A scored report
One composite score from 0 to 100, anchored to real calibration examples so a 70 means the same thing on every audit. No grading on a curve.
Coding agents do not fail randomly. They fail in predictable places: files too big to hold in context, names they cannot search for, missing docs that force them to guess, no tests to check their own work. The audit scores one repository across six weighted pillars, then hands you the exact, prioritized list of what to fix first.
Call fee credited toward your first engagement.
I take one audit at a time
Two-week turnaround from the day I get access. This is not a code-quality grade and it is not a security audit. The question is narrow and useful: how well can an AI agent work in this repo right now, and what is the fastest way to make it better.
One composite score from 0 to 100, anchored to real calibration examples so a 70 means the same thing on every audit. No grading on a curve.
Every pillar scored on its own rubric, with specific evidence: the files, the gaps, and exactly why each one slows an agent down.
Every fix ranked by leverage and tagged with pillar, impact, and effort. You get the short list that moves the number, not a transformation deck.
We go through the whole report together, live. You leave knowing what to do first, why it matters, and what it will change.
Documentation and Agent-Specific Readiness carry half the score between them, because they are the two things that most determine whether an agent starts a task oriented or lost.
README quality, architecture docs, and ADRs. Can an agent learn the project without reading every file? Quality over length: a lean top-level file that points to deeper docs beats one giant file that tries to hold everything.
The CLAUDE.md, nested guidance files, skills, and IDE configs an agent can actually use. We reward a thin harness with fat, judgment-carrying skills. Decorative files that exist only to look ready get discounted.
Linting configured and enforced, a CI pipeline, pre-commit hooks, code review, and clean secret hygiene. Zero guardrails scores near the floor; layered enforcement scores high.
Whether agents have a feedback loop to check their own work. Without tests, a human becomes the only reviewer, which kills the autonomy that makes agents worth using. We look at coverage and whether tests run in CI.
Whether code follows your language's conventions and can be found by natural-language search. Full terms beat abbreviations; descriptive filenames beat cryptic ones. Think code SEO for agents.
God files and average file size, weighted by whether they are the files agents actually need to modify. Lower-weighted on purpose: strong docs, tests, and guardrails can compensate for inherited size.
Most codebases that have never been tuned for agents land in the 40 to 65 range. That is not a failing grade. It means the foundations are there and the gaps are fixable, which is exactly what the roadmap is for.
Client details are fictional; the structure and the specificity are exactly what you receive.
Agents can do real work here, but slowly and with frequent correction. The code is well-organized and consistently named, so an agent can navigate it. What is missing is everything that lets an agent work on its own: no project-specific guidance file, thin test coverage (6% of source), and nothing running on a pull request. The gap is not the code. It is that the repo assumes a human is always in the loop. Close three gaps and this becomes a 76.
| Pillar | Weight | Score | Band |
|---|---|---|---|
| Documentation and Context | 25% | 61 | Can Work Here |
| Agent-Specific Readiness | 25% | 28 | Struggle |
| Guardrails and CI/CD | 15% | 44 | Survive |
| Testing and Verification | 15% | 35 | Struggle |
| Naming and Discoverability | 10% | 88 | Thrive |
| File Size and Complexity | 10% | 72 | Can Work Here |
| Composite | 100% | 58 | Agents Survive |
The full audit scores all six pillars in detail, lists every finding with evidence, and gives you the complete prioritized roadmap with effort tags. Projected after the roadmap: 76 out of 100, Agents Are Productive. Conservative.
The scoring engine was built metadata-first. It can run on your repository’s structure and configuration alone: file counts, directory layout, which config and doc files exist, and previews of the key ones. If granting source access is a blocker, the score runs on extracted metadata without me reading your proprietary source at all. I do not retain a copy after delivery, and I never delete tracked files, change permissions, or invent tooling. NDA-friendly either way.
It starts with a $300 readiness call, credited in full toward the audit if we go ahead. I take one audit at a time, so slots are genuinely limited: booking the call is how you hold one.
Call fee credited toward your first engagement.
Multi-repo and team-level engagements are scoped one to one. Every organization’s surface area is different, so there is no list price. We talk through what you have, what you are trying to unlock, and what a right-sized engagement looks like.
Not sure yet? Take the 10-minute readiness quiz for a free directional read before you book.
Take the quizStart with the readiness call: a hype-free conversation about where your codebase stands and whether the audit is worth it for you. If it is, the call fee comes off the price and we get to work.
Call fee credited toward your first engagement.
Tech Thaumaturge Weekly
Not ready to book? One useful letter a week on making codebases agent-ready. Practical, no hype.
Email only, one-click unsubscribe anytime. See our privacy page.