~/blog/guide/ai-quality
AI and code quality
Generation got cheap. Judgement did not. On tests, churn and the bar that stays where it was.
What this guide covers
Generating code now takes seconds. Judging whether it is correct, whether it fits the system, and whether it will hold up costs exactly what it always did. That gap is the subject of this cluster, because nearly every quality problem with AI code turns out to be a variant of it.
I write this from practice. I use agents daily and they produce real work. Which is exactly why the friction stands out: output grows, review capacity does not, and everything you skip judging comes back later with interest. Speed got cheap, judgement didn't works that through on what large-scale orchestration does to your review.
It starts before the code exists
Most bad AI code traces back to a bad brief. The prompt is not the spec is the whole argument: a vague intention is not requirements, and the model will happily fill the gaps with something plausible. The sequel is the spec you didn't read, where letting an agent write the spec is fine and forwarding it unread to a second agent is where it goes wrong.
The same holds for how things look. Leave the design brief open and the model picks the median. The AI tells on my site holds this site against Anthropic's own list of AI design defaults, and shows how a test can pin a default until it passes for a decision. Afterwards I measured the same thing in code: 23 models got the same open brief, and most of them added fields nobody asked for.
There is also a signal in how hard the agent is finding it. Your agent's suffering is your technical debt speaking: when a model keeps stumbling in one corner of your codebase, that corner was already confusing before it arrived.
Tests that prove something
The obvious reflex is to have the AI write the tests too, and let quality take care of itself. Almost. Coverage only proves a line was executed, never that anyone would notice it being wrong. And a model that makes a reasoning error in the code will cheerfully repeat that error in the test. In AI writes the tests, mutation testing checks if they work I use mutation testing to make that false confidence measurable: sabotage the code and see whether a single test goes red.
A suite can also stop proving anything one catch block at a time. Empty catch blocks in AI code is about the // ignore that gets an agent past ESLint's recommended config, the test step that can no longer fail because of it, and the lint rule that sees through the comment.
Review is the other gate, and it changes when the author is a model. How to review a pull request an AI wrote is what I look for, in what order, and which of your old review instincts stop firing.
What happens to your codebase
Quality reaches past the individual PR. At codebase level three patterns keep showing up.
The first you can measure: code churn, lines rewritten within two weeks because the first version was never really finished. The second is quieter: the lava layer, code nobody understands anymore and therefore nobody dares touch. The third is the one people laugh at until they grep for it: twenty ways to format money in one repo, because an agent with no memory of your helpers writes a fresh one every time.
Measuring AI code quality is how you put numbers on all three, and why three of the obvious instruments hand you a headline figure that flatters AI code instead of catching it. When the diagnosis is in, the consolidation pass is the treatment, and a cleanup only holds if the next session cannot re-add what you just removed. Quality ratchets are how you enforce a standard your repository already breaks everywhere: freeze what exists, fail on what gets added.
Every one of those instruments measures code that runs. What a dependency graph finds that your coding agent misses covers the other half: the exported function nothing imports, the safety net that was tested and never wired in, and why a defect living between files stays invisible to anything reading one file at a time.
The measuring instruments get asked a neighbouring question they cannot answer on the spot, which is whether the model you just swapped in is the reason things feel worse. Comparing coding models without a benchmark is the harness for that one, and the reason your instinct to re-run the prompt and diff it proves nothing.
The bar that stays where it was
Which is why this cluster keeps landing on the same standard. Never ship code you don't understand: if you cannot explain a piece of code to a colleague without pointing at the AI, it does not belong in your repo. And keep your reading muscles trained, because you can't spot the bug if you didn't write the code. Debugging instinct is built from thousands of rounds of writing and reading your own work, and it evaporates faster than you expect.
None of these rules are new. They held when code was typed by hand. AI made them more urgent, because the volume at which you can break them has exploded. Typing speed used to be a natural cap on how much poorly understood code could land in your repo per week. That brake is gone and nothing automatic took its place. The replacement has to come from your process, which is what the posts in this cluster try to supply.
A bar is only as steady as the person holding it, and that is the half of the story you cannot see in yourself. Eight reports, one bar puts a number on it: one public rubric, eight submissions read back to back, and a script afterwards showing that my own bar for clarity had slipped three quarters of a level along the way.
Where this touches the workflow
Quality starts upstream, in how you drive an agent. A proper spec, small steps and review as a fixed gate prevent more misery than any linter. The full working method lives in the agentic coding guide, under related topics below.
Below are three starting points, then every post in this cluster, newest first.
best entry points
- Why you should never ship code you don't understand
The core rule of this whole cluster. If you read one piece, read this.
- How to review a pull request an AI wrote
Judgement in practice. What to look for in a pull request an agent wrote, and why green CI says nothing about it.
- You can't spot the bug if you didn't write the code
What happens to your reading skill when AI takes over the writing.
all articles in this topic
Nobody asked for it: what coding agents add when you don't decide
Three days after taking the AI tells off this site, I wanted to know what they look like in code. So I gave 23 models the same three TypeScript tasks, 266 runs in five agent harnesses. Most of them add fields nobody asked for, every model line has a handwriting you can recognise, and two of the newest GPT models leave the data model as the prompt described it.
The AI tells on my site, and which ones I chose
A list of ten tells of AI-generated UI went round over the weekend, and item seven describes this site down to the fonts. Anthropic lists the rest of its look as a default in its own design skill. Asking which parts had a written reason separated my own choices from what an agent had added, and those additions went.
Eight reports, one bar: judging an AI safety sprint
I judged eight submissions at the Apart Research AI Incident Response Sprint. Keeping my own bar in the same place from the first report to the eighth took an apparatus, and the apparatus caught me dropping it three quarters of a level.
Empty catch blocks in AI code: the comment that silences your linter
AI-generated code swallows errors in catch blocks that hold nothing but a comment, and ESLint's recommended config waves them through because of that comment. What I found in my own repositories, why a model writes it, and the rule that sees through the comment.
What a dependency graph finds that your coding agent misses
I pointed a scanner at an internal app that agents had been working in for months. It took under two seconds and produced four tickets. Not one of the defects was in a file, which is exactly why nothing reading files had found them.
Quality ratchets for AI code: baselines that only tighten
A lint baseline lets you enforce a standard you cannot fix today. Existing violations are grandfathered, new ones fail the build, and the file can only shrink. How PHPStan, ESLint, detekt and Sonar do it, and the four ways a ratchet quietly stops working.
Did the model get worse? Comparing coding models without a benchmark
You swapped models, something feels worse, and you have nothing to point at. Why re-running the prompt and asking a judge model both fail, and the small boring harness that answers the question.
Measuring AI code quality: the dashboard beyond coverage and mutation testing
Coverage proves a line ran, mutation testing proves a test would catch a bug, and neither tells you if the code is quietly getting harder to change. What complexity, clone detection and architecture fitness functions actually catch.
Cleaning up AI code: the consolidation pass your codebase is waiting for
Churn, duplication and the lava layer are diagnoses. This is the treatment: how to clean up AI-generated technical debt with a scheduled consolidation pass, and which parts the agent can do itself.
Your codebase has twenty ways to format money
AI code duplication is quiet: an inherited codebase with twenty inconsistent money and date formatters. How near-clones form, how to spot them, and the CLAUDE.md rules that stop them.