AI writes the tests. Mutation testing checks if they work.
9m read time

AI writes the tests. Mutation testing checks if they work.

AI-generated tests hit high coverage in seconds, but coverage only proves a line ran. Mutation testing proves the tests would actually catch a bug, and Chaos-MCP puts that loop inside your agent.

Give Claude or ChatGPT a function and ask for tests. Seconds later you have a tidy suite, green across the board, with coverage numbers that would make an auditor smile.

It feels like free productivity.

There's a catch.

Coverage measures the wrong thing ​

Code coverage tells you a line ran. It does not tell you a test would notice if that line were wrong.

Those are different claims, and AI is unusually good at the first and unusually bad at the second. Models recognise patterns. The parrot writes tests that mirror the implementation in front of it, line for line.

The tests confirm the code is consistent with itself. Whether the code is correct is a question they never ask.

A green suite that proves nothing ​

Here is a function. Four lines, nothing exotic.

js
function shippingCost(weight) {
  if (weight > 10) {
    return 15
  }
  return 5
}

Ask an AI for tests and you get something like this:

js
test('heavy parcels cost more', () => {
  expect(shippingCost(20)).toBe(15)
  expect(shippingCost(2)).toBe(5)
})

Run coverage. 100%. Both branches executed, every line touched, the report is a wall of green.

Now change one character in the source: > becomes >=. At a weight of exactly 10, the original returns 5 and the mutant returns 15. That is a real bug, the classic off-by-one at a boundary.

Re-run the tests. They still pass.

Nothing in the suite ever calls shippingCost(10). The two tests picked 20 and 2, comfortably away from the edge, so the branch is "covered" while the boundary that the branch exists to enforce is never checked. The exact place a developer is most likely to get it wrong is the exact place the tests are silent.

That is the single most common shape of AI-generated test, not a contrived one: exercise both branches with values from the safe middle, hit 100% coverage, and never touch a boundary.

Worse, some generated tests assert almost nothing. Call the function, check the result is defined, move on. Coverage counts that line as covered. It has verified nothing at all.

Why AI misses the bugs that matter ​

AI has no intuition. It doesn't know which mistake a developer is likely to make. It doesn't know your business rules unless you spell them out. And it rarely invents the odd scenario on its own.

The classics:

  • a > where you meant >=
  • a flipped boolean
  • a forgotten null check
  • a missing validation
  • an edge case around empty collections or boundary values
  • an exception path nobody tests

The generated tests look great. They mostly walk the happy path.

This is the can't-spot-the-bug problem wearing a green checkmark. The instinct that says "test the value right on the boundary" comes from having been burned by that boundary before. The model was never burned. It has read a million tests and reproduces their average shape, and the average test does not probe the edge.

The same blind spot, twice ​

Here's what makes it worse.

The AI writes the production code. Then the same AI writes the tests.

Both come out of the same reasoning. Both make the same assumptions. If the model misread the requirement and used > when the rule was "10 and over", it will write the production code with > and then write tests that expect > behaviour. The bug and the test that should catch it are drawn from the same mistaken premise.

You end up with a green suite that guards a bug instead of catching it.

That's how a lava layer forms with a passing CI badge on top. Nobody revisits code that is green and covered. The wrong assumption sets, and the tests are the concrete poured over it.

Mutation testing asks a better question ​

Mutation testing flips the target. Instead of testing your code, it changes your code, on purpose, in small ways, and checks whether your tests scream.

A mutation tool makes edits like:

  • > becomes >=
  • == becomes !=
  • && becomes ||
  • + becomes -
  • true becomes false
  • a return value gets swapped

Each edit is a mutant: a copy of your code with one deliberate fault. After each one, the whole suite runs again.

If a test fails, the mutant is killed. Good. Your tests noticed the fault.

If no test fails, the mutant survived. That means your tests would not have caught that fault if a developer had introduced it by accident. A surviving mutant is a bug your suite is blind to, demonstrated in code you can read.

Run it on the shippingCost example and the > to >= mutant survives, pointing straight at line 2. The tool is telling you, precisely, that your boundary is untested.

What a mutation score actually buys you ​

Kill 11 mutants out of 12 and your mutation score is 92%. That single number carries more information than any coverage percentage.

Where coverage says "this code ran", mutation testing says "these tests are strong enough to catch a mistake". The second sentence is the one you actually care about.

100% coverage with a poor mutation score is entirely possible, and common. The reverse almost never happens: to kill mutants you have to write assertions that pin behaviour down, and those assertions drag coverage up as a side effect. Chase the mutation score and coverage comes along for free. Chase coverage and you get neither.

One honest caveat, because the number isn't magic. Some mutants are equivalent: the change produces code that behaves identically for every possible input, so no test can ever kill it. Those aren't gaps, and a good workflow lets you mark them as such so they stop dragging the score down. A mutation score below 100% is a to-do list rather than a verdict, and part of the work is deciding which survivors are real.

A perfect mutation score still says nothing about the function underneath the tests: whether it's needlessly nested, duplicated three other places, or reaching into a module it shouldn't. That's a separate dashboard, and it needs different instruments.

The tools already exist ​

Mutation testing isn't new, and you don't have to write the mutating yourself. Every major language has a mature engine.

StrykerJS is the one for TypeScript and JavaScript, and it's worth understanding how it works because the others follow the same pattern. Stryker parses your source into an abstract syntax tree, then walks the tree and generates a mutant for every place a mutator applies. Its mutators cover arithmetic (+ to -), equality (== to !=), boundary operators (> to >=), logical operators (&& to ||), booleans, string literals, array declarations, optional chaining, and more. For each mutant it runs your existing test runner, vitest, jest, mocha, whatever you already use, and records whether a test failed.

The clever part is how it stays affordable. Running your entire suite once per mutant would be brutal on a large project, so Stryker first collects coverage data and then, for each mutant, only runs the tests that actually reach that line. A mutant in a function nothing tests is reported as "no coverage" without wasting a run. That optimisation is the difference between mutation testing finishing over a coffee and running all night.

The other languages give you the same thing:

  • cosmic-ray for Python, driven from your pytest or unittest suite.
  • cargo-mutants for Rust, which can list every planned mutant without running a single test, so you know the size of the job up front.
  • Infection for PHP, on top of PHPUnit.

They all report the same core number, the mutation score, and they all point at the exact line and the exact mutation that survived.

So why isn't this already everywhere? Cost and friction. Running the suite per mutant is expensive, and wiring an engine into each project, learning its config format, and parsing its output is real work. That friction is the gap most teams never cross.

The workflow that actually works ​

AI is genuinely good at the first draft of a test suite. Let it do the repetitive part. Just don't let it be the final word.

A workflow that holds up:

  1. Let the AI generate the first tests.
  2. Read them. Do they assert behaviour, or just execute lines?
  3. Run mutation testing.
  4. Improve the tests until the mutants that matter are dead.
  5. Add the edge cases and business scenarios the model never considered.

That keeps AI as an accelerator and out of the quality gate. Steps 1 and 5 are where the model helps most, and step 3 is the one that keeps the whole thing honest. For how to set up such a loop end to end, see the guide on agentic coding.

The prompt was never the spec, and a green test run isn't one either.

Putting the loop inside the agent ​

Reviewing survivors by hand is the honest version. It's also slow, and the thing that generated the weak tests is sitting right there, able to fix them.

So I built a tool for that: Chaos-MCP.

It's an MCP server that removes exactly the friction I just described. Your agent calls one tool. Chaos-MCP detects the language, picks the right engine (StrykerJS, cosmic-ray, cargo-mutants or Infection), runs the mutants in an isolated sandbox (your real workspace is never touched), and parses the output back into a ranked to-do list. The survivors come back ordered by severity, each with a note on why the gap is dangerous and what kind of test would kill it.

The loop it enables:

  1. Audit a file. Chaos-MCP returns the surviving mutants and a runId.
  2. The agent writes tests aimed at those survivors.
  3. Verify with the same runId. It re-runs only the previous survivors and reports which are now dead.
  4. Repeat until the file is clean.

Same idea as the manual workflow, except the agent that wrote the tests is the one made to prove they catch bugs. There's a second tool, triage_test_coverage, that ranks a whole directory weakest-first so you know where to start, and a gate mode that returns a pass/fail against a minimum score for CI to block on.

One honest note: it's pre-release. Not on npm yet, install from source. The repo is public and the README walks through setup.

I ran it against a codebase this week and pushed a file to 100% coverage with every mutant killed. The coverage number was the easy half. Chaos-MCP is what told me the tests behind it actually meant something.

Because that is the whole point of a test.

Not to go green.

To stop a bug before your users find it.

(31 of 40)
1You don't have an AI problem. You have a process problem.2Why you should never ship code you don't understand3Stop copy-paste engineering4The lava layer: why AI code is slowly petrifying your codebase5The brilliant parrot problem: what AI actually does when it 'thinks'6The prompt is not the spec7The bureaucracy of bots: why we are checking the checker8The day Claude deleted my production database9The arms race for your trust: Mythos, Cyber and the security hype10Stop letting your agents write Markdown11Your agent's suffering is your technical debt speaking12You can't spot the bug if you didn't write the code13One in four: the security debt nobody's counting14Your 10x developer is gated by a 0.1x pipeline15Benchmarks said frontier. Developers said "dumb."16Caveman vs context-mode: small mouth, or smaller room?17Code churn is the lava you can still measure18The ceiling is made of concrete19The token-saver tax: walking back my Caveman advice20Even the malware is AI slop now21ThePrimeagen was right22Tokenmaxxing is what happens when you measure the wrong thing23They just asked the bot nicely: your support agent is the attack surface24Speed got cheap. Judgement didn't.25Your coding agent has no world model. You built it one.26The Ferrari has a limiter: a day with Claude Fable 527The off-switch was never yours28An open MCP server is worse than an open database29The most resilient job is eating its seed corn30The off-switch works both ways now31AI writes the tests. Mutation testing checks if they work.32How to get better at reading code: a practice routine33Learning to program in the age of AI: what I would learn first34Who is responsible for AI-generated code? You are, and 2026 wrote it down35When not to use AI for coding: the tasks I still do by hand36Hiring junior developers in 2026: the collapse is a hiring decision37Software estimation with AI: the typing was never the estimate38Slopsquatting: checking that the package exists is not a defence39Eight reports, one bar: judging an AI safety sprint40Nobody's agent went rogue