Nobody asked for it: what coding agents add when you don't decide
13m read time

Nobody asked for it: what coding agents add when you don't decide

Three days after taking the AI tells off this site, I wanted to know what they look like in code. So I gave 23 models the same three TypeScript tasks, 266 runs in five agent harnesses. Most of them add fields nobody asked for, every model line has a handwriting you can recognise, and two of the newest GPT models leave the data model as the prompt described it.

On Monday I took the AI tells off this site. The rule I ended up with was simple: a tell is a choice nobody made. The look was fine. The parts that nobody could give a reason for went.

One question stayed with me. Design gets screenshotted and passed around. Code mostly does not. If an agent fills every gap in a design brief with the median, what does it put in the gaps of a code brief, where nobody is looking?

So I measured it.

How I measured it ​

Flow of one run: prompt and credentials into the agent, then transcript, smoke test, labels and pairs

The idea was to give every model the same open brief and then count what it filled in. That sounds simple, and most of the work went into keeping it fair.

Three tasks with the gaps left open ​

Three small TypeScript tasks, each a few sentences long: a REST API for bookmarks, a CSV-to-JSON command-line tool and a client library for a flaky, rate-limited HTTP API. Each prompt names TypeScript and ends with "Include tests." Everything else is left open on purpose: framework, storage, test runner, argument parser, retry strategy. This is the whole API prompt:

text
Build a small REST API in TypeScript for storing bookmarks. A bookmark has a URL, a title
and optional tags. Support creating, listing, fetching one, updating and deleting
bookmarks, and persist them so they survive a restart. Include tests.

TypeScript is fixed because language choice is a known bias of its own. Left open, the first Gemini run I tried wrote Python.

Every agent as shipped, in a clean room ​

Five harnesses: Claude Code, Codex, Antigravity, OpenCode and Freebuff. 23 models in total, each at the settings its harness ships with. Antigravity puts the reasoning level in the model name, so there I had to pick one: medium, and low for the Pro model, which has no medium.

Each run got a fresh, empty directory and a fresh process, and nothing else. The agents normally read a lot from your home directory: config, rules, memory, plugins, MCP servers. So every harness ran from a clean home that held only its login. Without that, Antigravity loaded my own MCP servers and Claude Code my Gmail and calendar connectors, and the runs would have measured my setup instead of the model.

Four of the harnesses run headless and report in JSON. Freebuff only has a terminal interface, so a small driver types the prompt in tmux, reads the screen and waits until the agent is done. When an agent asked a question before building, it got the same answer every time: "Choose whatever you think is best and build it." Every run had 45 minutes.

A colleague (thanks, Jelle!) ran GPT-6 Astra and GPT-6.1 Sol on their own account, with the same harness and the same prompts, and sent back the results.

Checking which model actually answered ​

The name in a model picker is a claim. For every run I read the model that actually answered from the agent's own transcript, session file or database, and flagged any run where the two disagreed. That turned out to matter.

Does it work, and what did it choose ​

Share of runs that worked per model: 13 models at 100%, Solar Mini 4 and Nemotron 3.5 lowest at 22%

Every project then went through a smoke test, the way a person would try it: install it with the package manager it chose, build it if its start script needs a build, start it and call the API, convert a CSV with and without a header, run the library's own tests and type-check it with its own TypeScript. When a project failed, I read why. Several times the fault was in my test, and I fixed the test and reran it for every project.

I read the choices themselves from the source code alone, leaving the tests and the README aside: which framework is imported, where the data goes, which test runner is installed, how the retries back off. Each of those is an axis with a label per run. Whenever a rule mislabelled a real run, that run became a test case before the rule was fixed, and every run was labelled again.

Finally, the comparison. For each axis I took every pair of runs and counted how often the two made the same choice: pairs from the same model, and pairs from two different models. If a model only follows the trade, both numbers come out the same. If a model has its own habit, its own runs agree more often.

All 23 models ​

Median minutes and tokens per run per model on log scales, from Sonnet 5.5 at 0.9 min to Solar Mini 4 at 39

Per harness, each model with the runs that worked and the median time per run:

  • Claude Code: Haiku 4.5 (14/15, 2.9 min), Sonnet 5.5 (15/15, 0.9 min), Opus 5.5 (9/9, 2.1 min)
  • Codex: GPT-5.6 Luna (11/15, 2.4 min), GPT-5.6 Terra (15/15, 1.7 min), GPT-5.6 Sol (8/9, 6.9 min), GPT-6 Astra (9/9, 3.8 min), GPT-6.1 Sol (9/9, 4.6 min)
  • Antigravity: Gemini 3.8 Flash (15/15, 3.6 min), Gemini 3.7 Flash (15/15, 2.0 min), Gemini 3.6 Flash (9/9, 7.2 min), Gemini 3.1 Pro (3/9, 1.1 min)
  • OpenCode: Big Pickle (9/9, 20.1 min), Muse Spark 1.3 (8/9, 2.6 min), Nemotron 3 Ultra (8/9, 9.2 min), Nemotron 3.5 Lightning (2/9, 22.4 min)
  • Freebuff: GLM 5.3 Flash (9/9, 19.2 min), Solar Mini 4 (2/9, 39.1 min), MiMo 2.6 Flash (9/9, 8.4 min), Solar Pro 4 (3/9, 6.0 min), DeepSeek V4.1 Flash (15/15, 5.0 min), GPT-6 Luna (8/9, 3.0 min), MiMo 2.6 Pro (8/8, 7.3 min)

A run worked when it passed the smoke test. The chart above puts time next to tokens for the four harnesses that report them: input plus output for a whole run, mostly cached context that the harness reads again on every step. Freebuff only shows the size of its context.

Every harness ran at its own default reasoning level. In Freebuff that means max for GLM, high for DeepSeek and GPT-6 Luna, and none for the Solar and MiMo models. Only in Antigravity did I pick one, as described above.

The runs took place between 28 September and 1 October 2026, on Claude Code 2.1.285, Codex 0.154.0, Antigravity 1.2.13, OpenCode 1.18.29 and Freebuff 0.1.4. Seven models got five runs per task, the rest three. MiMo 2.6 Pro has one library run fewer: its paid hour ran out before the last run could start.

Four models are missing from this overview. Ling 3.0 Flash and the free mimo-v2.5 in OpenCode returned only errors, in all nine runs each. Space Bunny Alpha in Freebuff failed twelve times out of twelve. Muse Spark 1.2 turned out to be answered by a different model on Freebuff's free plan, so there was nothing of its own to measure.

The whole experiment was cheap: about $13 at list price for all Claude runs, six points of a weekly Codex limit, Freebuff's free daily allowance, and nothing for Antigravity or the free models in OpenCode.

18 runs did not count: the provider returned errors and left nothing behind. The rest is 248 runs.

The field nobody asked for ​

Of 83 API runs, 69 added createdAt, 32 a tag filter, 27 a health endpoint and 27 a SIGTERM handler

The prompt says a bookmark has a URL, a title and optional tags.

In 69 of the 83 API runs, the agent added createdAt or updatedAt anyway. 19 of the 23 models did it, nearly all of them in every run. Further down the same list: filtering by tag in 32 runs, a health endpoint in 27, a graceful shutdown on SIGTERM in 27. Nobody asked for any of these.

Each of these can be a good idea. If you want to know when a bookmark was saved, createdAt is exactly right, and an agent that adds it saves you a step. A health endpoint belongs in anything that runs behind a load balancer.

The trouble is that here nobody decided. The field is there because bookmark APIs usually have one, and it looks like diligence, so it gets through review. Once it is in, it stays: someone sorts on it, migrates it and argues about what it means. Created according to which clock? Does editing the tags count as an update?

That is the code version of a UI tell. Whether it belongs is your call, as long as you know it is there.

Who left the timestamps out is the interesting part. GPT-6 Astra and GPT-6.1 Sol: 0 of 6 runs. The three GPT-5.6 models from the same vendor: 13 of 13.

The newer two still add plenty: all six of their runs include a SIGTERM handler. They kept the operational extra and dropped the field in the data model, and the field is the one that is hard to take back. Gemini 3.1 Pro and Nemotron 3.5 Lightning skipped the timestamps as well, but both wrote very little code overall, so that says little.

The layer every model shares ​

Nearly every run: strict TypeScript 245/248, no linter 243/248, own retry 79/82, backoff 77/82, fetch 73/82

Below the additions sits a layer that hardly varies at all:

  • Strict TypeScript in 245 of 248 runs.
  • No linter in 243 of 248. Five runs set one up.
  • A hand-written retry loop in 79 of 82 library runs, with no retry package, usually with exponential backoff and plain fetch.

And yes, three retries, because that is what everyone uses. These are reasonable choices, the trade's median.

Research on models outside an agent points the same way. A study of eight models found they prioritise familiarity and popularity over suitability when they pick a library or a language. LangChoiceBench found that when models pick Python, the choice is mostly "automatic or driven primarily by ease". I fixed the language, so the ease shows up in everything else.

The missing linter is the one I would act on. Nobody decided against it either.

Every model line has a handwriting ​

Agreement within one model against between models, per axis: equal for tsStrict, far apart for storage and tests

Where the ecosystem itself is split, runs of the same model agree with each other far more often than runs of different models: a median of 17 percentage points across the axes I measured. To check that this was not luck from three runs, I gave seven of the models five runs per task.

It held:

ModelHabit
GPT-5.6 LunaPlain node:http instead of a framework, 5 of 5 API runs. GPT-6 Luna: 3 of 3
Sonnet 5.5node:util parseArgs for the CLI, 5 of 5
Gemini 3.7 and 3.8 FlashExpress and jitter on the retry delay in every run, vitest for the API and the library
Haiku 4.5Express with jest for the API, every run
Opus 5.5node:sqlite for storage, 3 of 3

Storage per model family: Gemini, Opus 5.5 and Big Pickle choose SQLite, almost all others a JSON file

Storage shows it most clearly. 60 of the 83 API runs keep bookmarks in a JSON file. SQLite comes almost entirely from three model lines: Gemini in 15 of its 16 runs, Opus in 3 of 3 and Big Pickle in 3 of 3, plus one run of GPT-6 Astra.

Haiku and Sonnet, in the same harness as Opus, chose JSON files. So that preference belongs to the model, whatever tool drives it.

Luna's habit survived a version bump. The stack you get depends on which model happened to be selected, and a new model can change it without anyone changing a line of the brief.

Who asked first ​

Freebuff runs that asked first: DeepSeek 6/15, MiMo 2.6 Pro 3/8, MiMo 2.6 Flash 2/9, Solar Pro 1/9, the rest none

Only Freebuff lets an agent stop and ask before it builds. The other four harnesses ran headless, where nobody is there to answer. In Freebuff, four of the seven models asked at least once: DeepSeek in 6 of its 15 runs, MiMo 2.6 Pro in all three of its API runs.

They asked about exactly the gaps I had left open: which framework, how to store the bookmarks, which test runner, how to stay under the rate limit. And the options often came with a favourite already marked. MiMo offered "Express + JSON file (Recommended)", the two most common choices in the whole experiment.

So the question was more of a confirmation. Every one of them got the same reply, "Choose whatever you think is best and build it.", and what they built next still counts as their own default.

What the name does not tell you ​

The picker and the transcript did not always agree. Freebuff's "MiMo 2.6 Flash" answered as mimo-v2.5 in every transcript. Muse Spark 1.2, on the free plan, was answered by a different model altogether. Two free models in OpenCode never returned a working response all night.

And then there was GPT-5.6 Luna. In 2 of its first 9 runs it wrote five files outside its own working directory, noticed, and ran find across my whole experiments folder to locate one of them. It saw file names and read nothing else, but it is the only model in 266 runs that left its directory. None of the agents had approval prompts in the way, so nothing stopped it.

What this does not show ​

  • Model or harness. I tried to run the same model in two harnesses and could not: the free MiMo endpoint returned only server errors, and Gemini CLI no longer serves the free tier. What I have are contrasts inside one harness, Opus against Haiku and Sonnet, and two versions of Luna in two harnesses.
  • Quality. The smoke test checks that the code works, not that it is good.
  • Anything general. Three small tasks, two to five runs per model per task, counts and no significance tests, all in an empty directory. A study of 26,760 pull requests written by agents found they rarely add a dependency, in 1.3% of PRs, and draw on a diverse set of libraries. Real projects look different.

My own harness was not flawless either. OpenCode takes its project directory from $PWD rather than from where it is started, and for its first two runs it worked in the root of my measuring rig. A check I had added for files written outside the run caught it, and every run since then lands where it should.

What to do with this ​

Most defaults are fine. The shared layer is what I would pick myself, and more often than not so is a createdAt. A default only becomes a problem when you did not know you had it.

Decide before the agent starts. Whatever you leave open, the model fills with its own median, and what you meant is not what you specified. Fields, storage, the linter, the test runner: if you care, put it in your CLAUDE.md or in the prompt.

Review for additions as well as mistakes. Hold the data model against the request. Every field the prompt did not ask for gets a reason or goes.

Expect the stack to drift when the model changes. The handwriting belongs to the model line. Switch models and the next feature arrives with a different framework, a different test runner and different storage, unless the brief pins them.

Add the linter yourself. Nobody else will.

On the design side, a tell was a pulsing badge nobody needed. In code it is a createdAt nobody asked for. Keep it if you want it, as long as keeping it is a decision.