AI Agents Build Code But Miss Spec Details
When the Agent Delivers More Than It Can Account For
Pei Yang and colleagues have described a benchmark called Zero2Repo in a paper on arXiv, submitted on 29 September 2026. [1] The setup is deliberately austere: an agent receives a product requirements document, an interface contract, and an empty workspace. It must deliver a complete repository in the project’s native ecosystem. No starter code. No half-finished scaffold. No human pair programmer glancing over its shoulder. The rules are strict in a way that matters for the result. The agent works in an isolated container with no route to the code-hosting sites, installing or importing the upstream implementation is blocked, and the acceptance tests stay sealed until the agent makes an explicit submission. The reward is binary — every hidden test must pass, and no language model judge is allowed an opinion. [1]
The task itself is validated before it is ever used. A reference implementation adapted from the upstream project must pass the full acceptance suite, and adversarial validation must show that the same suite rejects incorrect implementations. If the tests cannot tell a right repository from a wrong one, the task is sent back for revision. [1] Only then does the benchmark report a result that deserves slow reading: even on eleven tasks drawn from repositories that frontier models have very likely seen during training, the strongest of the three frontier model-harness setups — GPT-6 Astra — solves only ten of the eleven; Claude Opus 5.5 solves nine, and Grok 4.7 High solves seven. [1] Every failing submission passes between 89.9 and 99.4 percent of the hidden tests. And for the two strongest agents — GPT-6 Astra and Claude Opus 5.5 — 67 to 100 percent of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem. [1]
The agent built the house. It poured the foundation, framed the walls, wired the electricity, installed the plumbing. It missed one clause buried in the spec — a low-frequency rule, a corner case, a thing that happens once in a thousand runs. The house stands. The house fails inspection. The failure is not in the walls. It is in the reading.
The Specification as a Document Nobody Reads the Same Way Twice
Consider what a product requirements document actually is. It is not a blueprint. A blueprint tells you where the load-bearing wall goes and what gauge wire to run. A PRD tells you what the software should do when a user clicks a button at 3 a.m. on a leap day while the database is mid-failover. The information density is uneven by design. Some requirements are stated once and never repeated; others are implied by the shape of the interface contract; others are buried in a sentence that begins with “Note that” and ends with a semicolon. A human engineer reading that document brings a lifetime of pattern recognition about which clauses matter and which are throat-clearing. The agent brings something else — a capacity to hold the entire document in working memory without fatigue, and a corresponding inability to know which clause is the one that will break everything.

Zero2Repo’s authors built a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Every requirement has to correspond to runnable behavior of the pinned project, and each environment has to rebuild from its recipe. Public names and repository identifiers are neutralized before a task is released. The current release contains Python, TypeScript, Go, and C++ tasks. The pipeline and harness make no language-specific assumptions. That is the point: the failure mode is not a syntax error or a missing library. It is a misreading of intent. It did not understand which parts of the document were load-bearing.
Where the Failures Actually Live
The paper’s failure analysis is the part worth reading twice. Every failed hidden test was read individually and sorted into one of five modes. One family keeps recurring — the long-tail edge-case rule: a rule that is stated in the specification, whose surrounding feature works, but whose rare branch does not — empty input, an unknown server state, an option combination, a retention or exclusion rule. Another is a single omission that breaks a test group: one root cause, usually an interface or a convention shared with the wider ecosystem, that fails many tests at once. The shares shift with the configuration — in one of the three setups, single omissions accounted for most of the failures — but the same causes keep reappearing. Then a non-functional failure in which the core semantics of a whole component are wrong, environment-dependent behavior that hinges on file descriptors, terminals, encodings, TLS or the Git working tree, and, at the tail, the occasional uncaught crash. [1]

The shape of that list is diagnostic. Under all-or-nothing grading, none of these failures are dramatic: no agent timed out, none broke on the build, none failed because a package was missing. Familiarity with the source is visible in the code itself. On one parser task, 71 percent of the normalized lines in the Claude Opus 5.5 implementation are identical to the upstream project, and 26 of 34 long comments match verbatim. The agents know these repositories well enough to reproduce their main functionality. They fail on the last few percent. And the misses cluster more than the spread suggests: one task in the set is failed by all three configurations, with four of its hidden tests falling for every one of them.
The Role That Was Never in the Job Description
Here is what Zero2Repo describes. The agent can build a repository from scratch. It comes within a few percent of every hidden test. What it cannot do — what no current system does reliably — is exercise the kind of judgment that a senior engineer applies without noticing. The judgment that says: this clause is the one that will bite us in production. This convention matters because every consumer of the library depends on it.
That judgment is rarely written down explicitly, and the paper is careful about what its numbers mean. It does not read the failures as a lack of knowledge; the agents clearly know these codebases. It reads them as targets: each missed rule is a concrete, addressable gap. The rule was stated, and by the paper’s own account every failure was detectable before submission from the public specification alone — none of the agents checked for it. Building the artifact and knowing which part of the specification is load-bearing turn out to be different skills — and the benchmark’s numbers suggest the second is not closing as fast as the first.
