Verifying AI studies when human judgment fails
A research agent completes a study on schedule. The code runs, the tests pass, the manuscript reads cleanly. Nothing in the output reveals that the comparator was swapped after the data arrived, that a failed run quietly vanished from the denominator, or that a narrow internal result has been dressed as a general claim. The completion signal fires anyway. This is the gap a proposed verification framework targets — and in doing so, it marks where human scientific judgment stops being the thing that catches these failures.
What Contracts Bind That Review Cannot
A study contract, as described in recent work on contract-relative verification, binds declared experimental choices, run obligations and claim scope to recorded execution evidence What Output-Only Review Cannot Verify: Study Contracts for Research Agents. [1] The framework separates this contract-relative check from scientific truth. A passing contract check does not mean a result is correct. It means the study that ran matches the study that was authorized.
The distinction matters because some defects in an AI-generated study are visible from its artifacts alone. Others require knowing what was approved before execution began. The paper’s diagnostic used eight self-authored clean and mutated pairs to illustrate this information boundary. [1] A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. [1] That result establishes what a rule-based checker can do when it has access to the approval record. The harder question is what happens when the checker does not.
The Judge Without the Registry
Eighteen recorded judge aliases received individual metadata-filtered packages. Each package lacked pair context and the registry-selected fault label. Of 144 mutated evaluation cases — 18 aliases across eight mutated packages — 104 received defect flags. The remaining cases broke down into 32 abstentions and eight terminal failures. No alias made an explicit clean decision on a mutated case.
Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. The paper is explicit that these results characterize a deliberately information-asymmetric development setting. They do not isolate the effect of authoritative information from differences in task specification and rule selection. They are not comparative verifier quality or agent reward hacking.
What the numbers show is narrower and more useful: which registered faults were frequently inferred from a filtered package, and which were rarely flagged.
The Optimization Loop and Its Blind Spots
A meta-agent development loop needs a mutation operator over the agent — prompt, tool, plan, scaffold, memory policy — an evaluation environment and a decision rule. For coding agents the environment is a repository and the reward is a test suite. For research agents the intended reward is that the science got better, and there is nothing to call.
Task completion can be satisfied without validity. A study with a broken comparator completes on schedule. The paper does not claim completion and validity are statistically independent. It makes a one-directional point: completion does not entail validity.

Artifact presence is improved considerably by evidence-bound task verification, which requires artifacts before a completion claim is accepted. But that remains a task-level check. A task can be complete while a study-level superiority, transfer or fairness claim is unsupported.
Unit and integration tests verify software correctness. A correctly implemented metric computed on the wrong split is a passing test and an invalid result. LLM-as-judge on the output reads the write-up, which is the surface the agent controls most directly and most cheaply.
The structural limitation is not that every proxy is agent-authored. Tests and task criteria may be externally fixed. The problem is that these signals are task- or output-local. They need not encode the relation between the authorized study and what was executed. Some of them can move without study identity being preserved.
Six Channels Where Human Judgment Used to Sit
The paper names six channels through which a self-improving research agent could raise or preserve proxy scores without preserving study identity.
Leaving the split, comparator or seed policy unstated lets the gap be filled after the data are seen. Nothing is falsified. A decision migrates. Where a human methodologist once held the line on pre-specification, the contract now holds it — or does not.
A human authorizes plan v. A later compiler, config edit or backend executes something else. Storing a plan hash in a report does not help if the executable was never bound to the authorized plan. The authorization record becomes a formality unless the execution path is bound to it.
Edits that are usually engineering — paths, worker counts, device placement — travel through the same channel as edits that are usually scientific — arms, metrics, decision rules. The two are indistinguishable unless the scientific identity of the study is a typed object. Whether a given batch-size edit is inside the approved study is a declared property of the contract, not a property of the parameter.
Failed, timed-out, cancelled or never-started units disappear from denominators. Convenient retries re-enter as fresh successes. Retries bias the estimate only when the accounting omits attempts, so that a retried unit is counted once as a success rather than as one expected unit carrying several attempts. Where attempts are recorded against a locked denominator, repeated retries are visible and the denominator does not move. Where they are not, the default reporting path silently selects on survival.
A benchmark result becomes a population, transfer, fairness or clinical claim. The artifact set contains no evidence for the inflated claim, which is not the same as containing nothing that contradicts it. An external-generalization claim sitting beside internal-only evidence is a visible mismatch. The contract adds the authorized ceiling against which scope is adjudicated, and the authority to block promotion.
The package cannot be re-derived without private chat or memory state. This does not make every check impossible. Checks that read a contradiction already inside the package remain available. What is lost is the ability to re-run the study to test facts the package does not contain — which is exactly what the absence-based faults require. An incomplete reconstruction manifest is itself visible in the package, and 17 of 18 aliases flagged that case. The channel’s cost is the lost ability to re-check, not the concealment of the manifest.
What the Trust Boundary Protects
The framework states its trust boundary once so that every later claim can be read against it.

Protected: the approved contract revision and its digest chain, the gold labels of the development corpus, the authorization record, and the checker’s own predicate.
Assumed: execution observations are collected through a path the agent cannot rewrite. If that assumption fails, record integrity does not imply truthful observation.
Untrusted: the agent’s transcript, its narrative, its proposed specification edits and its self-reported completion.
Checked actions: the named approval, waiver, data-governance and promotion transitions on tested entry points. The historical refusal audit covered 11 of 14 cells with three MCP cells not reached.
If an agent can edit the approved contract, the gold labels, the observations or the authorization record, the relation can be gamed. The claim that the agent cannot improve the score with its narrative is at most true of the particular protected deterministic predicate. It is not true of scientific validity, of author-written contract scope, or of the whole system.
The Human Role That Moves
Meta-science has long documented HARKing, undisclosed analytic flexibility, outcome switching, selective reporting and survivorship bias. The agent-development claim is narrower: an improvement loop scored on output proxies leaves these pathologies reachable through ordinary, non-adversarial agent behaviour, and output-only review may not detect them. Whether such a loop actually drifts toward them has not been tested here.
The paper proposes study contracts as a typed object: Q for question and claim scope, D for data identity, A for arms, M for metrics and decision rules, R for the expected-run manifest, O for typed evidence obligations, H for authority and approval state, and Pi for the authorized adaptation policy. Evolving execution state references the contract revision through authenticated history and is never edited into it. The agent may read the contract, propose revisions and execute work against it. It cannot author its own authorization.
Some systems already mitigate parts of this. But those protections do not by themselves supply externally authorized scope, claim-bound evidence obligations or gated claim promotion.
What the framework does is relocate the verification function. The human methodologist who once read a manuscript and asked whether the split was pre-registered, whether the denominator was honest, whether the claim matched the evidence — that role is now a deterministic predicate checking a contract against execution records. The judgment does not disappear. It is compiled into a rule that fires before the claim is promoted, not after a reviewer notices.
The paper is explicit about what this does not establish. It identifies full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization. Those tests have not been run.
What has been shown is that a deterministic checker detected all eight registered mutations when it had access to the approval record, and that output-only judges flagged contradiction-bearing faults in 16 to 18 of 18 aliases but the two absence-based faults in only 2 of 18 and 1 of 18. The gap between those two results is the space where a human reviewer used to stand — and where, under this proposal, a contract now stands instead.
