🌿freegardner

Synapse

Harness Gains Mostly Come From Visibility Not Quality

09 Oct 2026 · via Rss.arxiv

Harness Gains Mostly Come From Visibility Not Quality
AI-generated image

Harness Gains Mostly Come From Visibility Not Quality

When a benchmark score goes up, the usual assumption is that the model got better. A preregistered study of harness search suggests the more useful question is whether the measurement got better, and whether anyone can tell the difference.

A harness is everything around a frozen model that changes what it is scored on: the prompts, reasoning switches, token budgets, retry rules and parsers. Harness search proposes a change to any of these and keeps it if it wins. The practice is widespread, and the gains usually look real. A gain that replicates on untouched rows shows that the selected harness beat its parent by more than selection noise. [1] That leaves three questions open, and a preregistered study has now answered them with uncomfortable precision.

The first question is visibility versus quality: did the answers get better, or did answers the parent never produced in readable form — because it ran out of tokens or never committed to one — become visible to the parser? The second is baseline adequacy: would a simple setting the search never tried do as well, and does the model beat trivial baselines at all? The third is evaluation validity: are the inputs, the gold labels, the split and the served model what they are said to be? Each question has a literature, and most of it works after the fact. What was missing was a study that fixes in advance, for a stated contrast, how much of a harness gain is visibility, tests the benchmark authors’ own protocol against searched harnesses with a margin, and injects known defects one at a time to ask whether replication reproduces them. [1].

The Experiment That Fixed Its Questions First

The study preregistered its design and froze it in advance. It uses three small models, three benchmark pools, a replication partition and a confirmatory test partition, an adaptive arm in which GEPA searches for harnesses, and six evaluation defects injected one at a time. All 47 primary endpoints, in four families, are reported with every deviation and addendum. Everything computed after the outputs existed is labelled as such. The results cover three small models, three pools and six defects, and nothing beyond them.

The literature search covered 2024-01 to 2026-10-03 and was capped at 200 web searches, and three of its six parts had to be resumed after a rate limit, so the statements are bounded. [1]. The study follows a case study of harness search around a frozen longevity language model. [1]. That earlier search scored every candidate through a caller that sent each benchmark row’s reference answer to the model as an earlier assistant turn. The rescue it selected replicated on a second shard through the same caller, so replication could not detect the defect: every call carried it. Trivial baselines, input audits and per-call prompt-token reconciliation did. After the caller was corrected, a preregistered re-evaluation found that the search’s starting configuration parsed none of its 556 prospective rows, and that a thinking-off setting the search had never tried was non-inferior to the rescue at less than half the token cost. [1].

This paper asks whether that one model, one benchmark and one defect is a pattern.

Where the Gain Actually Came From

Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells. In each of those cells, the gain came mostly from questions where the capped setting gave no readable answer. [1]. The study split the gain exactly into rows only the protocol parses, rows both parse, and rows only the capped configuration parses, and fixed a dominance statistic and a gain gate in advance. The split is not a post-hoc rationalization. It was the measurement plan.

The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons. It lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. The search found improvements over its own starting conditions. It did not find improvements over the simplest alternative. [1].

Harness Gains Mostly Come From Visibility Not Quality (Image 1)
AI-generated image

A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. The technique is not the study’s; it serves as an evaluation control, with a forced answer or a forced end of thinking. What this study adds is the test of the engine’s thinking-token budget as a harness setting against a capped configuration, with the comparison fixed in advance. [1].

The Defect That Replicated Itself

In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect’s distortion instead of revealing it. [1] The study injected six evaluation defects one at a time. For each defect and model it fixed one claim in advance: that replication through the same pipeline confirms the defect’s distortion. It then measured the false alarms of five audit checks. Other work injects faults with known ground truth and measures false alarms: tool bugs in agent benchmarks, tampering edits, reward-hacking channels and planted gains. [1].

On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline. That is the baseline adequacy question answered at its starkest: a model tuned for the domain cleared the bar that a constant label sets, and the others did not. [1].

What the Serving Stack Does at Fixed Weights

Serving stacks change outputs at fixed weights. Answers can flip while accuracy stays flat. A served route is not its model identifier. vLLM used to decide tying from the configuration alone and silently discarded a stored head; its fix is in the version the study used (vLLM 0.30.0), which loads the stored head. The study adds a measured instance of the mismatch, since the mechanism is known. [1].

The identity for a single configuration or between models is completion rate times conditional accuracy. Under a shared budget, thinking off beats capped thinking. Capped single-call evaluation of hybrid models has been called an artifact of protocol. Other work finds that unparsed answers hide correct ones, separates traces cut by a cap from traces that finished without an answer, reads a harness-evolution gain informally as completion rather than correctness on the twelve tasks that had been searched, and counts repaired abstentions for constrained decoding. The study adds a pre-specified dominance statistic and gain gate, applied prospectively to the contrast between a thinking-off protocol and a capped thinking configuration. [1].

The Baselines That Would Not Go Away

Several studies find, after the fact, that simple baselines match searched or evolved harnesses and complex agents. A preregistered non-inferiority test of a simple alternative exists for agent memory, and equivalence margins have been used for scaffolds and for compression. The study’s test fixes the margin in advance for the benchmark authors’ own protocol, against a derived rescue and four GEPA-selected harnesses. [1].

Harness and scaffold gains vanish against matched-budget test-time scaling and simple baselines. A gain can be completion rather than correctness. Thinking off can beat a thinking cap. Evaluation pipelines carry defects that move scores by tens of points. Each of these findings existed before this study. What did not exist was a single design that put them together, fixed the questions in advance, and reported every endpoint. [1].

What the Study Does Not Claim

Harness Gains Mostly Come From Visibility Not Quality (Image 2)
AI-generated image

The study does not claim that harness gains can be illusory, that thinking off can beat a thinking cap, or that agent and harness evaluations can be preregistered, since each has been shown or done before. It does not claim that harness search finds genuine quality gains. The scope is three small models, three pools and six defects. The findings do not extend beyond them. [1].

Prompt optimisers and harness search treat the text and control flow around fixed weights as the object of search. GEPA is the adaptive arm of the study. Earlier work has shown, after the fact, that such gains can be illusory: evolved harnesses and scaffolds do not beat matched-budget test-time scaling or simple baselines, a hardware-verification harness gains mostly completion, a search started from a defective seed fails, and self-improving loops tamper with or hack their evaluators. Liu et al. report that GEPA, started from its official GSM8K seed, which puts the answer before the reasoning, falls from 23. [1] The study’s question is narrower: where the gain of one protocol comes from, and whether a known defect would be seen, with the questions and decision rules fixed in advance.

The Image That Remains

A search that repairs its own broken starting points and still cannot beat the simplest setting it never tried is not a search that failed. It is a search that succeeded at the wrong thing. The harness improved. The score improved. The model did not. What improved was the pipeline’s ability to read what the model had already produced: answers that were there all along, waiting for a parser that could see them. [1].

The gain was real. It was just not the gain anyone was looking for.


Sources

  1. arXiv — Paper

Mentioned organisations (context, not sources)

← back to the garden