🌿freegardner

Synapse

Reading AI agent failures to choose harness or weights

11 Oct 2026 · via Rss.arxiv

Reading AI agent failures to choose harness or weights
AI-generated image

Reading AI agent failures to choose harness or weights

A long-horizon AI agent — one that plans a trip, books the parts, checks the result, and delivers a finished itinerary — can be made better in exactly two places. You can change the harness: the code and text wrapped around a frozen model, its system prompts, memory tools, checkpoints, output gates, and the tools it is permitted to call. Or you can change the weights, fine-tuning the model itself on trajectories of past behavior. Each lever has its own research literature. What has been missing is a rule for deciding which one to pull. A new line of work, described in a paper available as Improving a long-horizon LLM agent, builds that rule empirically and then tests it twice. [1] The paper’s central move is to stop treating failure as a single quantity. Instead, every failed trajectory is labeled by the first signal that fires. Unbacked assertions in the plan count as one class. The harness’s own interventions getting in the way count as another. Loops and exhausted step budgets count as a third. A delivered plan that is simply poor counts as a fourth. The first three are process failures — the system tripped over its own machinery. The last is a content failure — the system ran fine and produced something inadequate. In the paper’s reading, that distinction, computed deterministically from process signals and the scorer, is what tells you which lever to use.

The claim the paper tests is precise: harness evolution removes the process classes and leaves the content class in place. The behavior the harness instills can then be trained into the weights. And the content class is what weight training is for. The experiments run on DeepPlanning, a travel-planning benchmark with deterministic programmatic scoring, using 48 development tasks for every decision and 72 held-out tasks that no optimization loop ever saw.

What the harness can fix

The measurement protocol comes first, because the per-rollout score on this benchmark is noisy. Every configuration is reported as a four-rollout mean against fresh anchors, and a fresh anchor follows every accepted change. That discipline matters: without it, the difference between a real gain and a lucky run is invisible.

Against that baseline, a self-evolving harness loop — in which an LLM proposes edits to its own harness and a judge verifies them — lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30, and of Qwen3.5-9B from 0.32 to 0.44. [1] The gains are real, they survive held-out validation, and they were achieved without touching a single weight.

Where those gains land is the interesting part. For the 4B model, held-out delivery rises from 55% to 90% under the evolved harness, while loop failures fall from 64% to 37% of trajectories on the development tasks. In both cases, what the loop leaves behind is the content class: poor plans that were delivered on time and without process errors.

That rise is not a regression. It is arithmetic. When you remove the process failures that were masking the content failures, the content failures become a larger share of what remains. The harness did its job. The question is why it stops there.

Why the loop cannot see what it needs to fix

Component ablations answer that question with unusual directness. The compliance-enforcing components of the harness — the exit gate, the checkpoint, the hard blocks — are mostly within noise. The capability-granting components — a state ledger, a distilled lessons file, invented tools — carry the gain. The loop is not winning by forbidding bad behavior. It is winning by giving the model more to work with.

Reading AI agent failures to choose harness or weights (Image 1)
AI-generated image

But the evidence package that the proposal model reads leads with gate and checkpoint metrics and never aggregates the scorer’s failure composition. The loop optimizes what it can see, and what it can see does not include the cause of the content failures. This is, in the paper’s reading, a structural limit, not a tuning problem. The proposer is shown the wrong dashboard.

That finding has a wider implication for anyone building self-improving systems. A loop is only as good as the signal it is handed. If the signal describes process health but not output quality, the loop will converge on process health. In the paper’s view, this is not a flaw to be patched with better prompts. They treat it as a boundary that tells you when to switch levers.

What the weights can learn

LoRA adapters trained on trajectories collected under the evolved harness transfer to held-out tasks. Under the original, seed harness, the adapter alone adds 0.13 for both model sizes — matching or exceeding what harness evolution had delivered on its own. For the 9B model, the adapter alone reaches the evolved-harness level, meaning that at that size the two levers substitute for each other rather than compounding.

Each adapter moves its own failure class. The 4B adapter raises delivery and cuts loops. The 9B adapter cuts content failures (F4) from 23% of development trajectories and 28% of held-out trajectories to 7% and 5% respectively. [1] That is the class the rule assigns to weight training, and the adapter moves it.

The controls make the result harder to dismiss. A placebo adapter trained on answer-shuffled trajectories falls 0.16 below the base model on development and 0.23 below on held-out, which means the gain lives in the trajectory content rather than in the act of fine-tuning. The effect is not an artifact of adapter rank or training volume.

A rule applied twice

The paper’s contribution is not a single technique but a diagnostic procedure that runs in two stages. First, read the failure composition. If the dominant class is process failure, evolve the harness. If it is content failure, train the weights. Second, read what the accepted edits changed. Train in the gains that changed what the model writes. Keep the harness for the gains that changed what the model sees.

The WebArena-Lite transfer test shows why the second reading matters. The same loop transfers there, adding 0.09 across 117 unseen tasks, with one interface edit — keeping the two most recent full pages in context instead of one — accounting for most of the gain. But the gain lives in what the model sees — two pages of history instead of one — and adapters trained on those trajectories do not add to it. The harness changed the model’s input. There was nothing to internalize into the weights because the improvement was never about what the model produced.

In the paper’s reading, that is the cleanest demonstration of the rule. The same loop, the same training method, two benchmarks, two different answers about where the gain should live. On DeepPlanning, the harness fixed process failures and the weights fixed content failures. On WebArena-Lite, the harness fixed perception and the weights had nothing left to learn.

Generality across models and families

Reading AI agent failures to choose harness or weights (Image 2)
AI-generated image

The seed measurements generalize. Across nine evolution lines on eight models from six families and two benchmarks, the initial failure composition separates the five lines where the harness lever paid (+0.05 to +0.14) from the three where it paid nothing. The taxonomy does not require running the full loop to be useful. Reading which class of failure dominates tells you in advance whether harness evolution is worth attempting.

The paper also positions its taxonomy against the existing landscape. Those taxonomies are organized by the kind of error and require an LLM or human annotator. This one is organized by the lever that removes the error and is computed by a deterministic checker. The difference is, in the paper’s reading, operational: a taxonomy that tells you what went wrong is useful for post-mortems. A taxonomy that tells you which intervention will fix it is useful before you spend the compute.

What the result actually claims

It is worth being precise about the scope. The diagnose-then-intervene rule was validated on two benchmarks and eight models. The specific numbers — 0.16 to 0.30 for 4B and 0.32 to 0.44 for 9B on held-out tasks, +0.13 from the adapter alone under the seed harness — belong to those settings. The claim that survives the noise threshold is structural: the right lever can be read off the failure composition, and the composition is computable from signals the system already produces.

In the paper’s view, harness evolution is not always worth trying, and weight training does not always stack with it. On the 9B model, the two levers substitute. On WebArena-Lite, the adapter added nothing. Those are negative results within the paper’s own framework, and they are reported as such. The rule is not “do both.” The rule is “read first, then choose.” What makes the work notable is that it treats the choice between runtime and training as an empirical question with a measurable answer, rather than a matter of architectural preference. The harness and the weights are not competing philosophies. They are two instruments, and the failure composition tells you which one the system in front of you actually needs.


Sources

  1. arXiv — Paper

Mentioned organisations (context, not sources)

← back to the garden