🌿freegardner

Synapse

Correct Answers Invalid Traces Reveal Broken AI Reasoning

03 Oct 2026 · via Rss.arxiv

Correct Answers Invalid Traces Reveal Broken AI Reasoning
Image: Jenny.yip320 / Wikimedia Commons (CC BY-SA 4.0)

Correct Answers Invalid Traces Reveal Broken AI Reasoning

A grade-school math problem with a correct answer and a broken explanation is not a small thing. A model that reaches the right number through a reasoning chain that never happened has not reasoned at all — it has guessed well and dressed the guess in the grammar of logic. That gap between what a system claims and what it does is the subject of a 2026 paper by Ratish Puduppully and five co-authors, Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces. [1] The title is the finding. The finding is uncomfortable.

The Trace That Wasn’t There

Chain-of-thought prompting was sold as a window. Ask a model to show its work, the thinking went, and you get a readable path from question to answer — a way to check not just whether the machine arrived but how. That promise assumed the visible steps and the actual computation were the same thing. Puduppully’s team tested the assumption where it is easiest to test: arithmetic and word problems simple enough that every intermediate step can be verified by hand. The models produced correct answers. The traces that accompanied them, when checked step by step, did not hold up. Operations were skipped. Numbers appeared without derivation. A line of reasoning would jump from a premise to a conclusion that the premise could not support, and the final answer would still land. Over half of the invalid traces passed every syntactic and arithmetic check, failing only on semantic dependency checks.

This is not an occasional glitch. It is a structural feature of how these systems generate text. These systems do not solve the problem and then narrate the solution. They produce a fluent sequence of tokens that resembles a solution, and the answer emerges from a process the narration does not describe. When the answer happens to be right, the trace is a fiction that got lucky. When it is wrong, the trace is a fiction that did not. Either way, the window is painted on.

Plausibility Is Not Verification

The reason this matters beyond math is that the same mechanism operates everywhere a model is asked to explain itself. A medical summary that cites the right study for the wrong reason. A legal analysis that reaches a sound conclusion through a misreading of precedent. A financial recommendation that names the correct risk factor while actually responding to something else entirely. In each case the output passes the only test most users apply — does it look right? — while failing the test that matters — is it right for the stated reason? The grade-school setting strips away the camouflage. In a domain where the steps are checkable, the steps do not check out.

Correct Answers Invalid Traces Reveal Broken AI Reasoning (Image 1)
AI-generated image

The paper’s contribution is not that models sometimes produce bad reasoning. It is that the bad reasoning and the correct answer can coexist at high rates, which means the answer is not evidence for the reasoning and the reasoning is not evidence for the answer. They are two separate outputs from the same system, correlated loosely and misleadingly. A user who reads the trace and trusts the answer is trusting a relationship that does not exist.

The Training Data Problem

A separate paper, Generating Verifiable Chain of Thoughts from Execution-Traces, addresses a related root cause. Synthetic chain-of-thought training data, that paper’s authors write, “often consists of plausible-sounding explanations generated by teacher models, and not verifiable accounts of actual program behavior.” The models are trained on text that sounds like reasoning. They learn to produce text that sounds like reasoning. The sound is the target. The reasoning is not.

That paper’s approach is instructive because it shows what verification actually requires. The team instrumented code to capture real execution traces — what the program actually did, step by step — then narrated those traces into natural language and cross-checked each narration against the original trace. The result was 54,000 verified rationales, bidirectional, teaching models to reason both forward and backward. Models fine-tuned on this verified data improved substantially, with a peak gain of +26.6 on LiveCodeBench-Exec, +22.2 on CruxEval, and +19.5 on HumanEval. The improvement came not from more data or bigger models but from data where the explanation matched the computation. Verification quality, the authors demonstrate, directly determines both reasoning and code generation capabilities.

What the Math Papers Already Knew

The Answer That Reverses Everything

Here is the data point that should unsettle anyone who has been reassured by a model’s visible reasoning. In the grade-school setting, 31.6% of correct answers came with invalid traces on the hardest instances tested, making the trace unreliable as a verification tool. [1] Not occasionally. Not at the margins. At rates that make the whole enterprise of reading the reasoning suspect. The model does not hide its work. It is showing work it did not do. The explanation is generated by the same process that generates the answer, and neither one is accountable to the other.

Correct Answers Invalid Traces Reveal Broken AI Reasoning (Image 2)
AI-generated image

This reverses the usual worry. The concern about AI deception has focused on models that lie — that know the truth and say something else. That is not what is happening here. The model does not know the truth and does not know it is not telling it. It produces a correct answer and a false explanation with equal fluency, and it has no mechanism for noticing the discrepancy. The deception is not intentional. It is structural. The system cannot distinguish between reasoning and the appearance of reasoning, because it was trained on the appearance and rewarded for the outcome. The gap between what it claims and what it does is not a bug in the model. It is the model.

What the grade-school math reveals is that the window into a model’s thinking is not a window. It is a screen. The correct answer on the other side is real. The steps that appear to lead to it are a projection, generated after the fact, plausible and wrong. Anyone who has trusted a model because the reasoning looked sound now has a reason to check the reasoning itself — and to discover, as the paper did, that the check often fails. The answer was right. The trace was not. The difference is everything.


Sources

  1. arXiv — Paper

Mentioned organisations (context, not sources)

← back to the garden