AI Verifiers Trust Corrupted Rationales Over Evidence
The Message Between the Machines
A modern AI question-answering system rarely works alone. It is a relay race of specialists. One model retrieves the evidence, another reasons toward an answer, and a third judges whether that answer holds up. Each handoff is a message passed across a boundary, and each message carries assumptions nobody has fully audited.
The most interesting of these messages is the rationale. When a reasoner model proposes an answer, it often includes a short explanation of why that answer is correct. The verifier model then receives this explanation alongside the evidence and the candidate answer. The rationale is supposed to help the verifier do its job better — to check the reasoning, catch errors, and reach a more reliable final answer.
Research published on arXiv in July 2026 by Jiameng Zhang and Hongqiu Wu asks a deceptively simple question about this handoff: once an answer has already been proposed, what does the rationale message actually change? [1] The answer, it turns out, is not what the designers of these pipelines assume. The study ran on 400 examples from MuSiQue, HotpotQA, and 2WikiMultiHopQA, with DeepSeek serving as both generator and verifier.
The study introduces what the authors call a message-intervention diagnostic. The method is elegant in its restraint. It fixes the evidence and the candidate answer. It holds everything constant except one variable: the rationale passed from reasoner to verifier. Then it measures two outcomes separately — whether the final answer changes, and whether the verifier’s judgment about whether the answer is supported changes. Most evaluations collapse these two into a single accuracy number. This one pulls them apart, and what separates them is where the deception lives.
Two Things a Verifier Returns
The core insight of the diagnostic is that a verifier in a role-specialized pipeline does not return one thing. It returns two: an answer and a support judgment. When a rationale is passed along, it can influence either or both. A pipeline that only tracks final accuracy will never see the difference.
Zhang and Wu tested their protocol on 400 examples drawn from three multi-hop QA datasets — MuSiQue, HotpotQA, and 2WikiMultiHopQA — using DeepSeek as both generator and verifier. Each example was verified under five conditions: no rationale at all, the original reasoner-generated rationale, a harmless meaning-preserving paraphrase, an entity-swapped corruption, and an answer-conflicting corruption.
The results expose a gap between what rationales appear to do and what they actually do. Faithful rationales added almost no answer accuracy over passing no rationale at all. On MuSiQue, the original rationale improved exact match by only 0.5 points over evidence and answer alone. The explanation, in other words, contributed nearly nothing to getting the right answer.
But the same rationales exerted enormous influence on the support judgment. Under a blind verifier prompt — one that does not explicitly ask the verifier to check the rationale — harmless paraphrases shifted support in only 0 to 2.5 percent of examples. Corrupted rationales shifted support in 10 to 22 percent. When the system explicitly instructed the verifier to check rationale faithfulness, the channel grew stronger still: corrupted rationales changed support in 34 to 55 percent of examples, with strong statistical significance on every dataset.
The rationale is not helping the verifier find the truth. It is telling the verifier what to believe. And the verifier, more often than not, believes it.
The Corruption Overtrust Problem
The sharpest finding in the paper has a name: corruption overtrust. This is the case where the verifier keeps accepting a corrupted rationale — one that has been deliberately altered with entity swaps or conflicting claims — and treats the answer as supported anyway.
Human auditors were brought in to check. Of 42 valid corruptions, 16 were classified as corruption-overtrust cases. When blind human reviewers examined a sample of the corrupted rationales that the model had accepted, they rejected or marked unclear 9 out of 10 of them. The model was accepting explanations that a human, seeing the same text, would not trust.
This is the deception at the heart of the pipeline. The rationale presents itself as a justification. It reads like reasoning. It carries the surface features of an explanation — entities, relations, a causal-sounding chain. But its actual function in the system is not to justify. It is to persuade. And it persuades most effectively when it is wrong.
The paper also identifies a second failure mode: correct-answer penalty. Here a correct answer gets rejected because its accompanying rationale is corrupted. Of the 42 valid corruptions in the human audit, 13 were correct-answer penalty cases: the answer remained correct, but the verifier marked it unsupported because the rationale was corrupted. The rationale that was supposed to support the answer instead undermines it.
A third pattern, local answer dominance, occurs when the direct evidence is strong enough to preserve the answer despite a broken upstream rationale. In these cases the pipeline stumbles into correctness — not because the rationale helped, but because the evidence was too clear to override. The system appears to work. It is not working the way its designers think.
When the Channel Goes Quiet
Not every pipeline exhibits the same vulnerability. Zhang and Wu ran cross-model and task-boundary checks to map when the rationale channel is active, amplified, inert, or folded into the task label. The results describe several distinct receiver states, each with different implications.
An active channel changes support assessment. This is the dangerous state, the one that needs auditing for overtrust. A harmless channel leaves support stable under meaning-preserving rewording — the rationale is being read, but it is not distorting judgment. An inert channel discounts the external rationale entirely. The authors found this in a DeepSeek-R1 boundary check, where the model effectively ignored the passed rationale. In that case, the rationale adds interface complexity without communication value. The pipeline is paying a cost in latency and tokens for a message that does nothing.

The fourth state, task-coupled, occurs when the rationale is folded into the task label — when the verifier treats the explanation as part of the answer rather than as a separate claim to evaluate. This is perhaps the most insidious, because the distinction between explanation and answer collapses entirely.
The diagnostic turns a vague design question — should a pipeline pass rationales to later roles? — into a measurable one: what behavior does this message field change? The answer varies by model, by task, and by prompt. There is no universal rule. There is only the audit.
The Gap Between Explanation and Persuasion
What makes this research unsettling is not that AI systems make mistakes. It is that the mechanism designed to catch mistakes — the rationale passed to a verifier — is itself a source of error. The explanation is not a window into the model’s reasoning. It is a message with its own persuasive force, and that force operates independently of whether the message is true.
The paper is careful to distinguish its approach from traditional chain-of-thought faithfulness work. That line of research asks whether an explanation reflects the computation of the same model that produced an answer. Zhang and Wu ask something different: what happens when explanation-like text becomes a communication payload consumed by a later role? The rationale is not a trace of hidden reasoning. It is a message sent across a boundary, and messages can be corrupted, misread, or overtrusted.
The finding that faithful rationales add almost no answer accuracy is itself significant. It means the rationale is not doing the work it was designed to do. It is not improving answer selection. It is not helping the verifier find the truth more reliably. Its measurable effect is almost entirely on support assessment — the judgment about whether an answer is acceptable. The rationale is not a tool for verification. It is a tool for persuasion, and the verifier is its audience.
This is the gap the research exposes. The system claims to verify. It appears to verify. It produces a support judgment, a final answer, all the outputs of a functioning pipeline. But the mechanism that produces those outputs is not the mechanism the design assumes. The rationale is not a check on the answer. It is a nudge on the judge.
What the Numbers Hide
The paper reports that final answers move less than support judgments — between 2 and 30 percent across conditions, compared to 10 to 22 percent under a blind prompt and 34 to 55 percent under an explicit rationale-checking prompt for support. [1] And only 2.9 to 35.3 percent of corrupted support flips co-occur with answer changes. [1] This means that in 64.7 to 97.1 percent of cases where the rationale corrupts the verifier’s judgment, the final answer stays the same.
A pipeline evaluated only on final accuracy would see almost nothing wrong. The answer is correct. The system appears to work. But the support judgment — the thing that is supposed to tell you whether the answer is trustworthy — has been quietly compromised. The verifier is approving answers for the wrong reasons. It is right by accident.
This is the deception that accuracy metrics cannot see. The system produces the correct output through a process that is not the process it claims to use. The rationale is not supporting the answer. It is overriding the verifier’s independent judgment, and sometimes the override happens to land on the right answer. The authors are careful to note that not every support flip is a failure: if the system output includes a corrupted rationale, rejecting the bundle can be the correct behavior, because the answer may be right while the communicated explanation is not a supported claim.
The authors argue that rationale sharing should be evaluated as a verification-message mechanism, not merely as a route to higher answer accuracy. The distinction matters because the two evaluations produce different conclusions. By accuracy, the rationale is nearly useless. By influence on judgment, it is powerful and dangerous.
The Historical Parallel
There is a pattern here that extends beyond AI systems. Institutions have long relied on explanations to justify decisions, and those explanations have long been mistaken for the reasons behind the decisions. A report that explains why a decision is sound is read as analysis. It functions as assurance. The distinction between the two only becomes visible when the decision fails and the report is re-examined.
The rationale in a role-specialized QA pipeline occupies the same position. It is presented as reasoning. It is consumed as assurance. The verifier does not independently verify the rationale; it absorbs it. The explanation becomes the decision.
What Zhang and Wu have built is a way to measure that absorption. Their intervention protocol isolates the message and watches what it does. The results show that the message does a great deal — just not what it claims. The rationale does not communicate reasoning. It communicates confidence, and the verifier accepts the confidence without checking the reasoning.
The Scalability Trap
The appeal of role-specialized pipelines is scalability. You can swap models, adjust prompts, add roles, and the system keeps running. Each role has a defined input and output. The interfaces are clean. The architecture is modular.
But the rationale interface is not clean. It is a channel through which unverified claims flow from one role to another, and the receiving role has no reliable way to distinguish a faithful rationale from a corrupted one. The verifier cannot check the reasoning because the reasoning is not what it receives. It receives a summary, a narrative, a persuasive artifact.
The paper’s cross-model checks show that this vulnerability is not specific to one model family. It appears across configurations, though its strength varies. The channel can be active, amplified, inert, or task-coupled. The designers of a pipeline cannot know which state they are in without running the diagnostic.
This is the scalability trap. The architecture scales. The verification does not. Adding more roles and more messages does not add more checks. It adds more channels through which unverified claims can travel. The system grows more complex, and the complexity creates new surfaces for the same failure.

What the Diagnostic Reveals
The message-intervention protocol is not a fix. It is a measurement. It tells you whether your rationale channel is doing what you think it is doing. In most cases, the answer is no.
The three contributions the authors claim are modest in framing but significant in implication. First, a protocol for testing what a rationale changes after it is passed to a verifier. Second, paired metrics that separate answer selection from support assessment while holding evidence and candidate answers fixed. Third, evidence across MuSiQue, HotpotQA, and 2WikiMultiHopQA that harmless paraphrases rarely change verifier judgments while corrupted rationales substantially change support judgments.
The human audits add a dimension that metrics alone cannot capture. When blind humans look at the corrupted rationales the model accepts, they reject or mark unclear 9 out of 10. The model’s threshold for accepting an explanation is far lower than a human’s. The model is not verifying. It is complying.
This is the gap between what the system appears to do and what it actually does. The system appears to verify answers by checking rationales. It actually verifies answers by absorbing rationales. The difference is invisible in the final output. It is visible only when you intervene in the message and watch what changes.
The Time Horizon Problem
Progress and consequence rarely share the same clock. A pipeline that passes rationales between roles improves its architecture today. The cost of that improvement — the corruption overtrust, the correct-answer penalty, the support judgments that flip without answer changes — shows up later, in the accumulated unreliability of a system that appears to work.
The research does not propose a solution. It proposes a way of seeing. The rationale channel is not a bug to be fixed. It is a property of the architecture to be understood. Some pipelines will find their channel inert and remove the rationale entirely. Some will find it active and add auditing. Some will find it task-coupled and restructure the interface. The diagnostic does not tell you what to do. It tells you what is happening. The authors’ own design rule is narrow: before relying on a message field, test what receiver behavior it actually changes, and treat rationale messages as claims to verify against evidence rather than as traces of reasoning.
What is happening, in the systems the authors tested, is that the rationale is not a check. It is a message that carries persuasive force independent of its truth. The verifier receives it, weighs it, and often defers to it. The final answer may be correct. The support judgment may be wrong. The system does not know the difference, and neither does anyone reading its output.
The Explanation That Explains Nothing
The most disquieting finding is the one that sounds most benign: faithful rationales add almost no answer accuracy. The rationale, when it is correct, does not help. The rationale, when it is corrupted, does harm. The message that was supposed to improve verification is either inert or destructive.
This is the deception in its purest form. The rationale looks like reasoning. It reads like reasoning. It is passed between roles as if it were reasoning. But its measurable effect on the system’s ability to find correct answers is negligible. Its measurable effect on the system’s judgment about whether answers are supported is substantial. The rationale is not a tool for finding truth. It is a tool for manufacturing agreement.
The verifier agrees with the rationale. The rationale agrees with the answer. The answer may or may not be correct. The pipeline produces a result, and the result carries the appearance of verification. The appearance is the product. The verification is not.
Zhang and Wu’s diagnostic makes this visible. It holds the evidence fixed, holds the answer fixed, varies only the rationale, and watches what moves. What moves is the support judgment. What stays still is the answer accuracy. The rationale is not doing the work of verification. It is doing the work of persuasion, and the verifier is persuaded.
The gap between what the system claims — a reasoning pipeline with a verification step — and what it does — a persuasion pipeline with a compliance step — is the space this research opens. It is not a large space. It is a single message passed between two calls. But inside that message, the difference between checking and believing collapses, and the system does not know which one it is doing.
