Routing Probes Can Improve Without New Information
The Improvement That Wasn’t There
A probe is a small predictor bolted onto a trained vision transformer. It reads the model’s internal routing signals — expert gates, attention-residual weights, halting scores — and tries to guess whether the model got the answer right. When the probe with routing beats the probe without it, researchers tend to conclude that routing carries information about errors that the model’s outputs do not. That conclusion, according to a paper by Wenhao Liang and five co-authors, does not follow. [1] The improvement can appear even when routing contains nothing the outputs don’t already determine.
The study, titled “Routing Probes Can Improve Without New Information: An Exact-Null Audit of Uncertainty Beyond Model Outputs,” constructs a test where the answer is known in advance. [1]. [1] The authors keep the real pairs of outputs and routing signals from six trained vision transformers, but they redraw the correctness labels from a logistic function of the model’s confidence alone. The generator is fitted on one half of the data and frozen; on the other half, correctness is sampled from it. [1]. Routing is then uninformative by construction — it cannot carry any information about the labels beyond what the outputs already provide, because the labels were generated without it.
Under this null, a width-matched MLP comparison still reported a routing gain in 308 of 600 evaluations at the confidence-only view — 51.3 percent. A linear comparison reported none. The MLP was scikit-learn’s MLPClassifier with early stopping enabled, which selects its checkpoint by validation accuracy while every comparison is scored by log loss and Brier score. [1]. That mismatch, the authors found, is the cause. When they held each training trajectory fixed and read the same trajectory at the checkpoint preferred by validation log loss instead of validation accuracy, raw detections fell from 50 out of 120 to zero. [1] An independently implemented PyTorch probe showed the same contrast: 83 out of 120 against zero. Replacing routing with independent Gaussian features produced the same pattern — 80 out of 120 against zero — so the failure may not require routing structure at all.
The finding is narrow and specific, and that specificity is what makes it useful. The probe appeared to learn something about errors from routing. It had not. It had learned to exploit a selection rule that favored checkpoints by a metric different from the one used to score them.
The Distinction That Explains the Gap
The study’s Proposition 1 makes the underlying problem concrete. If correctness depends on a quantity O squared, adding R equals O squared helps a linear probe even though O already determines R. The gap persists with unlimited data and equal parameter counts after zero-padding. This is the difference between usable information and information content, a distinction established in earlier work by Xu and colleagues and by Hewitt and colleagues. A feature can improve a probe’s fit without carrying any information the probe did not already have access to.
What the audit measures is how often, and for what reasons, that inference fails for the probes, sample sizes, and controls used in practice, where routing and outputs are strongly dependent. The authors do not argue that routing never carries incremental information. They argue that the comparison as commonly run cannot tell the difference. Fitting a better probe and testing for incremental information are different problems, and each needs its own validation.
The shuffle controls tell a similar story. Across all four output views, MLP comparisons against globally and output-matched shuffled routing — controls intended to detect spurious gains — reported gains in 308 and 293 of 1,920 evaluations. A control that fires under the null is not a control. Applying log-loss selection to all 1,920 MLP evaluations of the null benchmark removed all 528 raw detections and most shuffle-control detections. The raw detection rate fell from 27.5 percent to zero observed detections.

What the Repaired Test Cannot Do
Removing false detections under the null does not make the comparison sensitive. In two matched synthetic settings with an implanted signal of about 0.005 nats, the repaired raw comparison detected it in zero out of 20 replicates in each setting, and its mean fitted gain stayed negative. The comparison was clean under the null but insensitive to a real signal of that size.
A conditional permutation test built on an estimated routing law performed differently. It detected the implanted signal in 11 out of 20 replicates in one setting and 10 out of 20 in the other, while rejecting rarely under the null — 59 times in 3,840 evaluations. With the true conditional law, such a test has guaranteed level. With an estimated law, its behavior must be measured, and the authors measured it. The contrast between the two approaches is the paper’s central methodological point: a test can be valid without being useful, and useful without being valid. The repaired comparison achieved validity at the cost of power. The conditional test retained power at the cost of needing an estimated law whose behavior under the null had to be verified empirically.
Where the Signal Actually Appears
Applied to real correctness labels across 22 routing families and 68 training runs, the conditional test found replicated conditional-assignment evidence in 12 of 176 family-probe-view cells. All 12 were in five DeiT attention-residual families. None passed an additional same-width noise criterion. Eight of the 12 cells kept their evidence under two specified variants of the conditional law.
This result constitutes model-relative evidence, not a general claim about routing. It appears in attention-residual families and not elsewhere in the panel. It survives two variants of the conditional law but fails a noise criterion. The authors report it as what it is: a signal in a specific architectural family that needs further validation before it can be read as incremental information about errors.
The Precedent That Proves This Is Not New
The pattern — a measurement that improves for reasons unrelated to the thing being measured — has a long history in statistics and machine learning. Selection by one metric while scoring by another is a classic source of apparent gains. [1]. The paper’s contribution is not the discovery of the pattern but its exact quantification in a setting where the null is known and the cause is isolated.
What makes the audit unusual is the construction of the null itself. Rather than arguing about whether routing might carry information, the authors build a world where it cannot and measure how often the standard comparison says otherwise. The 51.3 percent detection rate under the null is not an estimate or a bound. It is a count. The 308 out of 600 evaluations are not a simulation of what might happen. They are what happened when the labels were redrawn from an output-only generator. [1].
The six-model panel, the four output views, the 1,920 MLP evaluations, the 3,840 conditional test evaluations — these are the specific numbers from a specific experiment. [1]. They do not generalize to all probes or all routing signals. They describe what happened when this comparison was run this way on these models.

The Open Variable
The paper ends where the evidence ends. On real correctness labels, the conditional analysis yields model-relative evidence in five DeiT attention-residual families. In four it persists under two specified variants of the conditional law. No family passes an additional noise criterion. [1].
Whether routing carries information about errors beyond model outputs remains an open question. What the audit establishes is that the standard comparison cannot answer it. The improvement of a probe with routing over a probe without it is not evidence of incremental information. The checkpoint selection rule, the scoring metric, the control design, and the estimator’s capacity all shape the result in ways that can produce gains without information.
The last open variable is the conditional law itself. The conditional permutation test needs an estimated routing law to generate assignments. With the true law, the test has guaranteed level. With an estimated law, its behavior must be measured case by case. The paper measures it on this benchmark and finds it rejects rarely under the null while detecting an implanted signal in roughly half the replicates. Whether a better estimate of the conditional law would sharpen the test — or whether the signal in the DeiT families would survive a stricter noise criterion — is not settled here. Those are the questions the next study has to answer.
The audit’s contribution is a clean null, a specific cause, and a repaired comparison that is valid but not sensitive. It does not tell researchers what to conclude about routing. It tells them what they cannot conclude from the comparison they have been running. [1].
