Training Paths Decide AI Convention Preferences
A model trained on the same problems written under two incompatible conventions — both correct, neither privileged — ends up just as capable either way. What changes is which convention it commits to — and a benchmark that only counts right answers cannot see the difference. That is the finding at the center of a 2026 paper by Wenhui Chen of the University of Macau, and it matters less for what it says about machines than for what it says about a specific kind of human judgment: deciding which of two correct answers applies here. [1]
Two Correct Answers, One Silent Choice
Chen’s setup is deliberately plain. Take a corpus of problems, write them under two incompatible conventions — both correct, neither privileged — and train a model on them. Then vary the order in which the data arrives. Ten orderings of one corpus, one training budget, everything else held fixed.
The result is a 12.29-sigma arrangement effect on one coordinate (Chen). On the other coordinate — the one that measures capability — it is exactly zero. Same model, same data, same question. The difference is entirely in what you measure.
The coordinate that moves is called allocation (Chen): the share of the model’s behavior written under one convention versus the other. Across twelve arms, that share runs from 0.04 to 0.87 (Chen) — nearly the whole range. The coordinate that does not move is capability, measured as the sum of accuracy on convention A and accuracy on convention B. Across the same twelve arms, that sum stays constant to within 9.7 percent (Chen).
Put differently: the model can solve the problems either way. What the training path decides is which way it prefers. And no exact-match benchmark — the kind that scores a model by whether its output string matches a reference string — registers the shift at all, because under a convention-agnostic score the arrangement switch is exactly zero by construction.
The Knob Nobody Was Turning
The paper’s central theoretical claim is that the learning-rate schedule is not a background condition. It is the averaging operator, and it decides whether ordering effects survive training at all — a claim the paper proves as a bound, not merely asserts.
Chen proves a bound in which arrangement and schedule enter as separate multiplied factors (Chen). Arrangement enters only through the period of its alternation — how long each convention runs before switching. Schedule enters only through how much weight the endpoint can place on any single moment of the run. The two factors do not communicate.
The practical consequence is sharp. A decaying schedule, the kind that ends near zero, cannot place a large step size and an uncontracted remainder at the same moment. A constant schedule does exactly that at the final step. So under a constant learning rate, ordering effects bloom. Under a decaying one, they are averaged away by the time anyone scores the endpoint.
The measurement matches. Ten arrangements of one corpus, one budget, the same three seeds, the two families differing in lr_scheduler_type and in nothing else, span 2.20 contrast floors under a cosine schedule and 11.63 under a constant rate (Chen). Under the single cosine that essentially every published training run uses, the same ten arms collapse into two distinguishable states — where their own resolution would allow about ten. The constant-rate family’s intraclass correlation is 0.836, with a bootstrap interval of [0.315, 0.926] that lies entirely above zero (Chen).
Order matters and order does not matter are not competing findings in the literature. They are two settings of one knob — which is what the divided record on data ordering looks like from here.
What the Model Was Never Asked

Here is where the paper stops being about optimization and starts being about a specific kind of judgment.
The model, trained on conflicting but equally valid supervision, does not lose the ability to handle either convention. It loses the neutrality to choose between them. The training path writes a commitment into the parameters — a default, a preference, a habit of answering in one form rather than the other. And the standard evaluation apparatus, by construction, cannot see it.
This is a specific kind of blindness, and it maps onto a specific kind of human labor. When a person adjudicates between two correct-but-incompatible conventions — a legal standard that varies by jurisdiction, a medical coding scheme that differs between insurers, a notation that two engineering teams both consider canonical — the value they add is not the ability to compute the answer. It is the judgment of which answer applies here. That judgment is contextual, it is often undocumented, and it is exactly what an exact-match metric cannot score. The paper’s own construction makes this concrete: its natural corpus pits numeral answers against spelled-out answers, both correct, and the model trained on it commits to one form as a default.
Chen’s model has the capability. What the training path removes is the neutrality. And the benchmark, by construction, cannot tell the difference between a model that chose and a model that was chosen for — because the sum of accuracies on both conventions is conserved while the allocation share moves.
The Conservation That Hides the Loss
The paper’s most consequential number is not the 12.29 sigma. It is the 9.7 percent — the arm-to-arm conservation of the capability sum on the synthetic corpus, against 15 percent on the natural one.
Across twelve arms, the sum of accuracy on both conventions stays constant to within 9.7 percent while the allocation share runs from 0.04 to 0.87. Something is conserved — total problem-solving capacity, to within an eighth — and something else is wildly variable: which convention that capacity gets expressed through.
This is the shape of a problem that has been technically solved and socially unresolved. The question of which work it does, in which form, according to whose convention, has been absorbed into the training pipeline as an accident of data ordering and learning-rate schedule. Nobody decided. The path decided — and the paper’s own framing is that this is a stopping phase, a commitment device whose exact-match victory is its own objective’s defeat.
Chen reports one intervention that recovers the choice: marking the convention in the prompt collapses the switch and reaches 87.5 percent of the union ceiling (Chen) — recovering 74.9 percent of the gap to the additive ceiling. In other words, if you tell the model which convention applies, it can follow. The commitment is not permanent. But it is the default, and the default is what gets deployed.
The Role That Was Never Named
There is a job description buried in this paper, and it is not the job of training the model.
It is the job of deciding, before training, which conventions the corpus will contain and in what proportion. It is the job of deciding, before deployment, which convention the model should default to when the prompt is silent. It is the job of deciding, after deployment, whether a shift in allocation share from 0.04 to 0.87 is a bug or a feature — a question the paper’s own decomposition makes askable for the first time.
None of these decisions are made by the model. All of them are made by someone, or by no one. Chen’s paper shows that the training path makes them implicitly, through the interaction of data ordering and learning-rate schedule, and that the standard evaluation suite reports none of it — because the metric it uses cannot express the coordinate along which the decision moves.
The human role that AI makes superfluous is not the role of solving the problem. That was automated long ago. The role that is now at risk of disappearing without being replaced is the role of saying which correct answer counts — and the paper’s quiet demonstration is that this role can be vacated not by a decision but by a default, set by a hyperparameter and reported by no one.

A Benchmark That Cannot See the Choice
The paper does not argue that benchmarks are bad. It argues something narrower and harder to dismiss: that the choice of metric determines whether a real effect is visible or invisible — and that the field’s default metric is blind to the coordinate along which the effect moves.
A convention-agnostic score — one that asks only whether the model got the problem right, not which form it used — reports exactly zero for the arrangement switch that the paper measures at 12.29 sigma under exact match. Not a small number. Zero. The effect is not attenuated by the metric. It is orthogonal to it, because on a solved problem reallocating samples between the two conventions conserves their sum.
This is not a flaw in the benchmark. It is a definition of what the benchmark measures. Accuracy asks: can the model solve this? Allocation asks: which solution does the model prefer? These are different questions, and the second one has no standard home in the evaluation literature — which is what the paper’s Definition 2 is for.
The consequence for professional practice is direct. Any deployment that relies on a single accuracy number to certify a model is certifying capability while remaining blind to commitment. The model passes. The convention it will use when the prompt is ambiguous was set by the data ordering and the learning-rate schedule, and no one measured it — because the standard metric cannot express the question.
What the Path Writes
Chen’s paper is not about AI replacing people. It is about something more specific and more uncomfortable: AI absorbing a decision that people used to make explicitly and making it implicitly, in a place where no one is looking — the training pipeline itself.
The model does not choose between conventions the way a person chooses. It does not weigh jurisdiction, context, or consequence. It averages. The averaging operator is the learning-rate schedule, and the average it produces is a commitment that no exact-match benchmark can see — a commitment the paper measures as a movement of the allocation share at fixed capability sum.
The capability is conserved to within an eighth. The commitment is not. And the role that used to sit between those two facts — the role of deciding which right answer is the right one — has been quietly relocated into the training pipeline, where it is set by a hyperparameter and reported by no one.
The problem is technically solved. The model can do the work. What remains unresolved is who decides what the work means — and whether anyone will notice when the answer is no longer a person, but a learning-rate schedule.
