Behavior Matching Without Mechanism Recovery in LLMs
When a Model Matches Behavior Without Recovering the Mechanism
A system that produces the right answer for the wrong reason is not a system that understands. This is the oldest problem in evaluation, and it has acquired new urgency now that large language models are deployed as interactive agents and behavioral simulators. A recent study by researchers investigating sequential structure in LLMs puts the issue plainly: observed behavior does not uniquely determine the process that generated it. [1] Two action trajectories can look identical in their surface statistics while arising from entirely different latent mechanisms — one from independent sampling, another from history-conditioned rules. Matching aggregate behavior, the researchers argue, does not establish that a model has recovered the conditional structure that produced a sequence Sequential Structure Recovery in LLMs. [1] That distinction matters more than it might first appear. When an LLM simulates a user in a multi-turn tool-use task, or plays the role of a negotiating counterpart, or generates synthetic behavioral data for a sandbox environment, the value of the simulation depends on whether the model has captured the process behind the behavior. If it has only captured the proportions — how often each action appears — then the simulation is a statistical replica, not a behavioral one. It will look right in aggregate and fail in exactly the situations where sequential context matters most.
The research team designed a framework to separate these two capacities. Their approach decomposes model behavior into three distinct abilities: matching marginal distributions, recognizing latent strategies, and executing conditional rules. This decomposition is the study’s central methodological contribution. Instead of asking whether a model “understands” a game or “simulates” a player, it asks which specific component of sequential reasoning succeeds and which fails. The answers turn out to be more revealing than a single accuracy score would have been.
Rock-Paper-Scissors as a Controlled Probe
The experimental setting is deliberately minimal. Two-player Rock-Paper-Scissors provides an action space of exactly three choices — rock, paper, scissors — and a structure in which opponent responses can induce history-dependent behavior. The small action space is not a limitation but a design choice: it enables rule-level evaluation, where every possible conditional dependency can be enumerated and tested. Each player is governed by an unobserved strategy, and the model must infer the latent strategies from an observed trajectory, estimate action distributions where applicable, and produce a continuation consistent with what it has inferred.
The researchers populated this environment with players of increasing complexity. Statistical players ignore history entirely: static players always choose the same action, while distributional players sample independently from fixed categorical distributions over the three actions. These players test whether a model can recover marginal action statistics. Markov players, by contrast, choose actions according to deterministic rules over recent history — the opponent’s previous move, the previous joint state, or second-order histories. These players test whether a model can infer and follow sequential dependencies beyond marginal frequencies.
This progression — from fixed distributions to first-order Markov rules to higher-order conditional dependencies — forms what the researchers call an ability-based evaluation pyramid. Each level requires something more than the last. A model that clears the first level has learned to count. A model that clears the second has learned to condition on a single previous observation. A model that clears the third has learned to track a richer history. The pyramid structure allows the researchers to pinpoint exactly where the capacity breaks down, rather than reporting a single aggregate score that would obscure the pattern.
What Three Experiments Revealed

Across the RPS experiments and a one-player task, the results converged on a consistent finding: longer context does not reliably improve strategy recovery, correct identification does not guarantee faithful rule execution, and performance degrades most clearly as conditional dependencies become higher-order. [1] This degradation persisted even in the one-player setting, where RPS semantics and player interaction were removed entirely, replaced by a stochastic n-gram continuation task with varying dependency orders.
The one-player extension is significant because it eliminates alternative explanations. If the failure were specific to game-theoretic reasoning or opponent modeling, removing the opponent should have helped. It did not. If the failure were specific to the RPS action space, replacing it with a generic n-gram task should have changed the pattern. It did not. The bottleneck appears to be the recovery and sustained execution of higher-order conditional structure itself — a capacity that may be orthogonal to any particular domain.
One comparison stands out for what it says about the nature of the gap. A prefix-only empirical n-gram estimator — a simple count-based model that estimates transition probabilities directly from observed prefixes — remained comparatively stable at high order. This is not a story about LLMs being worse than classical methods at everything. It is a story about a specific architectural mismatch: the models can approximate the distribution but cannot reliably implement the conditional rule that generates it. The LLM, for all its flexibility, does not.
The Difference Between Recognizing and Reproducing
The study’s framework separates distribution matching from conditional rule following, and the results show that these two capacities dissociate in practice. A model can correctly identify a latent strategy — labeling it as, say, “win-stay, lose-shift” — and then fail to generate actions that actually follow that rule. Correct recognition, in other words, does not ensure faithful simulation. The model knows what the strategy is called but cannot reliably do what the strategy does.
This dissociation has direct implications for behavioral simulation. If a model is used to generate synthetic user behavior for training or evaluation, and if that model can name the behavioral pattern but not instantiate it, then the synthetic data will carry the label without the structure. Downstream systems trained on such data will inherit the mismatch. They will learn to recognize the pattern without learning to reproduce it — a failure mode that propagates silently, because the synthetic data looks plausible in aggregate.
The degradation at higher orders sharpens the point. First-order Markov rules — where the next action depends on the single previous observation — are handled better than second-order rules, where the dependency spans two rounds. This is not a gradual decline but a structural one. The models appear to have a working memory for sequential dependencies that is shallow enough to handle the simplest cases and too shallow for the rest. The prefix-only estimator, which has no such memory limit, does not show the same decline.
Why This Is a Gain, Not Just a Warning
It would be easy to read these results as a limitation story — LLMs cannot do this, so be careful. But the more useful reading is about what the framework makes visible. Before this study, the question “can an LLM simulate sequential behavior?” had no clean answer because there was no clean decomposition. A model that matched marginal frequencies would score well on aggregate metrics and be declared successful. The failure would only surface later, in deployment, when the sequential structure mattered and the model’s outputs diverged from the true generative process.
The ability-based evaluation pyramid changes that. It provides a diagnostic that separates three capacities that were previously conflated. A practitioner can now ask: does my model match the distribution? Does it identify the strategy? Does it execute the rule? And the answers can differ. A model might pass the first test and fail the second, or pass the second and fail the third. Each failure mode has different implications for different applications. If the goal is to generate realistic-looking aggregate data, marginal matching may suffice. If the goal is to simulate an agent that responds appropriately to context, rule execution is required.

The finding that a simple n-gram estimator outperforms LLMs at high-order conditional tasks is also a gain in a different sense. It identifies a concrete baseline that any proposed solution must beat. It also suggests where to look for improvement: the problem is not that the models lack the information — the prefix is available, the transition rules can be stated explicitly — but that they cannot sustain the conditional structure during generation. That is a more tractable problem than “the model does not understand sequences,” because it points to specific architectural or training interventions rather than a vague call for more data or larger models.
The Evaluator Faces the Same Problem
The study’s central insight — that apparent behavioral fidelity can mask incorrect generative mechanisms — applies not only to models but to the humans who evaluate them. When a model produces a plausible continuation, the evaluator faces the same problem the model faces: the output looks right, but the process behind it may be wrong. And the evaluator, like the model, has limited access to the latent mechanism.
This creates a structural asymmetry in how AI systems are assessed. The things that are easy to measure — aggregate accuracy, distributional match, surface plausibility — are precisely the things that can be achieved without recovering the underlying structure. The things that matter — whether the model has captured the conditional dependencies that generate the behavior — are harder to measure and easier to fake. A model that has learned to produce the right marginal statistics will pass the easy tests. It may fail the hard ones, but the hard ones are not the ones that get run by default. The source’s own numbers make this concrete: among 282 incorrectly recovered strategies, 41.8% still landed within a total-variation distance of 0.02 of the target marginal distribution, and 62.1% within 0.10. [1].
The researchers’ framework is a step toward closing that gap. By explicitly separating distribution matching from rule following, it makes the harder question askable. But the framework itself requires a controlled setting where the latent process is known exactly — a luxury that does not exist in most real-world applications. In the wild, the evaluator does not know the true strategy. The model’s output is the only evidence available, and that evidence is exactly the kind that can be produced by the wrong mechanism. The study’s own design is the proof: it works because the researchers built the players, so they knew the rule the model was supposed to recover. [1].
The Consequence No One Draws
If a model can match the distribution without recovering the rule, then every downstream system that relies on that model inherits the mismatch. The synthetic data looks right. The simulated agent behaves plausibly. The evaluation passes. And the conditional structure that was supposed to be captured — the sequential dependency that makes the behavior meaningful — is absent. The system has learned the surface and missed the substance, and there is no error signal to indicate the loss.
This is the consequence that is easy to state and hard to act on. It does not lead to a single fix. It leads to a question that every deployment of an LLM as a behavioral simulator must answer: not “does the output look right?” but “does the output come from the right process?” The framework exists now. The remaining question is whether anyone will use it before the synthetic data has already been generated, the downstream models already trained, and the mismatch already baked in.
