Code That Runs Is Not Code That Works
A Program That Executes Is Not a Program That Obeys
A benchmark called Spec2Game, described in a paper posted to arXiv, sets out to measure something that most evaluations of large language models quietly skip. [2] The paper’s authors put the problem plainly: generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language The authors generated 3,330 complete Pygame projects from detailed natural-language specifications across 14 language models, then checked not just whether the code ran, but whether it did what the specification asked. [2] Their central finding is blunt: high executability does not imply faithful specification realization. [2] A program can launch, render a window, accept keyboard input, and still violate the rules it was supposed to implement.
That distinction sounds academic until you consider what it means for anyone who has watched a coding assistant produce something that looks impressive and behaves incorrectly. The benchmark separates four dimensions of quality — Executability, Specification Realization, Code Quality, and User-Facing Quality — precisely because conflating them hides the problem. [2]. A model that scores near-perfect on getting code to run may score far lower on getting code to do the right thing.
The authors built Spec2Game around 15 game families, each with one canonical task and nine controlled rule variants, for 150 task instances total. [2]. The games span three implementation-complexity levels. Each specification follows a seven-part structure: game overview, scene and interface, objects and states, player interaction, core rules and mechanics, game flow and progression, and goal, feedback, and termination. This structure matters because it forces the model to handle not just objects but the relationships between them over time.
Where Models Actually Improve
The component-level analysis in the paper contains the most useful finding for anyone trying to understand where AI genuinely helps. Models perform substantially better on Game Element Modeling than on Rule and Mechanism Modeling or Goal and Termination Modeling. [2]. In plain terms: they can create the pieces — the player character, the obstacles, the score display — but they struggle to wire those pieces into a coherent system of rules that produces correct outcomes.
This is not a small gap. It is the difference between generating a sprite that moves when you press a key and generating a game where pressing that key at the right moment under the right conditions triggers a state transition that eventually leads to a win or loss condition. The first is a rendering problem. The second is a systems problem. The benchmark shows that current models are much better at the first than the second.
The authors note that games provide a suitable setting for studying this problem because they combine objects, user inputs, rules, state transitions, feedback, and termination conditions within a single interactive system. [2]. Each of those elements can be specified concretely, and each can be checked through controlled execution. That is what makes the benchmark useful: it isolates the specific places where generation fails.
The Measurement Problem
Before Spec2Game, evaluating whether a model could generate a complete interactive program meant looking mostly at whether the program ran. The authors point out that execution success or functional correctness alone cannot capture the different quality aspects of generated projects. [2]. A game may appear correct while violating its specified rules, or implement the rules correctly while failing to communicate its current state clearly to the player.
The benchmark addresses this through checkpoint-level evaluation. The authors extract checkable functional checkpoints from the specifications and provide the generation model with checkpoint-derived public behavioral requirements, while withholding evaluator-only metadata, detection methods, and scoring information. [2]. This means the model knows what behaviors are expected but not how they will be tested. The checkpoints are assigned to specific generation stages: requirements about state modeling go to the state.py stage, while requirements about rules and mechanics go to the game.py stage. [2].
The generation protocol itself is staged: state.py, then game.py, then rendering.py, then main.py. [2]. Each stage uses a shared generation contract and stage-specific instructions, issued as separate requests without shared chat history. Once generated, each module is locked and can only be used as fixed context by subsequent stages. The main.py stage receives only public API summaries of the preceding modules, not their full source code. This design forces the model to work with the interfaces it defined rather than the implementations it wrote, which is closer to how real software development works.
What the Numbers Show

Across 14 language models and 3,330 generated projects, the pattern held consistently. Executability was nearly saturated — most models could produce code that runs. Specification realization was substantially weaker. The authors identify this as a key capability limitation in complete game generation. [2] The gap is not marginal: among 630 candidate projects, 387 passed all six executability checks, yet 87 of those still scored below 60 on specification realization. [2].
The benchmark also includes two ablation studies and a human assessment of 585 generated projects, though the details of those are in the appendices. [2] The core result is that the gap between running and working is not a minor artifact of a few bad generations. It is systematic.
The authors position Spec2Game as complementary to existing benchmarks. HumanEval, MBPP, EvalPlus, and LiveCodeBench evaluate self-contained programming problems. [1] SWE-bench addresses repository-level issues. WebGen-Bench, WebCoderBench, and MiniAppBench study interactive application generation. V-GameGym evaluates repository-derived Pygame tasks using code, image, and video evidence. GameCraft-Bench evaluates complete Godot games. Spec2Game adds curated Pygame specifications, controlled rule variants within each game family, and separate evaluation of the four quality dimensions.
The controlled variants deserve attention. For each game family, the authors introduce nine rule variations that preserve core gameplay while changing specific rules or design factors. This is designed to discourage models from simply reproducing conventional implementations they may have seen during training. If a model generates a standard Pong clone from memory, it will fail the variant that changes the paddle physics or the scoring rules. The variants are provided as standalone tasks, and the affected specification content is updated accordingly.
The Complexity Framework
The authors also developed a mechanism-grounded complexity scoring framework. [2]. Implementation complexity, in their definition, characterizes the demands of translating a game specification into a complete program, not the challenge faced by a player. They draw on game-design research to define five dimensions: Operation Complexity, Object and Rule Complexity, Spatial and Scene Complexity, Adversarial Behavior Complexity, and Interface and Information Complexity.
The scoring formula sums 13 sub-dimensions, each rated from 1 to 3, producing a total score. [2]. Based on the distribution across 30 candidate games, the authors define three levels: Easy (score 18 or below), Medium (19 to 27), and Hard (28 or above). Each game family’s nine variants inherit the complexity label of its canonical task. The full rubric is in the appendix, but the framework itself is a contribution: it provides a consistent basis for organizing tasks and for future extension.
The 15 game families were selected based on implementation feasibility in Pygame, evaluation feasibility, and diversity of gameplay mechanics. [2]. The paper includes a table classifying them by type and implementation complexity, though the specific games are not named in the main text.
What This Means for How We Use These Systems
The practical implication is not that language models cannot generate working software. They can, and the benchmark confirms that executability is nearly saturated. The implication is that “it runs” is a much lower bar than most people assume, and that the gap between running and correct is where the real work lives.
For anyone using these systems to generate code, the finding suggests a specific workflow: generate, then verify against the specification, not just against the compiler. The benchmark’s checkpoint approach offers a model for how to do that verification systematically. Extract the specific behaviors the program should exhibit, then check each one independently rather than assuming that a successful run means a successful implementation.
The authors also note that functional behavior and visual presentation are distinct: a game may appear correct while violating its specified rules, or implement the rules correctly while failing to communicate its current state clearly. [2]. This means evaluation needs to look at both what the program does and what the user sees, which is why User-Facing Quality is a separate dimension from Specification Realization.
The benchmark does not perform semantic repair. Candidate programs are not executed during generation, and no runtime or evaluator feedback is returned to the model. [2]. A single technical retry is permitted only for request failures or outputs deemed incomplete, using the same prompt and no additional error feedback. This is a deliberate choice: the benchmark measures what the model can do in one shot, without the benefit of iterative debugging.
The Open Question

The paper’s conclusion is not that models fail. It is that they succeed at something easier than what we asked for. The 3,330 generated projects demonstrate that language models can produce executable code from natural-language specifications at scale. [2]. They also demonstrate that executable code is not the same as correct code, and that the difference is concentrated in specific areas: rules, mechanisms, goals, and termination logic.
What remains unclear is whether this gap is a fundamental limitation or a training artifact. The models tested were not fine-tuned on the benchmark’s checkpoint structure. They received the public behavioral requirements but not the detection methods or scoring rules. Whether a model trained specifically to close the gap between Game Element Modeling and Rule and Mechanism Modeling would perform differently is an open question the paper does not answer.
The authors frame Spec2Game as a benchmark for evaluating LLMs’ ability to generate complete interactive Pygame projects from detailed natural-language specifications, with an emphasis on the realization of specific requirements. [2]. The emphasis is theirs. The finding is that realization lags execution. For anyone who has ever shipped code that passed its tests and still broke in production, that lag will feel familiar. The contribution of this work is measuring it precisely, at scale, and in a setting where the specification is unambiguous enough to check.
The benchmark is available as a resource for future evaluation. The complexity framework is designed to support extension. The checkpoint methodology is described in enough detail to be reproduced. [2]. What the paper does not offer is a solution to the gap it identifies. It offers a way to see the gap clearly, which is the necessary first step. Whether the next step — closing it — is a matter of better training, better prompting, or better tools remains unresolved. The 3,330 projects suggest that the answer will not come from making models generate more code. It will come from making them generate code that does what it was asked to do.
