🌿freegardner

Synapse

World Models Trained on Games Miss Real Physics

28 Sep 2026 · via Wired

World Models Trained on Games Miss Real Physics
AI-generated image

World Models Trained on Games Miss Real Physics

A world model can watch a thousand hours of a forklift game and still crush the pallet. It has learned the sequence — approach, lower, lift — without learning the force. The distinction sounds academic until you put the same model in a warehouse. Then it becomes the whole problem.

The Cause-and-Consequence Gap

The AI systems that write fluent paragraphs and answer exam questions have never touched anything. They were trained on text, and text is a record of what people said about the world, not the world itself. That was fine as long as the task was producing more text. It stops being fine the moment the task is steering a vehicle or gripping an object with the right pressure.

This is the gap that a class of models called world models is meant to close. Instead of predicting the next word, a world model predicts what happens next in a physical space — if I push here, the object moves there. The ambition is to give machines a working intuition for cause and effect. The trouble is that intuition has to be built from data that shows both the action and its consequence, and that kind of data barely exists on the open internet.

“You need cause and consequence,” says Xiatian Zhu, an associate professor specializing in AI at the University of Surrey. “On the internet, we have very little of this type of data.” [1]

So the industry went looking for a substitute. It found one in an unlikely place: the thumbstick.

The Exhaust Product

Every time someone plays a 3D video game, the controller records a small stream of decisions — which direction, how hard, how long. Paired with what appears on screen, that stream is exactly the shape of data a world model needs. Visual input on one side, action on the other. A British startup called Worldmodeldata, advised by Yann LeCun, is betting that this byproduct of entertainment can be repackaged as training material for machines that are supposed to operate in reality.

The logic is seductive. Games are abundant, varied, and cheap to harvest. Players generate the data for free, and manual sensor-based data collection yields only a small amount of data by comparison. The startup says it packages controller inputs and other data collected by video game studios — an exhaust product available in massive quantities — into training datasets for world models. Its chief executive, Rhea Loucas, frames the pitch as a kind of rescue mission: “There are millions of great games, and they are more and more similar to the real world. Why don’t we take the vast, abundant, diverse experiences from video games, and teach AI?” [1]?”

The phrase to notice is “more and more similar.” Similar is not the same. And that distinction is where the deception begins — not a lie told by a machine, but a gap between what a training set appears to teach and what it actually teaches.

What the Controller Never Records

A video game character picks up an apple. On screen, the motion looks complete. Underneath, nothing happened. The engine did not simulate the friction between fingertip and skin, the slight rotation of the wrist to keep the fruit from rolling, the micro-adjustments that a real hand makes without conscious thought. The developer drew a hand near an apple and triggered an animation. The apple disappeared from the table and appeared in the inventory. Physics was implied, not computed.

A model trained on that footage learns the choreography and misses the mechanics. It sees a hand and an apple and a successful outcome, and it infers a rule that does not hold: that reaching is the same as grasping. In the game, that inference is harmless. In a kitchen, a factory, or a surgical suite, it is the difference between a task completed and an object destroyed.

World Models Trained on Games Miss Real Physics (Image 1)
AI-generated image

Ming-Yu Liu, who leads world model development at Nvidia, puts it plainly: models fed on game inputs are unlikely to handle tasks that demand fine motor control, because game physics is often eccentric and developers take shortcuts to create the illusion of realism. “I would be more conservative on using video game data for manipulation,” he says. “The physics for manipulation is much more involved.” [1]

Zhu reaches the same conclusion from the academic side. Games, he notes, are simulators with “some degree of physical grounding,” but that grounding is “very coarse, approximate.” [1] The model is not wrong about the world. It is fluent in a world that does not exist.

The Confidence Problem

Here is what makes this more than a technical footnote. The failure mode is not a model that refuses to act. It is a model that acts with the same fluency it showed in training, unaware that the rules have changed. In the game, the animation always completes. In reality, the apple slips.

That is the shape of AI deception worth worrying about — not a system that schemes, but a system whose competence is real within its training distribution and counterfeit outside it. The model has no way to signal the boundary. It does not know it has crossed from the world it learned into the world it was meant for. It simply keeps predicting, and the prediction is confident.

The industry’s own framing makes this harder to see. When a company says it has licensed a million hours of gameplay, the number sounds like coverage. But hours measure exposure, not fidelity. A million hours of a world where grip strength is a number in a config file does not add up to one hour of a world where grip strength is a physical fact. The dataset is large and thin at the same time.

The Bottleneck Nobody Can Buy Around

Some labs have tried to sidestep the problem by generating their own data. They attach sensors to humans and robots in controlled environments and record every motion. The approach is honest about physics but starved for scale. It produces a trickle where the models need a flood, and it captures the tidy scenarios — the pick-and-place tasks that a paid demonstrator can repeat — while missing the messy ones.

“You can pay people to demonstrate pick-and-place tasks,” says Nicole Fraenkel, a partner at Khosla Ventures, which has invested in General Intuition. “But repetition alone won’t capture the disorder of the world you’re asking a machine to operate in.” [1].”

The disorder is the point. The corner cases are where machines fail, and corner cases are precisely what a curated dataset, whether harvested from games or filmed in a lab, tends to exclude. Fraenkel is blunt about the stakes: “The cost of error with a car, plane, drone, factory forklift, or autonomous quadruped is very high.” [1].”

So the industry faces a choice between two incomplete teachers. One has seen everything but understands nothing. The other understands a little but has seen almost nothing. Neither can tell you which lessons will transfer.

The Promise and the Postponement

Loucas believes video game data will eventually make up the bulk of world model training, with task-specific data layered on top to refine the result. She calls it a possible “GPT moment” for world models. The phrase is doing a lot of work. No one has demonstrated that the same threshold exists for physical intuition, or that gameplay is the fuel that reaches it. [1].

Fraenkel, whose firm has money on the line, is more candid: “There are many paths to the promised land. The truth is, the jury is still out on which one is going to work best.” [1].”

World Models Trained on Games Miss Real Physics (Image 2)
AI-generated image

That is an honest sentence, and it is also the crux. The problem is not that video game data is useless. It may well be useful for generating hyperrealistic video or 3D environments, as Liu suggests. [1]. The problem is that “useful for something” and “sufficient for this” are different claims, and the marketing around world models tends to blur them. A model that can render a convincing kitchen is not a model that can cook in one.

Technically Solved, Socially Unresolved

The deeper issue is not which dataset wins. It is that we have built a class of systems whose failures are invisible until they act. A language model that hallucinates a citation can be checked. A world model that hallucinates friction cannot — not until the arm has already closed around the object and the object has already broken. The error is physical, and physical errors do not come with a confidence score.

This is where the technical and the social part ways. The engineering problem — how to train a model on cause and consequence — is being attacked from a dozen directions, and one of them may work. The human problem is that we have no good way to know which model has actually learned the world and which has merely learned a convincing simulation of it. Both will behave identically in the demo. Only one will behave correctly when the apple slips.

Until that distinction can be measured before deployment rather than discovered after, every world model is a claim rather than a guarantee. The thumbstick data may get us closer to machines that understand physics. It will not, by itself, get us closer to knowing when they do not. That part remains unsolved — not because the technology is immature, but because we have not decided what proof would even look like.


Sources

1. Wired — Quote source (original article)

Mentioned organisations (context, not sources)

- University of Surrey — Organisation (homepage)

- Worldmodeldata — Organisation (homepage)

- Nvidia — Organisation (homepage)

- Khosla Ventures — Organisation (homepage)

- General Intuition — Organisation (homepage)

← back to the garden