Physical AI replaces hidden human judgment in work
The question everyone asks about physical AI is when robots will become useful. That is the wrong question. The right question is what happens to the human judgment that currently sits inside every industrial process, construction site, and logistics operation. Unitree’s spectacular IPO collapse — from $66 billion to roughly half that value in a single week — was not a failure of hardware. It was the market realizing that a robot body without a brain is just an expensive sculpture. [1]. But the deeper truth is more uncomfortable: the brain being built for those bodies is not designed to replicate human skill. It is designed to replace the human decision-making that we never even noticed was there.
The Hidden Layer of Work
Every robot deployment story is actually a story about removing a layer of judgment that was never documented, never formalized, and never compensated as what it truly was. Consider the excavator operator at Bedrock, the company running autonomous digging machines. The obvious narrative is that the robot replaces the person holding the joystick. The less obvious narrative is that the robot replaces the person who looked at a pile of dirt and decided, based on years of experience, that the next scoop should come from the left side rather than the right. That person’s expertise was never written down. It was never encoded in a training manual. It existed only in the space between their eyes and their hands, refined over thousands of hours of watching how soil behaves when it is wet, dry, compacted, or loose.
The robotics industry calls this the data crisis. Foxglove, the company organizing the Actuate conference where 1,500 developers gathered to build AI brains for robots, literally had a booth at that event promising to solve “the robotics data crisis.” [2] The crisis is framed as a technical problem: there is not enough high-quality training data for AI models. But the crisis is actually an epistemological one. The knowledge that needs to be captured was never stored anywhere. It lives in the embodied judgment of people who have spent careers learning to read situations that no sensor can fully capture, no camera can fully frame, and no lidar scan can fully represent.
Antioch, a startup building simulation tools for model builders, suggests physical AI is in its “GPT-2 era” — the OpenAI model that predated ChatGPT by several years. [3] That framing is meant to be reassuring, suggesting that more data and more compute will eventually push the field over the hump. But GPT-2 was trained on text that already existed. The data for physical AI does not exist. It has to be created, captured, or — and this is the part that gets skipped — it has to be extracted from people who never knew they were generating it. The judgment of an experienced construction supervisor, a veteran warehouse manager, or a skilled machine operator is the raw material that this industry is mining, and the people who possess that judgment are not being compensated for the extraction.
The 80 Percent Trap
The hard numbers tell a story that the industry does not want to confront. General-purpose humanoid robots, the ones that promise to do anything a person can do, currently achieve an 80 percent success rate on basic tasks. That sounds almost good until you think about what it means in practice. An 80 percent success rate means that one out of every five laptop closures fails, one out of every five table cleanings leaves a mess, one out of every five pushes misses its target. No customer cares about a general-purpose robot that works at 80 percent success rate, as Théophile Gervet, the CEO of Genesis AI, put it bluntly at the Actuate conference. Genesis AI raised a $105 million seed round this year, which tells you how much money is flowing toward this problem. [4]
The 80 percent figure is not just a technical limitation. It is a judgment problem. The robot does not know which tasks it can do well and which tasks it cannot. It does not know when to ask for help, when to try a different approach, or when to stop and reassess. A human worker at 80 percent competence would be fired. A human worker at 80 percent competence would also know they were at 80 percent and would adjust their behavior accordingly, taking more time on difficult tasks, seeking assistance when needed, and communicating uncertainty to supervisors. The robot has none of that self-awareness. It simply attempts the task, succeeds or fails, and moves on to the next attempt without any recognition that its failure might have consequences.
This is where the judgment replacement becomes visible. The person who used to do the task was not just executing a sequence of motions. They were making continuous decisions about whether their current approach was working, whether conditions had changed, whether the task was worth completing at all. That metacognitive layer — the ability to judge your own performance while performing — is not something that end-to-end learning has cracked. The industry talks about “manipulation in the wild,” as Bedrock CTO Kevin Peterson describes his company’s excavation work, but the wild is not just a physical environment. [2] It is a judgment environment, full of ambiguous situations that require interpretation, not just execution.

The vertical versus general debate that is splitting the industry — companies like Gritt building solar farms, Agility deploying robots in industrial settings, and Bedrock operating excavators — is really a debate about where judgment lives. Vertical companies are betting that judgment can be embedded in the specific task, that the narrowness of the application will make the 80 percent problem manageable. General companies are betting that judgment will emerge from scale, that more data and more compute will eventually cross a threshold where the robot can handle novel situations. Both bets are about the same thing: who owns the judgment that used to belong to the worker.
The Invisible Transfer
The autonomous vehicle industry is furthest ahead in physical AI, and the reason is instructive. Self-driving cars can collect data from human-driven vehicles, capturing millions of miles of real-world driving without needing to build a single robot. But the deeper advantage is that the judgment being replaced — the decision to brake, to swerve, to yield — is relatively simple compared to the judgment of a construction supervisor or a warehouse manager. Driving is mostly about avoiding contact, not manipulating the physical environment. The judgment is reactive rather than generative, responsive rather than creative. That is why companies like Wayve and Uber are now launching robotics labs focused on humanoid form factors. They have solved the easy version of the judgment problem, and they are moving toward the hard version.
Wayve CEO Alex Kendall argues that it is too early to commit to any one hardware platform, that advances in sensors and components are coming quickly, and that a truly general model should be more agnostic. [6] That is a technical position, but it is also a judgment position. Kendall is saying that the hardware is not the bottleneck, that the value is in the model, and that the model is where the judgment will live. His company is licensing models to car makers, betting that a multibillion-dollar business in eyes-off autonomy for under $1,000 worth of hardware will fund the development of a truly general embodied AI model. The consumer moment he is waiting for — when you can buy a car that drives itself without supervision for the cost of a mid-range phone — is a moment when the judgment of every driver on the road becomes obsolete at once.
The transfer of judgment is not happening in a visible way. There is no ceremony, no announcement, no moment when a worker is told that their expertise has been extracted and encoded. It happens incrementally, through the accumulation of data, the refinement of models, and the gradual expansion of what robots can do without human intervention. The Foxglove announcement at Actuate — a new product built on top of Nvidia’s Cosmos open-weight world model that allows engineers to search visual and lidar data with natural language queries — is a tool for accelerating that transfer. [2] It lets model builders find the moments in their data where judgment was exercised, isolate those moments, and turn them into training examples. The expertise that took a human decades to acquire becomes a training set that can be replicated infinitely.
The Moment That Never Arrives
Sam
Altman recently said that the ChatGPT moment for physical AI is just a few years away. [7] That prediction assumes that physical AI will have a moment like ChatGPT had, a sudden inflection point where the technology becomes undeniably useful and captures the public imagination. But Foxglove CEO Adrian Macneil offers a different perspective. He points out that what made ChatGPT a moment was distribution — the company went from zero to a million active users in a week. Distribution in the physical world is harder. You cannot download a robot. You have to build it, ship it, install it, and maintain it. Macneil says he would be excited for the Apple II moment or the IBM PC moment in robotics, when you can buy a home robot that starts doing useful and fun stuff. That is a more modest vision, but it is also a more honest one.
The ChatGPT moment framing misses something important. ChatGPT did not replace judgment. It augmented it, giving people access to information and synthesis that would have taken hours to produce manually. Physical AI is different. It does not augment the judgment of the excavator operator or the warehouse worker or the construction supervisor. It replaces that judgment entirely. The robot does not help the operator dig better. It removes the operator from the loop. The judgment that was once the most valuable thing that person owned becomes a training dataset, and then becomes obsolete.
Alex Kendall’s vision of a robot that works out of the box, that you can talk to in natural language and have it perform any basic manipulation task with 80 percent plus reliability, is not a vision of augmentation. [6] It is a vision of replacement. The person who used to close the laptop, clean the table, or push the object is gone. The judgment that person exercised is now embedded in the model. The 80 percent threshold is not a technical milestone. It is the point where the replacement becomes economically rational, where the cost of the robot plus the cost of the model is less than the cost of the human plus the cost of the human’s judgment.
The final consequence that no one draws is this: the judgment being replaced was never just about the task. It was about the person. The excavator operator’s judgment about soil conditions was also judgment about safety, about when to stop, about when conditions were too dangerous to continue. The warehouse worker’s judgment about how to stack boxes was also judgment about how to protect their body from injury, how to work efficiently without exhausting themselves, how to pace themselves through a long shift. The construction supervisor’s judgment about how to sequence work was also judgment about how to manage people, how to motivate them, how to recognize when someone was struggling. None of that judgment is in the training data. None of it can be extracted from lidar scans and video feeds. And none of it will be replaced by a robot that achieves 80 percent success on manipulation tasks.

The industry is building machines that can do tasks. It is not building machines that can hold judgment. The distinction matters because the tasks are visible while the judgment is invisible. When Unitree’s valuation collapsed, the market was responding to the visible failure — robots that could not do value-creating work. But the invisible failure is more profound. The judgment that is being extracted from workers, encoded into models, and then used to make those workers obsolete is not being replaced. It is being lost. The simulation tools, the world models, the natural language search interfaces — all of these are ways of capturing the output of judgment without capturing the judgment itself. The result is a world where the tasks get done but the understanding disappears, where the 80 percent success rate is celebrated as progress while the 20 percent failure rate is where all the judgment used to live.
Sources
1. Unitree
2. Foxglove
3. Antioch
4. Genesis AI
6. Wayve
7. Nvidia
