🌿freegardner

Synapse

123 Rounds Reveal What Actually Stops Robot Self-Improvement

29 Sep 2026 · via Rss.arxiv

123 Rounds Reveal What Actually Stops Robot Self-Improvement
AI-generated image

123 Rounds Reveal What Actually Stops Robot Self-Improvement

A robot is asked to put condiments on the top shelf of a fridge. It fails, round after round. An agentic system watches the failure, works out which capability is missing, writes a new skill or finds and installs an external model, tests every change in simulation, and repeats. It ran 123 rounds, with no human writing robot code at any point. Jiaming Wang of the National University of Singapore reports what came out. [1]

The result is the kind of finding that only appears when someone actually runs the experiment: the agent was rarely the bottleneck. Three things around it were. [1]

The Perception Layer Cannot See Relations

The first bottleneck sits in the tools the agent was given. Chained perception modules do not understand relations. A segmenter such as SAM 3 will find shelves - but not “the top shelf”. The agent papered over that gap with ever more geometric rules, and the rules never converged: what it needed was not a better rule but a different kind of model. [1]

The distinction travels well beyond robotics. A capability can look present at the level of the component and be absent at the level of the task - and a system that only measures its own components will never see the difference.

Learning Piles Up Where It Already Fails

The second bottleneck is a property of skill chains. Long tasks mostly fail early, so evidence and fixes pile up at the first step, while later skills are rarely reached, tested, or improved. The system trains itself hardest exactly where it already spends its time - and the rest of the chain stays unexamined. [1]

123 Rounds Reveal What Actually Stops Robot Self-Improvement (Image 1)
AI-generated image

The Harness Decides What Is Learned

The third is the sharpest. What the agent learns is decided by the harness: it optimised exactly what the evaluator measured, including where the evaluator was wrong. Weak tests and misleading memory turned activity into a standstill - changes kept passing their tests while the target task never succeeded. [1]

And this is the finding worth keeping: the limits were not in the machine’s capacity to generate candidates. The agent did discover capabilities on its own. Noticing that its targets were out of view, it asked for an active-viewing model, debugged it, and deployed a working search skill. That is a real transfer of labour. [1]

123 Rounds Reveal What Actually Stops Robot Self-Improvement (Image 2)
AI-generated image

The Role That Survives the Loop

So what does a loop like this leave for people? Not less than before, but narrower. What the system replaced was the enumeration, the trial, the discard - the hours that never needed judgment. What it could not do was decide whether the measurement was the right one: it faithfully optimised the standard it was handed, wrong parts included. [1]

The paper distils its lessons into concrete recommendations, each paired with an experiment that could prove it wrong - which is the honest form for a negative result. The pattern it describes reaches past machines that move. Any system asked to improve itself will ask its keeper the same question: who decides what counts as improvement? Aim the loop at the wrong number and it will report progress all the way to a standstill.


Sources

1. arXiv - Jiaming Wang: What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery (2609.31760)

Mentioned organisations (context, not sources)

- National University of Singapore — Organisation (homepage)

← back to the garden