Skip to content
ShinyMetal
← Research explainers
Research explainerResearch

RPG lets a robot practice in simulation before it folds a towel. It is still research.

A new research system improved its reported manipulation results through simulated practice and then completed 30 controlled physical trials, but the work is not a home robot product.

What the paper saysOn held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds.Wang et al., arXiv
Illustration generated with AI. Not a photo of a real product.

The result

A research team from UC Berkeley, Amazon FAR, MIT, and the University of Chicago says it can improve a robot’s task system by having it practice related work in simulation before taking the revised system to hardware. The project is called Reconstruct, Practice, Go Real, or RPG. Its result is a research claim, not a product announcement: there is no robot for sale here, and the authors say their code is still coming soon.

The central reported number is substantial. In the paper, the authors say their system moved from 28.6% success after its first practice round to 95.0% after 15 rounds on held-out initializations of 22 simulated manipulation tasks. After calibration and hardware adaptation, they report 30 successful physical trials out of 30, spread across three tasks.

Those tasks matter because they resemble fragments of a home robot’s job: putting a ball into a drawer and closing it, folding a towel, and moving a bowl from one hand to the other. They are not a demonstration of an autonomous robot that can handle a whole room, a whole laundry load, or an unfamiliar kitchen.

How it works

RPG does not retrain its underlying model weights. Instead, it starts with an offline dataset, identifies a manipulation capability to work on, and constructs a related simulation task. During practice, it uses execution feedback, simulator state, and available videos to diagnose failures. It can add a reusable symbolic skill, refine an existing one, or revise the system prompt that coordinates perception and control.

The authors then test individual changes and combinations across tasks before keeping them in the shared skill library. That validation step is the interesting part. Robot demos often show a single polished behavior; this system is designed to reject a change if it helps one task while hurting the rest of the tested set.

The project site gives the physical-trial detail the paper summarizes. For the drawer task, success required the ball to be inside and the drawer fully closed. For the towel, the final footprint had to fall within a specified size range and have roughly parallel sides. For the bowl, the robot had to complete the transfer and hold it with its right hand for at least 3 seconds. Those are useful definitions, but they are still controlled criteria chosen by the researchers.

What it means for a home robot

The home-robot relevance is less about the three chores than the workflow. A useful household machine cannot be re-engineered from scratch every time it misses a drawer handle or changes its grip on a towel. If simulated practice can safely produce skills that survive transfer to hardware, it could reduce some of the costly human tuning behind each new task.

RPG’s physical examples also underline why task labels need care. “Fold towel” is a bounded test with a prepared robot, calibration, and a measurable endpoint. It is not evidence that a robot can sort a family’s mixed laundry, recover from a dropped garment, and put everything away without help. “Store ball in drawer” is likewise not evidence of general tidying.

The authors report that all methods received the same calibration and hardware-adaptation procedure. Their project page compares RPG with two CaP-Agent0 baselines and reports stronger completion results for RPG in these three trials. That is a reasonable experimental comparison, but it does not establish performance against every robot-learning system or every home layout.

Caveats

First, this is a preprint. It has not been presented here as a peer-reviewed consumer-robotics result. Second, the strongest 95.0% number is from 22 simulated tasks, not a multi-week home trial. The physical result covers 30 trials, ten each on three narrow tasks.

Third, the paper’s workflow relies on assets that ordinary homeowners do not have: an offline dataset, a simulator, calibration, hardware adaptation, and evaluation criteria. The authors explicitly describe a frozen system being deployed after those steps. That is a research pipeline, not a feature an owner can turn on.

Finally, no tracked consumer robot has adopted RPG. The project is useful evidence that simulation can help improve reusable manipulation skills without changing model weights. It is not evidence that today’s home humanoids are ready to practice their way through an unsupervised day of chores.

For now, the right reading is modest: this is a well-specified step toward more adaptable robot behavior, with real hardware tests worth watching. The gap between a 30-trial experiment and a dependable domestic appliance remains the whole story.

We labeled this story Research — A lab result, not a product. How we label claims

Sources

  1. Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents1 Oct 2026
  2. RPG: Reconstruct, Practice, Go Real project site3 Oct 2026

Read nextAll news