Skip to content
ShinyMetal
← Research explainers
Research explainerResearch

Figure's Helix 2.5 did chores in 30 homes it had never seen, and finished about half

A model pretrained on human video let Figure 03 tidy, fold towels and make beds in unfamiliar homes at 56% full-task success, up from 9%.

What the paper saysA person can walk into an unfamiliar home and start working immediately. This is currently not true for robots.Figure, Helix 2.5 announcement, 2026-09-17
Illustration generated with AI. Not a photo of a real product.

The result

Figure introduced Helix 2.5 on September 17, 2026, and the headline number is modest on purpose: 56%. That is the share of full chores its Figure 03 humanoid completed in 30 Bay Area homes it had never been inside, across three jobs: tidying a living room, folding towels and making a bed. A comparable model trained without Figure's new pretraining data managed 9% on the same test.

The scoring was strict. Figure says a tidy only counted if "all 13-15 toys scattered in the scene are picked and placed in the basket," a towel run only counted if every towel was folded and put away, and a bed only counted if both pillows and the comforter corners were at the top with the comforter pulled smooth. Half a job was a failure. If a person had to step in for safety, the run was aborted and marked as failed.

Figure also says Helix 2.5 matched the success rate of its January model, Helix 02, "while using half as much adaptation data," and did it without collecting any data in the test homes.

How it works (plain words)

Most robot learning so far has worked like private tutoring. You bring the robot to a place, show it the task many times in that place, and it gets good at that place. Move the furniture and the skill often falls apart.

Helix 2.5 changes the order of operations. First, Figure pretrains one large model on Index, a dataset of people doing ordinary physical tasks, recorded through a phone app. Figure took Index out of stealth on August 25, 2026, and says it has 264,000 app downloads in 108 countries, more than 16 million uploaded videos, and $15 million paid to the people who record them. Figure says that per 1,000 hours, the data covers 373 unique tasks, 1,146 unique objects and 116 unique environments.

Then Figure teaches each chore with a smaller batch of robot data collected elsewhere, and sends the same model into homes it has never seen. According to The AI Insider's write-up, no single evaluation task made up more than 1.9% of the pretraining data, so the model was not simply memorising towel folding from human video.

The bet is the same one language models made: watch enough of the world and you pick up the general shape of how things work, then a small amount of specific teaching goes a long way. Figure says it has committed $3.5B of compute to training Helix.

What it means for a home robot

This is the first result we have seen that tests the thing buyers actually care about: does the robot work in my house, not in the maker's showroom? Thirty homes is a small sample, but they were real homes with real layouts, and the model got no warm-up in any of them.

For a future owner, three things stand out.

  • Setup could get shorter. If a robot needs less teaching per home, the installer visit (or remote teaching session) that early home robots rely on could shrink.
  • Whole-body chores are on the table. Bed making and floor-level tidying mean walking, crouching and using both hands together. Those are exactly the jobs a wheeled arm on a counter cannot do.
  • The data flywheel is people, not robots. Index pays humans to film their chores. That scales faster than building robots to collect data, and it is why Figure can talk about variety across homes at all.

Caveats

A 56% full-task success rate is a research number, not a product number. Roughly one chore in two ended with something left undone, or with a person stepping in. You would not keep a dishwasher that finished one load in two.

Figure has not published per-task success rates, completion times or speeds in its announcement, so we cannot tell whether towels went better than beds, or whether a tidy took two minutes or twenty. The objects were chosen by Figure and checked to be absent from training, which is good practice, but the homes were all in one region. "Zero-shot" here means no data from the test home; each chore was still taught with robot data gathered elsewhere.

There is also a privacy question the announcement does not answer. Index is built from videos people record inside their own homes and workplaces, and Figure's Index page does not describe how that footage is stored or used beyond training.

Figure's own framing is fair: "The point is not that general humanoid robotics is solved." Figure 03 is not on sale to households, and Helix 2.5 is a research result. We are labelling this Research, and we will want to see the per-task breakdown and an outside evaluation before reading it as a buying signal.

We labeled this story Research — A lab result, not a product. How we label claims

Sources

  1. Helix 2.5: Zero-Shot 30-Home Generalization17 Sep 2026
  2. Introducing Index: Building The World's Largest and Most Diverse Physical Dataset25 Aug 2026
  3. Figure Unveils Helix 2.5 With Zero-Shot Humanoid Generalization Across 30 Homes17 Sep 2026
  4. Introducing Helix 02: Full-Body Autonomy27 Jan 2026

Robots in this story

Read nextAll news