XPENG says IronMind turns 10,000 hours of first-person video into better humanoid grasping. It is still a lab result.
The paper reports stronger manipulation results after camera-space pretraining, but it does not show a household robot ready to work unsupervised.
The 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining.IronMind paper, XPENG Robotics
The result
XPENG Robotics has posted IronMind, a research system for teaching a humanoid robot dexterous manipulation from a large collection of first-person human video and robot demonstrations. The claim is narrow but useful: rather than trying to translate a human body directly into a robot body, the system represents actions from the point of view of the head-mounted camera.
The accompanying paper, submitted September 30, says the team trained on more than 10,000 hours of egocentric video and heterogeneous robot data. In its reported real-robot evaluation, the largest pretraining run reached 55.0% success across six out-of-distribution manipulation tasks. The authors say smaller runs of up to 5,000 hours reached at most 11.7%, while a version without pretraining reached 5.0%.
Those are research results from the authors, not a consumer-product claim. IronMind ran on XPENG's IRON-R01 humanoid after a separate post-training stage with teleoperated robot demonstrations. The work does not announce a home robot, a sale date, or an independently audited benchmark.
How it works
First-person video is attractive robot-training material because people generate a great deal of it while opening drawers, picking up objects, and moving through ordinary rooms. It also has an awkward omission: a camera can see hands and objects without recording the wearer's full torso position. Conventional human-to-robot retargeting needs assumptions about that missing body geometry.
IronMind's answer is to keep both human and robot movements in camera space. The project page says robot trajectories are calibrated into the head camera's frame, while human and robot actions are aligned by meaning, including wrist pose and finger flexion. This avoids requiring a direct, joint-for-joint correspondence between a person and a robot.
The authors also say they filtered poor recordings, split videos into atomic tasks, re-captioned clips, adjusted motion speed, and weighted frames by quality. That matters because more raw video is not automatically better training data. A robot learning to grasp a mug needs demonstrations with a visible task and usable motion, not hours of someone looking around a kitchen.
What it means for a home robot
The most interesting part for a future household machine is the data question. Home work is varied: the same cabinet, towel, bottle, or countertop looks different across homes and changes position from one minute to the next. A method that gets useful robot training signal from human video could make it easier to cover that variety than collecting every example through teleoperation.
But the paper's result is not a dependable chore. A 55.0% aggregate success rate across six test tasks means many attempts still fail. The paper does not establish long-run reliability, safe operation around people, recovery after a dropped object, or performance in a stranger's home. Those are the things an owner would need before trusting a robot with daily work.
The result also depends on a target-robot adaptation step using teleoperated trajectories. In other words, first-person video is a useful starting point, not a replacement for robot-specific data and testing.
Caveats
IronMind is a preprint and project release from the team that built it. We found no independent replication or peer-reviewed publication. The reported task results come from the authors' own protocol, including a post-training stage on their target robot.
The project page shows fully autonomous demonstrations at normal speed, which is worth watching as evidence of what the lab tested. A demo is still different from a measured record of repeated operation in homes. The authors' own paper frames the work as a foundation for humanoid manipulation, not evidence that a consumer humanoid can complete household work reliably.
For now, IronMind is a data-and-training result: a credible attempt to make first-person human video more useful to humanoid hands. It narrows one gap between seeing a person perform a task and teaching a robot to attempt it. The larger gap is turning that attempt into safe, repeatable work in a real home.
We labeled this story Research — A lab result, not a product. How we label claims