HomeBody gives a Unitree G1 a memory of the kitchen. It is still a research demo.
A Stanford and Caltech project combines spatial memory with reusable skills so a humanoid can work across an unfamiliar room, while leaving the hard reliability questions open.
In a previously unseen kitchen, our system allows a Unitree G1 guided by GPT Astra to clean up across the room and retrieve a remembered object from an underspecified request, without environment-specific training data or additional policy learning.HomeBody project page, Stanford TML
The result
A humanoid can only do a household job if it can keep track of a room after its camera view changes. That is the problem addressed by HomeBody, a research system from Stanford and Caltech. The project connects a Unitree G1 to a frontier vision-language model, a persistent map of the room, and a small library of physical skills.
The team shows the G1 in a previously unseen kitchen completing two multi-step demonstrations. In one, it gathers coffee bags and discards specified cartons. In another, it finds medicine stored out of view, opens a drawer, hands the item over, and disposes of a carton. Those are research demonstrations, not evidence that a buyer can order a household robot to manage a kitchen.
The interesting part is where HomeBody puts the intelligence. Rather than using a learned vision-language-action policy to turn every observation into robot commands, it asks a high-level model to choose among reusable skills such as navigation, picking, placing, opening a drawer, and picking from a drawer. The project page calls this a replacement for the middle layer of the usual VLM-to-VLA-to-controller stack.
How it works
Before it receives the household request, the robot explores the room. HomeBody records camera observations, LiDAR and SLAM geometry, joint poses, and selected waypoints. It then uses that material to construct a digital twin in Isaac Sim. That stored spatial context gives the planner somewhere to look when an object or drawer is no longer visible from the robot's current position.
The division of labor matters. The model chooses the next skill and can revise its plan after a failure; the skill library and lower-level controller handle the movement. It is a practical answer to a recurring home-robot problem: long tasks require a robot to remember where it has been, but continuous low-level control is not the same job as deciding what to do next.
AI Weekly's report, published September 27, identifies five skills in the demonstration and describes the project as a direct challenge to the assumption that a learned action model must sit between a high-level model and the robot controller. That interpretation is fair, but it should not be confused with a product comparison. HomeBody evaluates one system in a constrained research setting.
What it means for a home robot
The demos are closer to a useful household task than a single scripted grasp. Cleaning up scattered objects and retrieving something from a drawer require navigation, object selection, arm choice, manipulation, and state tracking across a room. The medicine demonstration also turns on an ordinary household ambiguity: the item is not visible when the person asks for it.
That does not make the result a general home assistant. The project page does not present a public success-rate table across homes, task variants, or long unattended runs. It does not establish that the robot can safely deal with arbitrary clutter, pets, children, stairs, or objects outside its skill library. We also do not know the intervention rate when a grasp, map, or door interaction fails.
HomeBody is therefore best read as a research result about system architecture. It suggests that a robot may not need a custom action model for every new room if it can first build a dependable spatial memory and call skills that already work. The consumer question is harsher: whether that memory, those skills, and the hardware keep working after the kitchen stops looking like the test kitchen.
Caveats
The researchers are unusually clear that the approach has operational costs. Building the Real2Sim reconstruction takes setup time and uses model APIs. High-level model reasoning introduces pauses between skill executions. The accompanying report also notes finger-servo overheating during extended operation. Those are not cosmetic details. A household robot needs to work repeatedly, with acceptable speed, before a tidy-up demo becomes a home appliance.
The G1 in this work remains a research platform. HomeBody adds evidence that long-horizon household behavior can be assembled from memory and reusable skills, but it does not change the status of consumer home humanoids. We will look for repeatable evaluation across more homes, clearer failure data, and evidence that the stack can operate without a lab team nearby.
We labeled this story Research — A lab result, not a product. How we label claims
