Skip to content
ShinyMetal
← Research explainers
Research explainerResearch

InterEvolve lets a Unitree G1 revise task rewards at test time. It is still research.

A UIUC research project changes staged reward programs instead of retraining a humanoid controller, then shows a Unitree G1 perform a narrow physical demonstration.

What the paper saysEvolve the task, not the controller.InterEvolve project page
Illustration generated with AI. Not a photo of a real product.

A humanoid does not need a new neural network every time a box starts in a different spot. That is the premise behind InterEvolve, a University of Illinois Urbana-Champaign research project posted October 1. The system keeps its whole-body controller fixed and instead rewrites the staged reward program that tells the controller what to try. The authors show the resulting policy on a physical Unitree G1, but this is a research result, not a new G1 feature or a household robot product. The paper and its project page are the underlying sources.

The result

InterEvolve treats a task as a small program: stages, completion conditions, and numerical constants that define what counts as progress. An LLM agent proposes revisions to that program after looking at execution feedback. A numerical optimizer adjusts the constants. Candidates run in parallel simulation scenarios, and the system keeps a candidate only after comparing it with the current best program, according to the paper's method and evaluation description.

That is a different claim from a robot learning a new motor policy on its own. The behavioral controller is frozen. The authors argue that a broad controller may already contain useful motions, and that a better task specification can expose them. Their project page puts it more plainly: “Evolve the task, not the controller.”

The reported work includes simulated, contact-rich tasks such as handling boxes and arranging them over multiple steps. The project page also shows a physical Unitree G1 repeatedly kicking a box after it moves. It says the robot uses onboard perception except where noted; the page specifically notes motion capture for box pose in that kicking demonstration. That detail matters. A demo with an external object-position measurement is not the same as a robot finding every object and completing a home task under ordinary household conditions.

How it works

The system has two layers. First is an object-aware behavioral foundation model, which turns the reward program into whole-body motion. Second is the evolution loop, which has the LLM revise the structure of a reward program and an optimizer tune its values. Feedback from simulated rollouts informs the next proposal, while verified programs enter a skill library for later tasks. The paper says this allows prior programs to be reused or composed for later tasks without updating the controller weights.

In practical terms, that could reduce the work required to adapt a capable research robot to a variation of a task it already partly understands. A robot that can carry and place a box may need a different sequence of objectives when the box is smaller, placed elsewhere, or part of a longer arrangement. InterEvolve is an attempt to search those objectives rather than retrain the robot from scratch.

What it means for a home robot

The home-robot implication is modest but useful. Homes create endless small variations: a hamper is in a new place, a cupboard door is partly open, or an object is a different size. A controller that can safely reuse previous capability could be more practical than one that needs a new training cycle for each variation. The authors' examples show that kind of reuse across boxes, suitcases, arrangements, and grasping scenarios.

But the project does not establish reliable domestic autonomy. The paper's main experiments are simulation evaluations, and the physical demonstration is narrow. It does not report a long run of unsupervised chores in unfamiliar homes, recovery from household clutter, or safety performance around people and pets. It also does not make a claim about a product ship date, price, or consumer availability.

Caveats

InterEvolve is a preprint, not an independently reviewed product evaluation. Its reported results should be read as evidence that the method is promising within the authors' experimental setup, not evidence that a Unitree G1 can now take on open-ended chores. The system also spends computation at test time to search and validate reward programs; the paper does not turn that cost into a household deployment claim.

For now, the useful takeaway is narrower: researchers are testing ways to make fixed humanoid controllers more adaptable by changing the task description around them. The real G1 demonstration gives that idea some physical grounding. It does not close the gap between a box-kicking demonstration and a robot that can be trusted to work alone in a kitchen.

We labeled this story Research — A lab result, not a product. How we label claims

Sources

  1. InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation1 Oct 2026
  2. InterEvolve project page1 Oct 2026
  3. InterEvolve arXiv record1 Oct 2026

Robots in this story

Read nextAll news