HygieneRoboBench asks whether a household robot remembers what it touched. It is still research.
A new benchmark tests whether a planner can use a robot's contact history to choose a safe next household action, rather than merely finish the task.
A robot that has touched a contaminated tray should not handle fruit as if nothing happened. That sounds obvious to a person. It is less obvious to a task planner that mainly measures whether the fruit reached the plate. A research team has built HygieneRoboBench, a benchmark intended to test that missing piece: whether a household robot can remember contact history, update its plan when new contact information appears, and still respect the homeowner's priorities. It is a planning benchmark and a controlled demonstration, not a consumer robot that can safely manage a real kitchen.
The result
The authors describe 624 task instances in 134 task families, spanning five household activity groups and seven household areas. The benchmark changes one relevant condition at a time: what a gripper or object touched earlier, which resource a user wants conserved, or whether a new contact event appears during a task. The central question is not simply whether a system completes a request. It is whether its next action remains safe after the state of cleanliness has changed.
The paper also introduces Hygiene-NSP, a planner that combines language grounding, reconstructed contact history, and constraint solving. In the reported benchmark evaluation, the authors say it reached a 94.4 percent safe-resolution rate and a 90.4 percent optimal-safe-resolution rate. Those are results within the authors' own benchmark, not a measure of household reliability. The project page shows a real-robot case alongside simulation, but it does not establish that a robot can recognize every real-world contaminant, follow public-health guidance in an unstructured home, or clean safely without supervision.
How it works
HygieneRoboBench represents a home task with a contact history, hygiene rules, treatment options, costs, and user priorities. A planner must work out what is safe from that history before deciding whether to wash, replace a contact pad, put an object down, or continue. The project illustrates the point with a spoon and fruit: the same immediate scene can call for a different next move depending on whether the robot washed after touching a contaminated tray.
That framing matters because a camera snapshot cannot reliably answer every hygiene question. A robot may need a record of where its hands, tools, and shared objects have been. The benchmark then asks whether the planner changes course when it learns a new contact event, instead of treating the original plan as permanent.
The system treats hygiene as a formal planning problem. That makes it useful for comparing planners under the same rules, but it also means its contamination and treatment rules are modeled assumptions. Real homes involve uncertain surfaces, allergens, raw food, cleaning chemicals, pets, children, and incomplete observations. A benchmark can expose a planning failure without solving the sensing and verification problem that creates the record in the first place.
What it means for a home robot
A useful home robot will need more than a good grasp and a convincing task demo. It will need to know when an earlier action changes what is safe to do next. HygieneRoboBench gives researchers a more demanding way to test that reasoning than a simple success-or-failure chore score. It also makes user tradeoffs explicit: conserving water, time, cleaner, or replacement parts can lead to different acceptable plans.
For a buyer, this is a reminder to ask a plainer question than whether a robot can load a dishwasher on video: what does it do after touching something it should not carry to the next object? No current result here answers that question for a commercial home robot.
Caveats
This is research, submitted to arXiv on October 6. The reported rates come from a benchmark created by the authors and should be read as evidence about the tested setup, not as an independent certification. The project page includes a real-robot case, but the work is principally a benchmark and planner evaluation. It does not announce a product, customer deployment, price, or timeline.
The useful contribution is narrower and more durable: it gives household-robot researchers a way to penalize plans that reach the right destination by taking the wrong hygienic path. Before that becomes a consumer feature, robots will still need dependable perception, contact tracking, cleaning procedures, and a conservative way to handle what they do not know.
We labeled this story Research — A lab result, not a product. How we label claims