Skild's S1 learns a new task from one video of a person doing it
Skild AI says its S1 model picks up unseen tasks like repotting a plant from a single first-person video, with no retraining.
A single demonstration in context is worth roughly 380 post-training examplesSkild AI, S1 blog post, August 2026
The result
Skild AI says its new robot model, S1, can learn a chore it has never seen from a single video of a person doing it, with no retraining. Skild introduced S1 in August 2026, and NVIDIA wrote it up on September 10, 2026.
The tasks Skild showed are homey: repotting a plant, flipping pancakes, making pour-over coffee and assembling a kit. Skild says the tasks run up to 10 minutes and were never seen during pretraining. The prompt for each was "one egocentric human video demonstration," meaning a first-person recording of a person's hands doing the job.
The key numbers, as Skild reports them:
- On unseen tasks, "one example in context puts it at a 66% success rate." NVIDIA describes this as a per-step success rate, compared with 9% for a model prompted with words alone.
- "A single demonstration in context is worth roughly 380 post-training examples." NVIDIA puts that at 50 to 100 hours of manual data collection saved.
- For the plant-potting task, "the time from demonstration to autonomous execution was 11 minutes." Skild adds that most of its time went into moving furniture and setting up the scene.
How it works (plain words)
Today, teaching a robot a new job usually means fine-tuning: collect dozens or hundreds of robot demonstrations, run a training job, ship an updated model. That takes days and a data team.
S1 skips the training step. It treats the demonstration video the way a chatbot treats the text in your prompt. The video goes into the model's short-term context, the model's weights stay frozen, and the robot works out how to do the same job with its own body in the scene in front of it. Researchers call this in-context learning. Skild's pitch is simple: "Show it a video of a task, short or long, seen or unseen, and it executes."
The hard part hidden in that sentence is translation. A human hand is not a robot gripper, the camera angle is different, and the kitchen in the video may not match the kitchen the robot is standing in. S1 has to map what it sees a person do onto what its own arms can do. Skild says the benefit grows as it scales pretraining, and reports 96% success on seen tasks at scale.
What it means for a home robot
If this holds up, it changes who can teach a home robot. Every home has a few jobs nobody else has: the way you like the dog bowl rinsed, the odd latch on the pantry, the plants on the back step. A robot that needs a data-collection team for each of those jobs will never learn them. A robot that can watch you do it once, on your own phone, might.
It also changes the teaching model that early home humanoids rely on. 1X plans to use remote teleoperators to guide its NEO through chores it doesn't yet know. One-video learning points toward a future where the owner is the teacher and nobody at the company has to drive the robot through your home.
Caveats
This is a company blog result, not a peer-reviewed paper, and Skild's post does not describe the robot hardware it used or how many trials sit behind each number. We could not find failure cases in the post.
Read the 66% carefully. A per-step success rate is not the same as finishing the whole job. On a task with dozens of steps, a robot that gets two steps in three right will rarely finish on its own. That makes this a very different number from Figure's whole-chore 56% for Helix 2.5, and the two should not be compared directly.
Skild also sells into industry first. NVIDIA reports that Skild has 60 or more deployment partnerships across manufacturing, logistics, inspection, security and food preparation, and other coverage says S1 is going to a limited set of industrial partners first. There is no Skild home robot you can buy, and none has been announced.
So we are filing S1 under Research. The idea that matters for your home is teaching by showing, and on that, S1 is the most concrete claim we have seen this year.
We labeled this story Research — A lab result, not a product. How we label claims
Sources
- Introducing S1: In-Context Learning for RoboticsAug 2026
- Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video10 Sep 2026
- Skild AI unveils S1 robot foundation model that learns tasks from video demonstrations14 Sep 2026
- NEO humanoid designed for household use, available for preorder30 Oct 2025
- Helix 2.5: Zero-Shot 30-Home Generalization17 Sep 2026