DigestAI news desk
OpenAI board member warns company is not on track to prevent catastrophic AI loss of control OpenAI launches Agents API beta for long-running cloud agents OpenAI solves Navier-Stokes problem, sparking academic controversy over data use RTK Token Savings Debunked: Cost Benchmarks Disagree The Waymo effect: AI making research less collaborative Meta’s Muse AI Agent Seeks User Trust with Secure Architecture Thelio Mira AI Linux Workstation with 192 GB GPU Memory Claude No Longer Available to Minors
Agents & Tools updated 4 min read

Skild AI launches S1 foundation model that learns robot tasks from a single video

Skild AI unveiled its S1 robot foundation model, which can pick up previously unseen, long‑horizon tasks from just one video demonstration. Built on NVIDIA’s AI infrastructure, the model uses in‑context learning—no weight updates or task‑specific retraining—allowing an operator to record a short clip and have the robot execute the task autonomously. Demonstrations include plant potting, pancake…

1 source primary source

Key points

  • S1 learns new, long‑horizon robot tasks from a single video without weight updates.
  • In tests, S1 achieved ~66% success per step, over seven times better than a comparable AI system.
  • Skild AI reached a $100 M annual revenue run rate within ten months of its first deployment.

In internal tests, S1 achieved roughly 66% success at each manipulation step, more than seven times the performance of a comparable AI system that managed only 9%. The company reports that a single video can replace about 380 hand‑collected training examples, saving 50‑100 hours of manual data collection. Skild AI has already reached a $100 million annual revenue run rate within ten months of its first commercial deployment and counts over 60 partnerships across manufacturing, logistics, inspection, security, and food preparation.

The launch is backed by a broader collaboration with NVIDIA, leveraging tools such as Isaac Lab, Cosmos, Omniverse, and the Newton physics engine to generate synthetic data, simulate edge cases, and accelerate inference. Together they aim to move adaptable robot intelligence from research labs into dynamic factory floors, exemplified by a joint deployment with Foxconn for high‑precision assembly of NVIDIA Blackwell systems.

The story so far

7 episodes →
  1. Skild AI launches S1 foundation model that learns robot tasks from a single video this story
Full story from NVIDIA Blog · by Sasa Docca primary source Open source ↗

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

NVIDIA Blog · 10 September 2026

Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming.

Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses video as input to understand and execute the task without updating its weights or undergoing task-specific post-training — a technique called in-context learning.

Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment. The companies are working together to move adaptable robot intelligence from the lab into factories and other dynamic operating environments.

“Learning by experience, and not preprogramming, is the step change that has happened in robotics,” said Deepak Pathak, cofounder and CEO of Skild AI. “NVIDIA Isaac Lab and NVIDIA Cosmos technologies help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments.”

The launch comes as the company reached a $100 million annual revenue run rate 10 months after its first commercial deployment. In that time, Skild has built more than 60 deployment partnerships with work spanning manufacturing, logistics, inspection, security, food preparation and other applications.

Learning New Work From One Video

Most industrial robots are built for fixed jobs, so each new product, process or layout requires more data, retraining and validation.

S1 takes a different approach: An operator records a video of the desired task and provides it to the model as a prompt. It interprets the demonstrated intent, objects and sequence, then maps them into actions for the robot in front of it — with no retraining — and often for a task not covered by its pretraining dataset.

S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing and kit assembly. These tasks can span dozens of manipulation steps and require the robot to compose skills in sequences it hasn’t previously performed.

In one plant-potting test, the Skild AI team moved from recording the demonstration to autonomous execution on hardware in just 11 minutes. The model can also adjust when objects move, recover from errors and combine skills in sequences that weren’t explicitly programmed.

In Skild’s tests on new, multistep tasks, its S1 robot succeeded about 66% of the time at each step, compared with 9% for a similar AI system — a more than sevenfold improvement. Skild also estimates that showing the robot one short video example can be as useful as giving it roughly 380 hands-on training examples. A person collecting those examples manually could take 50-100 hours.

From Research to Factory Work

S1 breaks the cycle of needing to constantly retrain robots for new factors by letting operators demonstrate new tasks directly without requiring a new dataset or training run for every change. Where customer agreements permit, experience from Skild’s commercial deployments can inform the broader model and help accelerate future deployments.

That work is already in action on the factory floor. Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances across a multistep task. The work requires precise motion, contact-aware control, sequence tracking and recovery when the scene differs from the plan.

NVIDIA Technology Across the Development Cycle

NVIDIA accelerated computing gives Skild the scale to train its shared robot brain using simulation, human video, teleoperation and, where permitted, deployment data. NVIDIA Cosmos open world foundation models help diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale.

Skild is extensively using NVIDIA’s open simulation frameworks to train and validate its robot brain before real-world deployment. NVIDIA Omniverse libraries and the NVIDIA Isaac Sim framework provide physically based virtual environments for generating data, testing edge cases and validating behaviors.

Skild further strengthens the skills of its brain through reinforcement learning in Isaac Lab, an open modular robot learning framework. Powered by the Newton physics engine, Isaac Lab helps Skild’s engineers accurately model various physical parameters, such as forces, contact, collision and pressure, and reduce the simulation-to-reality gap.

Skild and NVIDIA are also jointly developing new GPU-accelerated simulation solvers that quickly and accurately model how robots physically touch, grip and manipulate solid objects. They’ll soon be made available to all developers as part of Newton.

As models move toward production, NVIDIA Nsight tools help engineers find performance bottlenecks during training, and the NVIDIA TensorRT software development kit optimizes inference so robots can respond quickly in the physical world. Together, these technologies connect the data, simulation, training and deployment stages instead of treating them as separate systems.

Read Skild AI’s S1 research and explore the NVIDIA Isaac robotics platform.

This text was published by NVIDIA Blog and written by Sasa Docca. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

Related stories