
Skild AI Uses Single Video to Teach Robots New Tasks
Skild AI partnered with Nvidia to deploy a foundation model that enables industrial robots to learn physical assembly chores from one video recording.
Umar Abubakar | 23 Sept. 2026 · 7 min read

Programming an industrial machine usually requires massive amounts of time and specialized engineering. When a factory changes a product line, the technical team must write entirely new instructions or spend days manually guiding a mechanical arm to record exact movements. This manual recording process, known as teleoperation, creates a severe bottleneck for industrial automation. Skild AI recently introduced a completely different approach that bypasses this massive delay. The startup, backed heavily by Nvidia, launched a foundation model that allows physical machines to learn complicated, multi-step actions simply by watching a single video recording of a human performing the job.
This development changes the underlying timeline of factory automation. Instead of spending weeks collecting a new dataset and running a long computational training cycle, operators can use a short video as a direct prompt. Skild AI built the S1 model to interpret the intent of the human operator, identify the relevant objects on the table, and execute the proper sequence of movements. The software does not require engineers to adjust the underlying mathematical weights of the program. It relies on in-context learning, mapping visual observations directly into physical commands that the hardware can understand instantly.
Moving Past the Teleoperation Bottleneck
To understand why this matters, we must look at how the technology sector trains software. Large language models became incredibly smart because developers scraped billions of text documents from the public internet. Robotics never had a comparable source of native data. Engineers tried to solve this by manually driving robots around sterile laboratories to collect exact motor torque sequences. That strategy fails to capture the unpredictable nature of real physical spaces. A human driving a robot cannot easily generate enough data to cover every possible mistake or unexpected physical collision.
Deepak Pathak, the chief executive officer of Skild AI, openly stated that learning by experience is the only way to advance the field. His team saw that the internet already holds billions of instructional videos. Humans naturally learn to complete physical jobs by watching others. We do not need someone to calculate the exact force required to lift a coffee cup. We watch a demonstration, understand the goal, and map those actions to our own hands. The S1 model attempts to replicate this exact biological process for machines.
The Nvidia Simulation Infrastructure
Translating a video of a human hand into commands for a metal claw is a massive technical challenge. Skild AI relies heavily on the computing infrastructure provided by Nvidia to bridge this gap. The software uses the Nvidia Cosmos platform to break the video down into structured descriptions. It analyzes the visual data to understand exactly what the human is attempting to achieve during the recording.
Before a machine attempts the job in the physical world, the software tests the actions inside a highly detailed virtual environment. Skild AI utilizes the Nvidia Isaac Sim framework to generate synthetic training scenarios. These virtual worlds simulate accurate physics, allowing the software to practice gripping objects, calculating surface friction, and predicting collisions without breaking expensive hardware. The Newton physics engine calculates the exact pressure required for the mechanical hand to hold an object without dropping or crushing it. This rigorous virtual testing drastically reduces the risk of deployment errors on the factory floor.
Assembling Hardware on the Foxconn Line
The companies are already deploying this technology in high stakes commercial environments. Skild AI and Nvidia recently placed the S1 software inside dual-arm machines operating at Foxconn manufacturing facilities. These particular machines are actively assembling the highly advanced Nvidia Blackwell server systems.
The required assembly sequence demands exact precision. The hardware must install a busbar component, place a limit block, and correctly fasten sixteen different screws. It must remember the correct execution order and quickly adjust if a component sits slightly out of place. If a physical disturbance occurs on the table, such as a screw rolling slightly out of reach, the software identifies the deviation visually and adjusts the mechanical arm to retrieve it. This ability to adapt to minor changes prevents the entire assembly line from halting over a minor physical error.
Previously, programming a machine to handle this level of detail required massive engineering teams. Now, an operator simply records the assembly process, feeds the video into the software, and allows the system to generate the physical commands. If a component specification changes next month, the factory does not need to hire an external programmer. They simply record a new video. This rapid adaptation directly attacks a massive labor shortage affecting the semiconductor supply chain. Factories globally lack the skilled human technicians required to build these advanced computer systems.
Measuring Success and Financial Growth
The internal testing metrics reveal exactly how capable this video learning process has become. Skild AI tested the S1 model on completely unseen, multi-step scenarios. The software achieved a success rate of 66 percent at each individual step. When the team ran the same test using a standard baseline program, the older software managed a success rate of only nine percent. Beyond heavy industrial manufacturing, the software demonstrates flexibility across everyday physical chores. The engineering team successfully used video prompts to teach the machines how to brew pour-over coffee, cook pancakes, and sort random objects into designated containers. The S1 model successfully handles continuous jobs lasting up to ten minutes, which involves linking dozens of separate physical movements together without human intervention.
The economic impact of this speed is massive. The startup estimates that analyzing a single short video demonstration provides the equivalent utility of roughly 380 manual teleoperation examples. In older workflows, collecting that much manual data required up to 100 hours of human labor. During one internal test involving planting seeds in a pot, the engineering team recorded a video and successfully transferred the instructions to an autonomous machine in exactly eleven minutes.
This rapid deployment speed directly fueled massive financial success for the startup. Skild AI reached an annual recurring revenue run rate of $100M just ten months after launching its first commercial product. Reaching that financial target so quickly proves that massive industrial clients are eager to pay for software that cuts their setup time from weeks to minutes. The company currently manages over sixty active deployment partnerships spanning warehouse logistics, factory inspection, physical security, and food preparation.
Redefining Commercial Automation
We must clarify exactly how this exact software operates. The machine does not spontaneously generate intelligence from a single video. The S1 model is already heavily pre-trained using massive amounts of general physical data, human demonstrations, and simulated environments. The single video simply acts as an incredibly detailed instruction manual for a system that already understands general physics and movement.
This distinction makes the technology highly appealing to corporate buyers. A warehouse manager does not need to understand machine learning architecture. They simply need a camera and a worker who knows how to perform the physical job. The software handles the complicated translation between human intent and mechanical execution. The machine can recover from simple errors, adjust to unexpected object placement, and combine its pre-trained skills in completely new ways to finish the assigned sequence.
The partnership between Skild AI and Nvidia will likely expand further. The engineering teams are currently building new hardware-accelerated simulation solvers designed exclusively to calculate physical contact and solid object manipulation faster than ever before. They plan to release these tools to external developers soon. This open approach will likely push the entire robotics industry away from manual programming and closer to visual observation.
The era of rigid, heavily programmed factory machines is slowly ending. The future belongs to software that can watch, understand, and adapt instantly. By proving that a single video can successfully instruct a machine to build an advanced computer server, Skild AI completely altered the economics of physical automation.
Read More on TechRobust:

Umar Abubakar
Umar Abubakar
Expertise:Editorial Leadership, Product Design (UI/UX), Digital Media Strategy, Technology Systems, Product Architecture
Award:TechRobust Visionary Leader of the Year 2025
Umar serves as Editor-In-Chief and CEO of TechRobust, combining editorial vision with senior product design expertise to shape how modern technology stories are built, packaged, and told. Overseeing all editorial verticals, he directs coverage across global and regional tech landscapes while applying deep design thinking to publication strategy and reader experience.