
TwelveLabs Pegasus 1.6 Translates Human Video For Robots
The Seoul based software builder released an updated video analysis tool designed to convert messy human footage into structured data, allowing robotics developers to train physical machines faster and more economically.
Inioluwa Ademidun | 6 Oct. 2026, 7:10 PM · 3 min read

Training physical machines to interact with the real world is incredibly expensive. Programmers usually have to physically maneuver a robot arm thousands of times to teach the software how to pick up a single object. A Seoul based software builder named TwelveLabs wants to eliminate this slow, manual process. The company recently released an upgraded software model named Pegasus 1.6. This specific software digests raw video footage recorded from a human perspective and converts those recordings into structured data that physical machines can instantly understand.
The underlying problem the startup fixes involves visual translation. A human wearing a camera can record themselves assembling a computer component. While a human watching that video understands exactly when the worker drops a screw or switches hands, a machine only sees a collection of colorful pixels. Pegasus translates those pixels into exact mathematical labels. It timestamps the exact second a hand reaches for a tool, labels the specific object being grabbed, and records whether the final task succeeded or failed.
Standardizing Video Data
Providing this structured video data gives robotics companies a huge shortcut. Instead of hiring engineers to build highly customized physical testing environments, they can simply strap cameras to human factory workers, record their daily shifts, and feed the video through the Pegasus system. The software breaks the human actions down into standardized training materials. This exact push to use video for mechanical instruction is accelerating rapidly. We tracked similar educational approaches when Skild and Nvidia trained robots using single video learning. The industry desperately needs cheaper ways to teach physical movement.
The startup built five distinct capabilities into the 1.6 update. The software segments actions into precise timestamps. It generates detailed text descriptions explaining the spatial relationships between the human hands and the objects they hold. It also scores the raw video footage, automatically rejecting clips that are blurry or useless before they reach the main training servers. Checking for visual accuracy saves researchers from paying for garbage data.
Filtering Private Information
Protecting corporate privacy is another main feature. The software automatically scans the raw footage and removes human faces or sensitive paper documents sitting on a desk. Removing sensitive information guarantees that a robotics laboratory does not accidentally ingest private employee records while training their machines. Gathering training material requires heavy capital, a reality we documented when Sequoia backed Mecka at a high valuation specifically for robot data collection. Companies want to buy clean, safe training material without violating local privacy laws.
TwelveLabs charges customers based on the exact amount of video processed. The current retail price sits at $1.75 per video hour analyzed, along with separate fees for generating the actual text descriptions. By renting access to their video analysis tool, the startup avoids the heavy physical costs associated with actually building the steel robots. They strictly supply the educational material. The market for physical hardware is highly competitive and brutally expensive, a reality exposed when Human Capital raised $100M just to finance physical machine construction. Selling the software layer is a much safer financial bet.
The success of Pegasus will depend entirely on its accuracy. If the software incorrectly labels a human hand movement, the robot learning from that data will drop its payload and shatter the object. Robotics developers will demand absolute proof that the video analysis accurately maps the physics of the real world. If TwelveLabs can consistently provide flawless video translation, they will become the primary textbook publisher for the entire mechanical engineering sector.
Read More on TechRobust:

Inioluwa Ademidun
Inioluwa Ademidun
Expertise:African Tech Ecosystem, Early-Stage Startups, Emerging Market Dynamics, Venture Capital & Tech Reporting, Product Management
Award:TechRobust Contributor of the Year 2025
Inioluwa is a Senior Product Manager by day and an investigative technology reporter by night, bridging the gap between scalable software architecture and high-impact journalism. She delivers deep-dive analysis on venture-backed founders, regulatory shifts, and grassroots tech ecosystems across Africa and global emerging markets.