
Pegasus provides temporal context, spatial reasoning, and judgement on whether a task was completed. Source: TwelveLabs
As the physical world rapidly digitalizes, AI teams are gathering vast amounts of complex video, but transforming raw footage into actionable models remains a critical challenge, according to TwelveLabs Inc. The company today released its Pegasus 1.6 model with added capabilities for understanding and navigating complex real-world environments.
“Our mission has always been to help machines understand how the world works through video,” stated Jae Lee, co-founder and CEO of TwelveLabs. “Physical AI is the next expression of that mission. Most of what people know about doing physical work, such as a changing grip or a recovery after something slips, has never been captured in a form a machine can learn from.”
“With our newest model release, we can now turn that footage into structured, reviewable knowledge, so robotics and physical AI teams can train on real human experience instead of starting from scratch,” he added.
Founded in 2021, TwelveLabs said it has created a “full-stack video intelligence platform” using its Marengo and Pegasus models, which it designed to see and understand video the way humans do in a fraction of the time. Developers can use this single system that gets smarter over time to access and act on all of their video content, said the Seoul-based company.
Pegasus 1.6 focuses on egocentric video
TwelveLabs said its latest version of Pegasus marks its expansion into physical AI, generating rich insights from real-world perspectives. It solves workflow-specific challenges so that machines such as robots, drones, and autonomous vehicles can perceive, reason, and safely act in the physical world like never before, the company claimed.
“We’re focused on the video understanding layer,” Jae Lee told The Robot Report. “You can’t just take millions of hours of video and feed it into a robotics model and expect useful training data to come out. You first need to understand what’s happening in the video, what action is taking place, when it happens, what the person is interacting with, how the behavior changes over time, and when other people or bystanders appear in the field of view.”
“Pegasus 1.6 provides that more precise understanding so robotics and data teams can then turn it into the training data they need,” he added.
TwelveLabs noted that Pegasus 1.6 is its first AI model built to understand egocentric video, which is shot from the point of view of the person doing the work. This could be someone cooking a meal, assembling parts on a factory line, or operating a robot remotely. The model does not require specific cameras, said Lee.
“We’re not requiring robotics teams to use a particular camera or proprietary hardware to capture that footage,” he said. “What matters is being able to capture the actions and interactions taking place from the operator’s point of view.”
In addition, Pegasus 1.6 can work with existing video data and allows for analysis of still images in addition to video.
“One of the reasons we’re focused on egocentric video is that it’s much easier to collect and scale than teleoperation data,” acknowledged Lee. “The goal is to take that footage and make it more useful for robotics teams by identifying the actions taking place, breaking them into precise time segments and capturing finer-grained details about the behavior.”

Head-mounted footage can break models trained on broadcast, instructional, and cinematic video. Source: TwelveLabs
TwelveLabs supports five workflows
Pegasus 1.6 currently supports five workflows powered by its video-native model. They include:
- Action segmentation and labeling: This feature accelerates model training with standardized datasets by automatically generating precise, time-stamped action labels for tasks, steps, objects, and hand-object interactions from raw video, mapped to the customer’s domain-specific taxonomy.
- Dense caption labeling: This enables natural language understanding for robots by producing rich, descriptive language for spatial relationships, scene context, and hand-object interactions to train advanced language-conditioned robot policies.
- Quality scoring: Users can save time by filtering out low-quality video by automatically evaluating and scoring video clips for action clarity, framing, and stability before sending footage to human reviewers.
- Search and curation: Customers can uncover critical edge cases by surfacing rare events, long-tail scenarios, and duplicate clips across an entire video repository using simple, natural-language search queries.
- Consent and compliance flagging: Privacy and compliance can be maintained by detecting and flagging faces, bystanders, and sensitive onscreen or paper data before video footage enters downstream development pipelines.
Editor’s note: Physical AI is among the session track topics at RoboBusiness 2026, which will be on Oct. 20 and 21 in Santa Clara, Calif. Register now to attend.
Customers to benefit from existing capabilities
Pegasus 1.6 builds on the video understanding capabilities that TwelveLabs developed for enterprises that maintain massive video libraries.
The new release extends Pegasus 1.5’s functionality that attracted several new customers. TwelveLabs cited Time-Based Metadata (TBM), which allows users to define a custom schema and automatically receive timestamped, structured metadata from video content. This is particularly useful in aiding contextual understanding, as people often narrate what they are doing in egocentric clips, it said.
Pegasus 1.6 has also improved entity recognition for more consistent tracking of hands, objects, and tools across clips. The model also offers faster, more cost-efficient processing for high-volume video workloads, according to TwelveLabs.
Lee said that TwelveLabs’ system adds video understanding to other sensor modalities.
“We’re focused on understanding the behavior we can observe in the video,” he said. “Pegasus 1.6 can identify the action taking place, segment it precisely in time, understand what the left and right hands are doing, and how the surrounding environment is changing.”
“We can also provide a coarse understanding of how the limbs are moving based on the video, but we’re not trying to replace the robot’s tactile or actuator-level sensing,” he added. “Our role is to give robotics teams a richer understanding of the behavior in the video that they can combine with their own sensor data to build more precise trajectories and learning policies.”
TwelveLabs said that Pegasus 1.6 expands on its existing collaborations with robotics developers.
“We’re working with robotics labs and data teams that are using video to help scale the data available for robot training,” said Lee. “A lot of the work is focused on dexterity and manipulation, tasks like packaging, assembly, and cleaning, as well as more specialized industrial applications such as semiconductor quality control. The broader goal is to help these teams make much larger amounts of human behavioral video useful for training, which is far more scalable than teleoperation data.”





Tell Us What You Think!