TwelveLabs Launches Pegasus 1.6 for Physical AI Video
TwelveLabs has released its Pegasus 1.6 model, expanding its full-stack video intelligence platform into physical AI by introducing capabilities to understand first-person video and support robotics workflows.

Expanding Video Understanding Into Physical AI
As the physical world rapidly digitalizes, AI teams are gathering vast amounts of complex video, but transforming raw footage into actionable models remains a critical challenge, according to TwelveLabs Inc. The Seoul-based company released its Pegasus 1.6 model with added capabilities for understanding and navigating complex real-world environments, marking its official expansion into physical AI.
“Our mission has always been to help machines understand how the world works through video,” stated Jae Lee, co-founder and CEO of TwelveLabs. “Physical AI is the next expression of that mission. Most of what people know about doing physical work, such as a changing grip or a recovery after something slips, has never been captured in a form a machine can learn from.”
Founded in 2021, the company has created a full-stack video intelligence platform using its Marengo and Pegasus models, which are designed to see and understand video the way humans do in a fraction of the time. Developers can utilize this single system to access and act on all of their video content.
Understanding Egocentric Video for Robotics
According to TwelveLabs, Pegasus 1.6 is its first AI model built to understand egocentric video, which is shot from the point of view of the person doing the work. This includes tasks such as cooking a meal, assembling parts on a factory line, or operating a robot remotely. The model does not require specific cameras or proprietary hardware.
“We’re not requiring robotics teams to use a particular camera or proprietary hardware to capture that footage,” Jae Lee said in an interview with The Robot Report. “What matters is being able to capture the actions and interactions taking place from the operator’s point of view.”
In addition to supporting head-mounted footage, Pegasus 1.6 can work with existing video data and allows for the analysis of still images. Lee noted that egocentric video is much easier to collect and scale than teleoperation data, making it a valuable resource for training advanced machine learning models.
Five Workflows Powered by Pegasus 1.6
Pegasus 1.6 currently supports five distinct workflows powered by its video-native model, designed to solve workflow-specific challenges so that machines such as robots, drones, and autonomous vehicles can perceive, reason, and act safely.
The supported workflows include action segmentation and labeling, which automatically generates precise time-stamped action labels for tasks, steps, and hand-object interactions. Another feature is dense caption labeling, which produces rich descriptive language for spatial relationships and scene context to train advanced language-conditioned robot policies.
Additional workflows feature quality scoring for filtering out low-quality video clips, search and curation for uncovering rare edge cases across entire repositories using natural-language queries, and consent and compliance flagging to detect sensitive data or bystanders before development pipelines utilize the footage.
Building on Existing Enterprise Capabilities
The new model builds upon video understanding capabilities previously developed for enterprises maintaining massive video libraries. Pegasus 1.6 extends the functionality of Pegasus 1.5, including Time-Based Metadata (TBM), which allows users to define custom schemas and automatically receive timestamped, structured metadata.
Furthermore, the updated model features improved entity recognition for more consistent tracking of hands, objects, and tools across video clips. TwelveLabs states that the model also offers faster, more cost-efficient processing for high-volume video workloads.
Lee emphasized that TwelveLabs focuses specifically on the video understanding layer rather than replacing tactile or actuator-level sensing. The goal is to provide robotics teams with a richer understanding of behavior that can be combined with their own sensor data.
Industry Applications and Future Outlook
TwelveLabs is actively working with robotics labs and data teams to help scale the data available for robot training. Much of this work focuses on dexterity and manipulation tasks such as packaging, assembly, and cleaning, as well as specialized industrial applications like semiconductor quality control.
As the physical AI landscape continues to evolve, industry events are highlighting these advancements. For instance, physical AI is among the session track topics scheduled for RoboBusiness 2026, which will take place on Oct. 20 and 21 in Santa Clara, California.
Sources
- The Robot ReportPegasus 1.6 brings video understanding to physical AI, says TwelveLabs
Continue chronologically




