TwelveLabs Brings Video Understanding to Physical AI With Latest Launch

Pegasus’ physical AI capabilities turn egocentric video into data for teaching machines to understand the real world

TwelveLabs, a leading video intelligence company, today announced the release of Pegasus 1.6, which adds new capabilities for understanding and navigating complex real-world environments. This latest model enables TwelveLabs expansion into physical AI, the domain of artificial intelligence that can perceive, reason about, and act in the physical world through machines like robots, drones, autonomous vehicles, and industrial equipment. Now TwelveLabs can solve key workflow-specific challenges and unlock the full potential for physical AI.

As the physical world rapidly digitalizes, physical AI teams are gathering vast amounts of complex video, but transforming raw footage into actionable model intelligence remains a critical challenge. TwelveLabs bridges this gap, automatically generating rich insights from real-world perspectives and enabling machines to see, understand, and safely interact with their environments like never before.

Marketing Technology News: MarTech Interview with Jema Birkmeyer, VP @ Attentive

“Our mission has always been to help machines understand how the world works through video,” said Jae Lee, CEO and co-founder of TwelveLabs. “Physical AI is the next expression of that mission. Most of what people know about doing physical work, such as a changing grip, or a recovery after something slips, has never been captured in a form a machine can learn from. With our newest model release, we can now turn that footage into structured, reviewable knowledge, so robotics and physical AI teams can train on real human experience instead of starting from scratch.”

Notably, TwelveLabs’ new Pegasus 1.6 model is the first TwelveLabs model built to understand egocentric video. Egocentric video is shot from the point of view of the person doing the work, whether that’s someone cooking a meal, assembling parts on a factory line, or operating a robot remotely.

Pegasus 1.6 currently supports five core workflows powered by this video-native model. They are:

  • Action segmentation and labeling: Accelerates model training with standardized datasets by automatically generating precise, timestamped action labels for tasks, steps, objects, and hand-object interactions from raw video, perfectly mapped to your domain-specific taxonomy.
  • Dense caption labeling: Enables natural language understanding for robots by producing rich, descriptive language for spatial relationships, scene context, and hand-object interactions to train advanced language-conditioned robot policies.
  • Quality scoring: Save time by filtering out low-quality video by automatically evaluating and scoring video clips for action clarity, framing, and stability before sending footage to human reviewers.
  • Search and curation: Uncover critical edge cases by effortlessly surfacing rare events, long-tail scenarios, and duplicate clips across your entire video repository using simple natural language search queries.
  • Consent and compliance flagging: Protect privacy and maintain compliance effortlessly by detecting and flagging faces, bystanders, and sensitive onscreen or paper data before video footage enters downstream development pipelines.

Marketing Technology News: How MarTech Is Enabling Autonomous Brand Engagement Across Channels?

Pegasus 1.6 builds on the industry-leading video understanding capabilities TwelveLabs developed for enterprises that maintain massive video libraries.

The new release extends Pegasus 1.5’s foundational functionality that attracted several new customers, particularly from Time-Based Metadata (TBM), which allows users to define a custom schema and automatically receive timestamped, structured metadata from video content. This is particularly useful in aiding contextual understanding, as people often narrate what they are doing in egocentric clips. Pegasus 1.6 also improved entity recognition for more consistent tracking of hands, objects, and tools across clips. The model also offers faster, more cost-efficient processing for high-volume video workloads.

The post TwelveLabs Brings Video Understanding to Physical AI With Latest Launch first appeared on PressReleaseCC.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top