TwelveLabs has announced a video intelligence platform and API that utilizes its Marengo and Pegasus models to enable natural language search and automated content analysis. The platform claims to index video at 60x real-time speeds and is currently used by organizations including NFL Media and MLSE.
The immediate implication is a drastic reduction in the latency between live capture and searchable archival data, allowing broadcasters to monetize highlights in near real-time. Within the streaming ecosystem, this technology shifts the burden of metadata generation from manual tagging to automated, multimodal reasoning that understands causation and narrative across two-hour timelines. This capability directly addresses the scalability issues faced by sports leagues and studios managing petabytes of legacy footage. As organizations like NFL Media integrate these APIs, the industry should watch for a shift in ad-tech toward scene-level brand safety filters that do not rely on static metadata or human review.
TwelveLabs enters a video AI market where established platform vendors are rapidly embedding intelligence directly into their pipelines. Bitmovin's 2026/2027 Video Developer Report found that 98 percent of 486 video professionals surveyed are already using AI or ML in their workflows, with audio transcription, translation, and foreign dubbing as the most common applications at 48 percent, followed by content recommendations at 34 percent and visual quality optimization at 30 percent. That near-universal adoption means TwelveLabs must differentiate not on whether AI belongs in video workflows but on depth of multimodal understanding and indexing speed.
Mux has moved aggressively to make video AI a first-party capability rather than a bring-your-own integration. In 2026, Mux launched Robots, an API that runs AI analysis natively alongside stored video assets, eliminating the need for customers to hold separate OpenAI or Hive API keys. The product evolved from an open-source TypeScript toolkit called @mux/ai, released in December 2025, which handled glue work between Mux assets and external AI providers. Mux also recently introduced Robots Directives for orchestrating multi-step workflows, signaling that the company views AI orchestration as a platform layer rather than a point feature. This positions Mux as a direct competitor to TwelveLabs for teams that want AI capabilities bundled with encoding, delivery, and analytics in a single vendor relationship.
The broader competitive landscape for video intelligence APIs includes several managed platforms shipping built-in AI in 2026. Mux offers Claude-powered auto-chaptering, semantic search, and an MCP server alongside GenAI clips planned for Q3 2026, while Cloudflare Stream provides per-title AI encoding, Hive moderation, and Whisper-based captions. AWS IVS pairs Bedrock and Rekognition for low-latency interactive use cases. TwelveLabs' claimed 60x real-time indexing speed and its Marengo-Pegasus architecture targeting natural language search across vision, audio, and text simultaneously represent a distinct technical approach compared to these platforms, which tend to bolt discrete AI features onto existing encoding or delivery stacks rather than building a unified multimodal embedding layer from the ground up.
TwelveLabs has launched a new video intelligence platform that processes multimodal data at 60x real-time speed, indexing one hour of footage in just 60 seconds. By utilizing the Marengo and Pegasus models, the system enables natural language search across vision, audio, and text, significantly accelerating content review and automated highlight creation.
The TwelveLabs platform indexes video at 60x real-time speed, which allows it to process one hour of footage in 60 seconds.
The platform is powered by the Marengo embedding model and the Pegasus language model, which work together to enable natural language search across vision, audio, and text.
Early adopters of the TwelveLabs platform include NFL Media and MLSE, who are using the technology for automated highlight creation and archive management.
Unlike platforms that bolt discrete AI features onto existing stacks, TwelveLabs uses a unified multimodal embedding layer designed from the ground up to support natural language search across vision, audio, and text simultaneously.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source