TwelveLabs launches Jockey platform to provide persistent video memory layer
TwelveLabs has unveiled 'Jockey,' a platform designed to create a persistent 'memory layer' for video ecosystems by preserving spatiotemporal relationships in video corpora. The system utilizes the company's Morango encoder and Pegasus language model to enable complex reasoning and querying across video collections for media, surveillance, and advertising applications.
Key Takeaways
- Jockey uses the Morango encoder and Pegasus language model to analyze video as a spatiotemporal volume rather than just a transcript.
- The platform features a queryable 'context graph' that links entities, relationships, and events across different files and camera angles.
- Three primary demos showcased Jockey's ability to track player performance in sports, detect safety incidents in city traffic, and identify ad-friendly hero shots.
- TwelveLabs released five core principles for the system, emphasizing that high-cost video understanding should be 'ingested once' and repurposed multiple times.
Why It Matters
TwelveLabs is addressing a fundamental bottleneck in AI-driven media management: the loss of temporal context in traditional frame-sampling methods. By creating a persistent memory layer, Jockey allows streaming and broadcast operators to query archives for complex cross-camera narratives rather than just keywords. Concretely, this shifts video from a passive storage cost into a queryable strategic asset that can automate content assembly and compliance at scale. This move places TwelveLabs in direct competition with multimodal offerings from hyperscalers like Amazon and Google, who are also racing to solve video-native reasoning. Watch for TwelveLabs to report adoption metrics from sports leagues and ad agencies as Jockey exits private beta and moves into general availability.
Additional Context
The launch of Jockey follows a significant capital infusion for TwelveLabs, which secured $100 million in Series B funding in July 2024 to accelerate its 'superintelligence' for video understanding, per SiliconAngle. This round included backing from New Enterprise Associates, Amazon Web Services, and NVIDIA, positioning the startup as a major challenger in the multimodal AI space. The company's expansion into strategic memory layers is supported by its recent achievement of the AWS AI Competency in June 2026, which TwelveLabs noted coincided with an 13.1% performance lead for its Pegasus model over rivals in specific multimodal tasks, according to their internal benchmarking Data.
Enterprise application of this technology has already yielded measurable operational efficiencies. TwelveLabs reported in June 2026 that a partnership with Maple Leaf Sports & Entertainment (MLSE) reduced the time required for video search and retrieval from 16 hours down to just nine minutes. This real-world validation highlights the industry's shift from basic metadata tagging toward deep temporal reasoning. Furthermore, the company expanded its ecosystem in April 2026 by launching a dedicated partner program, adding collaborators like Quickplay, Mux, and Vidispine to integrate its video-native models directly into established media production and distribution workflows, per PRWeb.
Read full article at startuphub.ai
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source