Mira Murati previews Thinking Machines' continuous-stream multimodal AI architecture
Mira Murati, co-founder and CEO of Thinking Machines Lab, presented their new 'interaction models' approach at Bloomberg Tech. These models process audio, video, and text as continuous, parallel streams for real-time, multimodal interaction, diverging from traditional turn-based systems. The approach requires low-latency streaming inference and integrated capture, posing significant compute and engineering challenges.
Key Takeaways
- Thinking Machines' architecture processes 200ms of input while generating 200ms of output to achieve sub-400ms total response latency.
- The lab's only shipping product as of June 2026 is Tinker, an API for fine-tuning open-source models that launched in October 2025.
- Murati confirmed the startup is utilizing 128K context windows and mixture-of-experts (MoE) backbones for its interaction models.
- The startup has raised $2 billion to date from investors including Andreessen Horowitz and Nvidia to fund compute-heavy streaming inference.
Why It Matters
Native multimodal streaming represents a shift from the 'bolted-on' interactivity of current voice assistants to an architecture where latency is a core model feature. For the streaming industry, this suggests a future where AI can monitor live video feeds—such as sports or security—and provide mid-stream feedback without the 1-2 second lag typical of batch processing. The challenge remains the extreme compute cost of maintaining persistent state for continuous video and audio inputs at under 200ms granularity. Success will depend on whether Thinking Machines can move beyond the 'Tinker' development wedge into production-ready deployments as rivals like OpenAI and Google iterate on their own low-latency voice and vision modes.
Additional Context
Since Murati’s departure from OpenAI in September 2024, the competitive landscape for low-latency multimodal interaction has narrowed to a high-stakes hardware race. Per Bloomberg (June 2026), Thinking Machines Lab reached a $12 billion valuation following a $2 billion seed round led by Andreessen Horowitz. To support the heavy compute demands of parallel audio-video streams, the startup secured a multiyear supply agreement with Nvidia for its Vera Rubin GPU architecture in March 2026, according to Crypto Briefing. This hardware commitment is critical as the lab aims to scale its TML-Interaction-Small model for public release by late 2026. While Thinking Machines emphasizes transparency, it faces significant internal volatility. Per The Next Web (June 2026), the lab has navigated a series of researcher departures to Meta and OpenAI despite its aggressive hiring from those same firms during its 2025 launch phase. Observers such as Puck News (May 2026) noted that until the Bloomberg appearance, the company had been 'eerily quiet,' relying on Tinker—a service for fine-tuning open-weights models like Llama 3—to maintain developer mindshare while focusing internally on the higher-complexity streaming inference stack. Technically, the 'interaction model' approach mirrors a broader industry transition toward full-duplex systems. Competitors such as Kyutai’s Moshi model previously demonstrated theoretical latencies as low as 160ms for audio, according to Unitlab (January 2026). However, Murati’s vision adds continuous video processing to the mix, requiring significantly higher throughput. As OpenAI prepares for a reported IPO and Anthropic targets a $1 trillion valuation, Thinking Machines is positioning its 'human-AI collaboration' thesis as a strategic alternative to the increasingly closed ecosystems of incumbent frontier labs.
Read full article at letsdatascience.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source