Mira Murati’s Thinking Machines Lab launches Inkling, an 975B-parameter open-weight model
Thinking Machines Lab has released Inkling, a 975-billion-parameter open-weight mixture-of-experts model trained on 45 trillion multimodal tokens. The startup aims to provide enterprise customers with a customizable alternative to proprietary, closed-loop AI models for specific production and development tasks.
Key Takeaways
- Inkling utilizes a Mixture-of-Experts (MoE) architecture, activating only 41 billion of its 975 billion parameters per task to optimize speed.
- The model was trained on 45 trillion multimodal tokens, including text, code, audio, and video, though outputs remain limited to text and structured data.
- Benchmarks show Inkling achieves coding performance comparable to Nvidia’s Nemotron 3 Ultra while consuming approximately 66% fewer tokens.
- Enterprises can fine-tune the model using Tinker, the company’s infrastructure platform, ensuring proprietary business logic remains within private environments.
Why It Matters
The release of Inkling signals a shift toward local, open-weight deployment for high-volume enterprise workloads where data sovereignty is paramount. By providing a customizable starting point rather than a finished consumer product, Thinking Machines targets a gap left by proprietary 'black box' models. For the streaming industry, this suggests a move toward specialized, vertically-integrated AI stacks that manage content metadata and code generation without exposing internal libraries to third-party labs. As open-weight performance converges with closed frontier models, the industry's focus is shifting from raw intelligence to operational control and IP protection. To track next, watch for verified SWE-bench performance gains following the first wave of enterprise fine-tuning on the Tinker platform.
Additional Context
Thinking Machines Lab, founded in February 2025, quickly became a focal point for AI investment, closing a record $2 billion seed round at a $12 billion valuation in July 2025 per Reuters. The lab, which includes senior alumni from OpenAI and Anthropic such as Chief Scientist John Schulman, initially focused on building 'interaction models' designed for real-time collaboration. In March 2026, the company secured a strategic agreement for a gigawatt of Nvidia compute capacity on the Vera Rubin architecture to support its training clusters, according to Wikipedia and The Next Web. This hardware advantage allowed the lab to develop Inkling in roughly nine months by leveraging a practice known as distillation, which uses outputs from existing models like Moonshot’s Kimi K2.5 to accelerate learning. The launch coincides with a documented rise in enterprise skepticism regarding closed-loop APIs. Per Microsoft’s CEO Satya Nadella in July 2026, companies are increasingly concerned about the 'Reverse Information Paradox,' where they pay for AI using both subscription fees and the proprietary knowledge embedded in their prompts. This trend is bore out by data from Vercel’s July 2026 AI Gateway Production Index, which found that open-weight models now process 29% of all production AI tokens. While open models accounted for nearly a third of token volume, they represented less than 4% of total cloud spending, highlighting the aggressive cost advantages Thinking Machines is pursuing with its open-weight strategy. This shift suggests a maturing market where enterprises prioritize building their own 'trust boundaries' over renting general-purpose intelligence from frontier labs.
Read full article at techcrunch.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source