Nvidia to acquire Hugging Face for $12.9 billion amid record revenue
Nvidia has agreed to acquire Hugging Face for $12.9 billion while reporting record quarterly revenue of $96.2 billion. Additionally, Google introduced Gemini 3.5 Transcribe for streaming audio and Gemini Omni 1.1 Flash for generative video production, alongside various other AI model and hardware updates.
Key Takeaways
- Nvidia's quarterly revenue doubled year-over-year to $96.2 billion, with data centers representing 93% of total sales.
- Google launched Gemini 3.5 Transcribe for streaming audio, featuring sub-second latency and 4.0% word-error rates.
- Apple debuted Mac Studio configurations with M5 Ultra chips supporting 512GB of unified memory for local AI workloads.
- OpenAI's new Jalapeño inference chip demonstrated up to 3.6 times lower latency than Nvidia GB200 systems.
Why It Matters
Nvidia's move to own the primary repository for open-source models effectively vertically integrates the AI stack from silicon to software distribution. For the streaming industry, Google's release of Gemini 3.5 Transcribe and Omni 1.1 Flash provides immediate tools for low-latency live captioning and generative video production. This consolidation occurs as hardware costs rise, with Nvidia passing 17% price increases to customers due to memory supply constraints. The ecosystem is now bifurcating between massive cloud-based models and high-end local hardware like Apple's M5 Ultra for private inference. Watch for whether OpenAI's Jalapeño chip benchmarks lead to a reduction in their API pricing to compete with Nvidia's hardware dominance.
Additional Context
Hugging Face has become the dominant distribution layer for open-source AI models, hosting more than 1.5 million models and serving as the default hub for researchers and enterprises deploying machine learning pipelines. The platform's strategic value was underscored when Nvidia confirmed its $12.9 billion acquisition of Hugging Face in August 2026, a deal that places the open-source repository under the same roof as the world's leading AI chipmaker. Prior to this deal, Hugging Face had raised $235 million in a Series D round at a $4.5 billion valuation in 2023, with investors including Salesforce, Google, Amazon, and Nvidia itself. The acquisition gives Nvidia direct control over the primary channel through which developers discover, fine-tune, and deploy models, consolidating influence from silicon through software distribution.
On the regulatory and business front, the deal arrives amid intensifying scrutiny of AI market concentration. The European Commission's Digital Markets Act framework already designates large platforms as gatekeepers, and EU competition officials signaled in mid-2026 that they would examine whether vertical integration of AI infrastructure and model distribution raises antitrust concerns. Meanwhile, Nvidia's own pricing power is under pressure from the opposite direction: the company disclosed that memory supply constraints would drive approximately 17% price increases on high-end server racks, a move that could accelerate enterprise interest in alternative inference hardware. Google, for its part, continues to position Gemini as a multi-modal platform play rather than a hardware consolidation strategy, releasing Gemini 3.5 Transcribe and Gemini Omni 1.1 Flash as API-first products aimed at developers building real-time applications.
For streaming and video applications specifically, Google's Gemini 3.5 Transcribe targets low-latency speech-to-text for live captioning workflows, while Gemini Omni 1.1 Flash extends generative video capabilities with faster inference times suitable for near-real-time production pipelines. Google announced Gemini 3.5 Transcribe in August 2026 with support for streaming audio input and multi-language output, positioning it as a direct competitor to Whisper-based transcription services that many streaming platforms currently rely on. The combination of Nvidia's hardware dominance and Hugging Face acquisition targets software distribution creates a vertically integrated stack that could pressure independent inference providers, while Google's API-first approach offers streaming companies an alternative path that avoids capital expenditure on GPU clusters. Apple's M5 Ultra chip, which supports on-device inference for models up to 500 billion parameters, represents a third vector for companies seeking to reduce cloud dependency for private video processing workloads.
Read full article at patmcguinness.substack.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source