Google's Gemma 4 brings multimodal reasoning and 256K context to open weights
Google released the Gemma 4 family of open multimodal models, featuring dense and Mixture-of-Experts architectures with support for text, image, and (in smaller models) audio inputs, along with expanded context windows up to 256K tokens and native system prompt support.
Key Takeaways
- Five model sizes: E2B (2.3B effective), E4B (4.5B effective), 12B, 26B A4B MoE (3.8B active of 25.2B total), and 31B Dense (30.7B parameters)
- Gemma 4 31B scores 89.2% on AIME 2026 (no tools), up from Gemma 3 27B's 20.8% — a 68-point jump
- The 26B A4B MoE activates only 3.8B of 25.2B total parameters, routing across 8 active experts out of 128 total plus 1 shared
- E2B and E4B support text, image, and audio input; the 12B, 26B, and 31B models support text and image only
- Context windows: 128K on E2B/E4B, 256K on 12B/26B/31B — all up from Gemma 3's limits
Why It Matters
Gemma 4's Apache 2.0 licensing and multimodal capabilities make it a viable open-weight option for teams building video and content analysis pipelines without cloud API dependency. The family spans from 2.3B edge models that run on phones to a 31B dense model ranking #3 on Arena AI's text leaderboard, giving media companies a single model family for both on-device and server-side AI tasks. Watch whether the 26B A4B MoE variant — which activates only 3.8B parameters per token — gains traction as a cost-efficient inference choice for high-volume content processing workloads.
Additional Context
Google announced Gemma 4 on April 2, 2026, positioning it as the company's most capable open model family, built from the same research and technology as Gemini 3 (per Google Blog, April 2026). The release came under Apache 2.0 licensing — a shift from prior Gemma generations — removing usage restrictions that had prompted developer feedback. Google reported over 400 million Gemma downloads to date across more than 100,000 community-built variants. The 31B dense model ranks #3 and the 26B MoE ranks #6 on Arena AI's text leaderboard as of early April 2026. The family expanded after launch. Google released the 12B Unified model on June 3, 2026 — an encoder-free architecture that projects raw image patches and audio waveforms directly into the LLM's embedding space, reducing multimodal latency (per Google AI for Developers, June 2026). Quantization-Aware Training variants followed in June 2026, with 4-bit and 2-bit optimized checkpoints for mobile and consumer GPU deployment. Google collaborated with Qualcomm Technologies and MediaTek on edge optimizations for the E2B and E4B models. Gemma 4 enters a competitive open-weight landscape. Meta's Llama 4 Scout offers a 10M token context window — roughly 40x larger than Gemma 4's maximum — but requires approximately 70GB VRAM and carries a 700M MAU license restriction. Alibaba's Qwen 3.5 edges ahead on MMLU Pro (86.1% vs Gemma 4's 85.2%) and offers native video input. Gemma 4 differentiates on math reasoning (89.2% on AIME 2026 vs Llama 4 Scout's approximately 36%) and is the only family with native audio input at the edge model tier (per CloudInsight, May 2026; Lushbinary, June 2026). For streaming and video applications, Gemma 4's multimodal support includes variable image resolution with configurable token budgets ranging from 70 to 1,120 tokens per image. Google's documentation notes that lower visual token budgets suit video understanding workloads where processing many frames at speed outweighs fine-grained detail extraction, while higher budgets benefit OCR and document parsing tasks (per Ollama, June 2026; Google Blog, April 2026).
Read full article at ollama.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source