Liquid AI launches LFM2.5 encoders optimized for CPU-based media classification
Liquid AI has released its LFM2.5-Encoder-230M and 350M models designed for long-context tasks up to 8,192 tokens. These bidirectional models are optimized for CPU-based inference, targeting edge and on-premise streaming or media utility applications such as classification, safety filtering, and intent routing.
Key Takeaways
- New models support 8,192-token context windows, roughly 15 pages of text, for document-scale media analysis.
- LFM2.5-Encoder-230M performs 3.7x faster than ModernBERT-base on CPU at maximum context length.
- Architecture replaces causal attention with bidirectional masks and non-causal convolutions to improve local token representation.
- Models are released as open-weights on Hugging Face, specifically targeting low-latency edge AI and sensitive on-premise workloads.
Why It Matters
This release shifts document-scale AI analysis from expensive GPU clusters to standard CPUs, providing a viable path for private, local media processing. By outperforming established models like ModernBERT on throughput, Liquid AI addresses the high-volume, cost-sensitive pipelines common in streaming content moderation and metadata tagging. The move enables streaming providers to run complex intent routing and safety filters at the edge, reducing cloud latency and operational expenses. Watch for deployment metrics in automotive and consumer electronics as Liquid AI leverages its AMD partnership to scale LFM2.5 integration across non-GPU hardware stacks.
Additional Context
Liquid AI, a startup spun out of MIT’s CSAIL, has rapidly expanded its product line since raising $250 million in Series A funding led by AMD Ventures in December 2024. Per Reuters and VentureBeat, that round valued the company at roughly $2.35 billion, positioning it as a primary challenger to transformer-based architectures. The company’s Liquid Foundation Models (LFMs) rely on dynamical systems and signal processing theory to achieve higher memory efficiency than traditional Large Language Models (LLMs), reportedly using up to 90% less memory in certain configurations.
In July 2025, Liquid AI introduced the LFM2 series, which the company claims delivers 2x faster decode performance than Alibaba’s Qwen3 on CPU hardware. This development coincides with a broader industry shift toward Small Language Models (SLMs) and edge AI. Per MarketsandMarkets (November 2025), the SLM market is projected to reach $5.45 billion by 2032 as enterprises prioritize performance-per-watt and data sovereignty over raw parameter counts. Liquid AI has already established partnerships with G42 and Capgemini to deploy these efficient models for on-premise and sovereign AI solutions.
The streaming video and media sector is increasingly adopting these compact encoders for real-time insights. According to Streaming Media (January 2025), AI-driven workflow automation is now a baseline requirement for major platforms to manage hyper-relevant recommendations and ad-tier safety. By offering an encoder that stays performant at long context lengths, Liquid AI targeting B2B applications where analyzing full transcripts or legal contracts previously required significant local inference strategies cloud-based GPU overhead.
Read full article at liquid.ai
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source