H Company launches NeoMME multimodal encoders to optimize visual document retrieval
H Company has released NeoMME, a family of open-source multimodal encoders designed for visual document retrieval. The models, available in 260M and 800M parameter sizes, utilize a shared bidirectional Transformer architecture to improve efficiency on the ViDoRe v3 benchmark by eliminating the need for causal decoders.
Key Takeaways
- NeoMME-Retriever-260M achieves an nDCG@10 of 0.523, matching ColQwen2.5 performance with 14 times fewer parameters
- Architecture supports 16,384 tokens, allowing the processing of two 4K UHD images simultaneously without OCR preprocessing
- Compression techniques including hierarchical token pooling and asymmetric quantization reduce storage requirements by up to 255 times
- Encoding throughput reaches 51 pages per second on an NVIDIA L40S, doubling the speed of ColModernVBERT
Why It Matters
The release of these open-source encoders provides a more efficient path for streaming and media enterprises to index high-resolution visual assets without the heavy compute costs of traditional generative vision-language models. By bypassing OCR and preserving document layouts, the technology streamlines the retrieval of complex metadata from charts and tables within large media libraries. This shift toward smaller, specialized bidirectional Transformers challenges the industry's reliance on massive causal decoders for non-generative tasks. As multimodal search becomes a standard feature for content management systems, watch for the adoption rate of these Apache 2.0 licensed models within Hugging Face Transformers to gauge their impact on production-scale retrieval-augmented generation pipelines.
Additional Context
The release of NeoMME arrives amid intensifying competition in open-source multimodal retrieval models. Hugging Face, which hosts NeoMME and the ViDoRe benchmark, has become the primary distribution channel for vision-language retrieval models, with ColPali and its derivatives establishing a new baseline for late-interaction visual document retrieval since mid-2025. The ViDoRe benchmark itself has evolved through multiple versions to test increasingly complex document layouts, and NeoMME's bidirectional architecture represents a deliberate departure from the causal-decoder approach that ColPali and similar models inherited from generative vision-language models. The business case for smaller, specialized encoders is gaining traction across the AI infrastructure stack. Nokia announced in June 2026 that its Autonomous Network Fabric would run on AWS, combining agentic AI with intent-based networking for cross-domain automation, illustrating how enterprises are shifting from monolithic foundation models toward modular, task-specific AI components that can be orchestrated across cloud environments. Similarly, Ericsson launched its AI in RAN commercial software subscription on June 11, claiming up to 20% higher downlink throughput across more than 15 live deployments, demonstrating that specialized AI models deployed at scale can outperform general-purpose systems on narrow tasks. This pattern of domain-specific efficiency over brute-force scale mirrors the architectural bet NeoMME makes for document retrieval. On the technical front, the bidirectional Transformer approach that NeoMME employs draws on the ModernBERT lineage, which demonstrated that encoder-only architectures can match or exceed decoder-based models on retrieval and classification tasks at a fraction of the compute cost. Ericsson's agentic AI blueprint places a Telco DataOps Platform at the center of its OSS/BSS architecture, using streaming data pipelines to clean and contextualize information before agents make decisions, a pattern that parallels how NeoMME's encoders are designed to feed structured embeddings into retrieval-augmented generation systems. The Apache 2.0 licensing under which H Company released NeoMME removes commercial restrictions, positioning the models for integration into proprietary content management and media asset platforms where permissive licensing is a prerequisite for production deployment.
Read full article at unite.ai
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source