Zilliz launches Milvus 3.0 to index data directly from object storage
Zilliz has launched Milvus 3.0, an update to its open-source vector database that introduces a lake-native architecture for direct indexing and retrieval from object storage. The release includes features like external collections, a new storage engine called Loon, and enhanced retrieval capabilities to support complex AI and RAG workflows.
Key Takeaways
- External Collections allow Milvus to index and query data in Lance, Iceberg, Parquet, or Vortex without moving source files.
- Loon storage engine reduces read amplification for point access, cutting measured I/O significantly compared to standard Parquet baselines.
- StructList enables native multi-vector retrieval for complex entities like multi-patch images and documents using ColBERT or ColPali models.
- Optimized sparse indexing in version 3.0 is three times smaller than previous iterations while maintaining comparable recall levels.
Why It Matters
For streaming platforms managing massive video metadata and frame embeddings, Milvus 3.0 addresses the 'infrastructure tax' of syncing data lakes with vector databases. By indexing directly on S3 or Azure Blob storage, engineering teams can build RAG and recommendation systems without maintaining mirrored datasets. This pivot toward lake-native architecture signals a shift away from isolated vector silos toward unified data foundations that support both real-time retrieval and batch analytics. Watch for whether competitors like Pinecone or Weaviate introduce similar direct-lake indexing to counter Zilliz's reduced operational overhead narrative.
Additional Context
The push toward lake-native vector databases arrives as enterprise unstructured data volumes reach a tipping point. Per Komprise (December 2025), 74% of enterprises now manage over 5 petabytes of unstructured data, a 57% increase since 2024. This explosion in rich media and AI telemetry is making traditional ETL pipelines—where data is copied from storage into specialized databases—operationally unsustainable due to synchronization lag and mounting cloud egress costs.
Market analysts at BARC U.S. (July 2026) noted that Zilliz’s architectural shift aligns with a broader industry convergence where transactional and analytical workloads are merging. Competitors such as Databricks and Snowflake have recently introduced features to unify vector search with their existing lakehouse environments. Specifically, the integration with open table formats like Apache Iceberg is becoming a prerequisite for enterprise adoption, as it prevents vendor lock-in and allows standard tools like Spark to interact with AI datasets.
In the specialized vector market, the rivalry remains intense. While Zilliz positions Milvus for massive-scale horizontal growth, Pinecone continues to lead in managed serverless simplicity, and Turbopuffer—which raised seed funding in late 2025—is gaining traction with a similar object-storage-first philosophy. AI spending on data readiness (April 2026) suggests that enterprise agentic AI adoption will increase sevenfold by 2029, pressuring infrastructure providers to prove they can handle the real-time analytics for agentic AI workloads.
Read full article at businesswire.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source