Kioxia launches GP1 SSD hitting 10 million IOPS for AI infrastructure
Kioxia has announced its GP1 SSD, designed as a high-performance memory extension tier for AI infrastructure to alleviate memory bottlenecks. The drive utilizes PCIe 6.0 and second-generation XL-FLASH to reach 10 million IOPS with sub-5μs latency, targeting hardware integrators serving AI-intensive storage architectures.
Key Takeaways
- Delivers 10 million random read IOPS and sub-5μs latency, rivaling discontinued Intel Optane performance specifications.
- Uses second-generation XL-FLASH with 1-bit-per-cell (SLC) format and supports the latest PCIe 6.0 and NVMe 2.2 standards.
- Targeting AI storage architectures including Nvidia’s CMX scheme for Key-Value (KV) cache memory extension.
- Available in multiple form factors (E3.S, E1.S) with specific support for cold-plate liquid cooling to handle dense AI rack thermal loads.
Why It Matters
The GP1 addresses the critical bottleneck between High Bandwidth Memory (HBM) and standard storage, allowing AI clusters to process massive datasets without the prohibitive cost of saturating nodes with HBM. For the streaming ecosystem, this facilitates more efficient real-time recommendation engines and personalized ad-insertion at scale by accelerating the retrieval of large inference context windows. As Intel's Optane exit left a vacuum in the storage-class memory (SCM) tier, Kioxia is positioning itself as the primary hardware supplier for high-performance metadata and cache layers. Watch for testing results from VAST Data, which previously qualified Kioxia’s FL6 as its primary Optane alternative.
Additional Context
The launch of the GP1 follows a broader industry pivot toward PCIe 6.0 to support the data-heavy demands of generative AI. Per Samsung Electronics (July 2026), the company recently began mass production of its own PCIe 6.0 PM1763 enterprise SSD, which achieved roughly 6.92 million IOPS in initial validation. While Samsung’s entry focuses on high-capacity TLC flash for training and general HPC, Kioxia’s GP1 specifically targets the latency-sensitive 'memory extension' niche previously dominated by Intel's 3D XPoint technology. Intel officially ceased shipments of its Optane Persistent Memory 200-series in late 2025 (per Tom's Hardware), leaving specialized AI infrastructure providers like VAST Data and Weka seeking high-endurance, low-latency NAND alternatives. Integration with the Nvidia ecosystem is also a primary driver for these high-performance drives. At GTC (March 2026), Nvidia introduced the BlueField-4 DPU and the CMX (Context Memory Storage) architecture, which offloads Key-Value (KV) cache management from the CPU to a dedicated storage tier. According to Nvidia, this architecture is designed to improve token throughput by up to 5x for long-context inference. Kioxia's GP1 is engineered to plug directly into this reference architecture, allowing infrastructure builders like Supermicro and Quanta Cloud Technology to build JBOF (Just a Bunch of Flash) systems that function as an externalized pool of shared GPU memory. Market adoption of PCIe 6.0 is expected to ramp significantly through the second half of 2026 as server platforms based on AMD’s 'Venice' EPYC and Nvidia’s 'Vera' CPUs reach general availability. Per ServeTheHome (August 2026), while PCIe 5.0 remains the standard for current mainstream AI clusters, the transition to Gen6 is critical for eliminating the I/O stalls that occur when context windows grow into the hundreds of thousands of tokens. The GP1’s 50 Drive Writes Per Day (DWPD) rating also signals a strategic trade-off, offering higher endurance than standard TLC drives (which typically offer 1–3 DWPD) to handle the frequent read-write cycles of ephemeral AI cache data.
Read full article at blocksandfiles.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source