Infinity raises $15M to breach Nvidia’s CUDA moat with automated code
Inference startup Infinity has raised $15 million at a $100 million valuation to develop its Ignition agent, which automatically writes and optimizes kernel code for AI inference across diverse chip architectures. The software aims to reduce reliance on the Nvidia CUDA ecosystem by enabling high-performance AI deployment on alternative hardware.
Key Takeaways
- Seed round led by Touring Capital with participation from OpenAI and Anthropic researchers.
- Ignition agent uses a self-optimizing feedback loop to write, test, and debug low-level chip code.
- Early case study showed the agent reached 92% of d-Matrix’s Corsair chip peak performance in 10 hours.
- Business model avoids upfront licensing, instead taking a cut of cost savings and performance gains.
- Revenue is already being generated through a ship-design partnership with hardware challenger d-Matrix.
Why It Matters
The standard barrier for non-Nvidia hardware has always been the software gap, as writing high-performance kernels for new chips can take years of human engineering. Infinity’s automated approach collapses this timeline to days, potentially commoditizing AI hardware by making migration between chips friction-free. For the streaming industry, which is pivotally shifting from model training to high-volume inference, this provides a path toward significantly lower token costs and reduced dependence on tight GPU supply chains. Watch for rival silicon providers to adopt similar automated stacks to accelerate time-to-market for their specialized inference accelerators.
Additional Context
The funding for Infinity arrives as the AI sector undergoes a structural shift from model training to large-scale inference. According to industry estimates cited by Business Wire in July 2026, inference workloads are projected to represent roughly two-thirds of all AI compute spending this year. This shift has intensified the search for alternatives to Nvidia’s dominant H100 and H200 GPUs. Per research from TrendForce in June 2026, shipments for custom AI Application-Specific Integrated Circuits (ASICs) are expected to grow 44.6% in 2026, significantly outpacing the 16.1% growth projected for general-purpose GPUs. Hardware challengers are already moving toward full-scale deployment to meet this demand. For instance, d-Matrix, a key partner for Infinity, announced in June 2026 that its Corsair inference platform has entered volume production. D-Matrix claims its Corsair chips deliver 10x faster performance and 5x better energy efficiency for Large Language Model (LLM) inference than traditional GPUs, specifically targeting the latency-sensitive needs of real-time voice and agentic AI applications. Per AI Multiple, d-Matrix reached a $2 billion valuation following a $275 million Series C round led by Temasek and Microsoft's M12 in late 2025. While hardware performance improves, the software remains the primary bottleneck for widespread adoption. Per Million Miner in July 2026, Nvidia's CUDA ecosystem still maintains approximately six million developers and 18 years of deeply integrated libraries. Competitors like AMD have tried to close this gap with the ROCm platform, and its MI355X chip reportedly matches Nvidia's Blackwell on raw compute. However, the introduction of automated agents like Infinity’s Ignition marks a new strategy: using AI itself to bypass the manual labor of software porting, which has historically been the most effective protector of Nvidia's 80% market share.
Read full article at techcrunch.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source