Positron AI funding reaches $875M to challenge Nvidia inference dominance
Positron AI has raised $875 million in a funding round to develop its Titan inference appliance, which utilizes LPDDR5X memory to bypass HBM supply constraints. The system features custom Asimov silicon designed to achieve high hardware utilization for large-scale AI model deployment.
Key Takeaways
- Titan appliances utilize LPDDR5X memory to achieve 90% throughput, bypassing current HBM supply chain constraints.
- Custom Asimov chips are designed to process 26 times more tokens per dollar than Nvidia Blackwell GB300 NVL72 systems.
- A single Titan system supports large language models with 32 trillion parameters and 10 billion token context windows.
- Mass production is scheduled for the second half of 2027 using TSMC three-nanometer manufacturing nodes.
Why It Matters
The massive capital injection into Positron AI highlights a strategic pivot toward hardware that circumvents the HBM supply bottleneck currently slowing AI deployment. By utilizing LPDDR5X memory typically found in smartphones, the company aims to lower the cost of large-scale inference for streaming and media enterprises managing massive datasets. This approach challenges the premium pricing model of Nvidia by focusing on hardware utilization efficiency rather than raw theoretical bandwidth. If Positron successfully tapes out its Asimov silicon at TSMC this year, it could provide a viable alternative for high-parameter model hosting. Watch for the first simulation-to-silicon performance benchmarks when the Asimov chip enters production in late 2027.
Additional Context
Positron AI enters a crowded field of inference-focused chip startups seeking to undercut Nvidia's dominance in AI serving. The company's Titan appliance, which pairs custom Asimov silicon with LPDDR5X memory, represents a bet that memory bandwidth efficiency matters more than raw HBM throughput for production inference workloads. Light Reading reported in June 2026 that Ericsson and Nokia are diverging sharply on AI-RAN architecture, with Nokia committing its entire Layer 1 RAN stack to Nvidia's CUDA platform and GPUs, underscoring how deeply Nvidia's ecosystem has penetrated adjacent compute markets. That same GPU-centric dependency is precisely the supply constraint Positron aims to exploit by sourcing commodity memory from established DRAM suppliers rather than competing for limited HBM allocations from SK Hynix and Samsung.
The competitive landscape for inference hardware has intensified considerably in 2026. IEEE ComSoc's Technology Blog documented a cluster of announcements in June 2026 showing telcos transitioning from isolated AI pilots to production-grade AI operations deployed across live networks, creating massive new demand for inference capacity at the edge. SK Telecom announced a gigawatt-scale AI Cloud built on Nvidia DGX SuperPOD architecture, while MTN Group detailed plans to convert 18,000 African tower locations into a distributed AI inference grid. These deployments illustrate the scale of inference demand that Positron's cost-optimized approach targets, particularly for operators and media companies that need to run large models without absorbing HBM premium pricing.
Nokia's recent infrastructure partnerships reveal the broader ecosystem dynamics Positron must navigate. Nokia announced at DTW Ignite 2026 that it is combining with AWS and Databricks to build a unified data and control layer for autonomous networks, claiming operators are already achieving automation rates above 90 percent and service delivery times under four hours. Nokia also partnered with Google Cloud to deploy six specialized Gemini-powered AI agents for network troubleshooting, targeting 50 to 80 percent reductions in problem-solving times. These agentic AI deployments require substantial inference compute at the edge, representing the exact workload profile where Positron's LPDDR5X-based Titan could offer meaningful cost advantages over GPU-based alternatives. The company's $5 billion valuation reflects investor confidence that inference, not training, will dominate AI compute spending as enterprises scale production deployments across streaming, telecom, and media verticals.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source