StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← Streaming Platforms
PlatformsProduct LaunchJuly 27, 2026

Dell and AMD target cloud token costs with modular AI inference

SiliconANGLE

Dell Technologies and AMD have launched a modular AI platform designed to help enterprises scale generative AI workloads on-premises. The system uses AMD Instinct GPUs and EPYC CPUs to provide organizations with a scalable, pre-validated infrastructure alternative to token-based cloud pricing.

Key Takeaways

  • Modular architecture utilizes AMD Instinct MI300X and MI350P GPUs to allow scaling without infrastructure rearchitecting
  • Pre-validated software stack includes AMD ROCm and open ecosystem frameworks to simplify the transition from proof-of-concept to production
  • On-premises deployment enables enterprises to bypass token-based cloud pricing for high-volume inference tasks
  • Infrastructure combines Dell PowerEdge nodes with EPYC CPUs to handle both GPU-heavy inference and CPU-driven orchestration

Why It Matters

Enterprises are hitting a ceiling with cloud GPU costs as AI workloads mature. By shifting to on-premises 'token generators,' organizations can stabilize budgets while maintaining data sovereignty and governance. This launch intensifies the competition for the enterprise AI stack, positioning AMD as a cost-effective alternative to Nvidia-dominated data centers. The ecosystem movement suggests a growing B2B preference for hybrid deployments that balance cloud flexibility with the TCO advantages of owned hardware. Watch for whether AMD's ROCm software can maintain performance parity with CUDA as businesses move deeper into complex RAG pipelines.

Additional Context

The push toward on-premises AI comes as enterprise inference costs begin to dominate tech budgets. According to Deloitte Tech Trends 2026, high-volume AI inference is placing unprecedented strain on cloud strategies, prompting a deliberate shift toward hybrid architectures that prioritize cost and data sovereignty. Industry data from DreamFactory in July 2026 suggests that local execution can reduce response times to under 40 milliseconds—a 97% improvement over cloud-based APIs—while reducing the cost per million tokens significantly for high-volume users. AMD is aggressively positioning itself as the primary alternative to Nvidia's market dominance. During the July 2026 Advancing AI event, AMD reported that its data center revenue reached $5.8 billion in Q1 2026, driven by demand for the MI350 series. While Nvidia continues to hold roughly 80% of the AI accelerator market, AMD's Instinct GPUs have gained traction through massive deployments at Meta and Microsoft Azure. Per SemiAnalysis in July 2026, AMD’s strategy includes high-performance memory configurations that offer structural advantages for large-scale enterprise inference. Dell's expanded partnership with AMD follows its long-standing collaboration with Nvidia under the 'AI Factory' banner. In April 2026, Dell confirmed the general availability of the Dell Lightning File System, an ultra-high-performance storage layer designed to eliminate I/O bottlenecks for GPUs during continuous inference runs. By offering both AMD and Nvidia configurations, Dell is positioning its hardware as the agnostic foundation for a projected $200 billion AI accelerator market where, according to AMD CEO Lisa Su, 60% of compute power will be dedicated purely to inference by late 2026.


Read full article at siliconangle.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

TV[BE]urope: Leyra Bets on ‘One Stack’ OTT SaaS for Growth
TV[BE]urope: Bitmovin bets on vertical video: faster ads, smaller web player
TV[BE]urope: ChannelSurfer brings “cable mode” to YouTube—no algorithm required
Broadcast: DAZN eyes the sky: live sports streaming goes inflight

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 2 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 2 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →