StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoProduct LaunchJuly 19, 2026

Perplexity launches WANDR benchmark to stress-test high-volume AI research agents

Perplexity launches WANDR benchmark to stress-test high-volume AI research agents
MarkTechPost

Perplexity AI has released WANDR, an open-source evaluation benchmark designed to test the efficacy of AI research agents in performing complex, multi-step knowledge gathering tasks. The benchmark consists of 500 tasks, requiring agents to discover entities and verify them against credible, evidence-backed source material.

Key Takeaways

  • WANDR tasks require a median of 245 records per task, utilizing a hierarchy to validate qualifying companies, employees, and supporting URLs.
  • Perplexity’s Search as Code (SaC) system led the benchmark with a 0.363 Soft F1 score, while Anthropic trailed at 0.249.
  • The benchmark reveals high costs for accuracy, with task expenses ranging from a $0.03 baseline to $324.83 for high-effort reasoning.
  • Evaluation involves re-fetching cited pages to verify that excerpts truly exist and support the specific claims made by the agent.

Why It Matters

This release shifts the evaluation of AI agents from simple question-answering toward the high-volume verification required for competitive intelligence and due diligence. For the streaming industry, such tools are critical for mapping fragmented licensing rights and global market data, but current results show a significant gap between production readiness and perfect accuracy. No existing system, including OpenAI or Anthropic, currently solves the benchmark, highlighting the persistent hallucination and coverage risks in automated research. Watch for whether OpenAI releases a comparable multi-step search benchmark for its 'SearchGPT' prototypes to challenge Perplexity’s technical lead in this niche.

Additional Context

The release of WANDR follows a period of heightened competition in the AI search and research sector. Per Reuters in July 2024, OpenAI officially entered this space with SearchGPT, a temporary prototype designed to combine its AI models with real-time web information. This move directly challenged Perplexity’s established model of using large language models as a primary interface for web discovery. While OpenAI has focused on consumer-facing search, Perplexity’s focus on benchmarks like DRACO and WANDR signals a strategic pivot toward proving the reliability of its systems for enterprise-grade research and data extraction. Institutional adoption of these agents remains cautious due to ongoing legal and technical frictions. Per The Verge in June 2024, Perplexity faced scrutiny over its web crawling practices, with some publishers claiming the company bypassed the Robots Exclusion Protocol (robots.txt). By open-sourcing WANDR on GitHub, Perplexity is attempting to formalize the 'Search as Code' category and establish industry standards for how research agents should be audited. This transparency is likely a response to the need for verifiable accuracy in B2B applications where a single citation error can invalidate a due diligence report. Technical performance across the sector remains varied, as shown by the cost-to-accuracy trade-offs identified in WANDR's initial results. According to reporting from TechCrunch in early 2026, the cost of running high-inference research agents has become a primary bottleneck for widespread enterprise deployment. Perplexity’s data showing a range from pennies to over $300 per task reflects the immense compute power required for 'deep research' modes. As streaming platforms increasingly use AI to automate metadata enrichment and competitor price monitoring, the efficiency metrics established by WANDR will become a key procurement benchmark for media tech stacks.


Read full article at marktechpost.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

WeRSM (We are Social Media): Google morphs Flow Music Spaces into end-to-end AI production studio
Tech Times: Black Forest Labs launches FLUX 3 multimodal model for video and robotics
BigGo: YouTube Ads engineers detail staged evaluation framework for LLM agents

Newest

about 22 hours ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
about 22 hours ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
about 22 hours ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
about 23 hours ago
Investing.com: TF1 Digital Revenues Jump 17% as Netflix Partnership Exceeds Growth Targets
about 23 hours ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
2 days ago
Ealing Times: YouTube debuts UK Shopping Affiliate Programme with M&S and Currys
2 days ago
Investing.com: AMD and Cerebras debut disaggregated architecture to slash AI inference latency
2 days ago
MediaPost: Sports leagues explore non-exclusive local rights as RSN model collapses
2 days ago
YouTube: Blackmagic Design details GPU optimization protocols for DaVinci Resolve workflows
2 days ago
Startup Fortune: AI data centers threaten US grid stability and freeze cloud pipelines
2 days ago
TechRadar: OpenAI joins coalition lobbying against strict open-weight AI model regulations
2 days ago
Startup Fortune: SPAN and Nvidia board residential homes with 16-GPU Blackwell compute nodes
2 days ago
Digital Applied: Google faces €890M EU fine as Digital Markets Act enforcement accelerates
2 days ago
iZOOlogic: Ultra Clean Android App Masquerades as Utility to Host Malware-Grade Adware
2 days ago
SiliconANGLE: HPE and AMD converge supercomputing and AI via liquid-cooled GX5000
2 days ago
MarketBeat: AMD data center revenue surges 38% to $10.25B on AI demand
2 days ago
PPC Land: Acast revenue per listen jumps 26% despite flat audience growth

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.Tech Times60
  4. 4.YouTube59
  5. 5.AdExchanger57
  6. 6.TechCrunch54
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →

Newest

about 22 hours ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
about 22 hours ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
about 22 hours ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
about 23 hours ago
Investing.com: TF1 Digital Revenues Jump 17% as Netflix Partnership Exceeds Growth Targets
about 23 hours ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
2 days ago
Ealing Times: YouTube debuts UK Shopping Affiliate Programme with M&S and Currys
2 days ago
Investing.com: AMD and Cerebras debut disaggregated architecture to slash AI inference latency
2 days ago
MediaPost: Sports leagues explore non-exclusive local rights as RSN model collapses
2 days ago
YouTube: Blackmagic Design details GPU optimization protocols for DaVinci Resolve workflows
2 days ago
Startup Fortune: AI data centers threaten US grid stability and freeze cloud pipelines
2 days ago
TechRadar: OpenAI joins coalition lobbying against strict open-weight AI model regulations
2 days ago
Startup Fortune: SPAN and Nvidia board residential homes with 16-GPU Blackwell compute nodes
2 days ago
Digital Applied: Google faces €890M EU fine as Digital Markets Act enforcement accelerates
2 days ago
iZOOlogic: Ultra Clean Android App Masquerades as Utility to Host Malware-Grade Adware
2 days ago
SiliconANGLE: HPE and AMD converge supercomputing and AI via liquid-cooled GX5000
2 days ago
MarketBeat: AMD data center revenue surges 38% to $10.25B on AI demand
2 days ago
PPC Land: Acast revenue per listen jumps 26% despite flat audience growth

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.Tech Times60
  4. 4.YouTube59
  5. 5.AdExchanger57
  6. 6.TechCrunch54
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →