Perplexity launches WANDR benchmark to stress-test high-volume AI research agents
Perplexity AI has released WANDR, an open-source evaluation benchmark designed to test the efficacy of AI research agents in performing complex, multi-step knowledge gathering tasks. The benchmark consists of 500 tasks, requiring agents to discover entities and verify them against credible, evidence-backed source material.
Key Takeaways
- WANDR tasks require a median of 245 records per task, utilizing a hierarchy to validate qualifying companies, employees, and supporting URLs.
- Perplexity’s Search as Code (SaC) system led the benchmark with a 0.363 Soft F1 score, while Anthropic trailed at 0.249.
- The benchmark reveals high costs for accuracy, with task expenses ranging from a $0.03 baseline to $324.83 for high-effort reasoning.
- Evaluation involves re-fetching cited pages to verify that excerpts truly exist and support the specific claims made by the agent.
Why It Matters
This release shifts the evaluation of AI agents from simple question-answering toward the high-volume verification required for competitive intelligence and due diligence. For the streaming industry, such tools are critical for mapping fragmented licensing rights and global market data, but current results show a significant gap between production readiness and perfect accuracy. No existing system, including OpenAI or Anthropic, currently solves the benchmark, highlighting the persistent hallucination and coverage risks in automated research. Watch for whether OpenAI releases a comparable multi-step search benchmark for its 'SearchGPT' prototypes to challenge Perplexity’s technical lead in this niche.
Additional Context
The release of WANDR follows a period of heightened competition in the AI search and research sector. Per Reuters in July 2024, OpenAI officially entered this space with SearchGPT, a temporary prototype designed to combine its AI models with real-time web information. This move directly challenged Perplexity’s established model of using large language models as a primary interface for web discovery. While OpenAI has focused on consumer-facing search, Perplexity’s focus on benchmarks like DRACO and WANDR signals a strategic pivot toward proving the reliability of its systems for enterprise-grade research and data extraction. Institutional adoption of these agents remains cautious due to ongoing legal and technical frictions. Per The Verge in June 2024, Perplexity faced scrutiny over its web crawling practices, with some publishers claiming the company bypassed the Robots Exclusion Protocol (robots.txt). By open-sourcing WANDR on GitHub, Perplexity is attempting to formalize the 'Search as Code' category and establish industry standards for how research agents should be audited. This transparency is likely a response to the need for verifiable accuracy in B2B applications where a single citation error can invalidate a due diligence report. Technical performance across the sector remains varied, as shown by the cost-to-accuracy trade-offs identified in WANDR's initial results. According to reporting from TechCrunch in early 2026, the cost of running high-inference research agents has become a primary bottleneck for widespread enterprise deployment. Perplexity’s data showing a range from pennies to over $300 per task reflects the immense compute power required for 'deep research' modes. As streaming platforms increasingly use AI to automate metadata enrichment and competitor price monitoring, the efficiency metrics established by WANDR will become a key procurement benchmark for media tech stacks.
Read full article at marktechpost.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source