parallelquant
July 19, 2026 · MarkTechPost

Perplexity releases WANDR, a benchmark for wide-and-deep research agents

Perplexity's WANDR is an open benchmark of 500 evidence-heavy tasks testing whether research agents can find many qualifying entities and back each with citable, re-verifiable evidence. Perplexity's own 'Search as Code' system currently leads, scoring 0.363 soft F1 and 0.133 hard F1.

Why it matters: The low absolute scores, well under half on the lenient metric and far lower on the strict one, show current research agents are still weak at exhaustive, verifiable search. That's a meaningful gap for a product category increasingly marketed as an 'AI research assistant,' where completeness and citability matter more than a single plausible-sounding answer.

Related updates