wandr

Tag

Cards List
#wandr

WANDR Benchmark: Evaluating Research Agents That Must Search Wide and Deep (15 minute read)

TLDR AI · 2026-07-15 Cached

Perplexity releases WANDR, an open benchmark and evaluation harness for research agents, consisting of 500 realistic data-collection tasks that require both wide discovery and deep verification. Initial results show even the strongest systems achieve low scores, highlighting that wide-and-deep research remains a challenging open problem.

0 favorites 0 likes
← Back to home

Submit Feedback