Tag
Perplexity releases WANDR, an open benchmark and evaluation harness for research agents, consisting of 500 realistic data-collection tasks that require both wide discovery and deep verification. Initial results show even the strongest systems achieve low scores, highlighting that wide-and-deep research remains a challenging open problem.