Tag
WANDR is a benchmark for evaluating AI agents on wide and deep research tasks, focusing on high-volume data collection with verifiable accuracy. It includes 500 tasks and an evaluation harness to stress-test current systems.