wide-and-deep

Tag

Cards List
#wide-and-deep

WANDR: A Benchmark for Wide and Deep Research

arXiv cs.LG · 2026-08-18 Cached

WANDR is a benchmark for evaluating AI agents on wide and deep research tasks, focusing on high-volume data collection with verifiable accuracy. It includes 500 tasks and an evaluation harness to stress-test current systems.

0 favorites 0 likes
← Back to home

Submit Feedback