Tag
Q2D-Web is a large-scale benchmark and leaderboard for evaluating retrieval models on web search, featuring 190 million documents and queries in ten languages with methods to minimize bias and reduce evaluation costs.
Introduces Agent Retrieval Bench, a file-level benchmark evaluating how well coding agents retrieve relevant repository files during the context-acquisition stage. The benchmark includes 427 samples across 25 repositories and evaluates various retrieval methods, finding no single family dominates.
OBLIQ-Bench introduces a suite of five oblique search tasks that expose a gap between retrieval and verification: reasoning LLMs easily recognize relevant documents once surfaced, but even state-of-the-art retrievers fail to surface them, highlighting overlooked bottlenecks in modern retrieval systems.