retrieval-benchmark

Tag

Cards List
#retrieval-benchmark

Q2D-Web: Evaluating First-Stage Retrievers at Scale (11 minute read)

TLDR AI · 3d ago

Q2D-Web is a large-scale benchmark and leaderboard for evaluating retrieval models on web search, featuring 190 million documents and queries in ten languages with methods to minimize bias and reduce evaluation costs.

0 favorites 0 likes
#retrieval-benchmark

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Hugging Face Daily Papers · 2026-07-27 Cached

Introduces Agent Retrieval Bench, a file-level benchmark evaluating how well coding agents retrieve relevant repository files during the context-acquisition stage. The benchmark includes 427 samples across 25 repositories and evaluates various retrieval methods, finding no single family dominates.

0 favorites 0 likes
#retrieval-benchmark

@lateinteraction: https://x.com/lateinteraction/status/2061285488166949254

X AI KOLs Timeline · 2026-06-01 Cached

OBLIQ-Bench introduces a suite of five oblique search tasks that expose a gap between retrieval and verification: reasoning LLMs easily recognize relevant documents once surfaced, but even state-of-the-art retrievers fail to surface them, highlighting overlooked bottlenecks in modern retrieval systems.

0 favorites 0 likes
← Back to home

Submit Feedback