Tag
AACL-IJCNLP 2026 Findings paper showing that classic lexical retrieval (BM25 plus a small cross-encoder) matches or beats agentic skill retrieval on agent skill selection across 3 of 4 settings, at roughly half the cost, arguing agentic retrieval should be the fallback rather than the default.
SWE-Explore introduces a benchmark for evaluating coding agents' repository exploration capabilities, requiring ranked lists of relevant code regions within line budgets. Experiments show agentic exploration outperforms traditional retrieval, and line-level coverage remains a key differentiator.