Beyond Retrieval: A Multitask Benchmark and Model for Code Search
Summary
This paper introduces CoREB, a contamination-limited multitask benchmark for code search that evaluates text-to-code, code-to-text, and code-to-code retrieval with fine-tuned reranking capabilities.
View Cached Full Text
Cached at: 05/11/26, 07:21 AM
Paper page - Beyond Retrieval: A Multitask Benchmark and Model for Code Search
Source: https://huggingface.co/papers/2605.04615
Abstract
A new code search benchmark called CoREB is introduced that addresses limitations of existing datasets by providing contamination-limited, multitask evaluation across text-to-code, code-to-text, and code-to-code retrieval tasks with fine-tuned reranking capabilities.
Code searchhas usually been evaluated as first-stageretrieval, even though production systems rely on broader pipelines withrerankingand developer-style queries. Existing benchmarks also suffer from data contamination, label noise, and degenerate binary relevance. In this paper, we introduceCoREB, a contamination-limited, multitask coderetrievalandrerankingbenchmark, together with a fine-tuned code reranker, that goes beyondretrievalto cover the fullcode searchpipeline.CoREBis built from counterfactually rewritten LiveCodeBench problems in five programming languages and delivered as timed releases with graded relevance judgments. We benchmark elevenembedding modelsand fivererankersacross three tasks:text-to-code,code-to-text, and code-to-code. Our experiments reveal that: \circone code-specialised embeddings dominatecode-to-code retrieval({sim}2{times} over general encoders), yet no single model wins all three tasks; \circtwo short keyword queries, the format closest to real developer search, collapse every model to near-zero nDCG@10; \circthree off-the-shelfrerankersare task-asymmetric, with a 12-point swing on code-to-code and no baseline net-positive across all tasks; \circfour our fine-tunedCoREB-Reranker is the first to achieve consistent gains across all three tasks. The data and model are released.
View arXiv pageView PDFProject pageGitHub0Add to collection
Models citing this paper1
#### hq-bench/coreb-code-reranker Text Classification• 4B• Updated4 days ago • 62 • 3
Datasets citing this paper1
#### hq-bench/coreb Viewer• Updatedabout 5 hours ago • 34.2k • 55 • 5
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.04615 in a Space README.md to link it from this page.
Collections including this paper3
Similar Articles
Recall Before Rerank: Benchmarking Deep Learning Models for Large-Scale Code-to-Code Retrieval
This paper benchmarks 17 deep learning models for first-stage recall in large-scale code-to-code retrieval, evaluating their precision, efficiency, and scalability across multiple programming languages and datasets. It introduces LLM-based code normalization and query rewriting schemes that improve precision for lower-performing models.
Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Introduces Agent Retrieval Bench, a file-level benchmark evaluating how well coding agents retrieve relevant repository files during the context-acquisition stage. The benchmark includes 427 samples across 25 repositories and evaluates various retrieval methods, finding no single family dominates.
ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval
ExecRetrieval introduces a benchmark to measure the functional-correctness gap in code-embedding retrieval, showing that top retrievers frequently rank buggy near-clone implementations above correct ones.
Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers
Researchers introduce CSR-L and CS-MTEB benchmarks showing that code-switching queries degrade IR system performance by up to 27%, revealing embedding-space divergence that current multilingual techniques cannot fix.
SWE Context Bench just proved something I think a lot of coding agent users already feel
A new benchmark paper 'SWE Context Bench' tests whether coding agents can reuse knowledge across tasks, highlighting a gap in existing benchmarks that only evaluate isolated problem-solving. The author discusses solutions like external memory and mentions tools such as langmem, mem0, supermemory, and Greplica.