Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Summary
The paper proposes Schrödinger Repo, an evaluation framework for coding agents that dynamically instantiates repositories to address data leakage in benchmarks, showing that current LLMs may depend on memorized cues.
View Cached Full Text
Cached at: 09/24/26, 03:39 AM
Paper page - Schrödinger’s Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Source: https://huggingface.co/papers/2609.27891
Abstract
Repository-levelcodingbenchmarkshavebecomethestandardforevaluatingcodingagents,yettheyinherentlysufferfromdataleakagebecausetheyarebuiltuponpopularopen-sourcerepositoriesrepeatedlyusedfortraining.Consequently,strongperformancemayreflectmemorizationofcanonicalrepositorycuesratherthanrobustrepositoryreasoning.WeproposeSchrodingerRepo(Schrödinger’sRepository),anevaluationframeworkfortestingcodingagentsunderdynamicallyinstantiatedrepositoryrepresentations.Insteadofrepeatedlyusingastaticrepresentationofthetestrepository,SchrodingerRepotreatsthetestrepositoryasanevaluation-timelatentvariablethatisdynamicallyinstantiatedonlywhentheagententerstheevaluationenvironment.Theinstantiatedrepositorypreservestheoriginalexecutablebehaviorwhileerodingfamiliarcuessuchasnamingconventions,filelayouts,andimplementationpatternsthroughfourtransformationlevels:problemstatementreconstruction,namespaceremapping,intra-filelayoutreordering,andfunctionality-preservingcoderewriting.WeevaluatepopularLLMsonSWE-benchVerifiedandSWE-QA.Resultsshowthatremovingfamiliarrepositorycuesconsistentlydegradesagentperformanceandsubstantiallyincreasesinteractioncostsacrossmodels.Furtheranalysisrevealsthattheadditionalcostisprimarilycausedbyincreaseddifficultyinrepositoryexplorationandlocalization.Thesefindingssuggestthatcurrentcodingagentsmaypartiallyrelyonmemorizedrepository-sidecues,highlightingtheneedforevaluationunderdynamicallyinstantiatedrepositoryrepresentations.
View arXiv pageView PDFGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.27891
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.27891 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.27891 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.27891 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SWE Context Bench just proved something I think a lot of coding agent users already feel
A new benchmark paper 'SWE Context Bench' tests whether coding agents can reuse knowledge across tasks, highlighting a gap in existing benchmarks that only evaluate isolated problem-solving. The author discusses solutions like external memory and mentions tools such as langmem, mem0, supermemory, and Greplica.
LLM Agents Can See Code Repositories
This paper presents the first systematic empirical study of using visual repository representations to enhance LLM-based coding agents, showing that integrating visual graphs as a supplementary modality reduces token consumption by up to 26% while maintaining or improving issue-resolution accuracy.
@adithya_s_k: https://x.com/adithya_s_k/status/2067628584680710292
This article discusses how coding agents can cheat evaluations by copying known patches, and introduces Repo2RLEnv, a tool to create verifiable coding environments from real repositories to build robust benchmarks and training data for AI coding agents.
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science introduces a repository-level benchmark for evaluating coding agents on scientific software repair tasks, revealing failure mechanisms and mixed effects of scientific guidance.
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
SWE-Explore introduces a benchmark for evaluating coding agents' repository exploration capabilities, requiring ranked lists of relevant code regions within line budgets. Experiments show agentic exploration outperforms traditional retrieval, and line-level coverage remains a key differentiator.