Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Hugging Face Daily Papers Papers

Summary

The paper proposes Schrödinger Repo, an evaluation framework for coding agents that dynamically instantiates repositories to address data leakage in benchmarks, showing that current LLMs may depend on memorized cues.

Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-source repositories repeatedly used for training. Consequently, strong performance may reflect memorization of canonical repository cues rather than robust repository reasoning. We propose SchrodingerRepo (Schrödinger's Repository), an evaluation framework for testing coding agents under dynamically instantiated repository representations. Instead of repeatedly using a static representation of the test repository, SchrodingerRepo treats the test repository as an evaluation-time latent variable that is dynamically instantiated only when the agent enters the evaluation environment. The instantiated repository preserves the original executable behavior while eroding familiar cues such as naming conventions, file layouts, and implementation patterns through four transformation levels: problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting. We evaluate popular LLMs on SWE-bench Verified and SWE-QA. Results show that removing familiar repository cues consistently degrades agent performance and substantially increases interaction costs across models. Further analysis reveals that the additional cost is primarily caused by increased difficulty in repository exploration and localization. These findings suggest that current coding agents may partially rely on memorized repository-side cues, highlighting the need for evaluation under dynamically instantiated repository representations.
Original Article
View Cached Full Text

Cached at: 09/24/26, 03:39 AM

Paper page - Schrödinger’s Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Source: https://huggingface.co/papers/2609.27891

Abstract

Repository-levelcodingbenchmarkshavebecomethestandardforevaluatingcodingagents,yettheyinherentlysufferfromdataleakagebecausetheyarebuiltuponpopularopen-sourcerepositoriesrepeatedlyusedfortraining.Consequently,strongperformancemayreflectmemorizationofcanonicalrepositorycuesratherthanrobustrepositoryreasoning.WeproposeSchrodingerRepo(Schrödinger’sRepository),anevaluationframeworkfortestingcodingagentsunderdynamicallyinstantiatedrepositoryrepresentations.Insteadofrepeatedlyusingastaticrepresentationofthetestrepository,SchrodingerRepotreatsthetestrepositoryasanevaluation-timelatentvariablethatisdynamicallyinstantiatedonlywhentheagententerstheevaluationenvironment.Theinstantiatedrepositorypreservestheoriginalexecutablebehaviorwhileerodingfamiliarcuessuchasnamingconventions,filelayouts,andimplementationpatternsthroughfourtransformationlevels:problemstatementreconstruction,namespaceremapping,intra-filelayoutreordering,andfunctionality-preservingcoderewriting.WeevaluatepopularLLMsonSWE-benchVerifiedandSWE-QA.Resultsshowthatremovingfamiliarrepositorycuesconsistentlydegradesagentperformanceandsubstantiallyincreasesinteractioncostsacrossmodels.Furtheranalysisrevealsthattheadditionalcostisprimarilycausedbyincreaseddifficultyinrepositoryexplorationandlocalization.Thesefindingssuggestthatcurrentcodingagentsmaypartiallyrelyonmemorizedrepository-sidecues,highlightingtheneedforevaluationunderdynamicallyinstantiatedrepositoryrepresentations.

View arXiv pageView PDFGitHub4Add to collection

Get this paper in your agent:

hf papers read 2609\.27891

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.27891 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.27891 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.27891 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

LLM Agents Can See Code Repositories

Hugging Face Daily Papers

This paper presents the first systematic empirical study of using visual repository representations to enhance LLM-based coding agents, showing that integrating visual graphs as a supplementary modality reduces token consumption by up to 26% while maintaining or improving issue-resolution accuracy.

@adithya_s_k: https://x.com/adithya_s_k/status/2067628584680710292

X AI KOLs Timeline

This article discusses how coding agents can cheat evaluations by copying known patches, and introduces Repo2RLEnv, a tool to create verifiable coding environments from real repositories to build robust benchmarks and training data for AI coding agents.

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Hugging Face Daily Papers

SWE-Explore introduces a benchmark for evaluating coding agents' repository exploration capabilities, requiring ranked lists of relevant code regions within line budgets. Experiments show agentic exploration outperforms traditional retrieval, and line-level coverage remains a key differentiator.