Scaling Test-Time Compute for Agentic Coding
Summary
A test-time scaling framework for agentic coding that compresses rollout trajectories into structured summaries and uses recursive voting/PDR to boost Claude-4.5-Opus to 77.6% on SWE-Bench Verified.
View Cached Full Text
Cached at: 04/23/26, 07:47 AM
Paper page - Scaling Test-Time Compute for Agentic Coding
Source: https://huggingface.co/papers/2604.16529 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Test-time scaling framework for agentic coding uses compact trajectory representations and recursive voting/parallel-distill-refine methods to improve long-horizon task performance.
Test-time scalinghas become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that can be directly compared, ranked or refined. Long-horizon coding agents violate this premise: each attempt produces an extended trajectory of actions, observations, errors, and partial progress taken by the agent. In this setting, the main challenge is no longer generating more attempts, but representing prior experience in a form that can be effectively selected from and reused. We propose atest-time scalingframework foragentic codingbased on compact representations ofrollout trajectories. Our framework converts each rollout into a structured summary that preserves its salient hypotheses, progress, and failure modes while discarding low-signal trace details. This representation enables two complementary forms of inference-time scaling. For parallel scaling, we introduceRecursive Tournament Voting(RTV), which recursively narrows a population of rollout summaries through small-group comparisons. For sequential scaling, we adaptParallel-Distill-Refine(PDR) to the agentic setting by conditioning new rollouts on summaries distilled from prior attempts. Our method consistently improves the performance of frontier coding agents acrossSWE-Bench VerifiedandTerminal-Bench v2.0. For example, by using our method Claude-4.5-Opus improves from 70.9% to 77.6% onSWE-Bench Verified(mini-SWE-agent) and 46.9% to 59.1% onTerminal-Bench v2.0(Terminus 1). Our results suggest thattest-time scalingfor long-horizon agents is fundamentally a problem of representation, selection, and reuse.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.16529 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.16529 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.16529 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Agentic Test-Time Scaling (GitHub Repo)
AutoTTS is an open-source tool that uses agentic discovery to automatically find optimal test-time scaling strategies for LLMs, significantly reducing token usage and cost through replay-based evaluation.
Can a coding agent use 57–85% less fresh model traffic without losing task success? I open-sourced my experiment
The author open-sources an execution and context layer for coding agents that cuts fresh model traffic by 57-85% while preserving task success in paired smoke tests on GPT-5.6 and Claude Opus 5, and seeks independent evaluation and sponsorship.
BTL-3 27B agentic coding and tool-use model from Bad Theory Labs (fits in 8.39GB)
Bad Theory Labs released BTL-3, a 27B open-weight agentic coding and tool-use model that fits in 8.39GB via custom quantization, achieving 92.2% performance retention and strong benchmarks including HumanEval 95.12% pass@1.
CogScale: Scalable Benchmark for Sequence Processing
CogScale is a benchmark of 14 scalable synthetic tasks designed to isolate and evaluate cognitive and memory abilities in sequence processing models. It provides a lightweight framework for rapid architectural validation and includes evaluations of seven architectures under strict parameter budgets.
I built a coding agent that gets 87% on benchmarks with a 4B parameter model, here's how
The author built SmallCode, a coding agent optimized for small local models, achieving 87% benchmark success with a 4B parameter model using techniques like compound tools, improvement loops, and token budgeting.