AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Hugging Face Daily Papers Papers

Summary

AgSpec 是一个检索式推测解码框架,通过从会话、工作区和全局语料中检索草稿并动态调整草稿长度,在多智能体编码基准上实现了最高 4.37 倍(batch size 1)和 4.76 倍(batch size 16)的生成吞吐提升,优于五种现有检索式草稿器和 EAGLE-3。

Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37times at batch size 1 and 4.76times at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
Original Article
View Cached Full Text

Cached at: 10/02/26, 08:27 AM

Paper page - AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Source: https://huggingface.co/papers/2610.01108

Abstract

Retrieval-basedspeculativedecoding(SD)draftstokensbycopyingcontinuationsfromexistingtext,whichsuitscodingagentsthatrepeatedlyreproducecode,logs,andearlierattempts.Yetexistingmethodsfallshortinagentpipelines:muchofthereusabletextismissingfromtheircorporaorstoredinaformthatdiffersfromwhattheagentemits,andtheirdraftlengthsignorethatacceptlengthvariesacrossagentsanddriftsoverturns.WepresentAgSpec,aframeworkthatsuppliesthecorpusanddraft-lengthpoliciesthatexistingretrievalengineslackincoding-agentpipelines.AgSpecretrievesfromsession,workspace,andglobalcorpora,retainingtheongoingsessiontrajectoryandindexingopenedfilesintheagent’semissionformat.Itboundseachagent’sdraftlengthwithanoffline-profiledcapandadaptsthelengthonlinefromverificationfeedback.Ontworepository-levelmulti-agentcodingbenchmarks,AgSpecoutperformsfiveretrieval-baseddraftersandEAGLE-3inmostevaluatedsettings,raisinggenerationthroughputoverautoregressivedecodingupto4.37timesatbatchsize1and4.76timesatbatchsize16.AgSpecalsoremainseffectiveonbenchmarkswithoutarepositoryoramulti-agentpipeline,showingthatitsgainsgeneralizetocodingagentsbroadly.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2610\.01108

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.01108 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.01108 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.01108 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

arXiv cs.CL

AngelSpec introduces a unified training and inference framework for speculative decoding that jointly optimizes autoregressive multi-token prediction and block-parallel diffusion drafters to handle heterogeneous real-world workloads. Experiments on the Hy3 model series show up to 2.4x speedup over autoregressive decoding and 11.8% higher throughput than DFlash.

Speculative Decoding with a Speculative Vocabulary

arXiv cs.CL

This paper proposes SpecVocab, a method for selecting a per-step vocabulary subset for the draft model in speculative decoding, achieving higher acceptance length and up to 8.1% throughput improvement over EAGLE-3.

What is Speculative Decoding? (trending on paperswithco.de) [R]

Reddit r/MachineLearning

Speculative decoding is an inference optimization technique that uses a fast draft model to propose future tokens verified in parallel by a larger model, improving LLM generation speed. The article highlights its trending status on Papers with Code and a recent SGLang blog post about state-of-the-art latencies using DFlash models.