DEdit: Iterative Draft Editing for Speculative Decoding

arXiv cs.CL Papers

Summary

DEdit introduces a diffusion-based drafter for speculative decoding that iteratively edits drafts via token-to-token predictions, using a ProposalMix training scheme to repair errors while preserving correct tokens. It achieves macro-average speedups of 5.72× and 5.97× on Qwen3-4B and Qwen3-8B across seven benchmarks.

arXiv:2609.38510v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:44 AM

# DEdit: Iterative Draft Editing for Speculative Decoding
Source: [https://arxiv.org/html/2609.38510](https://arxiv.org/html/2609.38510)
Longxuan YuAffiliation:University of California, RiversideAffiliation:Amazon Web ServicesWork done during an internship at Amazon Web ServicesEqual contributionBingsen ChenAffiliation:Amazon Web ServicesAffiliation:New York UniversityWork done during an internship at Amazon Web ServicesEqual contributionDongkyu LeeAffiliation:Amazon Web ServicesYi XiangAffiliation:Amazon Web ServicesHideo KobayashiAffiliation:Amazon Web ServicesSheng ZhangAffiliation:Amazon Web ServicesShuaichen ChangAffiliation:Amazon Web ServicesXing NiuAffiliation:Amazon Web ServicesZhuoyan XuAffiliation:Amazon Web ServicesGreg Ver SteegAffiliation:University of California, RiversideJiarong JiangAffiliation:Amazon Web Services

###### Abstract

Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel\. Diffusion\-based drafters further reduce drafting latency by proposing multiple tokens at once\. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions\. We introduce DEdit, a diffusion\-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token\-to\-token predictions\. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix\. To teach the model to repair errors while preserving correct predictions, we proposeProposalMix, a training scheme that mixes draft predictions with ground\-truth tokens based on first\-pass confidence during training\. Across seven benchmarks on Qwen3\-4B and Qwen3\-8B, DEdit achieves the highest macro\-average token acceptance and speedup among the evaluated drafters, reaching macro\-average speedups of5\.72×5\.72\\timesand5\.97×5\.97\\timesover autoregressive generation under greedy decoding, respectively\. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and thatProposalMixhalves harmful edits that shorten the accepted prefix\. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains\.

## 1Introduction

Figure 1:Bidirectional editing before verification extends the accepted prefix\.Autoregressive large language models \(LLMs\) generate text one token at a time, creating a sequential bottleneck for inference speed\. Speculative decoding mitigates this bottleneck by using a lightweight drafter to propose multiple tokens that the target model verifies in a single forward pass\([Leviathan et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib1);[Chen et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib2)\)\. Its efficiency depends on both the number of accepted tokens and the cost of producing each proposal\. However, autoregressive drafters themselves generate proposals sequentially, retaining the overhead of repeated forward passes\. To reduce this overhead, parallel drafting, particularly diffusion\-based drafting, has been widely adopted in recent work on speculative decoding\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10);[Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16);[Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17);[Wang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib18)\)\.

This shift to parallel drafting introduces a different challenge\. Because multiple draft positions are predicted without conditioning on one another, proposal quality tends to degrade at later positions, causing acceptance rates to decay\. Existing methods address this limitation by reintroducing causal dependencies among draft tokens to improve proposal quality\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16);[Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17);[Wang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib18)\)\. We take a different view\. The lack of causal dependence also means that an error at an early position does not necessarily invalidate later predictions\. In fact, later candidates can still be correct even beyond the first rejected token\. However, standard prefix verification discards this informative suffix entirely\. We therefore ask whether these otherwise wasted parallel predictions can instead be used to revise earlier errors and extend the accepted prefix before target verification\.

We introduceDEdit, a diffusion\-based drafter that generates an initial proposal from masked inputs and then iteratively refines all positions in parallel using bidirectional context\. This allows later predictions to help repair earlier errors before verification, after which the target model verifies only the final proposal using the standard procedure\. Figure[1](https://arxiv.org/html/2609.38510#S1.F1)illustrates how editing can correct both an early mismatch and the downstream continuation, extending the accepted prefix\.

Training DEdit requires inputs that resemble imperfect first\-pass drafts while retaining useful predictions that should not be overwritten\. We thus design a novel training algorithm,ProposalMix, which constructs training proposals by mixing first\-pass draft predictions with ground\-truth tokens according to confidence\. This creates partially correct proposals with both reliable context and realistic errors, teaching the editor when to preserve a prediction and when to revise it\.

Our analyses show that the drafter improves token acceptance through inference\-time edits beyond the single supervised editing step\. Editing also makes wider proposal windows more useful: extending the window from 16 to 32 tokens raises acceptance by 10\.3% after editing, compared with 2\.1% for DFlash\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\)\.ProposalMixteaches the editor to avoid harmful edits, halving the rounds in which editing shortens the accepted prefix\. Controlled comparisons indicate that editing gains rely on future proposal context, which causal correctors such as Domino\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16)\), restricted to prefix context, cannot use: hiding later tokens from the editor reduces acceptance on every task we evaluate\. This benefit is largest on predictable outputs: on reasoning, summary, and synthetic copy\-heavy tasks, DEdit’s acceptance advantage over Domino exceeds that on the standard seven\-benchmark suite and is largest on the copy\-heavy task\.

Across seven benchmarks under greedy and stochastic decoding, DEdit achieves the highest macro\-average token acceptance among the evaluated drafters on both Qwen3\-4B and Qwen3\-8B\. Editing trades additional drafter computation for longer accepted prefixes, so its speed benefit depends on how much drafting cost the execution backend exposes\. With CUDA Graph drafting and eager target verification, where the CPU dispatches target operations while the drafter runs, DEdit also achieves the highest macro\-average speedup, reaching5\.72×5\.72\\timeson Qwen3\-4B and5\.97×5\.97\\timeson Qwen3\-8B under greedy decoding\.

Our contributions are:

- •We introduce DEdit, a bidirectional drafter that generates and iteratively edits complete proposals before verification\. Its acceptance improves with additional inference\-time edits and continues to increase as proposal windows expand from 16 to 32 tokens\.
- •We develop joint proposal\-and\-edit training withProposalMix, which teaches the drafter to preserve correct tokens and revise incorrect ones in a way that aligns editing with the prefix\-acceptance objective of speculative decoding\.
- •We provide controlled evidence that access to future proposal context is a key source of the gains from bidirectional editing, revealing a capability unavailable to causal correction methods restricted to prefix context\.

## 2Preliminaries

### 2\.1Speculative Decoding

Speculative decoding accelerates generation from a frozen autoregressive target modelppusing an efficient drafterqq\([Leviathan et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib1);[Chen et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib2)\)\. Each decoding round uses a proposal window of widthWW, comprising one committed anchor tokeny0y\_\{0\}andW−1W\-1unverified candidate tokens\. The anchor is the last token of the current output prefix and remains fixed during drafting and editing\. The target evaluates the window in a single forward pass and accepts the longest candidate prefix consistent with its own predictions; the first mismatch terminates acceptance, and the remaining candidates are discarded regardless of matches at later positions\. Ifaacandidates are accepted, verification appends them and one additional target token, which becomes the anchor for the next round, advancing generation bya\+1a\+1tokens\. We denote the expected token acceptance per round byτ=𝔼⁡\[a\+1\]\\tau=\\mathbb\{E\}\[a\+1\]\. Because DEdit changes only proposal generation and retains this verification protocol, its improvements appear directly as a higherτ\\tau\.

### 2\.2Parallel Diffusion Drafting

Parallel diffusion drafters predict candidate blocks from masked inputs\. We consider the target\-conditioned, single\-pass formulation exemplified by DFlash\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\)\. Given decoded context𝒙\{\\bm\{x\}\}and target features𝒉⁡\(𝒙\)\{\\bm\{h\}\}\(\{\\bm\{x\}\}\), the drafter predicts all masked positions in𝒛\(0\)=\[y0,mask,…,mask\]\{\\bm\{z\}\}^\{\(0\)\}=\[y\_\{0\},\\text\{\{mask\}\},\\ldots,\\text\{\{mask\}\}\]simultaneously:

qθ,i\(1\)\(v\):=qθ\(Yi=v∣𝒙,𝒉\(𝒙\),𝒛\(0\)\),y^i\(1\)=argmaxv∈𝒱qθ,i\(1\)\(v\),i=1,…,W−1\.q\_\{\\theta,i\}^\{\(1\)\}\(v\):=q\_\{\\theta\}\\\!\\left\(Y\_\{i\}=v\\mid\{\\bm\{x\}\},\{\\bm\{h\}\}\(\{\\bm\{x\}\}\),\{\\bm\{z\}\}^\{\(0\)\}\\right\),\\qquad\\hat\{y\}\_\{i\}^\{\(1\)\}=\\arg\\max\_\{v\\in\\mathcal\{V\}\}q\_\{\\theta,i\}^\{\(1\)\}\(v\),\\quad i=1,\\ldots,W\-1\.The predictions form the initial proposal𝒛\(1\)=\[y0,y^1\(1\),…,y^W−1\(1\)\]\{\\bm\{z\}\}^\{\(1\)\}=\[y\_\{0\},\\hat\{y\}\_\{1\}^\{\(1\)\},\\ldots,\\hat\{y\}\_\{W\-1\}^\{\(1\)\}\]\. DEdit uses this proposal as the starting point for bidirectional editing before target verification\.

## 3DEdit

### 3\.1Overview

DEdit uses a diffusion\-based drafter to not only generate, but alsoiteratively edit proposalsbefore target verification\. We first describe the iterative editing procedure, explaining how each pass uses the preceding proposal as bidirectional context to update candidate tokens \(Section[3\.2](https://arxiv.org/html/2609.38510#S3.SS2)\)\. We then introduce joint proposal\-and\-edit training andProposalMix, which constructs editing inputs that teach the drafter to revise imperfect proposals while preserving correct predictions \(Section[3\.3](https://arxiv.org/html/2609.38510#S3.SS3)\)\. Finally, we describe how the trained editor is adapted to larger proposal windows through continued training \(Section[3\.4](https://arxiv.org/html/2609.38510#S3.SS4)\)\. Figure[2](https://arxiv.org/html/2609.38510#S3.F2)summarizes the inference and training pipelines\.

Figure 2:DEdit inference and training\.\(a\)At inference time, drafter generates and edits its proposal tokens throughKKpasses, sending only the final proposal to the target model for verification\.\(b\)ProposalMixconstructs P2 inputs for joint proposal\-and\-edit training\.
### 3\.2Iterative Proposal Editing

As illustrated in Figure[2](https://arxiv.org/html/2609.38510#S3.F2)\(a\), DEdit appliesK−1K\-1editing passes to the initial proposal𝒛\(1\)\{\\bm\{z\}\}^\{\(1\)\}using the same drafter, whereKKdenotes the total number of drafting passes, including initial proposal generation\. Throughout editing, the decoded context𝒙\{\\bm\{x\}\}, target features𝒉⁡\(𝒙\)\{\\bm\{h\}\}\(\{\\bm\{x\}\}\), and anchored tokeny0y\_\{0\}remain fixed\. At each passk∈\{2,…,K\}k\\in\\\{2,\\ldots,K\\\}, the drafter takes the complete preceding proposal𝒛\(k−1\)\{\\bm\{z\}\}^\{\(k\-1\)\}as input and predicts:

qθ,i\(k\)\(v\):=qθ\(Yi=v∣𝒙,𝒉\(𝒙\),𝒛\(k−1\)\),y^i\(k\)=argmaxv∈𝒱qθ,i\(k\)\(v\),i=1,…,W−1\.q\_\{\\theta,i\}^\{\(k\)\}\(v\):=q\_\{\\theta\}\\\!\\left\(Y\_\{i\}=v\\mid\{\\bm\{x\}\},\{\\bm\{h\}\}\(\{\\bm\{x\}\}\),\{\\bm\{z\}\}^\{\(k\-1\)\}\\right\),\\qquad\\hat\{y\}\_\{i\}^\{\(k\)\}=\\arg\\max\_\{v\\in\\mathcal\{V\}\}q\_\{\\theta,i\}^\{\(k\)\}\(v\),\\quad i=1,\\ldots,W\-1\.These predictions form the updated proposal𝒛\(k\)=\[y0,y^1\(k\),…,y^W−1\(k\)\]\{\\bm\{z\}\}^\{\(k\)\}=\[y\_\{0\},\\hat\{y\}\_\{1\}^\{\(k\)\},\\ldots,\\hat\{y\}\_\{W\-1\}^\{\(k\)\}\]\.

All candidate positions are updated simultaneously from the same preceding proposal\. With bidirectional attention, positioniican condition on both earlier and later candidate tokens, including those at positionsj\>ij\>i\. Later predictions can therefore provide context for correcting earlier errors and extending the accepted prefix\. Every candidate position is eligible for revision at each pass, without confidence gating\.

AfterKKpasses, only the final proposal𝒛\(K\)\{\\bm\{z\}\}^\{\(K\)\}is submitted to the target verifier\. All passes share the same parameters, so each additional edit requires one drafter forward pass but introduces neither additional model parameters nor an intermediate target query\.

### 3\.3Joint Training withProposalMix

We train a single network for both a proposal pass \(P1\) and an editing pass \(P2\), as illustrated in Figure[2](https://arxiv.org/html/2609.38510#S3.F2)\(b\)\. The proposal objective trains generation from masked inputs, while the editing objective teaches the network to preserve correct tokens and revise errors\. To support editing supervision,ProposalMixsupplies ground\-truth tokens at positions selected using first\-pass confidence and retains draft predictions elsewhere\. As the drafter’s predictions change during training, confidence\-based selection adapts the editing inputs to its evolving error patterns \(Appendix[B\.3](https://arxiv.org/html/2609.38510#A2.SS3)\)\.

For each positioni∈\{1,…,W−1\}i\\in\\\{1,\\ldots,W\-1\\\}, we define:

ci=maxv∈𝒱qθ,i\(1\)\(v\),mi=𝕀\[ci\>η\]\.c\_\{i\}=\\max\_\{v\\in\\mathcal\{V\}\}q\_\{\\theta,i\}^\{\(1\)\}\(v\),\\qquad m\_\{i\}=\\mathbb\{I\}\[c\_\{i\}\>\\eta\]\.Given the ground\-truth target tokenyi⋆y\_\{i\}^\{\\star\}, the mixed input is

z~i\(1\)=\{yi⋆,mi=1,y^i\(1\),mi=0\.\\tilde\{z\}^\{\(1\)\}\_\{i\}=\\begin\{cases\}y\_\{i\}^\{\\star\},&m\_\{i\}=1,\\\\ \\hat\{y\}\_\{i\}^\{\(1\)\},&m\_\{i\}=0\.\\end\{cases\}The anchored tokeny0y\_\{0\}remains fixed, and we setη=0\.5\\eta=0\.5across all experiments\. At high\-confidence positions \(mi=1m\_\{i\}=1\), the ground\-truth token provides correct context for editing\. This leaves correct first\-pass predictions unchanged and replaces high\-confidence errors\. At the remaining positions \(mi=0m\_\{i\}=0\), the original predictions are retained, exposing the editor to model\-generated errors as well as correct predictions that it should preserve\.

Ground\-truth tokens are used as proposal inputs only during training\. The first\-pass argmax predictions and confidence masks used to construct𝒛~\(1\)\\tilde\{\{\\bm\{z\}\}\}^\{\(1\)\}are detached from the computational graph\. During training, the second\-pass distribution is

qθ,i\(2\)\(v\):=qθ\(Yi=v∣𝒙,𝒉\(𝒙\),detach\(𝒛~\(1\)\)\),i=1,…,W−1\.q\_\{\\theta,i\}^\{\(2\)\}\(v\):=q\_\{\\theta\}\\\!\\left\(Y\_\{i\}=v\\mid\{\\bm\{x\}\},\{\\bm\{h\}\}\(\{\\bm\{x\}\}\),\\mathrm\{detach\}\(\\tilde\{\{\\bm\{z\}\}\}^\{\(1\)\}\)\\right\),\\qquad i=1,\\ldots,W\-1\.The editing loss therefore does not backpropagate through the first\-pass predictions or confidence decisions used to construct its input\.

Both passes are supervised across all valid future token positions𝒯⊆\{1,…,W−1\}\\mathcal\{T\}\\subseteq\\\{1,\\ldots,W\-1\\\}\. We use position\-decay weightswi=exp\(−\(i−1\)/γ\)w\_\{i\}=\\exp\(\-\(i\-1\)/\\gamma\), whereγ\>0\\gamma\>0controls the decay\. These weights emphasize earlier positions because an early mismatch prevents later candidates from being accepted, while retaining supervision for every valid position:

ℒCE\(k\)=−∑i∈𝒯wi​log⁡qθ,i\(k\)​\(yi⋆\)∑i∈𝒯wi,k∈\{1,2\}\.\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(k\)\}=\-\\frac\{\\sum\_\{i\\in\\mathcal\{T\}\}w\_\{i\}\\log q\_\{\\theta,i\}^\{\(k\)\}\(y\_\{i\}^\{\\star\}\)\}\{\\sum\_\{i\\in\\mathcal\{T\}\}w\_\{i\}\},\\qquad k\\in\\\{1,2\\\}\.We writeℒunmask=ℒCE\(1\)\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}=\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(1\)\}for the proposal pass andℒedit=ℒCE\(2\)\\mathcal\{L\}\_\{\\mathrm\{edit\}\}=\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(2\)\}for the editing pass, and train the shared parameters with the sum of these separately normalized losses:

ℒ=ℒunmask\+ℒedit\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}\+\\mathcal\{L\}\_\{\\mathrm\{edit\}\}\.We refer to minimizingℒ\\mathcal\{L\}over shared parametersθ\\thetaas*joint training*\. It contrasts with*two\-stage training*, in which a separate drafter is first trained withℒunmask\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}and frozen, and an editor is then trained withℒedit\\mathcal\{L\}\_\{\\mathrm\{edit\}\}on the drafter’s proposals\. Joint training avoids a second model at inference, and it may expose the editing pass to a wider range of errors, since the errors of the proposal pass change as it improves during training\.ProposalMixchanges only the input to P2; both passes use the same target labels, valid positions, and loss weights\. Editing supervision covers both positions supplied with ground\-truth tokens and those retaining draft predictions\. At inference time, every editing pass consumes the preceding model\-generated proposal, and passesk≥3k\\geq 3reuse the learned editing operation without additional training\. Appendix[B\.2](https://arxiv.org/html/2609.38510#A2.SS2)provides the training hyperparameters\.

### 3\.4Window Expansion

Bidirectional editing allows earlier positions to use later proposal tokens as context\. Expanding the proposal window therefore provides both more candidates for verification and additional future context for editing\. However, parallel diffusion drafters can exhibit lower conditional token acceptance rates at later positions\([Sandler et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib8);[Inco AI, 2026](https://arxiv.org/html/2609.38510#bib.bib11)\)\. Prior evaluations of DFlash further show that increasing the draft budget can reduce both acceptance length and end\-to\-end speedup\([Hu et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib12)\)\. To benefit from a larger window, the drafter must produce useful predictions at the additional positions and use them when revising earlier ones\. A drafter trained atW=16W=16, however, has never been supervised at these positions, so we continue training it at larger windows to adapt both proposal generation and editing to the expanded context\.

We initialize independent training runs atW=24W=24andW=32W=32from the same converged checkpoint trained atW=16W=16, and continue each for 3,000 optimization steps\. We increase the position\-decay parameter fromγ=7\\gamma=7atW=16W=16toγ=11\\gamma=11atW=24W=24andγ=15\\gamma=15atW=32W=32, slowing the decay of loss weights at more distant positions\.

## 4Experiments

### 4\.1Setup

Models and Benchmarks\.We evaluate DEdit on Qwen3\-4B and Qwen3\-8B\([Yang et al\., 2025](https://arxiv.org/html/2609.38510#bib.bib19)\), the targets used by DFlash and DSpark, across seven benchmarks: GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38510#bib.bib20)\), MATH\-500\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.38510#bib.bib21);[Lightman et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib22)\), and AIME25\([Mathematical Association of America, 2025](https://arxiv.org/html/2609.38510#bib.bib23)\)\(Math\); HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2609.38510#bib.bib24)\), MBPP\([Austin et al\., 2021](https://arxiv.org/html/2609.38510#bib.bib25)\), and LiveCodeBench\([Jain et al\., 2024](https://arxiv.org/html/2609.38510#bib.bib26)\)\(Code\); and MT\-Bench\([Zheng et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib27)\)\(Chat\)\. Main results \(§[4\.2](https://arxiv.org/html/2609.38510#S4.SS2)\) use an approximately 800K target\-aligned corpus of Nemotron and CodeAlpaca prompts\([Nathawani et al\., 2025](https://arxiv.org/html/2609.38510#bib.bib13);[Chaudhary, 2023](https://arxiv.org/html/2609.38510#bib.bib14)\); training\-design ablations \(§[4\.3](https://arxiv.org/html/2609.38510#S4.SS3)\) use a separately prepared 100K corpus\. Appendix[B\.1](https://arxiv.org/html/2609.38510#A2.SS1)reports their sources and exact statistics\.

Baselines and Drafter Configurations\.We compare against DFlash\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\), a single\-pass block\-diffusion drafter; Domino\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16)\), which adds lightweight causal correction; and DSpark\([Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17)\), which adds a sequential module and confidence\-scheduled verification\. AtW=16W=16, all methods are trained on the same corpus for the same number of epochs, use a five\-layer drafter conditioned on the same target layers, and share the evaluation protocol\. DFlash and DEdit have identical architectures; Domino adds a lightweight GRU corrector, and DSpark adds a low\-rank Markov head and a confidence head, which it uses to verify only a confidence\-selected prefix of its draft\. Following their original configurations, Domino and DSpark draftWWcandidates per window rather thanW−1W\-1\. Main results compare DEdit atW=32W=32with two editing passes \(K=3K=3, P3\) against baselines at their nativeW=16W=16\. Section[4\.3](https://arxiv.org/html/2609.38510#S4.SS3)reports window\-matched comparisons, and Appendix[C\.1](https://arxiv.org/html/2609.38510#A3.SS1)the trade\-off across editing passes\.

Evaluation Protocol\.We evaluate greedy \(Tp=0T\_\{p\}=0\) and stochastic \(Tp=1T\_\{p\}=1\) decoding with a maximum of 2,048 generated tokens, following the verification protocol in Section[2](https://arxiv.org/html/2609.38510#S2)\. Under greedy decoding, verification accepts candidates matching the target’s top\-1 predictions; under stochastic decoding, we follow the DFlash implementation\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\), in which the drafter proposes argmax tokens and verification accepts candidates matching tokens sampled from the target\. Thinking mode is disabled for both target models, both when generating training responses and during evaluation\. We report the mean token acceptance per roundτ\\tauand the end\-to\-end speedup over autoregressive decoding, measured on a single NVIDIA H100 GPU at batch size 1\. Drafting uses CUDA Graphs and target verification runs in eager mode for all methods, including the autoregressive baseline; Appendix[C\.2](https://arxiv.org/html/2609.38510#A3.SS2)details this execution setting and its effect on measured latency\. Full training configurations are provided in Appendix[B\.2](https://arxiv.org/html/2609.38510#A2.SS2)\.

### 4\.2Main Results

Table[1](https://arxiv.org/html/2609.38510#S4.T1)reportsτ\\tauand end\-to\-end speedup for both target models under greedy and stochastic decoding\. We draw three conclusions\.

Table 1:End\-to\-end speedup and mean token acceptance \(τ\\tau\) under the evaluation protocol in Section[4\.1](https://arxiv.org/html/2609.38510#S4.SS1); Avg\. is the seven\-suite macro average\. Baselines operate atW=16W=16\(Domino and DSpark draft 16 candidates; DFlash drafts 15\), while DEdit uses theW=32W=32P3 model after a 3K\-step continuation fromW=16W=16\. Bold marks the best value in each model/temperature block\.DEdit achieves the highest macro\-averageτ\\tauin all four target and decoding settings\. Under greedy decoding, it reachesτ=7\.58\\tau=7\.58on Qwen3\-4B and7\.757\.75on Qwen3\-8B,9\.7%9\.7\\%and7\.6%7\.6\\%higher than DSpark, the baseline with the highestτ\\tau; under stochastic decoding, the advantage narrows to6\.1%6\.1\\%and4\.5%4\.5\\%\. This gain comes from editing rather than from a stronger initial proposal: DEdit shares DFlash’s architecture, and its P1 acceptance is on par with DFlash’s \(Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(a\)\)\.

The acceptance gain varies across benchmarks\. On MATH\-500, DEdit improvesτ\\tauover the strongest baseline by1212–18%18\\%across the four settings, and on GSM8K and MT\-Bench by55–13%13\\%\. On HumanEval and MBPP, DEdit and DSpark are within4%4\\%of each other, and DSpark leads in four of the eight model–decoding combinations\. The benefit of editing is therefore task\-dependent; Section[5\.2](https://arxiv.org/html/2609.38510#S5.SS2)examines how it varies with the future context available in a proposal\.

In our execution setting \(Section[4\.1](https://arxiv.org/html/2609.38510#S4.SS1)\), the acceptance advantage also yields the highest macro\-average speedup, but by a smaller margin\. DEdit reaches5\.72×5\.72\\timesand5\.97×5\.97\\timesunder greedy decoding and4\.88×4\.88\\timesand4\.90×4\.90\\timesunder stochastic decoding,7\.07\.0–8\.9%8\.9\\%faster than Domino, the fastest baseline\. On Qwen3\-4B under greedy decoding, DEdit’sτ\\tauis14\.8%14\.8\\%higher than Domino’s, whereas its speedup is7\.9%7\.9\\%higher, because each editing pass adds a drafter forward\. Editing thus trades per\-round drafting cost for longer accepted prefixes, and whether this trade pays off depends on how much drafting cost the execution backend exposes \(Appendix[C\.2](https://arxiv.org/html/2609.38510#A3.SS2)\)\.

### 4\.3Ablation Studies

DEdit combines three design choices: iterative editing \(Section[3\.2](https://arxiv.org/html/2609.38510#S3.SS2)\), joint training withProposalMix\(Section[3\.3](https://arxiv.org/html/2609.38510#S3.SS3)\), and window expansion \(Section[3\.4](https://arxiv.org/html/2609.38510#S3.SS4)\)\. We ablate them on Qwen3\-4B under greedy decoding and report macro\-averageτ\\tauover the seven benchmarks in Table[2](https://arxiv.org/html/2609.38510#S4.T2)\. Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(a\) varies the window width and the number of editing passes to test whether wider windows help and whether this depends on editing; Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\) removesProposalMix, joint training, or both to measure how each contributes to editing quality\.

Table 2:Window scaling and training ablations on Qwen3\-4B\. Both panels report macro\-averageτ\\tauacross seven benchmarks\.\(a\)Proposal width and editing passes atTp=0T\_\{p\}=0; Domino draftsWWcandidates and the other methodsW−1W\-1\.\(b\)Training components in a controlled 100K setting: a frozen DFlash drafter supplies P1, and a separately trained editor is evaluated at P2 withW=16W=16;Δ​τ\\Delta\\tauis relative to Full\.\(a\)Window scaling

\(b\)Training ablations

Wider windows raise acceptance mainly through editing\. Without editing, extendingW=16W=16toW=32W=32barely changes DEdit’sτ\\tau\(5\.82→\\to5\.93\), and DFlash and Domino gain only2\.1%2\.1\\%and5\.0%5\.0\\%\. With two editing passes, the same extension raisesτ\\tauby10\.3%10\.3\\%\(6\.87→\\to7\.58\) and, in a separate timing run, speedup from5\.55×5\.55\\timesto5\.80×5\.80\\times\. Each editing pass also contributes more at wider windows: the second editing pass \(P3\) adds 0\.24 atW=16W=16but 0\.46 atW=32W=32, although it is never supervised during training\. The additional positions of a wider window are therefore more valuable as context for editing than as extra candidates for verification\.

ProposalMixand joint training both improve editing, withProposalMixcontributing more\. To isolate the training objective, Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\) holds the proposal fixed: a frozen DFlash drafter supplies P1 for all variants, and a separate editor is trained for each\. With joint training, this editor is trained with bothℒunmask\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}andℒedit\\mathcal\{L\}\_\{\\mathrm\{edit\}\}; without it, the editor is trained withℒedit\\mathcal\{L\}\_\{\\mathrm\{edit\}\}only, as in two\-stage training \(Section[3\.3](https://arxiv.org/html/2609.38510#S3.SS3)\)\. WithoutProposalMix, the editor receives unmodified DFlash proposals\. RemovingProposalMixlowersτ\\tauby 0\.144 and removing joint training by 0\.085, while removing both lowers it by 0\.257\. Joint training helps even though the editor never generates P1 at inference, suggesting that learning to draft from masked inputs also improves editing\.

## 5Analysis

Section[4](https://arxiv.org/html/2609.38510#S4)shows that DEdit’s acceptance gains come from editing, and that they depend on the training objective and grow with window width\. These aggregate results do not show how the editor achieves them, which raises two questions\. First,ProposalMixraisesτ\\tau\(Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\)\), but what behavior does it actually teach the editor \(Section[5\.1](https://arxiv.org/html/2609.38510#S5.SS1)\)? Second, wider windows help mainly through editing \(Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(a\)\), suggesting that the editor uses later proposal tokens as context; does it, and when does this help most \(Section[5\.2](https://arxiv.org/html/2609.38510#S5.SS2)\)?

Metric\.We reportAA, the number of accepted draft tokens per round, excluding the additional target token\. To compare drafters with different window widths, all methods in each analysis are scored on the same number of leading positions\.

### 5\.1What DoesProposalMixTeach the Editor?

\(a\) Effect ofProposalMix

\(b\) Future\-context benefit

Figure 3:Editing outcomes and future context on Qwen3\-4B\.\(a\)Edit\-gain distributionAP2−AP1A\_\{\\mathrm\{P2\}\}\-A\_\{\\mathrm\{P1\}\}for the Full and w/oProposalMixeditors of Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\)\. Percentages use all rounds; zero\-gain rounds are omitted \(60\.0% withoutProposalMix, 66\.5% with it\)\.\(b\)W=16W=16DEdit P2\-minus\-Domino gain versus correct future\-token count\. Points and bands show response\-cluster means and 95% cluster\-bootstrap CIs; gray bars show the number of rounds per bin on a log scale, from 56k to 127\. Fully accepted rounds \(7\.7%\) are excluded\.ProposalMixis designed to teach the editor when to preserve a prediction and when to revise it \(Section[3\.3](https://arxiv.org/html/2609.38510#S3.SS3)\)\. To examine what the editor actually learns, Figure[3](https://arxiv.org/html/2609.38510#S5.F3)\(a\) compares the Full and w/oProposalMixeditors of Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\), which edit proposals from the same frozen DFlash drafter, and measures the change inAAfrom P1 to P2\.

The editor trained withProposalMixpreserves more and revises better\. It shortens the accepted prefix in half as many rounds \(9\.9% to 4\.7%\)\. The share of rounds in which it extends the prefix stays nearly the same \(30\.1% to 28\.8%\), but more of these are large repairs: rounds in which editing extends the prefix by five or more tokens rise from 5\.4% to 6\.7%\. Together, these raise the mean edit gain from 0\.687 to 0\.797\.ProposalMixtherefore improves acceptance by avoiding harmful edits rather than by editing more often\. This matters under prefix verification, where a single harmful edit can invalidate an otherwise accepted prefix\.

### 5\.2How Does Editing Use Future Context?

Prefix verification discards correct predictions after the first mismatch, which we call*correct future tokens*\(the trailing C’s in\[C,C,W,C,C\], with C correct and W wrong\)\. A bidirectional editor can attend to them, whereas a causal corrector cannot\. We ask whether editing uses these tokens, and on which tasks this helps most\.

##### Does editing use future context?

DEdit’s advantage over a causal corrector grows with the number of correct future tokens\. We apply Domino’s causal corrector\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16)\)and DEdit’s P2 editor to identicalW=16W=16proposals from Domino’s backbone; the DEdit editor comes from the main experiments and has never seen Domino’s proposals\. Figure[3](https://arxiv.org/html/2609.38510#S5.F3)\(b\) shows that the advantage is negligible at zero correct future tokens and clearly positive from two onward\. After stratifying by accepted\-prefix length, each additional correct future token is associated with a 0\.140\-token increase in the advantage \(95% CI\[0\.130,0\.150\]\[0\.130,0\.150\]\)\. Since both correct the same proposals, this growth points to future context as the source of the advantage\.

Hiding future tokens from the same editor reduces acceptance\. Keeping DEdit’s weights and P1 proposals fixed, we replace bidirectional attention within the proposal by a lower\-triangular mask\. On General, Table[3](https://arxiv.org/html/2609.38510#S5.T3)shows that bidirectional attention adds 0\.222 accepted tokens at P2 and 0\.290 at P3\. It both repairs the first rejected token more often and corrupts the accepted prefix less often \(Appendix[A\.4](https://arxiv.org/html/2609.38510#A1.SS4)\), so future context supports repair and preservation alike\. Domino, a trained causal corrector, provides a reference: at P2, causal DEdit falls below Domino on General, Reasoning, and Summary, whereas bidirectional DEdit exceeds it on every slice\. Because the editor was trained with bidirectional attention, part of this gap may reflect the change in attention pattern at inference\.

Table 3:Fixed\-weight visibility intervention\. Values areAAover the first 16 candidate positions, averaged over rounds within each slice; General averages the seven suite\-level means\. Causal and bidirectional DEdit share aW=32W=32checkpoint and P1 proposal\. Domino \(W=16W=16\) takes the first token of this proposal and drafts the remaining 15 with its own backbone and corrector\.Δ\\Delta: bidirectional minus causal; underlined: below Domino\. Details in Appendix[A\.3](https://arxiv.org/html/2609.38510#A1.SS3)\.
##### When does future context help most?

The benefit of future context is consistent across tasks and largest when outputs are highly predictable\. Besides General, we evaluate Reasoning and Summary slices of a visible solution and its summary, and a synthetic Copy\-heavy task that copies an input with one specified change \(Appendix[A\.2](https://arxiv.org/html/2609.38510#A1.SS2)\)\. Bidirectional attention improves acceptance on every slice at both P2 and P3\. On Reasoning and Summary, it adds about 0\.20 tokens at P2 and 0\.25 at P3, comparable to General\. On Copy\-heavy, it adds 0\.724 and 0\.897, about three times as much\. When most of the output is determined by the input, proposals likely contain more correct future tokens for editing to use\.

\\FloatBarrier

## 6Related Work

##### Drafters for speculative decoding\.

Speculative decoding uses a lightweight drafter to propose tokens that the target model verifies in parallel\([Stern et al\., 2018](https://arxiv.org/html/2609.38510#bib.bib3);[Leviathan et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib1);[Chen et al\., 2023](https://arxiv.org/html/2609.38510#bib.bib2)\)\. Drafters include multi\-head predictors\([Cai et al\., 2024](https://arxiv.org/html/2609.38510#bib.bib4);[Ankner et al\., 2024](https://arxiv.org/html/2609.38510#bib.bib34)\), feature\-conditioned autoregressive drafters such as EAGLE\([Li et al\., 2024b](https://arxiv.org/html/2609.38510#bib.bib32);[Li et al\., 2024a](https://arxiv.org/html/2609.38510#bib.bib33);[Li et al\., 2025b](https://arxiv.org/html/2609.38510#bib.bib5)\), and diffusion drafters that denoise all positions in parallel\([Christopher et al\., 2025](https://arxiv.org/html/2609.38510#bib.bib6);[Sandler et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib8);[Li et al\., 2025a](https://arxiv.org/html/2609.38510#bib.bib7);[Cheng et al\., 2025](https://arxiv.org/html/2609.38510#bib.bib9);[Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10);[Zhang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib15)\)\. Diffusion drafters are fast because they predict a whole block in one forward pass, but each position is predicted without conditioning on the others\. Multi\-token prediction\([Gloeckle et al\., 2024](https://arxiv.org/html/2609.38510#bib.bib40);[DeepSeek\-AI, 2024](https://arxiv.org/html/2609.38510#bib.bib41)\)instead trains the target itself to predict several future tokens, and its prediction heads can be reused as drafters\. These methods improve how the initial proposal is generated; DEdit instead edits the realized proposal before verification\.

##### Dependency recovery for parallel drafts\.

Predicting draft positions independently causes acceptance to degrade at later positions\. Recent methods restore dependencies with lightweight causal correction\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16);[Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17);[Zheng and Li, 2026](https://arxiv.org/html/2609.38510#bib.bib36);[Wang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib18)\), extend such correction to draft trees\([Li et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib38)\), or select a coherent path among candidates\([Rusanovsky et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib37)\)\. In causal correction, each position conditions only on earlier tokens, so correct predictions later in the draft cannot inform earlier ones\. DEdit instead revises each position using proposal tokens on both sides\.

##### Iterative refinement and training for editing\.

Iterative refinement is well established in non\-autoregressive generation, from Mask\-Predict\([Ghazvininejad et al\., 2019](https://arxiv.org/html/2609.38510#bib.bib35)\)to diffusion language models that revise generated tokens\([Bie et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib28);[Chen et al\., 2026c](https://arxiv.org/html/2609.38510#bib.bib29)\), and Speculative Correction\([Chen et al\., 2026a](https://arxiv.org/html/2609.38510#bib.bib39)\)applies bidirectional refinement to complete drafts\. Using refinement for speculative drafting further requires training, since edits must extend the accepted prefix rather than merely improve token\-level accuracy\. PARD\-2\([An et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib30)\)and VAT\([Gu et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib31)\)align drafter training with target acceptance\. DEdit brings iterative refinement to speculative drafting, andProposalMixtrains the editor on partially correct proposals to preserve reliable predictions under prefix verification\.

## 7Conclusion

We propose DEdit, a diffusion\-based drafter that generates proposal tokens and iteratively edits them before the target model verifies them\. To train this drafter, we designProposalMix, a training scheme that teaches the model to repair errors in its proposals while keeping the correct ones\. Across seven benchmarks under greedy and stochastic decoding, DEdit achieves the highest macro\-average token acceptance among the evaluated drafters and, in our execution setting, the highest macro\-average speedup\. Our analyses show that editing uses correct future tokens as context, so the additional positions of a wider window are more valuable for editing than as extra candidates\.

### AI Use Statement

In this work, generative AI tools assisted with manuscript drafting and language editing and with implementing parts of the experimental code\. The training corpora include target\-model\-generated responses, as described in Appendix[B\.1](https://arxiv.org/html/2609.38510#A2.SS1)\. The authors take full responsibility for the final manuscript, experimental code, and reported results\.

#### Reproducibility Statement

All model architectures, training hyperparameters, loss formulations, and evaluation protocols are detailed in Section[4](https://arxiv.org/html/2609.38510#S4)and Appendix[B](https://arxiv.org/html/2609.38510#A2)\. The evaluation datasets used across all seven benchmarks are publicly accessible\. Code, checkpoints, and evaluation scripts will be made publicly available to facilitate reproducibility\.

## References

- Anet al\.\(2026\)Z\. An, T\. Liu, Z\. Liu, D\. Li, R\. Liu, and E\. BarsoumPARD\-2: target\-aligned parallel draft model for dual\-mode speculative decoding\.Note:arXiv preprint arXiv:2605\.08632External Links:2605\.08632,[Link](https://arxiv.org/abs/2605.08632)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Ankneret al\.\(2024\)Z\. Ankner, R\. Parthasarathy, A\. Nrusimha, C\. Rinard, J\. Ragan\-Kelley, and W\. BrandonHydra: sequentially\-dependent draft heads for Medusa decoding\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=FbhjirzvJG)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Austinet al\.\(2021\)J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. SuttonProgram synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Bieet al\.\(2026\)T\. Bie, M\. Cao, X\. Cao, B\. Chen, F\. Chen, K\. Chen, L\. Du, D\. Feng, H\. Feng, M\. Gong, Z\. Gong, Y\. Gu, J\. Guan, K\. Guan, H\. He, Z\. Huang, J\. Jiang, Z\. Jiang, Z\. Lan, C\. Li, J\. Li, Z\. Li, H\. Liu, L\. Liu, G\. Lu, Y\. Lu, Y\. Ma, X\. Mou, Z\. Pan, K\. Qiu, Y\. Ren, J\. Tan, Y\. Tian, Z\. Wang, L\. Wei, T\. Wu, Y\. Xing, W\. Ye, L\. Zha, T\. Zhang, X\. Zhang, J\. Zhao, D\. Zheng, H\. Zhong, W\. Zhong, J\. Zhou, J\. Zhou, L\. Zhu, M\. Zhu, and Y\. ZhuangLLaDA2\.1: speeding up text diffusion via token editing\.Note:arXiv preprint arXiv:2602\.08676External Links:2602\.08676,[Link](https://arxiv.org/abs/2602.08676)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Caiet al\.\(2024\)T\. Cai, Y\. Li, Z\. Geng, H\. Peng, J\. D\. Lee, D\. Chen, and T\. DaoMedusa: simple LLM inference acceleration framework with multiple decoding heads\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 5209–5235\.External Links:[Link](https://proceedings.mlr.press/v235/cai24b.html)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Chaudhary \(2023\)S\. ChaudharyCode Alpaca: an instruction\-following LLaMA model for code generation\.Note:GitHub repositoryExternal Links:[Link](https://github.com/sahil280114/codealpaca)Cited by:[§B\.1](https://arxiv.org/html/2609.38510#A2.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Chenet al\.\(2026a\)B\. K\. Chen, C\. Wu, and K\. KawaguchiSpeculative correction: draft\-then\-refine decoding for diffusion language models\.Note:arXiv preprint arXiv:2608\.02625External Links:2608\.02625,[Link](https://arxiv.org/abs/2608.02625)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Chenet al\.\(2023\)C\. Chen, S\. Borgeaud, G\. Irving, J\. Lespiau, L\. Sifre, and J\. JumperAccelerating large language model decoding with speculative sampling\.arXiv preprint arXiv:2302\.01318\.Cited by:[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38510#S2.SS1.p1.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2026b\)J\. Chen, Y\. Liang, and Z\. LiuDFlash: block diffusion for flash speculative decoding\.InProceedings of the 43rd International Conference on Machine Learning,Note:arXiv:2602\.06036Cited by:[§B\.2](https://arxiv.org/html/2609.38510#A2.SS2.p1.1),[§C\.2](https://arxiv.org/html/2609.38510#A3.SS2.p2.1),[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§1](https://arxiv.org/html/2609.38510#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.38510#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p3.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. d\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Chenet al\.\(2026c\)Z\. Chen, G\. Fang, X\. Ma, R\. Yu, and X\. WangDMax: aggressive parallel decoding for dLLMs\.Note:arXiv preprint arXiv:2604\.08302External Links:2604\.08302,[Link](https://arxiv.org/abs/2604.08302)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Chenget al\.\(2026\)X\. Cheng, X\. Yu, C\. Shao, J\. Li, Y\. Xiong,et al\.DSpark: confidence\-scheduled speculative decoding with semi\-autoregressive generation\.arXiv preprint arXiv:2607\.05147\.Cited by:[§B\.2](https://arxiv.org/html/2609.38510#A2.SS2.p2.1),[§C\.2](https://arxiv.org/html/2609.38510#A3.SS2.p2.1),[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§1](https://arxiv.org/html/2609.38510#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p2.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Chenget al\.\(2025\)Z\. Cheng, G\. Yang, J\. Li, Z\. Deng, M\. Guo, and S\. HuDEER: draft with diffusion, verify with autoregressive models\.arXiv preprint arXiv:2512\.15176\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Christopheret al\.\(2025\)J\. K\. Christopher, B\. R\. Bartoldson, T\. Ben\-Nun, M\. Cardei, B\. Kailkhura, and F\. FiorettoSpeculative diffusion decoding: accelerating language generation through diffusion\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 12042–12059\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- DeepSeek\-AI \(2024\)DeepSeek\-AIDeepSeek\-V3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Ghazvininejadet al\.\(2019\)M\. Ghazvininejad, O\. Levy, Y\. Liu, and L\. ZettlemoyerMask\-Predict: parallel decoding of conditional masked language models\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 6112–6121\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1633),[Link](https://aclanthology.org/D19-1633/)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Gloeckleet al\.\(2024\)F\. Gloeckle, B\. Youbi Idrissi, B\. Rozière, D\. Lopez\-Paz, and G\. SynnaeveBetter & faster large language models via multi\-token prediction\.InProceedings of the 41st International Conference on Machine Learning,Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Guet al\.\(2026\)G\. Gu, B\. Heo, H\. Jun, Y\. Kang, S\. Lee, S\. Yun, and D\. HanVerification\-aware training for speculative decoding\.Note:arXiv preprint arXiv:2608\.30135External Links:2608\.30135,[Link](https://arxiv.org/abs/2608.30135)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px3.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the MATH dataset\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Vol\.1\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Huet al\.\(2026\)L\. Hu, Z\. Feng, Y\. Wu, H\. Yuan, Y\. Zhao, Y\. Qian, B\. Wang, P\. Zhao, D\. Jiang, Y\. Zhu, T\. Rosing, and H\. ZhangJetSpec: breaking the scaling ceiling of speculative decoding with parallel tree drafting\.arXiv preprint arXiv:2606\.18394\.External Links:[Link](https://arxiv.org/abs/2606.18394)Cited by:[§3\.4](https://arxiv.org/html/2609.38510#S3.SS4.p1.1)\.
- Huanget al\.\(2026\)J\. Huang, Y\. Zhang, Q\. Zhang, H\. Lin, H\. Xu, and L\. ZhangDomino: decoupling causal modeling from autoregressive drafting in speculative decoding\.arXiv preprint arXiv:2605\.29707\.Cited by:[§B\.2](https://arxiv.org/html/2609.38510#A2.SS2.p2.1),[§C\.2](https://arxiv.org/html/2609.38510#A3.SS2.p2.1),[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§1](https://arxiv.org/html/2609.38510#S1.p2.1),[§1](https://arxiv.org/html/2609.38510#S1.p5.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p2.1),[§5\.2](https://arxiv.org/html/2609.38510#S5.SS2.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Inco AI \(2026\)Inco AIDFlash 2: Keep Drafting Parallel\.External Links:[Link](https://inco.ai/blog/dflash2/)Cited by:[§3\.4](https://arxiv.org/html/2609.38510#S3.SS4.p1.1)\.
- Jainet al\.\(2024\)N\. Jain, K\. Han, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. StoicaLiveCodeBench: holistic and contamination free evaluation of large language models for code\.arXiv preprint arXiv:2403\.07974\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Leviathanet al\.\(2023\)Y\. Leviathan, M\. Kalman, and Y\. MatiasFast inference from transformers via speculative decoding\.InProceedings of the 40th International Conference on Machine Learning,pp\. 19274–19286\.Cited by:[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38510#S2.SS1.p1.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2025a\)G\. Li, Z\. Fu, M\. Fang, Q\. Zhao, M\. Tang, C\. Yuan, and J\. WangDiffuSpec: unlocking diffusion language models for speculative decoding\.arXiv preprint arXiv:2510\.02358\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2026\)T\. Li, Y\. Luo, X\. Shang, and Z\. ShenDARTree: speculative diffusion decoding with autoregressive draft trees\.Note:arXiv preprint arXiv:2608\.13524External Links:2608\.13524,[Link](https://arxiv.org/abs/2608.13524)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2024a\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE\-2: faster inference of language models with dynamic draft trees\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2024b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE: speculative sampling requires rethinking feature uncertainty\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 28935–28948\.External Links:[Link](https://proceedings.mlr.press/v235/li24bt.html)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2025b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE\-3: scaling up inference acceleration of large language models via training\-time test\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.52202/085713-4562),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/c7b5a35ea98b62512a869c19ea7b03cb-Abstract-Conference.html)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Lightmanet al\.\(2023\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s verify step by step\.arXiv preprint arXiv:2305\.20050\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Mathematical Association of America \(2025\)Mathematical Association of AmericaAmerican invitational mathematics examination\.External Links:[Link](https://maa.org/maa-invitational-competitions/)Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Nathawaniet al\.\(2025\)D\. Nathawani, S\. Ding, V\. Lavrukhin, I\. Gitman, S\. Majumdar, E\. Bakhturina, B\. Ginsburg, and J\. Polak ScowcroftNemotron\-Post\-Training\-Dataset\-v2\.Note:Hugging Face datasetExternal Links:[Link](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)Cited by:[§B\.1](https://arxiv.org/html/2609.38510#A2.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Rusanovskyet al\.\(2026\)M\. Rusanovsky, Y\. Miron, R\. Uziel, O\. Belhasin, H\. Guo, R\. Zilberstein, M\. Ashkenazi, and M\. EladLiLiCorr: lightweight likelihood correlation of parallel drafts for speculative decoding\.Note:arXiv preprint arXiv:2608\.20530External Links:2608\.20530,[Link](https://arxiv.org/abs/2608.20530)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Sandleret al\.\(2026\)J\. Sandler, J\. K\. Christopher, T\. Hartvigsen, and F\. FiorettoSpecDiff\-2: scaling diffusion drafter alignment for faster speculative decoding\.InProceedings of Machine Learning and Systems,Vol\.8\.External Links:[Link](https://proceedings.mlsys.org/paper_files/paper/2026/hash/041dad5ed2191b44ba3ed0e00cdc3187-Abstract-Conference.html)Cited by:[§3\.4](https://arxiv.org/html/2609.38510#S3.SS4.p1.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Sternet al\.\(2018\)M\. Stern, N\. Shazeer, and J\. UszkoreitBlockwise parallel decoding for deep autoregressive models\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2026\)Z\. Wang, D\. Wertheimer, Y\. C\. F\. Lim, M\. Srivatsa, R\. K\. Ganti, M\. Zhang, and N\. WangxPress: parallel refinement for diffusion drafters in speculative decoding\.arXiv preprint arXiv:2608\.02438\.Cited by:[§1](https://arxiv.org/html/2609.38510#S1.p1.1),[§1](https://arxiv.org/html/2609.38510#S1.p2.1),[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.
- Zhanget al\.\(2026\)J\. Zhang, Z\. Yu, S\. Liu, E\. J\. Yu, Z\. Li, D\. Zhu, J\. Duo, W\. Xiong, Y\. Song, G\. Yu, J\. Zhu, and S\. LiDFlare: scaling up draft capacity for block diffusion speculative decoding\.arXiv preprint arXiv:2606\.02091\.Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px1.p1.1)\.
- Zheng and Li \(2026\)H\. Zheng and P\. LiDeLS\-Spec: decoupled long\-short contexts for parallel speculative drafting\.Note:arXiv preprint arXiv:2607\.07409External Links:2607\.07409,[Link](https://arxiv.org/abs/2607.07409)Cited by:[§6](https://arxiv.org/html/2609.38510#S6.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 46595–46623\.Cited by:[§4\.1](https://arxiv.org/html/2609.38510#S4.SS1.p1.1)\.

## Appendix AMechanism and Task Analysis

### A\.1Qualitative Editing Trace

The trace in Figure[4](https://arxiv.org/html/2609.38510#A1.F4)uses a fixed\-width proposal, so “insertion” and “deletion” patterns refer to repairing token shifts under sequence alignment; the tensor length stays fixed\.

Figure 4:A fullW=32W=32editing trace\.On an MBPP prompt, successive passes extend the accepted prefix from 8 to 18 to 30 of 31 proposal tokens\. Blue cells already match the target but remain blocked by an earlier error\.Figure[4](https://arxiv.org/html/2609.38510#A1.F4)shows the complete first verification round for theW=32W=32DEdit checkpoint used in the main results\. The P1 proposal first fails at slot 9 even though slots 10–14 already match the target\. P2 repairs the prefix through slot 18 and leaves another correct run at slots 20–27 after its first rejection\. P3 repairs the prefix through slot 30 but proposesurrencerather than the target token\_occat slot 31\. The verifier therefore accepts 30 proposal tokens and emits\_occas the next\-round anchor\. The reported counts are contiguous accepted\-prefix lengths, not token\-level accuracies\.

##### Observational audit details\.

TheW=16W=16comparison in Section[5\.2](https://arxiv.org/html/2609.38510#S5.SS2)contains 176,823 decoding rounds from 2,830 response clusters\. DEdit outperforms Domino on 17\.91% of rounds, matches on 69\.03%, and loses on 13\.06%\. For example,\[C,C,W,C,C\]has two correct future tokens after the first rejection\. At a future\-token count of 14, the observed gain reaches\+2\.654\+2\.654, although such high\-count rounds are rare\.

\\FloatBarrier

### A\.2Controlled Task Definitions

We use three reasoning–summary constructions and one copy\-heavy task to control how much useful future information is already visible\. These are mechanism diagnostics, not new task benchmarks\. All reasoning prompts are derived from frozen GSM8K examples, use visible model output withenable\_thinking=false, and are selected without using drafter performance\. The single\-response construction permits separate Reasoning and Summary measurements; the two\-turn and frozen\-prefix forms isolate the same summary\-continuation effect under alternative boundaries\.

##### Shared reasoning example\.

The three reasoning–summary formats below use this representative question:

> Amber, Micah, and Ahito ran 52 miles in total\. Amber ran 8 miles\. Micah ran 3\.5 times what Amber ran\. How many miles did Ahito run?

##### Two\-turn visible reasoning to summary\.

The first turn requests an explicit solution:

> Write a detailed, checkable solution\. Show the essential equations and intermediate calculations\. End with a separate line in exactly this format:Reasoning conclusion: <answer\>\.

The complete target response is then retained in the conversation, and the second turn asks:

> Using the solution above, write a concise self\-contained solution in 3 to 5 sentences\. Keep only the essential calculation steps\. End with a separate line in exactly this format:Final answer: <answer\>\. Do not mention the previous response, a reasoning trace, or these instructions\.

Thus the Summary continuation can condition on the exact visible Reasoning generated in the preceding turn\.

##### Single\-response reasoning and summary\.

This construction asks for both stages in one assistant response:

> Solve the problem and return exactly two plain\-text sections in this order\. Do not write any text before the first section label\. Reasoning: Write a detailed, checkable solution with the essential equations and intermediate calculations\. Summary: Write a concise self\-contained solution in 3 to 5 sentences\. Keep only the essential calculation steps and end with a separate line in exactly this format:Final answer: <answer\>\.

A representative frozen target response is:

> Reasoning: The total distance is 52 miles\. Amber ran 8 miles, and Micah ran3\.5×8=283\.5\\times 8=28miles\. Together they ran8\+28=368\+28=36miles, so Ahito ran52−36=1652\-36=16miles\. Summary: Amber ran 8 miles, and Micah ran3\.5×8=283\.5\\times 8=28miles\. Together they ran 36 miles\. Ahito therefore ran52−36=1652\-36=16miles\. Final answer: 16

The stage boundary is the literalSummary:delimiter, so proposal statistics can be accumulated separately before and after it\.

##### Frozen\-reasoning summary continuation\.

To isolate Summary generation from variation in the preceding Reasoning, we freeze the target\-produced prefix and begin evaluation after the delimiter:

> Reasoning: The total distance run by Amber, Micah, and Ahito is 52 miles\. Amber ran 8 miles\. Micah ran3\.5×8=283\.5\\times 8=28miles\. Amber and Micah therefore ran8\+28=368\+28=36miles, so Ahito ran52−36=1652\-36=16miles\. Summary:

The expected continuation is:

> Amber ran 8 miles, and Micah ran3\.5×8=283\.5\\times 8=28miles\. Together, Amber and Micah ran 36 miles\. Ahito ran52−36=1652\-36=16miles\. Final answer: 16

Only the Summary suffix is drafted and measured in this construction\.

##### Copy\-heavy continuation\.

The CodeXGLUE\-derived task exposes the exact old\-to\-new replacement and asks for the complete normalized method:

> Apply the requested bug fix to the normalized Java method\. Replace the exact token sequence<OLD\>VAR\_1</OLD\>with<NEW\>VAR\_3</NEW\>\. Change nothing else\. Return the complete corrected method as a single line with one space between tokens\. Return only the method, with no Markdown, code fences, labels, or explanation\. private boolean METHOD\_1 \( java\.lang\.String VAR\_1 \) \{ boolean VAR\_2 = false ; try \{ java\.lang\.Boolean \. METHOD\_2 \( VAR\_3 \) ; VAR\_2 = true ; \} catch \( TYPE\_1 error \) \{ VAR\_4 \. METHOD\_3 \( STRING\_1 \) ; \} return VAR\_2 ; \}

The expected response changes only the supplied token:

> private boolean METHOD\_1 \( java\.lang\.String VAR\_3 \) \{ boolean VAR\_2 = false ; try \{ java\.lang\.Boolean \. METHOD\_2 \( VAR\_3 \) ; VAR\_2 = true ; \} catch \( TYPE\_1 error \) \{ VAR\_4 \. METHOD\_3 \( STRING\_1 \) ; \} return VAR\_2 ; \}

Because the gold replacement is present in the prompt, this experiment measures lexical preservation and local correction, not standard CodeXGLUE program\-repair quality\.

\\FloatBarrier

### A\.3Task\-Slice Evaluation Details

Table[3](https://arxiv.org/html/2609.38510#S5.T3)evaluates the visibility intervention on the seven\-suite general workload and the three controlled slices defined above: a visible reasoning section, its conditioned summary, and a copy\-heavy continuation\. All three configurations receive the same target context\. The two DEdit configurations use the sameW=32W=32checkpoint, P1 proposal, and attention implementation, and edit all 31 candidate positions\. Domino drafts 16 candidates atW=16W=16\(Section[4\.1](https://arxiv.org/html/2609.38510#S4.SS1)\); to align its input with DEdit, we fix its first candidate to the first token of the DEdit P1 proposal, and its backbone and corrector generate the remaining 15\. Domino is a reference rather than part of the fixed\-weight intervention\.

For all configurations, we score the accepted prefix, after P2 and P3 for DEdit, over the same first 16 candidate positions and exclude rounds in which fewer than 16 tokens of the target continuation remain\. Reasoning and Summary use a fixed 1,000\-prompt GSM8K\-derived set; 943 responses satisfy the required section markers and contribute 26,572 Reasoning and 8,919 Summary rounds\. Copy\-heavy uses a fixed 1,000\-prompt set and contributes 5,874 rounds\. General is the macro average over seven suite\-level round means\.

### A\.4Fixed\-Weight Future\-Visibility Counterfactual

##### Intervention\.

We directly remove future proposal visibility from the Qwen3\-4BW=32W=32DEdit editor used in the main results while holding its checkpoint, target hidden states, P1 proposal, verifier, and target\-reference trajectory fixed\. Both configurations pass an explicit attention mask over the completeW=32W=32proposal\. The bidirectional configuration exposes all proposal positions, while the causal configuration replaces only intra\-proposal visibility with a lower\-triangular mask; all target\-prefix hidden states remain visible\. At P2, both configurations therefore receive exactly the same P1 input\. P3 iteratively edits the proposal produced by its own P2 configuration\. Following the task\-slice protocol, reported acceptance statistics evaluate the first 16 proposal tokens and exclude rounds in which fewer than 16 tokens of the target continuation remain\.

Figure 5:Fixed\-weight intervention on future proposal visibility\. Lines report response\-cluster mean bidirectional\-minus\-causal accepted tokens with 95% response\-bootstrap intervals; gray bars report each bin’s share of 181,205 eligible rounds\. P2 compares both masks on the same P1 proposal, while P3 applies the next edit to each configuration\. The hatched bar denotes the 10,382 rounds in which all 16 shared proposal tokens are accepted\.
##### Effect by available future information\.

Averaged over response clusters rather than rounds as in Table[3](https://arxiv.org/html/2609.38510#S5.T3), bidirectional visibility adds 0\.268 accepted tokens at P2, with a 95% interval of\[0\.253,0\.283\]\[0\.253,0\.283\], and 0\.361 at P3, with an interval of\[0\.343,0\.380\]\[0\.343,0\.380\]\. The effect is positive even when no correct token follows the first rejection:\+0\.041\+0\.041at P2 and\+0\.038\+0\.038at P3\. It then generally increases with available correct future tokens, reaching\+0\.750\+0\.750and\+1\.016\+1\.016at a count of six\.

Task\-level results after one and two editing passes are reported in Table[3](https://arxiv.org/html/2609.38510#S5.T3)\.

##### Repair and preservation\.

On General, bidirectional and causal visibility repair the first rejected token in 27\.85% and 27\.11% of P2 rounds, respectively, while corrupting the original accepted prefix in 2\.82% and 4\.75%\. At P3, the repair rates are 29\.68% and 27\.10%, and the prefix\-corruption rates are 3\.14% and 4\.69%\. Thus the editor’s bidirectional advantage is visible in both more frequent repair and better preservation, and its gains are largest when more correct future tokens follow the first rejection \(Figure[5](https://arxiv.org/html/2609.38510#A1.F5)\)\.

\\FloatBarrier

## Appendix BTraining Details

### B\.1Training Data

Table[4](https://arxiv.org/html/2609.38510#A2.T4)reports the source composition of the target\-aligned corpora\. Each record pairs a user prompt with a response generated by the corresponding Qwen3 target with thinking disabled\. Counts refer to released records before training\-time tokenization and sequence\-length handling; percentages use each corpus’s actual row count\.

##### Main training corpora \(800K\)\.

The main experiments use prompts from the chat, code, and math partitions of Nemotron Post\-Training Dataset v2\([Nathawani et al\., 2025](https://arxiv.org/html/2609.38510#bib.bib13)\)and CodeAlpaca\-20k\([Chaudhary, 2023](https://arxiv.org/html/2609.38510#bib.bib14)\), paired with target\-specific responses\. The Qwen3\-4B and Qwen3\-8B corpora contain 799,864 and 800,000 records, respectively\. Their source counts differ by 136 chat records; code, math, and CodeAlpaca counts are identical\.

##### Controlled training corpus \(100K\)\.

The Qwen3\-4B training\-design ablations use a 99,987\-record corpus\. Its recorded source families arenemotron,openr1\_math,opencodeinstruct, andevol\_codealpaca\. Thenemotronrecords span chat, code, math, and STEM partitions\. Its sources and proportions differ from those of the 800K corpora\.

Table 4:Training\-corpus composition: records \(percentage of each corpus\)\. Counts are computed from the releasedsourcefield and, for 100K, the jointsource/splitfields\. Nemotron rows denote v2 partitions for 800K and the mirror’snemotronlabels for 100K\. A dash denotes no records with that source label\.
##### Window\-expansion data and settings\.

TheW=24W=24andW=32W=32continuations use a fixed 200,000\-record sample from the corresponding target’s 800K corpus, selected uniformly without replacement with seed 42 and retained in source order\. This continuation sample is separate from the 100K training\-design corpus\. Each continuation runs 3,000 steps with learning rate10−410^\{\-4\}, block sizeWW, decayγ=11\\gamma=11and 256 anchors atW=24W=24, andγ=15\\gamma=15and 192 anchors atW=32W=32; other settings follow Table[5](https://arxiv.org/html/2609.38510#A2.T5)\. DFlash, Domino, and DEdit use the same continuation settings\.

### B\.2Training Implementation

Our implementation adopts the DFlash\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\)backbone and target\-feature conditioning scheme\. The target model remains frozen throughout drafter training\. Attention over the decoded context is causal, while attention across candidate positions is bidirectional\. The selected target layers and other training settings are listed in Table[5](https://arxiv.org/html/2609.38510#A2.T5)\.

Within each scale, DFlash, Domino\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16)\), DSpark\([Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17)\), and DEdit are trained on the same corpus for the same number of epochs\. Each baseline keeps its original training objective; DSpark, for example, combines CE and L1 losses with weights 0\.1 and 0\.9\. DEdit applies position\-weighted token CE to every valid position in both trained passes, in the main and the 100K settings\. Training supervises P1/P2; later passes iteratively reuse the editor at inference without additional training\.

Table 5:Training configurations for the controlled ablations and main results\. Shared entries span both columns\.
### B\.3ProposalMixInput Statistics

Table[6](https://arxiv.org/html/2609.38510#A2.T6)shows howProposalMixinputs change as the proposal pass improves\. We replay intermediate checkpoints of the Full model in Table[2](https://arxiv.org/html/2609.38510#S4.T2)\(b\) on a fixed probe set, using each checkpoint as both P1 and P2, withη=0\.5\\eta=0\.5\. Ground\-truth selection includes positions whose first\-pass prediction already matches the target; overwrite counts only positions whose token changes\.

Table 6:ProposalMixinput rates \(%;η=0\.5\\eta=0\.5\) when each checkpoint generates its own P1 on a fixed probe set\. GT selection marks positions chosen for ground\-truth conditioning; overwrite counts those whose tokens change\.
### B\.4Position\-Weighted Two\-Pass CE

Let𝒯\\mathcal\{T\}contain every valid non\-anchor token in the sampled training blocks\. Ifri∈\{1,…,W−1\}r\_\{i\}\\in\\\{1,\\ldots,W\-1\\\}is tokenii’s proposal offset, its weight is

wi=exp⁡\(−ri−1γ\),Z=∑i∈𝒯wi\.w\_\{i\}=\\exp\\\!\\left\(\-\\frac\{r\_\{i\}\-1\}\{\\gamma\}\\right\),\\qquad Z=\\sum\_\{i\\in\\mathcal\{T\}\}w\_\{i\}\.For each trained passk∈\{1,2\}k\\in\\\{1,2\\\}, the implementation computes a separately normalized cross\-entropy,

ℒCE\(k\)=−1Z∑i∈𝒯wilogqθ,i\(k\)\(yi⋆\),ℒ=ℒunmask\+ℒedit,\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(k\)\}=\-\\frac\{1\}\{Z\}\\sum\_\{i\\in\\mathcal\{T\}\}w\_\{i\}\\log q\_\{\\theta,i\}^\{\(k\)\}\(y\_\{i\}^\{\\star\}\),\\qquad\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}\+\\mathcal\{L\}\_\{\\mathrm\{edit\}\},whereℒunmask=ℒCE\(1\)\\mathcal\{L\}\_\{\\mathrm\{unmask\}\}=\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(1\)\}andℒedit=ℒCE\(2\)\\mathcal\{L\}\_\{\\mathrm\{edit\}\}=\\mathcal\{L\}\_\{\\mathrm\{CE\}\}^\{\(2\)\}\.W=16W=16,W=24W=24, andW=32W=32useγ=7,11,15\\gamma=7,11,15, respectively\. The valid\-token mask and positional weights are the same for both passes; only the P2 input differs\. In particular,ProposalMixdoes not narrow P2 supervision to uncertain positions\.

The positional decay reflects prefix verification: errors near the beginning of a proposal prevent later tokens from being accepted in the same round\. Supervising the full block still teaches later tokens that can become useful after an earlier edit\. Appendix[C\.1](https://arxiv.org/html/2609.38510#A3.SS1)reports how acceptance and speedup vary with the number of editing passes\.

\\FloatBarrier

## Appendix CInference and Systems Analysis

### C\.1Editing\-Pass Trade\-off

Figure 6:Trade\-off across editing passes for the sameW=32W=32checkpoint atTp=0T\_\{p\}=0\. Each configuration is evaluated twice with a shared autoregressive baseline\. Drafting uses CUDA Graphs and target verification remains eager\. Training supervises P1 and P2; P3 and P4 reuse the trained editor at inference\.As Figure[6](https://arxiv.org/html/2609.38510#A3.F6)shows, adding a second editing pass \(from P2 to P3\) raisesτ\\taufrom 7\.123 to 7\.577 and speedup from5\.478×5\.478\\timesto5\.724×5\.724\\times\. P4 further increasesτ\\tauto 7\.770, while speedup rises by only 0\.19% to5\.735×5\.735\\times\. Although P4 is the fastest measured configuration, we use P3 as the default because it retains 99\.8% of the observed P4 speedup with one fewer drafter forward per round\.

### C\.2Execution Setting

We evaluate single\-request decoding with CUDA Graph\-optimized drafting and eager target verification\. CUDA Graph replay reduces host submission overhead by launching a captured sequence of GPU operations through one submission\. We preallocate fixed\-shape input, position, attention, and cache buffers for each drafter\.

DFlash\([Chen et al\., 2026b](https://arxiv.org/html/2609.38510#bib.bib10)\)captures its single parallel draft forward\. Domino\([Huang et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib16)\)captures its backbone, output head, and complete GRU correction loop\. DEdit captures allKKpasses and proposal construction between passes\. DSpark\([Cheng et al\., 2026](https://arxiv.org/html/2609.38510#bib.bib17)\)captures its fixed\-shape backbone, Markov correction, and confidence kernels; its native confidence readback and dynamic width decision remain in the measured host path\. Target verification remains eager for every method, as does the autoregressive baseline\.

Draft and target GPU operations execute sequentially on the same CUDA stream\. During draft graph execution, however, the CPU can enqueue the subsequent eager target operations\. This enqueue\-ahead window can reduce gaps between target operations, partially hiding the additional drafting overhead\. Consequently, independently measured draft and target latencies need not sum to the observed round latency\. Additional passes still incur GPU computation, and their runtime benefit depends on the execution backend and workload\.

#### C\.2\.1Enqueue\-Ahead Diagnostic

Figure 7:Host enqueue\-ahead during CUDA Graph replay\.\(a\) The CPU submits eager target operations while the GPU executes the captured drafter; draft and target GPU operations remain sequential\. A forced host barrier delays target submission\. \(b\) Nsight traces decompose each round into the pre\-target draft interval, target compute, and gaps between target operations\.In Figure[7](https://arxiv.org/html/2609.38510#A3.F7), all seven profiled settings execute the same 2,294 target GPU operations\. Target compute remains nearly constant, while longer draft execution is accompanied by shorter inter\-operation gaps\. The forced\-barrier control delays target dispatch without changing the draft graph or arithmetic, supporting the role of host enqueue\-ahead\. Because Nsight instrumentation perturbs latency, these traces illustrate the mechanism rather than quantify its contribution to the main\-table speedups\.

\\FloatBarrier

## Appendix DLimitations and Future Work

##### Limitations\.

Our wall\-clock gains are measured in a single\-request setting with CUDA Graph\-optimized drafting and eager target verification\. CPU enqueue\-ahead partially hides drafting overhead in this setup, so these results do not directly establish speedups in production serving with continuous batching\. Additional editing passes still incur computation, and their exposed cost can outweigh the benefit of longer accepted prefixes\. We evaluate targets of up to 8B parameters\. Because every drafter is trained from scratch for each target under a matched budget, extending this controlled comparison to larger or newer targets is left to future work\.

##### Future Work\.

Realizing these acceptance gains in production requires joint algorithm and scheduling design\. A direct opportunity is to skip editing passes that do not help: about two\-thirds of editing passes leave the accepted prefix unchanged \(Figure[3](https://arxiv.org/html/2609.38510#S5.F3)\(a\)\), yet each costs a drafter forward\. Predicting such rounds, for example from proposal confidence or from whether predictions stop changing across passes, could reduce drafting cost with little loss in acceptance\. The same signals could inform how to adapt proposal width and verification length, and drafting and verification could be interleaved across requests\. Evaluating whether these signals support effective compute allocation under realistic serving workloads remains future work\.

Similar Articles

Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks

arXiv cs.AI

Introduces Speculative Refinement (SpecRef), a training-free hybrid decoding strategy that warm-starts a masked diffusion language model from an autoregressive draft using entropy-guided selective masking. Evaluated across six benchmarks, it reveals that code benchmarks conflate structural discovery with logical correctness, identifies a refinement tension phenomenon, and shows that evaluation protocols can produce different model rankings.

What is Speculative Decoding? (trending on paperswithco.de) [R]

Reddit r/MachineLearning

Speculative decoding is an inference optimization technique that uses a fast draft model to propose future tokens verified in parallel by a larger model, improving LLM generation speed. The article highlights its trending status on Papers with Code and a recent SGLang blog post about state-of-the-art latencies using DFlash models.

Teaching Diffusion to Speculate Left-to-Right

arXiv cs.CL

This paper proposes three training-time interventions (positional weighting, first-error focal loss, and chain loss) to align diffusion-based draft models with autoregressive verification in speculative decoding, improving accepted prefix length by 21–76% without extra inference cost.