并行生成草稿,通过深度进行条件化:用于推测解码的相邻因果注入
摘要
提出DSpine,一种通过网络深度注入相邻因果条件的并行推测解码drafter,在SGLang中实现高效并行执行,较DFlash在Qwen3-8B上平均接受长度提升27.8%、吞吐量提升23.3%。
arXiv:2609.36173v1 Announce Type: new
Abstract: Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.
查看缓存全文
缓存时间: 2026/09/30 09:46
# Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
Source: [https://arxiv.org/html/2609.36173](https://arxiv.org/html/2609.36173)
Keyu Chen2,†Haocheng Sun3Weibo Gu2Ruizhi Qiao2Xing Sun2Bo Jiang1,∗Affiliation:1Shanghai Jiao Tong University2Tencent YouTu Lab3Xiamen University
###### Abstract
Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix\. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors\. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier\. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor’s predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel\. A unified transfer space built from the target model’s output embeddings unifies layer\-wise injection with predecessor\-conditioned decoding, and layer\-wise output\-embedding supervision promotes the formation of predicted features in shallow layers\. Fused kernels and a transition cache execute both efficiently in parallel within SGLang\. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3\-4B and Qwen3\-8B\. At temperature zero on Qwen3\-8B, it raises the seven\-benchmark mean from DFlash’s 3\.77 to 4\.82 \(\+27\.8%\); in SGLang serving tests, it delivers 23\.3% higher throughput than DFlash on average\.
Figure 1:Mean acceptance length across seven math, code, and chat benchmarks\. DSpine leads all compared parallel drafters on Qwen3\-8B and Qwen3\-4B at temperatures 0 and 1; the annotations report its margin over the strongest baseline in each setting\.
## 1Introduction
Autoregressive decoding is a major source of latency in large language models\. Speculative decoding accelerates it without changing the target model: a lightweight drafter proposes multiple candidates, and the target model verifies them in a single forward pass\([Leviathan et al\., 2023](https://arxiv.org/html/2609.36173#bib.bib1)\)\. Its speedup depends on how many candidates are accepted per round and how cheaply they are drafted\. Autoregressive drafters such as EAGLE\-3 reuse target\-model features to improve predictions but still require one sequential drafting step per draft position\([Li et al\., 2025](https://arxiv.org/html/2609.36173#bib.bib2)\)\. Inspired by diffusion language modeling, DFlash instead predicts a masked block in parallel within one forward pass conditioned on target context features, so a longer draft adds no sequential drafting steps\([Chen et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib3)\)\.
This shift also changes how token dependencies are established\. Autoregressive drafting conditions each successor on its already selected predecessors, whereas positions in a parallel draft must form predictions before their predecessors commit to tokens, so individually plausible candidates can be inconsistent when combined\. Because the target accepts candidates in prefix order, a single inconsistency truncates the rest of the round’s draft, leaving much of the cheaply drafted block unaccepted\. Reducing drafting cost must therefore go hand in hand with establishing effective predecessor conditioning inside parallel computation\.
Recent work improves token consistency in parallel drafts through lightweight conditioning modules\. Inside the draft backbone, DFlash2 mixes adjacent positions through local convolutions, but these blend neighboring hidden states without separating the predecessor’s predictive information from the rest of its state\([DFlash2 Team, 2026](https://arxiv.org/html/2609.36173#bib.bib4)\)\. Domino, DSpark, and DFlash2’s path selector instead condition on the predecessor after the backbone to adjust token distributions or candidate selection\([Huang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib5);[Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6);[DFlash2 Team, 2026](https://arxiv.org/html/2609.36173#bib.bib4)\)\. Conditional decoding in these methods is thus left to a lightweight module after the backbone, which limits the flow of predecessor information to successors\. Moreover, the predecessor information these modules pass is confined to the final\-layer vocabulary space, so it can neither be combined with in\-layer features nor exploit the rich intermediate features formed across the backbone\.
Our empirical study yields two key insights\.\(1\) Earlier positions form predictions at shallower depth:layer\-wise readouts already recover substantial token\-predictive information at early positions in shallow layers \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(a\)\), so the prefix could condition successors whose computation is still in progress\.\(2\) Successor prediction depends on both the content of the adjacent predecessor and when it enters the network:even with correct earlier context, replacing only the adjacent predecessor with its draft prediction sharply lowers successor accuracy, and injecting the same correct predecessor features earlier clearly improves it \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\)\)\.
Guided by these insights, we proposeDSpine, a drafter with causal conditioning injection throughout the backbone\. At every layer, a gated adjacent injection module writes each predecessor’s predicted feature into its successor, so that the causal conditioning chain unfolds over network depth while all positions update in parallel\. We further build a unified transfer space from the output embeddings of the target model, which unifies layer\-wise injection with the information transfer in predecessor\-conditioned decoding\. We also design layer\-wise output\-embedding supervision that aligns the predicted features with this space, promoting their formation in shallow layers and reducing their discrepancy from token embeddings\. Finally, we introduce a transition cache that precomputes the scores of all adjacent candidate pairs in parallel, turning sequential conditional decoding into parallel computation\. Our contributions are:
- •We identify recoverable shallow prefix predictions and the effect of predecessor\-conditioning timing, motivating dependency formation across network depth\.
- •We propose DSpine, which injects causal conditioning throughout the backbone and links it to predecessor\-conditioned decoding through a unified output\-embedding space with layer\-wise supervision\.
- •We demonstrate that DSpine surpasses baselines trained on the same data in both acceptance length and SGLang throughput: on Qwen3\-8B, it raises the mean acceptance length from 3\.77 for DFlash to 4\.82 \([Figure1](https://arxiv.org/html/2609.36173#S0.F1)\), and its SGLang throughput exceeds that of DFlash by 23\.3%\.
## 2Preliminaries
### 2\.1Speculative Decoding
Given a confirmed contextx≤0x\_\{\\leq 0\}, a drafter proposesMMcandidatesx^1:M\\hat\{x\}\_\{1:M\}, which the target model verifies in one forward pass\([Leviathan et al\., 2023](https://arxiv.org/html/2609.36173#bib.bib1)\)\. Letqdraft,tq\_\{\\mathrm\{draft\},t\}be the actual proposal distribution given this context and preceding candidates; both it and the target distributionptp\_\{t\}include their respective sampling transformations\. Conditional on acceptance of all preceding candidates, candidatettis accepted with probability
min\(1,pt\(x^t\)qdraft,t\(x^t\)\),pt\(v\)=ptarget\(v∣x≤0,x^1:t−1\)\.\\min\\\!\\left\(1,\\frac\{p\_\{t\}\(\\hat\{x\}\_\{t\}\)\}\{q\_\{\\mathrm\{draft\},t\}\(\\hat\{x\}\_\{t\}\)\}\\right\),\\qquad p\_\{t\}\(v\)=p\_\{\\mathrm\{target\}\}\(v\\mid x\_\{\\leq 0\},\\hat\{x\}\_\{1:t\-1\}\)\.\(1\)At the first rejection, the remaining draft is discarded and a replacement is sampled from the corrected distribution \([Equation11](https://arxiv.org/html/2609.36173#A3.E11)\); if all candidates pass, the target supplies an additional token\. Greedy decoding accepts candidates matching the target’s choice under the same prefix\.
The average acceptance lengthτ=𝔼\[Nacc\]\+1\\tau=\\mathbb\{E\}\[N\_\{\\mathrm\{acc\}\}\]\+1is the number of tokens advanced per round, whereNaccN\_\{\\mathrm\{acc\}\}counts consecutively accepted draft tokens and the extra one is the correction or additional target token\. For a fixed runtime configuration, ignoring prefill and termination costs, the average time per output token \(TPOT\) is approximately
TPOTspec≈Tdraft\+Tverifyτ,\\mathrm\{TPOT\}\_\{\\mathrm\{spec\}\}\\approx\\frac\{T\_\{\\mathrm\{draft\}\}\+T\_\{\\mathrm\{verify\}\}\}\{\\tau\},\(2\)whereTdraftT\_\{\\mathrm\{draft\}\}andTverifyT\_\{\\mathrm\{verify\}\}are mean per\-round costs, including sampling and cache management\. Later candidates contribute only when the preceding prefix is accepted, and acceptance gains must compensate for added drafting cost\.
### 2\.2Parallel Drafting
DFlash\-style parallel drafters process a block ofM\+1M\+1positions, the known anchorx0x\_\{0\}followed byMMmasked candidate positions\([Chen et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib3);[Zhang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib7)\)\. In DFlash, fused target\-model features overx<0x\_\{<0\}supply K/V at every layer\. The backbone updates all positions in one pass with bidirectional block attention, and the LM head shared with the target produces candidate distributions\.
An autoregressive drafter generates each token conditioned on the predecessors it has drafted, whereas a DFlash\-style drafter predicts every position from the confirmed context alone, so the block proposal factorizes over positions:
qdraft\(x^1:M∣x≤0\)=∏t=1Mqdraft,t\(x^t∣x≤0\)\.q\_\{\\mathrm\{draft\}\}\(\\hat\{x\}\_\{1:M\}\\mid x\_\{\\leq 0\}\)=\\prod\_\{t=1\}^\{M\}q\_\{\\mathrm\{draft\},t\}\(\\hat\{x\}\_\{t\}\\mid x\_\{\\leq 0\}\)\.\(3\)Each position thus marginalizes over its possible predecessors instead of conditioning on the one actually drafted, and individually likely tokens can form inconsistent combinations\. Because verification accepts candidates in prefix order, such inconsistencies cut the accepted prefix, and acceptance decays rapidly at later positions\([Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6)\)\.
## 3When and How Prefix Predictions Condition Parallel Drafting
This section asks whether evolving prefix predictions can condition successor computation in parallel drafting\. We first examine when predictive information becomes recoverable and when reliable prefix conditions are most useful, then turn to the immediate predecessor: its contribution beyond earlier history and the timing of its use\. We use the frozen official Qwen3\-8B\-DFlash\-b16 with greedy Qwen3\-8B continuations as the correct tokens; datasets, metrics, and full results are in[AppendixB](https://arxiv.org/html/2609.36173#A2)\.
Figure 2:\(a\) Readout accuracy of each layer, grouped by block position\. \(b\) Gain in consecutive correct tokens when the drafted prefix is replaced by the correct prefix at the filled \(orange\) layers; top: from one layer onward; bottom: at two layers, moving the earlier one\. \(c\) Gain in successor accuracy when, at the specified layers, the adjacent predecessor’s features come from the correct token rather than its draft prediction \(L1–L5: all layers\); line color marks whether the history before the predecessor is correct \(orange\) or drafted \(blue\)\.### 3\.1Early Predictive Cues and the Timing of Conditioning
To examine when predictive information emerges, we read out predictions from each layer\. Intermediate DFlash states lie outside the input space of the target LM head and cannot be decoded directly, so for each layer we fit, on separate data, an equal\-capacity lightweight low\-rank projection that maps its frozen states into the space decoded by the frozen LM head\. At the same shallow depth, early positions are read out much more accurately than distant ones, while all position groups improve with depth \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(a\)\)\. Although positions update in parallel, predictive information is not equally available: prefixes already carry recoverable cues while successor predictions are still forming, so prefixes could condition successor computation before selecting their own tokens\.
We next examine when more accurate prefix conditions are most useful\. With a correct token as the anchor, we fill the preceding prefix with draft predictions or correct tokens, which the target model encodes for DFlash’s existing context pathway\. Sustained access to correct\-prefix features from an early layer yields a 6\.7% relative gain in consecutive correct tokens over the all\-draft\-prefix baseline, and delaying it reduces the gain \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(b\), sustained conditioning\)\.
Earlier access, however, also means that more layers receive correct features\. To isolate timing, we fix the total number of injection layers and vary only their placement: earlier injection raises the relative gain from 1\.2% to 4\.3% \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(b\), timing control\)\. Thus, even at a fixed number of uses, earlier access to reliable conditions improves successor prediction: a reliable prefix is valuable not only for more accurate content but also for timely use as successor representations form\.
### 3\.2Adjacent Predecessors Shape Successor Predictions
We next examine how the immediate predecessor affects successor prediction\. Keeping the target\-encoded earlier history correct, we use the predecessor as the anchor of a new draft block: replacing only the correct predecessor with its draft prediction lowers successor accuracy by 31\.8 percentage points on average over three block positions \(L1–L5 in[Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\)\)\. Correct earlier history thus does not fully compensate for an inaccurate predecessor: the adjacent token supplies additional predictive cues for its successor\.
The predecessor’s effect also depends on when it enters successor computation\. Keeping the draft\-predicted anchor, we replace its K/V at selected layers with the same\-layer K/V cached from a forward pass with the correct predecessor\. With correct earlier history, the same number of injection layers, and the same predecessor K/V at the final layer, earlier injection improves successor accuracy over later injection by 14\.8 points on average \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\)\), indicating that adjacent conditions affect how successor representations form\.
These two observations are complementary: shallow prefix states already carry recoverable predictive cues, and accurate predecessor features help successors more when supplied earlier\. Motivated by this, we explore transmitting predecessor prediction features to successors from shallow layers, letting token dependencies build up progressively through network depth \([Section4](https://arxiv.org/html/2609.36173#S4)\)\.
## 4DSpine: Adjacent Causal Injection and Output\-Embedding Supervision
In parallel drafting, every position predicts before its predecessors commit to tokens, so a successor cannot know which token its predecessor will take \([Section2\.2](https://arxiv.org/html/2609.36173#S2.SS2)\)\. DSpine gives each successor what its immediate predecessor is predicting\. The observations in[Section3](https://arxiv.org/html/2609.36173#S3)determine where this information comes from and when it arrives: from the adjacent predecessor, which supplies cues that earlier history cannot replace, and from the first layer on, since shallow prefix states already carry recoverable cues and accurate predecessor features help more when supplied earlier\.
[Figure3](https://arxiv.org/html/2609.36173#S4.F3)\(a\) outlines the resulting forward pass\. After every Transformer layer, each position receives the feature its predecessor currently predicts, and a gate decides how much of it to absorb; all positions still update in parallel within a single forward pass\. After the final layer, each position proposes its top\-KKcandidates, and tokens are selected from left to right: once a predecessor’s token is selected, it is written into its successor through the same injection, which yields the successor’s final scores\. We describe what is passed \([Section4\.1](https://arxiv.org/html/2609.36173#S4.SS1)\), how it is passed inside the backbone \([Section4\.2](https://arxiv.org/html/2609.36173#S4.SS2)\), how the same pathway completes decoding \([Section4\.3](https://arxiv.org/html/2609.36173#S4.SS3)\), and how the passed features are supervised and the drafter is trained \([Sections4\.4](https://arxiv.org/html/2609.36173#S4.SS4)and[4\.5](https://arxiv.org/html/2609.36173#S4.SS5)\)\.
Figure 3:Overview of DSpine\. \(a\) Injection passes predecessor features to successors at every layer \(green\); after the final layer, transition\-cached decoding selects each token conditioned on its selected predecessor \(orange\)\. \(b\) One injection layer: predecessor and successor features jointly gate how much of the predecessor’s message enters the successor’s residual stream; training adds cosine alignment toCCand, with probabilitypp, substitutesCyt−1C\_\{y\_\{t\-1\}\}for the predecessor feature\.### 4\.1Unified Transfer Space
Existing parallel drafters pass predecessor information to a successor in two ways: inside the backbone, as hidden states mixed across positions by attention or similar operations; after the backbone, as the selected predecessor token fed to an additional module\. The former mixes the predecessor’s prediction with the context it has gathered, leaving the successor to infer which token the predecessor favors; the latter provides a decided token but only after all layers have been computed\. DSpine passes the predecessor’s prediction itself from the first layer on, expressed in a space where vectors name tokens\.
The target LM head scores tokenvvasEv⊤hE\_\{v\}^\{\\top\}h, so its output embeddings naturally encode which token a state predicts\. We build the unified transfer space from them: weℓ2\\ell\_\{2\}\-normalize the rows of the LM head, apply PCA whitening, keeprrprincipal directions, and normalize each row again to obtain a fixed embedding tableCC\([AppendixA](https://arxiv.org/html/2609.36173#A1)\); whitening removes the component shared by all tokens, so that directions inCCdistinguish tokens\. Two kinds of information enter this space\. A predecessor whose token is still undetermined contributes a predicted feature: each layer projects its pre\-injection stateuℓ,tu\_\{\\ell,t\}, the output of theℓ\\ell\-th Transformer layer, tozℓ,t=fℓ\(uℓ,t\)=rnormalize\(Rℓuℓ,t\)z\_\{\\ell,t\}=f\_\{\\ell\}\(u\_\{\\ell,t\}\)=\\sqrt\{r\}\\,\\normalize\(R\_\{\\ell\}u\_\{\\ell,t\}\)without intermediate vocabulary decoding, whereRℓR\_\{\\ell\}is a layer\-specific learned projection torrdimensions andnormalize\\normalizedenotesℓ2\\ell\_\{2\}normalization\. A known tokenxx, namely the anchor or a decoded predecessor, enters directly as its embeddingrCx\\sqrt\{r\}\\,C\_\{x\}\. Because both take the same form, a single pathway can carry either a prediction or a decided token: DSpine passes predicted features inside the backbone and decoded tokens in predecessor\-conditioned decoding\.
### 4\.2Adjacent Causal Injection
After theℓ\\ell\-th Transformer layer, positionttreceives a message from its immediate predecessor only\. Shallow predecessor features are still forming and vary in reliability, so the successor gates how much of the message to absorb: the receiver featurezℓ,tz\_\{\\ell,t\}and the predecessor featurezℓ,t−1z\_\{\\ell,t\-1\}jointly determine the gate, while the predecessor feature is projected into a message that is gated and written back to the residual stream:
gℓ,t\\displaystyle g\_\{\\ell,t\}=σ\(Gqzℓ,t\+Gpzℓ,t−1\+b\),\\displaystyle=\\sigma\\\!\\left\(G\_\{q\}z\_\{\\ell,t\}\+G\_\{p\}z\_\{\\ell,t\-1\}\+b\\right\),\(4\)mℓ,t\\displaystyle m\_\{\\ell,t\}=Azℓ,t−1,\\displaystyle=Az\_\{\\ell,t\-1\},hℓ,t\\displaystyle h\_\{\\ell,t\}=uℓ,t\+RMS\(uℓ,t\)Wℓ\(gℓ,t⊙mℓ,t\)\.\\displaystyle=u\_\{\\ell,t\}\+\\RMS\(u\_\{\\ell,t\}\)\\,W\_\{\\ell\}\\left\(g\_\{\\ell,t\}\\odot m\_\{\\ell,t\}\\right\)\.\(5\)Hereσ\\sigmais the sigmoid function;AA,GqG\_\{q\},GpG\_\{p\}, andbbare shared across layers, and each layer has its own write matrixWℓW\_\{\\ell\}\. All messages use pre\-update features, so injection within a layer runs in parallel across positions; the anchor remains unchanged\.
#### Block\-causal attention\.
Unlike DFlash and other parallel drafters with bidirectional block attention, DSpine uses block\-causal attention in allLLbackbone layers: positionttattends only to target\-context K/V and block positions0,…,t0,\\ldots,t\. Attention and injection thus share the predecessor\-to\-successor direction, and the correct information injected into a position during training cannot flow back to earlier positions \([Section4\.5](https://arxiv.org/html/2609.36173#S4.SS5)\)\. Drafting remains a single, block\-parallel forward pass\.
#### Propagation through depth\.
Because earlier access to accurate predecessor features benefits successor prediction \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\)\), injection starts at the first layer and its effects propagate through depth: messages received at one layer enter the next layer’s computation, shaping the predicted features that successors send onward\. For example, information thatt−1t\-1receives fromt−2t\-2at layerℓ\\ellcan influence the message sent tottat layerℓ\+1\\ell\+1\([Figure3](https://arxiv.org/html/2609.36173#S4.F3)\(a\)\)\. The causal conditioning chain thus unfolds over network depth while positions within each layer update in parallel; earlier context remains accessible through block\-causal attention\.
#### Fused layer\-wise injection\.
Layer\-wise injection consists of small normalization, projection, gating, and residual\-write operations, which a direct implementation executes as many kernel launches with intermediate tensor traffic\. We merge the three projections into one matrix multiplication and use two fused kernels, one for the gated message and one for the residual write together with the next layer’s input normalization, all executed within the draft backbone’s CUDA graph\.
### 4\.3Predecessor\-Conditioned Decoding
Figure 4:Transition cache\. \(a\) Serial conditional scoring\. \(b\) Parallel transition precomputation followed by sequential selection; greedy implementation shown\.Predecessor\-conditioned decoding conditions each position on the token actually decoded at its predecessor, rather than on its predicted feature, and it needs no new module\. The final\-layer outputhL,th\_\{L,t\}first passes through the final RMSNorm and shared LM head to determine each position’s initial top\-KKcandidates\. A decoded predecessor is a known token, so it enters the unified transfer space directly and passes through the same final\-layer injection: for a tokenssat positiont−1t\-1, we replace the predicted featurezL,t−1z\_\{L,t\-1\}withrCs\\sqrt\{r\}\\,C\_\{s\}, changing both gate and message, and redo the final\-layer write \([Equation5](https://arxiv.org/html/2609.36173#S4.E5)\) from the cached pre\-injection stateuL,tu\_\{L,t\}\. We denote the output byhto\(s\)h\_\{t\}^\{\\mathrm\{o\}\}\(s\)and call this step*last\-write refinement*; the shared LM head maps it to conditional scoresSt\(s,v\)S\_\{t\}\(s,v\)over the initial candidates\.
#### Transition\-cached decoding\.
Because each position’s scores depend on the predecessor’s choice, computing them directly proceeds sequentially across positions\. The transition cache removes this sequential computation: once the backbone has fixed each position’s pre\-injection stateuL,tu\_\{L,t\}and top\-KKcandidate set𝒱t=\{vt,1,…,vt,K\}\\mathcal\{V\}\_\{t\}=\\\{v\_\{t,1\},\\ldots,v\_\{t,K\}\\\}, last\-write refinement depends only on the predecessor’s token, so before selection we compute the refined final\-layer features for all predecessor candidates in parallel and score the successor candidates, organized asTt\[i,j\]=St\(vt−1,i,vt,j\)T\_\{t\}\[i,j\]=S\_\{t\}\(v\_\{t\-1,i\},v\_\{t,j\}\)\([Figure4](https://arxiv.org/html/2609.36173#S4.F4)\)\. Enumerating allK2K^\{2\}pairs per transition trades additional scoring work for a shorter sequential dependency chain; in practice,K=16K=16suffices at a small additional cost \([Section5\.4](https://arxiv.org/html/2609.36173#S5.SS4)\)\. Selection then proceeds through the cache from the first position, without any sequential matrix operations\.
### 4\.4Layer\-Wise Output\-Embedding Supervision
The predicted features that injection passes serve two roles, and the final\-layer loss shapes them only indirectly through later computation\. First, they condition successors from the first layer on, yet in shallow layers they carry recoverable cues but are still forming \([Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(a\)\)\. Second, the injection receives predicted features inside the backbone but token embeddings in predecessor\-conditioned decoding, so the two should lie close in the unified transfer space\. DSpine therefore aligns the predicted features at every layer directly with the correct token’s embedding, promoting their formation in shallow layers and narrowing their gap to the token embeddings used in predecessor\-conditioned decoding\.
#### Directional alignment\.
With the projectionfℓf\_\{\\ell\}used for injection, we readeℓ,i=fℓ\(hℓ,i\)e\_\{\\ell,i\}=f\_\{\\ell\}\(h\_\{\\ell,i\}\)from post\-injection states and align it with the correct token’s embeddingCyiC\_\{y\_\{i\}\}:
ℒemb=∑ℓ=1L∑iwi\[1−cos\(eℓ,i,Cyi\)\]\.\\mathcal\{L\}\_\{\\mathrm\{emb\}\}=\\sum\_\{\\ell=1\}^\{L\}\\sum\_\{i\}w\_\{i\}\\left\[1\-\\cos\\\!\\left\(e\_\{\\ell,i\},C\_\{y\_\{i\}\}\\right\)\\right\]\.\(6\)wherewi=exp\(−\(ti−1\)/γ\)w\_\{i\}=\\exp\(\-\(t\_\{i\}\-1\)/\\gamma\)is the decay weight for block positiontit\_\{i\}\. Because the readout follows injection, this loss trains both the formation of each feature and the successor’s use of the received message; it constrains only the direction of a low\-dimensional projection, leaving the full hidden state free for contextual computation\.
### 4\.5Training
We sample draft blocks from training sequences and train the drafter with the target model, shared input embeddings, and LM head frozen\. Training supervises two outputs\. The first is the ordinary final\-layer output, which predicts over the full vocabulary and determines the top\-KKcandidates\. The second is the output of last\-write refinement: as at inference, it replaces the predicted predecessor feature with a token embedding and redoes the final\-layer write, except that the correct predecessor token is used, and predicts only over those candidates\. Both receive the same loss below, jointly training candidate generation and predecessor\-conditioned decoding\.
Letqiq\_\{i\}andpip\_\{i\}denote the draft and target distributions at positioniiunder the correct prefix\. Following the loss design of DSpark\([Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6)\), the loss is a weighted mix of token cross\-entropy and probabilityℓ1\\ell\_\{1\}distance:
ℒdist=∑iwi\[−αlogqi\(yi\)\+\(1−α\)‖qi−pi‖1\],\\mathcal\{L\}\_\{\\mathrm\{dist\}\}=\\sum\_\{i\}w\_\{i\}\\Bigl\[\-\\alpha\\log q\_\{i\}\(y\_\{i\}\)\+\(1\-\\alpha\)\\,\\\|q\_\{i\}\-p\_\{i\}\\\|\_\{1\}\\Bigr\],\(7\)whereα\\alphais the mixing weight\. The output of last\-write refinement is also included inℒdist\\mathcal\{L\}\_\{\\mathrm\{dist\}\}, except that its distributions are restricted to the top\-KKcandidates, with the target renormalized to that set\.
Early in training, predicted features carry little reliable information, which gives the injection little to learn from\. We therefore replace shallow\-layer predecessor features with correct\-token embeddings with some probability, keeping receiver features model\-derived\([Bengio et al\., 2015](https://arxiv.org/html/2609.36173#bib.bib8)\), and gradually phase this replacement out; the unified transfer space makes this substitution direct\. Last\-write refinement always uses correct predecessors\. The overall training objective weights layer\-wise supervision byλ\\lambda\(coefficients and schedules in[AppendixA](https://arxiv.org/html/2609.36173#A1)\):
ℒ=ℒdist\+λℒemb\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{dist\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{emb\}\}\.\(8\)
## 5Experiments
### 5\.1Experimental Setup
#### Models and evaluations\.
We conduct experiments on Qwen3\-4B and Qwen3\-8B\([Yang et al\., 2025](https://arxiv.org/html/2609.36173#bib.bib9)\)with thinking mode disabled, covering three task categories:*Math*: GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.36173#bib.bib10)\)and MATH\-500\([Lightman et al\., 2023](https://arxiv.org/html/2609.36173#bib.bib11)\);*Code*: HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2609.36173#bib.bib12)\), MBPP\([Austin et al\., 2021](https://arxiv.org/html/2609.36173#bib.bib13)\), and LiveCodeBench \(LCB\)\([Jain et al\., 2024](https://arxiv.org/html/2609.36173#bib.bib14)\);*Chat*: MT\-Bench\([Zheng et al\., 2023a](https://arxiv.org/html/2609.36173#bib.bib15)\)and Arena\-Hard\([Li et al\., 2024a](https://arxiv.org/html/2609.36173#bib.bib16)\)\. On all seven benchmarks, we measure the average acceptance lengthτ\\tau\([Section2\.1](https://arxiv.org/html/2609.36173#S2.SS1)\) in SGLang\([Zheng et al\., 2023b](https://arxiv.org/html/2609.36173#bib.bib17)\)at temperatures 0 and 1; serving throughput is measured on GSM8K, MATH\-500, HumanEval, and MBPP at temperature 0 and concurrency 2–32, with the same fixed verification budget for all methods \([AppendixC](https://arxiv.org/html/2609.36173#A3)\)\.
#### Baselines\.
We compare with four representative state\-of\-the\-art parallel drafters: DFlash\([Chen et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib3)\), Domino\([Huang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib5)\), DSpark\([Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6)\), and DFlash2\([DFlash2 Team, 2026](https://arxiv.org/html/2609.36173#bib.bib4)\)\. DFlash selects tokens independently across positions; the other three build on the DFlash backbone to model intra\-block dependencies\.
#### Implementation\.
For each target model, all drafters are retrained on the same ShareGPT data, with responses regenerated by the target; backbones use five layers and a block size of 16 and are trained for six epochs with the same global batch size and learning\-rate schedule \([AppendixA](https://arxiv.org/html/2609.36173#A1)\)\.
### 5\.2Main Results
Table 1:Average acceptance lengthτ\\tauon Qwen3 models\. Each cell gives temperature 0 / 1; bold marks the best result for each model and temperature, and underline marks the second best\.As shown in[Table1](https://arxiv.org/html/2609.36173#S5.T1), DSpine achieves the longest acceptance length on all benchmarks for both models at both temperatures\. On Qwen3\-8B, it raises the averageτ\\tauof DFlash from 3\.77 to 4\.82 and leads the strongest baseline by 10\.0% with greedy decoding \(DFlash2\) and by 12\.8% under sampling \(DSpark\); on Qwen3\-4B, its lead over the strongest baseline, DFlash2, is 5\.4% and 10\.1%, respectively\. Domino, DSpark, and DFlash2, which condition on predecessors after the backbone, already improve substantially over DFlash, while DSpine, which conditions on predecessors at every layer and again at token selection, extends the gain further, with a larger lead under sampling \(see[AppendixC](https://arxiv.org/html/2609.36173#A3)for each method’s proposal policy at temperature one\)\.
### 5\.3Serving Throughput
Figure 5:Serving throughput gain over DFlash on SGLang, in thousands of tokens per second \(temperature zero\)\. The dashed zero line is DFlash; shading and the strip above each panel give the lead of DSpine over DFlash2, the strongest baseline\.As shown in[Figure5](https://arxiv.org/html/2609.36173#S5.F5), DSpine achieves the highest throughput at every task and concurrency level on both models: on Qwen3\-8B it outperforms the strongest baseline, DFlash2, by 11\.1% on average and reaches up to 4\.0×\\timesspeedup over autoregressive decoding \([Table8](https://arxiv.org/html/2609.36173#A3.T8)\); on Qwen3\-4B it leads DFlash2 by 6\.5% on average\. The gain in acceptance length thus carries over to serving, and the extra cost of layer\-wise injection and transition\-cached decoding \([Section5\.4](https://arxiv.org/html/2609.36173#S5.SS4)\) does not offset it\.
### 5\.4Decoding Latency Breakdown
Figure 6:Module\-level breakdown of Qwen3\-8B decoding time on SGLang\.To examine where decoding time goes, we time each decoding round of Qwen3\-8B in SGLang, split into a draft stage and a verification stage\. As shown in[Figure6](https://arxiv.org/html/2609.36173#S5.F6), verification takes 72–76% of round time and costs about the same for all methods\. DSpine’s parallel injection adds at most 0\.15 ms per round across all five layers, and its predecessor\-conditioned decoding, which scores candidates and selects through the transition cache \([Section4\.3](https://arxiv.org/html/2609.36173#S4.SS3)\), costs less than half as much as the post\-backbone heads of DSpark and Domino\. Overall, a DSpine round costs about as much as a DFlash2 round \(7\.14 vs\. 7\.19 ms\), and because more tokens are accepted per round, DSpine spends 10\.9% less decoding time per generated token than DFlash2\.
### 5\.5Ablation Study
Table 2:Ablation of module
Table 3:Ablation of supervision loss
#### Effect of each component\.
We train each variant in[Table3](https://arxiv.org/html/2609.36173#S5.T3)independently\. Removing last\-write refinement reducesτ\\tauby 6\.6–7\.1%, and further removing parallel injection enlarges the drop to 8\.8–10\.7%\. Both components improve acceptance length, and their contributions are complementary\.
#### Supervision target\.
[Table3](https://arxiv.org/html/2609.36173#S5.T3)compares intermediate\-layer supervision: directional supervision toward compact output embeddings improves acceptance length, whereas per\-layer full\-vocabulary cross\-entropy falls below the unsupervised baseline\. This supports the design of[Section4\.4](https://arxiv.org/html/2609.36173#S4.SS4), as forcing intermediate states to classify tokens may interfere with later layers’ contextual computation\.
## 6Related Work
Speculative decoding lets a lightweight drafter propose tokens that the target model verifies losslessly\([Leviathan et al\., 2023](https://arxiv.org/html/2609.36173#bib.bib1);[Chen et al\., 2023](https://arxiv.org/html/2609.36173#bib.bib18)\)\. Follow\-ups mainly improve the drafter: Medusa predicts several positions with independent heads, Hydra makes these heads depend on drafted tokens\([Cai et al\., 2024](https://arxiv.org/html/2609.36173#bib.bib19);[Ankner et al\., 2024](https://arxiv.org/html/2609.36173#bib.bib20)\), the EAGLE series drafts autoregressively from target features\([Li et al\., 2024b](https://arxiv.org/html/2609.36173#bib.bib21);[Li et al\., 2024c](https://arxiv.org/html/2609.36173#bib.bib22);[Li et al\., 2025](https://arxiv.org/html/2609.36173#bib.bib2)\), and tree verification checks multiple continuations per round\([Miao et al\., 2024](https://arxiv.org/html/2609.36173#bib.bib23);[Li et al\., 2024c](https://arxiv.org/html/2609.36173#bib.bib22)\)\. Autoregressive drafting conditions each token on its predecessors but costs one sequential step per token\.
Parallel decoding methods predict multiple positions at once\([Stern et al\., 2018](https://arxiv.org/html/2609.36173#bib.bib24);[An et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib25);[Liu et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib26);[Christopher et al\., 2025](https://arxiv.org/html/2609.36173#bib.bib27);[Li et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib28);[Chen et al\., 2025](https://arxiv.org/html/2609.36173#bib.bib29)\)\. DFlash drafts a block in one pass of a lightweight block\-diffusion model conditioned on target features\([Chen et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib3)\); follow\-ups add layer\-wise target conditioning, position\-weighted losses, or draft trees\([Zhang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib7);[Wu et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib30);[Ringel and Romano, 2026](https://arxiv.org/html/2609.36173#bib.bib31)\)\. Since positions predict before their predecessors commit, the resulting tokens can be inconsistent\([Stern et al\., 2018](https://arxiv.org/html/2609.36173#bib.bib24);[Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6)\)\. Most remedies restore this dependency after the backbone, through a causal correction branch \(Domino\), a sequential head \(DSpark\), or a path selector over top\-kkcandidates \(DFlash2\)\([Huang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib5);[Cheng et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib6);[DFlash2 Team, 2026](https://arxiv.org/html/2609.36173#bib.bib4)\), as well as previous\-token conditioning, block refinement, or adjacent\-candidate scoring\([Rheinboldt et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib32);[Wang et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib33);[Rusanovsky et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib34)\)\. Inside the backbone, DFlash2 adds causal two\-tap convolutions, and DART and JetSpec use causal block attention\([DFlash2 Team, 2026](https://arxiv.org/html/2609.36173#bib.bib4);[Liu et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib26);[Hu et al\., 2026](https://arxiv.org/html/2609.36173#bib.bib35)\)\. DSpine also uses block\-causal attention but additionally injects each predecessor’s predicted feature, supervised in a unified transfer space, into its successor at every layer; the same injection carries decoded tokens in predecessor\-conditioned decoding\.
## 7Conclusion
DSpine injects causal conditioning throughout the backbone, building intra\-block token dependencies over network depth within a single parallel drafting pass\. Gated adjacent injection passes predecessor features to successors at every layer, and a unified transfer space with layer\-wise supervision links this injection to predecessor\-conditioned decoding\. A transition cache precomputes the scores of all adjacent candidate pairs in parallel, turning sequential conditional decoding into parallel computation\. Across math, code, and chat tasks, DSpine improves acceptance length over representative parallel drafters and translates these gains into higher SGLang serving throughput\. These results motivate further study of in\-layer information transfer, including more flexible communication ranges and interactions in other parallel decoding architectures\.
## References
- Leviathan et al\. \(2023\)Yaniv Leviathan, Matan Kalman, and Yossi Matias\.Fast inference from transformers via speculative decoding\.In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors,*Proceedings of the 40th International Conference on Machine Learning*, volume 202 of*Proceedings of Machine Learning Research*, pages 19274–19286\. PMLR, 2023\.
- Li et al\. \(2025\)Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang\.EAGLE\-3: Scaling up inference acceleration of large language models via training\-time test\.In*Advances in Neural Information Processing Systems*, volume 38, 2025\.[10\.52202/085713\-4562](https://doi.org/10.52202/085713-4562)\.URL[https://papers\.nips\.cc/paper\_files/paper/2025/hash/c7b5a35ea98b62512a869c19ea7b03cb\-Abstract\-Conference\.html](https://papers.nips.cc/paper_files/paper/2025/hash/c7b5a35ea98b62512a869c19ea7b03cb-Abstract-Conference.html)\.
- Chen et al\. \(2026\)Jian Chen, Yesheng Liang, and Zhijian Liu\.DFlash: Block diffusion for flash speculative decoding\.In*Proceedings of the 43rd International Conference on Machine Learning*, volume 306 of*Proceedings of Machine Learning Research*\. PMLR, 2026\.
- DFlash2 Team \(2026\)DFlash2 Team\.DFlash 2: Keep drafting parallel\.Blog post, August 2026\.URL[https://inco\.ai/blog/dflash2/](https://inco.ai/blog/dflash2/)\.
- Huang et al\. \(2026\)Jianuo Huang, Yaojie Zhang, Qituan Zhang, Hao Lin, Hanlin Xu, and Linfeng Zhang\.Domino: Decoupling causal modeling from autoregressive drafting in speculative decoding\.arXiv preprint arXiv:2605\.29707, 2026\.URL[https://arxiv\.org/abs/2605\.29707](https://arxiv.org/abs/2605.29707)\.
- Cheng et al\. \(2026\)Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, and Wenfeng Liang\.DSpark: Confidence\-scheduled speculative decoding with semi\-autoregressive generation\.arXiv preprint arXiv:2607\.05147, 2026\.URL[https://arxiv\.org/abs/2607\.05147](https://arxiv.org/abs/2607.05147)\.
- Zhang et al\. \(2026\)Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J\. Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, Jianchen Zhu, and Sujian Li\.DFlare: Scaling up draft capacity for block diffusion speculative decoding\.arXiv preprint arXiv:2606\.02091, 2026\.URL[https://arxiv\.org/abs/2606\.02091](https://arxiv.org/abs/2606.02091)\.
- Bengio et al\. \(2015\)Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer\.Scheduled sampling for sequence prediction with recurrent neural networks\.In C\. Cortes, N\. Lawrence, D\. Lee, M\. Sugiyama, and R\. Garnett, editors,*Advances in Neural Information Processing Systems*, volume 28\. Curran Associates, Inc\., 2015\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report, 2025\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\.Training verifiers to solve math word problems, 2021\.
- Lightman et al\. \(2023\)Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step, 2023\.
- Chen et al\. \(2021\)Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al\.Evaluating large language models trained on code, 2021\.
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton\.Program synthesis with large language models, 2021\.
- Jain et al\. \(2024\)Naman Jain, King Han, Alex Gu, Wen\-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar\-Lezama, Koushik Sen, and Ion Stoica\.LiveCodeBench: Holistic and contamination free evaluation of large language models for code, 2024\.
- Zheng et al\. \(2023a\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P\. Xing, et al\.Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena, 2023a\.
- Li et al\. \(2024a\)Tianle Li, Wei\-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E\. Gonzalez, and Ion Stoica\.From crowdsourced data to high\-quality benchmarks: Arena\-Hard and BenchBuilder pipeline, 2024a\.
- Zheng et al\. \(2023b\)Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E\. Gonzalez, Clark Barrett, and Ying Sheng\.SGLang: Efficient execution of structured language model programs, 2023b\.
- Chen et al\. \(2023\)Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean\-Baptiste Lespiau, Laurent Sifre, and John Jumper\.Accelerating large language model decoding with speculative sampling\.arXiv preprint arXiv:2302\.01318, 2023\.URL[https://arxiv\.org/abs/2302\.01318](https://arxiv.org/abs/2302.01318)\.
- Cai et al\. \(2024\)Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D\. Lee, Deming Chen, and Tri Dao\.Medusa: Simple LLM inference acceleration framework with multiple decoding heads\.In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pages 5209–5235\. PMLR, 2024\.
- Ankner et al\. \(2024\)Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan\-Kelley, and William Brandon\.Hydra: Sequentially\-dependent draft heads for Medusa decoding\.In*First Conference on Language Modeling*, 2024\.URL[https://openreview\.net/forum?id=FbhjirzvJG](https://openreview.net/forum?id=FbhjirzvJG)\.
- Li et al\. \(2024b\)Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang\.EAGLE: Speculative sampling requires rethinking feature uncertainty\.In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pages 28935–28948\. PMLR, 2024b\.
- Li et al\. \(2024c\)Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang\.EAGLE\-2: Faster inference of language models with dynamic draft trees\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen, editors,*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 7421–7432, Miami, Florida, USA, November 2024c\. Association for Computational Linguistics\.[10\.18653/v1/2024\.emnlp\-main\.422](https://doi.org/10.18653/v1/2024.emnlp-main.422)\.URL[https://aclanthology\.org/2024\.emnlp\-main\.422/](https://aclanthology.org/2024.emnlp-main.422/)\.
- Miao et al\. \(2024\)Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia\.SpecInfer: Accelerating large language model serving with tree\-based speculative inference and verification\.In*Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3*, ASPLOS ’24, pages 932–949\. ACM, 2024\.[10\.1145/3620666\.3651335](https://doi.org/10.1145/3620666.3651335)\.
- Stern et al\. \(2018\)Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit\.Blockwise parallel decoding for deep autoregressive models\.In S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett, editors,*Advances in Neural Information Processing Systems*, volume 31\. Curran Associates, Inc\., 2018\.
- An et al\. \(2026\)Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, and Emad Barsoum\.PARD: Accelerating LLM inference with low\-cost PARallel draft model adaptation\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=XbOyv7iVGL](https://openreview.net/forum?id=XbOyv7iVGL)\.
- Liu et al\. \(2026\)Fuliang Liu, Xue Li, Ketai Zhao, Yinxi Gao, Ziyan Zhou, Zhonghui Zhang, Zhibin Wang, Wanchun Dou, Sheng Zhong, and Chen Tian\.DART: Diffusion\-inspired speculative decoding for fast LLM inference\.arXiv preprint arXiv:2601\.19278, 2026\.URL[https://arxiv\.org/abs/2601\.19278](https://arxiv.org/abs/2601.19278)\.
- Christopher et al\. \(2025\)Jacob K Christopher, Brian R\. Bartoldson, Tal Ben\-Nun, Michael Cardei, Bhavya Kailkhura, and Ferdinando Fioretto\.Speculative diffusion decoding: Accelerating language generation through diffusion\.In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors,*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 12042–12059, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.[10\.18653/v1/2025\.naacl\-long\.601](https://doi.org/10.18653/v1/2025.naacl-long.601)\.URL[https://aclanthology\.org/2025\.naacl\-long\.601/](https://aclanthology.org/2025.naacl-long.601/)\.
- Li et al\. \(2026\)Guanghao Li, Zhihui Fu, Min Fang, Qibin Zhao, Ming Tang, Chun Yuan, and Jun Wang\.DiffuSpec: Unlocking diffusion language models for speculative decoding\.In Maria Liakata, Viviane P\. Moreira, Jiajun Zhang, and David Jurgens, editors,*Findings of the Association for Computational Linguistics: ACL 2026*, pages 20896–20910, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.[10\.18653/v1/2026\.findings\-acl\.1048](https://doi.org/10.18653/v1/2026.findings-acl.1048)\.URL[https://aclanthology\.org/2026\.findings\-acl\.1048/](https://aclanthology.org/2026.findings-acl.1048/)\.
- Chen et al\. \(2025\)Keyu Chen, Zhifeng Shen, Daohai Yu, Haoqian Wu, Wei Wen, Jianfeng He, Ruizhi Qiao, and Xing Sun\.ASPD: Unlocking adaptive serial\-parallel decoding by exploring intrinsic parallelism in LLMs\.arXiv preprint arXiv:2508\.08895, 2025\.URL[https://arxiv\.org/abs/2508\.08895](https://arxiv.org/abs/2508.08895)\.
- Wu et al\. \(2026\)Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, and Yilun Du\.D\-PACE: Dynamic position\-aware cross\-entropy for parallel speculative drafting\.arXiv preprint arXiv:2605\.18810, 2026\.URL[https://arxiv\.org/abs/2605\.18810](https://arxiv.org/abs/2605.18810)\.
- Ringel and Romano \(2026\)Liran Ringel and Yaniv Romano\.Accelerating speculative decoding with block diffusion draft trees\.arXiv preprint arXiv:2604\.12989, 2026\.URL[https://arxiv\.org/abs/2604\.12989](https://arxiv.org/abs/2604.12989)\.
- Rheinboldt et al\. \(2026\)Peer Rheinboldt, Frédéric Berdoz, and Roger Wattenhofer\.TreeFlash: Parallel AR\-approximation for faster speculative decoding\.arXiv preprint arXiv:2606\.03819, 2026\.URL[https://arxiv\.org/abs/2606\.03819](https://arxiv.org/abs/2606.03819)\.
- Wang et al\. \(2026\)Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K\. Ganti, Minjia Zhang, and Naigang Wang\.xPress: Parallel refinement for diffusion drafters in speculative decoding\.arXiv preprint arXiv:2608\.02438, 2026\.URL[https://arxiv\.org/abs/2608\.02438](https://arxiv.org/abs/2608.02438)\.
- Rusanovsky et al\. \(2026\)Matan Rusanovsky, Yoav Miron, Roy Uziel, Omer Belhasin, Hao Guo, Ran Zilberstein, Maor Ashkenazi, and Michael Elad\.LiLiCorr: Lightweight likelihood correlation of parallel drafts for speculative decoding\.arXiv preprint arXiv:2608\.20530, 2026\.URL[https://arxiv\.org/abs/2608\.20530](https://arxiv.org/abs/2608.20530)\.
- Hu et al\. \(2026\)Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao, Yu\-Yang Qian, Bojun Wang, Peng Zhao, Daxin Jiang, Yibo Zhu, Tajana Rosing, and Hao Zhang\.JetSpec: Breaking the scaling ceiling of speculative decoding with parallel tree drafting\.arXiv preprint arXiv:2606\.18394, 2026\.URL[https://arxiv\.org/abs/2606\.18394](https://arxiv.org/abs/2606.18394)\.
- Shah et al\. \(2024\)Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao\.FlashAttention\-3: Fast and accurate attention with asynchrony and low\-precision, 2024\.
## Appendix AImplementation Details of DSpine
#### Architecture\.
DSpine builds on the DFlash drafter for Qwen3\-8B \([Table4](https://arxiv.org/html/2609.36173#A1.T4)\): five Transformer layers with block\-causal attention, target context features fused from five target layers and supplied as K/V to every layer, and the target’s input embedding and LM head shared and frozen\. Injection adds about 33M parameters, roughly 3% of the backbone\.
#### Unified transfer space\.
LetE¯v\\bar\{E\}\_\{v\}be theℓ2\\ell\_\{2\}\-normalized LM\-head row of tokenvvandμ\\muits mean over the vocabulary, and letPrP\_\{r\}andΛr\\Lambda\_\{r\}hold the toprreigenvectors and eigenvalues of the centered Gram matrix∑v\(E¯v−μ\)\(E¯v−μ\)⊤\\sum\_\{v\}\(\\bar\{E\}\_\{v\}\-\\mu\)\(\\bar\{E\}\_\{v\}\-\\mu\)^\{\\top\}\. The table isCv=normalize\(Λr−1/2Pr⊤\(E¯v−μ\)\)C\_\{v\}=\\normalize\\bigl\(\\Lambda\_\{r\}^\{\-1/2\}P\_\{r\}^\{\\top\}\(\\bar\{E\}\_\{v\}\-\\mu\)\\bigr\), fixed and shared by all layers\.
#### Initialization\.
Each readout matrixRℓR\_\{\\ell\}is initialized withPr⊤P\_\{r\}^\{\\top\}; the write matricesWℓW\_\{\\ell\}and the predecessor gate matrixGpG\_\{p\}are zero\-initialized withb=0b=0, so every injection starts as the identity map\.
#### Losses\.
Both the backbone output and the last\-write refinement output useα=0\.1\\alpha=0\.1in[Equation7](https://arxiv.org/html/2609.36173#S4.E7)and are summed with equal weight\. The layer\-wise supervision weight isλ=βtsg\(CEbb\)\\lambda=\\beta\_\{t\}\\,\\sg\(\\mathrm\{CE\}\_\{\\mathrm\{bb\}\}\), whereCEbb\\mathrm\{CE\}\_\{\\mathrm\{bb\}\}is the token cross\-entropy of the backbone output andβt\\beta\_\{t\}rises linearly from 0 to 0\.5 over the first 146 steps\. For last\-write refinement, the cross\-entropy term is computed only where the correct token lies among theK=16K=16candidates, and theℓ1\\ell\_\{1\}term only where the candidates hold more than10−410^\{\-4\}of the target probability, against the target distribution renormalized to the candidates\.
#### Correct\-predecessor curriculum\.
Each draft block is selected with probabilityppby one draw shared by layers 1–3, which then userCyt−1\\sqrt\{r\}\\,C\_\{y\_\{t\-1\}\}as predecessor input\.ppis0\.50\.5for the first sixth of training and decays linearly to zero at one third\.
#### Training setup\.
All drafters share the same training configuration: the same ShareGPT data with responses regenerated by the target model with thinking disabled, the same draft block of one anchor and 15 candidates, a global batch of 112 sequences, AdamW without weight decay at a peak learning rate of6×10−46\\times 10^\{\-4\}with 4% linear warmup \(146 steps\) and cosine decay to zero, a maximum training sequence length of 3,072 tokens, and six epochs \(3,655 steps\)\.[Table4](https://arxiv.org/html/2609.36173#A1.T4)lists the remaining DSpine hyperparameters\.
Table 4:Remaining DSpine hyperparameters for Qwen3\-8B\.
## Appendix BDetails of the Conditioning Experiments
### B\.1Setup, Readouts, and Metrics
We use the GSM8K test set and HumanEval with the frozen official Qwen3\-8B\-DFlash\-b16 drafter; greedy Qwen3\-8B continuations \(thinking disabled, at most 1,024 new tokens\) serve as correct tokens\. Each block contains one anchor and 15 candidates\.
For layerℓ\\ell, the readout is
qℓ=softmax\(E\[n\(hℓ\)\+UℓVℓn\(hℓ\)\+bℓ\]\),q\_\{\\ell\}=\\softmax\\\!\\left\(E\[n\(h\_\{\\ell\}\)\+U\_\{\\ell\}V\_\{\\ell\}n\(h\_\{\\ell\}\)\+b\_\{\\ell\}\]\\right\),\(9\)whereEEis the frozen target LM head,nnis RMS normalization without a trainable scale, andUℓVℓU\_\{\\ell\}V\_\{\\ell\}has rank 64\. Each readout is fitted with cross\-entropy on 500 GSM8K training questions and selected on another 100, both disjoint from the test data \(AdamW, three epochs, batch size 256, learning rate10−310^\{\-3\}, weight decay10−410^\{\-4\}, 5% warmup, cosine decay\)\.
ForTTevaluated positions, the number of consecutive correct tokens is
KT=∑j=1T∏t=1j𝕀\[y^t=yt\],K\_\{T\}=\\sum\_\{j=1\}^\{T\}\\prod\_\{t=1\}^\{j\}\\mathbb\{I\}\[\\hat\{y\}\_\{t\}=y\_\{t\}\],\(10\)which, unlikeτ\\tau, compares predictions with fixed correct continuations\. We average blocks within each task and then over tasks; panels \(b\) and \(c\) of[Figure2](https://arxiv.org/html/2609.36173#S3.F2)average GSM8K and HumanEval with equal weight\.
### B\.2Prefix Content and Timing
We fix the correct tokenx7x\_\{7\}as a new anchor and fillx1,…,x6x\_\{1\},\\ldots,x\_\{6\}with either draft predictions \(P\) or correct tokens \(G\), keeping earlier context unchanged\. The target model encodes the context up tox6x\_\{6\}, and its fused features enter each draft layer as K/V; layer 1 always receives P features, while layers 2–5 follow the chosen schedule\. We evaluatex8,…,x15x\_\{8\},\\ldots,x\_\{15\}and report the relative gain inK8K\_\{8\}over the all\-P schedule\.[Table5](https://arxiv.org/html/2609.36173#A2.T5)lists all schedules, and[Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(b\) shows the average of the two datasets\.
Table 5:Prefix\-timing conditions\. Schedules list layers 1–5; P/G denotes predicted/correct prefix features\. Gains inK8K\_\{8\}are relative to the all\-P baseline\.
### B\.3Adjacent Predecessor versus Earlier History
For original\-block positions 4, 8, and 12, we use the immediately preceding token as a new anchor; the earlier history and the anchor independently take draft predictions or correct tokens\. The history is encoded by the target model and the predecessor enters the drafter through its input embedding, so each evaluated token is the first candidate of its new block\.[Table6](https://arxiv.org/html/2609.36173#A2.T6)reports the gain from correcting only the predecessor; the L1–L5 points in[Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\) average it over the three positions\. Correcting the predecessor helps in every setting; at position 12 with correct history, accuracy rises from 51\.00% to 93\.50% on GSM8K and from 48\.78% to 91\.31% on HumanEval\.
Table 6:Accuracy gain \(percentage points\) from correcting only the predecessor\. History refers to tokens before the predecessor; positions refer to the original block\.
### B\.4Timing of Predecessor Features
We cache the predecessor’s K/V at every layer in two forward passes, one with the predicted and one with the correct predecessor, and substitute the correct K/V only at selected layers\. The comparison supplies correct features twice: at layer 5 and at one of layers 2–4\. The all\-predicted and all\-correct schedules reproduce the two native forward passes, so the L1–L5 points in[Figure2](https://arxiv.org/html/2609.36173#S3.F2)\(c\) equal the content effect in[SectionB\.3](https://arxiv.org/html/2609.36173#A2.SS3)\. Averaged over the three positions, L2\+L5 improves accuracy over L4\+L5 by 14\.33 and 15\.24 points on GSM8K and HumanEval with correct history \([Table7](https://arxiv.org/html/2609.36173#A2.T7)\), and by 9\.63 and 8\.99 points with predicted history\.
Table 7:L2\+L5 minus L4\+L5 with correct earlier history\. Accuracy differences are in percentage points;K4K\_\{4\}counts at most four consecutive correct candidates, excluding the anchor\.
## Appendix CExperimental Details
Table 8:SGLang throughput \(tokens/s;[Figure5](https://arxiv.org/html/2609.36173#S5.F5)\)\. Green: speedup over AR; bold: best speculative method\.#### Evaluation\.
Each benchmark uses its complete test set \(1,319 GSM8K, 500 MATH\-500, 164 HumanEval, 500 MBPP, 1,055 LCB, 80 MT\-Bench, and 500 Arena\-Hard prompts; MT\-Bench runs both turns\)\. All methods receive identical prompts without a system message and generate at most 2,048 new tokens\. Acceptance length in[Table1](https://arxiv.org/html/2609.36173#S5.T1)is measured in SGLang with eight single\-GPU instances, each serving up to 16 concurrent requests; at temperature one, tokens are sampled from the full target distribution\. Every method proposes 15 draft tokens per round\.τ\\tauis averaged over the responses of each benchmark, and Avg\. is the unweighted mean over benchmarks\.
#### Lossless verification\.
If the first rejection occurs at positionjj, we keepx^1:j−1\\hat\{x\}\_\{1:j\-1\}and sample a replacement from
pjcorr\(v\)=\[pj\(v\)−qdraft,j\(v\)\]\+∑u∈𝒱\[pj\(u\)−qdraft,j\(u\)\]\+,\[a\]\+=max\(a,0\),p\_\{j\}^\{\\mathrm\{corr\}\}\(v\)=\\frac\{\[p\_\{j\}\(v\)\-q\_\{\\mathrm\{draft\},j\}\(v\)\]\_\{\+\}\}\{\\sum\_\{u\\in\\mathcal\{V\}\}\[p\_\{j\}\(u\)\-q\_\{\\mathrm\{draft\},j\}\(u\)\]\_\{\+\}\},\\qquad\[a\]\_\{\+\}=\\max\(a,0\),\(11\)where𝒱\\mathcal\{V\}is the vocabulary; if allMMcandidates are accepted, one additional token is sampled from the target\. Together with[Equation1](https://arxiv.org/html/2609.36173#S2.E1), this preserves the target distribution\.
#### Decoding at temperature one\.
All methods use SGLang’s lossless verification\. DSpark, DFlash2, and DSpine sample their proposals, the latter two from the softmax of their selector or transition\-cache scores over the top\-16 candidates of each position, and apply rejection sampling with these proposal probabilities; DFlash and Domino propose greedily and accept a candidate when it matches the token sampled from the target\.
#### Serving\.
Each instance uses one GPU \(tensor parallelism one, BF16, FlashAttention\-3\([Shah et al\., 2024](https://arxiv.org/html/2609.36173#bib.bib36)\), CUDA graphs\) at temperature zero with at most 2,048 output tokens on the complete GSM8K, MATH\-500, HumanEval, and MBPP test sets, with identical prompts in the same order\. Throughput is the total number of output tokens divided by the wall\-clock time of the whole workload, including prefill and scheduling\. All methods use the same static budget of 15 draft tokens per round\.
#### Supervision targets\.
The variants in[Table3](https://arxiv.org/html/2609.36173#S5.T3)use the same training data, disable injection, and are trained for three epochs \(1,827 steps\); token CE decodes every layer through the shared final RMSNorm and the frozen target LM head\. All variants are evaluated in SGLang at temperature zero\.相似文章
DFlash 2:保持并行起草
DFlash 2 通过并行预测令牌改进推测解码,在最小延迟下实现每次验证通过时输出增加超过20%,并集成到主要推理引擎如 SGLang 和 vLLM 中。
什么是推测性解码?(在paperswithco.de上热门)[R]
推测性解码是一种推理优化技术,它使用快速草稿模型提出未来 token,并由较大模型并行验证,从而提高 LLM 的生成速度。文章强调了它在 Papers with Code 上的热门状态,以及最近的 SGLang 博客文章,该文章介绍了使用 DFlash 模型实现的最先进延迟。
DeLS-Spec: 解耦的长短上下文用于并行推测性草拟
DeLS-Spec通过在DFlash上添加轻量级局部头,将推测性解码中的长上下文和短上下文建模解耦,无需完全重新训练即可实现一致的加速。它仅需要对局部头进行标准的下一个词元预测训练,并在Qwen3基准测试中提高了接受长度。
当并行起草器遇上并行推测解码
DPara是一个并行推测解码框架,通过预计算草稿表示来消除概率性回退,在Qwen3模型上相比自回归解码实现了平均3.21倍到3.52倍的速度提升。
@dzhulgakov:来自 @deepseek_ai 的 DSpark 巧妙融合了多种投机解码思路,将吞吐量提升 1.5 到 5 倍…
来自 DeepSeek AI 的 DSpark 集成了投机解码思路,在生产系统中实现 1.5 到 5 倍的吞吐量提升。本推文从基础开始讲解了 10 个关键思路。