Recovering Off-Policy Supervision for Speculative Decoding

arXiv cs.CL Papers

Summary

This paper proposes a rollout-based training framework (Anchor-Label Relabelling and In-Rollout Anchors) to recover full supervision for speculative decoding block drafters trained on off-policy corpora, boosting greedy accepted length by up to 36.5% over DFlash without modifying the training text.

arXiv:2609.38795v1 Announce Type: new Abstract: Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:45 AM

# Recovering Off-Policy Supervision for Speculative Decoding
Source: [https://arxiv.org/abs/2609.38795](https://arxiv.org/abs/2609.38795)
[View PDF](https://arxiv.org/pdf/2609.38795)

> Abstract:Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off\-policy token invalidates supervision for all subsequent slots in a block\. Existing approaches discard these divergent slots, resulting in severe supervision loss\. To resolve this problem while preserving the training corpus, we propose a rollout\-based training framework that recovers full supervision through two complementary components\. The first component, Anchor\-Label Relabelling \(ALR\), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots\. The second component, In\-Rollout Anchors \(IRA\), places draft blocks directly inside these rollouts to expose the drafter to target\-generated context, reusing precomputed rollout features at no additional target cost\. Across fixed vision\-language and text corpora, our framework increases greedy accepted length by up to 36\.5% over DFlash and consistently outperforms erasing baselines\. Notably, a single epoch of our method surpasses the best erase schedules\. After three epochs, it matches the acceptance length of training on target\-regenerated responses\. These results show that our framework provides an effective and compute\-efficient approach for training speculative drafters on fixed corpora without modifying the original text\. Code is available at[this https URL](https://github.com/js-lee-AI/ALR-IRA)\.

## Submission history

From: Jungseob Lee \[[view email](https://arxiv.org/show-email/4cb22465/2609.38795)\] **\[v1\]**Wed, 30 Sep 2026 02:27:37 UTC \(155 KB\)

Similar Articles

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Hugging Face Daily Papers

Draft-OPD introduces on-policy distillation with target-assisted rollouts and error replay to overcome the offline-to-inference mismatch in training draft models for speculative decoding, achieving over 5x lossless acceleration and improving upon EAGLE-3 and DFlash by 23% and 13% respectively.

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv cs.LG

This paper proposes an adaptive supervised anchoring framework for on-policy self-distillation, addressing the problem of rollout-conditioned signal degradation in language model training. The method separates rollout-conditioned distribution matching from canonical-context supervision, improving task acquisition while preserving general capabilities.

Teaching Diffusion to Speculate Left-to-Right

arXiv cs.CL

This paper proposes three training-time interventions (positional weighting, first-error focal loss, and chain loss) to align diffusion-based draft models with autoregressive verification in speculative decoding, improving accepted prefix length by 21–76% without extra inference cost.

Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks

arXiv cs.AI

Introduces Speculative Refinement (SpecRef), a training-free hybrid decoding strategy that warm-starts a masked diffusion language model from an autoregressive draft using entropy-guided selective masking. Evaluated across six benchmarks, it reveals that code benchmarks conflate structural discovery with logical correctness, identifies a refinement tension phenomenon, and shows that evaluation protocols can produce different model rankings.