Recovering Off-Policy Supervision for Speculative Decoding
Summary
This paper proposes a rollout-based training framework (Anchor-Label Relabelling and In-Rollout Anchors) to recover full supervision for speculative decoding block drafters trained on off-policy corpora, boosting greedy accepted length by up to 36.5% over DFlash without modifying the training text.
View Cached Full Text
Cached at: 10/01/26, 09:45 AM
# Recovering Off-Policy Supervision for Speculative Decoding Source: [https://arxiv.org/abs/2609.38795](https://arxiv.org/abs/2609.38795) [View PDF](https://arxiv.org/pdf/2609.38795) > Abstract:Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off\-policy token invalidates supervision for all subsequent slots in a block\. Existing approaches discard these divergent slots, resulting in severe supervision loss\. To resolve this problem while preserving the training corpus, we propose a rollout\-based training framework that recovers full supervision through two complementary components\. The first component, Anchor\-Label Relabelling \(ALR\), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots\. The second component, In\-Rollout Anchors \(IRA\), places draft blocks directly inside these rollouts to expose the drafter to target\-generated context, reusing precomputed rollout features at no additional target cost\. Across fixed vision\-language and text corpora, our framework increases greedy accepted length by up to 36\.5% over DFlash and consistently outperforms erasing baselines\. Notably, a single epoch of our method surpasses the best erase schedules\. After three epochs, it matches the acceptance length of training on target\-regenerated responses\. These results show that our framework provides an effective and compute\-efficient approach for training speculative drafters on fixed corpora without modifying the original text\. Code is available at[this https URL](https://github.com/js-lee-AI/ALR-IRA)\. ## Submission history From: Jungseob Lee \[[view email](https://arxiv.org/show-email/4cb22465/2609.38795)\] **\[v1\]**Wed, 30 Sep 2026 02:27:37 UTC \(155 KB\)
Similar Articles
Draft-OPD: On-Policy Distillation for Speculative Draft Models
Draft-OPD introduces on-policy distillation with target-assisted rollouts and error replay to overcome the offline-to-inference mismatch in training draft models for speculative decoding, achieving over 5x lossless acceleration and improving upon EAGLE-3 and DFlash by 23% and 13% respectively.
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
AdaptiveSpec is a training-free per-step speculative decoding method that adaptively adjusts token verification and draft tree shape to enhance LLM inference throughput, improving performance by up to 56% while maintaining high accuracy across benchmarks.
Adaptive Supervised Anchoring for On-Policy Self-Distillation
This paper proposes an adaptive supervised anchoring framework for on-policy self-distillation, addressing the problem of rollout-conditioned signal degradation in language model training. The method separates rollout-conditioned distribution matching from canonical-context supervision, improving task acquisition while preserving general capabilities.
Teaching Diffusion to Speculate Left-to-Right
This paper proposes three training-time interventions (positional weighting, first-error focal loss, and chain loss) to align diffusion-based draft models with autoregressive verification in speculative decoding, improving accepted prefix length by 21–76% without extra inference cost.
Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks
Introduces Speculative Refinement (SpecRef), a training-free hybrid decoding strategy that warm-starts a masked diffusion language model from an autoregressive draft using entropy-guided selective masking. Evaluated across six benchmarks, it reveals that code benchmarks conflate structural discovery with logical correctness, identifies a refinement tension phenomenon, and shows that evaluation protocols can produce different model rankings.