为推测解码恢复离策略监督信号

arXiv cs.CL 论文

摘要

本文提出一种基于 rollout 的训练框架(Anchor-Label Relabelling 与 In-Rollout Anchors),用于在离策略语料上训练的推测解码块草稿模型中恢复完整的监督信号,在不改动训练文本的情况下,将贪心解码下的接受长度较 DFlash 最多提升 36.5%。

arXiv:2609.38795v1 Announce Type: new Abstract: Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:45

# Recovering Off-Policy Supervision for Speculative Decoding
Source: [https://arxiv.org/abs/2609.38795](https://arxiv.org/abs/2609.38795)
[View PDF](https://arxiv.org/pdf/2609.38795)

> Abstract:Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off\-policy token invalidates supervision for all subsequent slots in a block\. Existing approaches discard these divergent slots, resulting in severe supervision loss\. To resolve this problem while preserving the training corpus, we propose a rollout\-based training framework that recovers full supervision through two complementary components\. The first component, Anchor\-Label Relabelling \(ALR\), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots\. The second component, In\-Rollout Anchors \(IRA\), places draft blocks directly inside these rollouts to expose the drafter to target\-generated context, reusing precomputed rollout features at no additional target cost\. Across fixed vision\-language and text corpora, our framework increases greedy accepted length by up to 36\.5% over DFlash and consistently outperforms erasing baselines\. Notably, a single epoch of our method surpasses the best erase schedules\. After three epochs, it matches the acceptance length of training on target\-regenerated responses\. These results show that our framework provides an effective and compute\-efficient approach for training speculative drafters on fixed corpora without modifying the original text\. Code is available at[this https URL](https://github.com/js-lee-AI/ALR-IRA)\.

## Submission history

From: Jungseob Lee \[[view email](https://arxiv.org/show-email/4cb22465/2609.38795)\] **\[v1\]**Wed, 30 Sep 2026 02:27:37 UTC \(155 KB\)

相似文章

Draft-OPD:面向推测式草稿模型的在线策略蒸馏

Hugging Face Daily Papers

Draft-OPD 引入在线策略蒸馏,结合目标辅助展开和错误重放,克服了训练用于推测解码的草稿模型时存在的离线到推理不匹配问题,实现了超过5倍的无损加速,相较于EAGLE-3和DFlash分别提升了23%和13%。

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv cs.LG

This paper proposes an adaptive supervised anchoring framework for on-policy self-distillation, addressing the problem of rollout-conditioned signal degradation in language model training. The method separates rollout-conditioned distribution matching from canonical-context supervision, improving task acquisition while preserving general capabilities.

训练扩散模型进行从左到右推测

arXiv cs.CL

本文提出了三种训练时干预方法(位置加权、首次错误焦点损失和链损失),用于在推测解码中将基于扩散的草稿模型与自回归验证对齐,使接受前缀长度提升21-76%,且不增加推理开销。

Speculative Refinement: 一种混合自回归扩散解码策略及其在不同基准测试中的行为表现

arXiv cs.AI

介绍了 Speculative Refinement (SpecRef),一种无需训练的混合解码策略,它通过熵引导的选择性掩码,从自回归草稿中热启动掩码扩散语言模型。在六个基准测试上的评估表明,代码基准测试混淆了结构发现与逻辑正确性,识别出了一种精炼张力现象,并显示评估协议可能产生不同的模型排名。