Tag
The paper introduces GLANCE, a one-pass block drafting method for lossless speculative decoding in vision-language models, achieving up to 2.93x faster generation without changing output.
This preprint introduces DBLast, a dependent block drafter for stochastic speculative decoding, using a low-rank latent mixture over token positions and an acceptance-oriented training objective to improve accepted draft length in higher-entropy decoding regimes. Experiments with Qwen3-4B and Qwen3-8B show consistent improvements over independent block sampling.