draft-augmented

Tag

Cards List
#draft-augmented

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

arXiv cs.CL ↗ · 2026-06-25 Cached

Dustin introduces a sparse verification framework for speculative decoding that leverages draft model signals and sparse attention head scoring to overcome the KV cache verification bottleneck, achieving up to 27.85x speedup in self-attention and 9.17x end-to-end decoding speedup on long-context tasks with negligible accuracy loss.

0 favorites 0 likes
← Back to home

Submit Feedback