block-parallel

Tag

Cards List
#block-parallel

DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting

arXiv cs.CL · 2026-07-09 Cached

DeLS-Spec decouples long- and short-context modeling in speculative decoding by adding a lightweight local head to DFlash, achieving consistent speedups without full retraining. It requires only standard next-token prediction training for the local head and improves acceptance length on Qwen3 benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback