copy-optimization

Tag

Cards List
#copy-optimization

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

arXiv cs.CL ↗ · 2026-09-18 Cached

SwitchSD is an adaptive framework for speculative decoding in LLMs that uses intrinsic model signals to dynamically switch between neural drafting and context-based copying, achieving up to 15% throughput gains over baselines like EAGLE3.

0 favorites 0 likes
← Back to home

Submit Feedback