stateful-models

Tag

Cards List
#stateful-models

SpecLA: Efficient Speculative Decoding for Linear-Attention Models

arXiv cs.CL · 2026-07-21 Cached

SpecLA proposes a speculative decoding runtime tailored for stateful linear-attention models, achieving up to 1.70x end-to-end speedup over autoregressive decoding on an NVIDIA H100 with a GDN-1.3B target.

0 favorites 0 likes
← Back to home

Submit Feedback