batched

Tag

Cards List
#batched

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

arXiv cs.LG · 6d ago Cached

ASPIRE is an asynchronous batched self-speculative decoding framework that enhances long-context LLM inference by enabling independent request scheduling and reducing attention staleness, achieving 1.70-4.58× speedup over baselines.

0 favorites 0 likes
← Back to home

Submit Feedback