Tag
ASPIRE is an asynchronous batched self-speculative decoding framework that enhances long-context LLM inference by enabling independent request scheduling and reducing attention staleness, achieving 1.70-4.58× speedup over baselines.