runtime-termination

Tag

Cards List
#runtime-termination

ART: Attention Run-time Termination for Efficient Large Language Model Decoding

arXiv cs.CL · 2026-06-02 Cached

This paper proposes ART, a lightweight run-time mechanism that tracks accumulated attention outputs during LLM decoding and terminates unnecessary KV block accesses, achieving 20% higher generation throughput with comparable accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback