Tag
This paper investigates the importance of temporal aggregation and ranking preservation in decoding-time KV cache compression for AI models, introducing InertiaKV methods to improve decode throughput.