Tag
This paper introduces a diversity-aware, layer-wise scoring method for KV cache eviction in large language models, incorporating attention dispersion and redundancy to improve performance on LongBench datasets.