long-context-models

Tag

Cards List
#long-context-models

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

arXiv cs.AI · 2026-09-03 Cached

HeadWiseKV is a training-free framework that compresses KV caches in hybrid long-context language models, reducing GPU memory usage and extending context lengths while maintaining quality.

0 favorites 0 likes
← Back to home

Submit Feedback