chunk-reuse

Tag

Cards List
#chunk-reuse

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

arXiv cs.AI · 4d ago Cached

KVBoost is a chunk-level key-value cache reuse system for efficient large language model inference that achieves high cache hit rates and significant speedup in time-to-first-token without quality loss, using dual-hash keying and deviation-guided recomputation.

0 favorites 0 likes
← Back to home

Submit Feedback