@akshay_pachaar: 90% of your KV cache never gets reused. (prompt caching was never meant to fix it) if your system prompt and tool defin…
Summary
CacheBlend, a EuroSys 2025 Best Paper, solves the problem that 90% of KV cache is never reused due to rigid prefix-matching in prompt caching. By selectively recomputing only boundary tokens between documents, it achieves 2-4x faster multi-document processing without quality loss, implemented in the open-source LMCache layer.
View Cached Full Text
Cached at: 07/20/26, 03:33 PM
90% of your KV cache never gets reused.
(prompt caching was never meant to fix it)
if your system prompt and tool definitions are stable, prompt caching is the single highest-leverage optimization available today. cached input tokens get up to 90% cheaper, and hit rates of 60 to 85% are realistic.
but it comes with one rigid rule. the cached portion must be an exact, byte-for-byte prefix of the new request. change a single character in that region and you get a full cache miss.
that rule breaks in three situations you hit constantly:
→ RAG with multiple documents. you cached document A alone and document B alone. a query now needs both. document B’s cached state was computed without any awareness of A, so it’s invalid and gets recomputed from scratch.
→ document order changes. the same three documents appear in a different order across requests. every permutation is a cache miss, even though the content is identical.
→ growing conversation history. each new turn changes everything after the stable prefix, so earlier cached states beyond it become useless.
Alibaba Cloud’s production data shows how bad this gets: 10% of KV cache blocks serve 77% of all hits. the rest sits in storage and never gets reused, because prefix matching won’t allow it.
CacheBlend, a research paper from the LMCache team (EuroSys 2025 Best Paper Award), attacks exactly this. the insight is that in modern transformers, tokens overwhelmingly attend to their own local context. only a small fraction of tokens carry real connections across document boundaries.
so instead of recomputing everything after the first cached document, CacheBlend reuses every document’s cache as-is and selectively recomputes just those few boundary tokens. those are the small orange fixes between documents in the diagram, and they are the entire cost. the result is 2 to 4x faster processing on multi-document queries with no quality loss.
the order problem disappears with it. shuffle the same documents however you like, and every permutation stays cached, where prefix caching recomputes all of them every time. the bottom of the diagram shows that side by side.
that’s the real shift: from caching prefixes to caching knowledge. every document in your knowledge base becomes a reusable cached asset, regardless of what order it appears in or what sits next to it.
CacheBlend ships inside LMCache, the open-source cache management layer that runs outside the inference engine and integrates with vLLM, SGLang, and TensorRT-LLM, on both NVIDIA and AMD GPUs.
check it out on GitHub: https://github.com/LMCache/LMCache
(don’t forget to star )
i wrote the full breakdown of the architecture, including why cache management should never live inside your inference engine. the article is quoted below.
stay tuned for more on this!
LMCache/LMCache
Source: https://github.com/LMCache/LMCache
A KV Cache Management Layer for Scalable LLM Inference
Blog | Documentation | Join Slack | Community Meeting | Roadmap
⭐ If LMCache helps you serve LLMs faster and cheaper, give us a star — it helps more teams discover the project.
Updates
- [2026/05] 🔥 Agentic workload benchmark on AMD MI300X (blog).
- [2026/04] 🔥 LMCache’s new multiprocess (MP) architecture release (blog).
- [2026/03] LMCache at GTC 2026 (post).
- [2026/01] LMCache multi-node P2P CPU memory sharing, from experimental feature to production (blog).
More
- [2025/11] LMCache x CoreWeave accelerate efficient LLM inference for Cohere (blog).
- [2025/10] LMCache joins the PyTorch Foundation and Tensormesh unveiled (blog, PyTorch).
- [2025/09] NVIDIA Dynamo integrates LMCache, accelerating LLM inference (blog).
- [2025/08] 🎉 LMCache hits 5,000+ GitHub stars (blog).
- [2025/08] LMCache supports gpt-oss (20B/120B) on day 1 (blog).
- [2025/07] Get faster LLM inference and cheaper responses with LMCache and Redis (Redis blog).
- [2025/07] LMCache extends its turbo-boost to multimodal models in vLLM V1 (blog).
- [2025/06] LLM Production Stack goes cross-hardware: AMD, Arm and Ascend (blog).
About
LMCache is a KV cache management layer for LLM inference. It turns KV cache from a temporary state into reusable AI-native knowledge that can be stored persistently, reused across multiple serving engines, monitored with an observability stack, and transformed for better generation quality. As a result, LMCache reduces TTFT (time-to-first-token) and improves throughput, especially for long-context agentic, multi-turn conversation, and knowledge-augmented workloads (e.g., RAG).
LMCache is vendor-neutral. It can be used as a KV cache layer for a range of mainstream open-source serving engines, inference frameworks, hardware vendors, storage systems, and infrastructure providers. The vendor neutrality allows users to freely switch between serving engines and storage vendors, while reusing the stored KV caches.
Key features
-
Engine-independent deployment: LMCache, as a standalone daemon process, manages KV cache independently from the inference engine process, so that KV cache will not be lost even if the inference engine crashes (i.e., no fate-sharing with engines).
-
Persistent, tiered KV cache offloading and reuse: Move KV caches out of GPU memory into a tiered storage hierarchy spanning CPU memory, local storage, and remote backends, enabling reuse across requests, sessions, and engine instances to reduce repeated prefill computation and improve TTFT.
-
Production-level KV cache observability: LMCache provides a rich set of KV cache observability metrics, including typical Kubernetes metrics (health monitoring, performance diagnostics), KV-cache-specific metrics (request-level and token-level prefix cache hits, lifecycle, request-level KV cache performance), management metrics (user-specific usage), and more.
-
Pluggable storage and transport backends: Easily integrate remote storage and KV transfer backends through a unified interface, enabling KV cache offloading and sharing across storage providers. Through this interface, LMCache supports storage backends including CPU RAM, local disk (SSD), Redis/Valkey, Mooncake, InfiniStore, S3-compatible object storage, NIXL, and GDS.
-
Non-prefix KV reuse: Extend KV reuse beyond prefix caching by reusing cached KV blocks at any position in the prompt. This leverages CacheBlend to selectively recompute tokens for quality recovery.
-
PD disaggregation and KV transfer: Support KV cache transfer from prefill workers to decode workers over NVLink, RDMA, or TCP through transport layers such as NIXL.
-
Pluggable KV transformation: A simple interface for researchers to write compression, token dropping, and custom serialization through a flexible SERDE interface.
LMCache is becoming an integral layer in the LLM inference ecosystem, with community-driven integration with serving engines, inference frameworks, hardware vendors, storage systems, and infrastructure providers:
Getting Started
To use LMCache, simply install lmcache from your package manager, e.g. pip:
pip install lmcache
For more setup options and examples, see:
Contributing
We welcome and value contributions and collaborations. Join us in improving LMCache. Check out the Contributing Guide or join our Slack community to get started.
Adoption and Partnerships
LMCache has a growing community of developers, researchers, industry adopters, and partners building the next generation of efficient LLM inference systems.
As an independent open-source project, LMCache is becoming the de-facto standard for KV Cache management in LLM inference. Its continued development and community work are supported in part by Tensormesh.
Citation
LMCache builds on research in KV cache management, including cache reuse, offloading, compression, and serving optimization. If you use LMCache in your research, please cite the LMCache paper and related work.
@article{cheng2025lmcache,
title={LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference},
author={Cheng, Yihua and Liu, Yuhan and Yao, Jiayi and An, Yuwei and Chen, Xiaokun and Feng, Shaoting and Huang, Yuyang and Shen, Samuel and Du, Kuntai and Jiang, Junchen},
journal={arXiv preprint arXiv:2510.09665},
year={2025}
}
Related papers
@inproceedings{liu2024cachegen,
title={Cachegen: Kv cache compression and streaming for fast large language model serving},
author={Liu, Yuhan and Li, Hanchen and Cheng, Yihua and Ray, Siddhant and Huang, Yuyang and Zhang, Qizheng and Du, Kuntai and Yao, Jiayi and Lu, Shan and Ananthanarayanan, Ganesh and others},
booktitle={Proceedings of the ACM SIGCOMM 2024 Conference},
pages={38--56},
year={2024}
}
@inproceedings{yao2025cacheblend,
title={Cacheblend: Fast large language model serving for rag with cached knowledge fusion},
author={Yao, Jiayi and Li, Hanchen and Liu, Yuhan and Ray, Siddhant and Cheng, Yihua and Zhang, Qizheng and Du, Kuntai and Lu, Shan and Jiang, Junchen},
booktitle={Proceedings of the twentieth European conference on computer systems},
pages={94--109},
year={2025}
}
License
The LMCache codebase is licensed under Apache License 2.0. See the LICENSE file for details.
Similar Articles
@akshay_pachaar: https://x.com/akshay_pachaar/status/2074502882812952666
A practitioner's guide to KV cache management, introducing the open-source LMCache architecture that cuts input token costs by 90% and speeds up LLM inference by up to 14x by eliminating redundant context processing in agentic workflows.
Probing the Prompt KV Cache: Where It Becomes Dispensable
This paper systematically investigates when and which parts of the prompt KV cache become dispensable during LLM decoding, showing that redundancy primarily involves chat template scaffolding rather than task content, and replacement with neutral filler preserves accuracy.
Explains how prompt caching works in LLMs, using Claude as a case study, detailing the transformer's KV cache mechanism and the cost benefits of caching static prefixes in agentic workflows.
Explains how prompt caching works in LLMs, using Claude as a case study, detailing the transformer's KV cache mechanism and the cost benefits of caching static prefixes in agentic workflows.
Enabling KV Caching of Shared Prefix for Diffusion Language Models
This paper proposes BiCache, a novel KV caching technique for shared prefixes in diffusion language models, which avoids accuracy collapse by dynamically reusing cached keys and values in shallow layers and achieves 36.3%–98.3% throughput improvement.
Prompt Caching In Agents
The article explains how prompt caching works in large language model agents, covering KV cache mechanics, prefill and decode phases, and the impact on latency, cost, and agent design.