When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

Hugging Face Daily Papers Papers

Summary

Proposes SOLAR, a learning-augmented framework for semantic cache replacement in LLM agents, outperforming classical heuristics by using regret-based modification timing and Bayesian online learning. Achieves constant competitive ratio and significant improvements over FIFO.

LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain largely ad-hoc. We formalize this as an online semantic cache replacement problem with switching costs, where items are matched by embedding similarity and hit quality is continuous rather than binary. Through experiments on two datasets from MemoryBench-Full (LoCoMo, DialSim) with 8 replacement policies, we reveal a surprising finding: classic heuristics (LRU, LFU) consistently underperform the naive FIFO baseline on semantic workloads, due to the absence of temporal locality and frequency concentration. We propose SOLAR, a learning-augmented framework that derives modification timing from regret accumulation (achieving sim17\% modification rate) and content selection from Bayesian online learning over implicit retrieval feedback. We prove SOLAR achieves a constant competitive ratio leq 3, independent of cache size and horizon (vs.\ Ω(K) for FIFO), and eviction regret O(KTlog T), matching the Ω(KT) lower bound up to logarithmic factors. Experiments demonstrate 5--75\% relative improvement over FIFO at tight cache sizes, with a clearly characterized phase transition at the working set boundary. Synthetic experiments with 5000-item pools further reveal an inverted-U relationship between pool size and retrieval quality, justifying capacity constraints as a retrieval noise phenomenon rather than a storage limitation.
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:55 AM

Paper page - When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

Source: https://huggingface.co/papers/2607.00394

Abstract

Research addresses cache management for LLM agents by formalizing semantic cache replacement as an online problem with switching costs, proposing SOLAR—a learning-augmented framework that outperforms traditional heuristics through regret-based modification timing and Bayesian content selection.

LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain largely ad-hoc. We formalize this as anonline semantic cache replacementproblem withswitching costs, where items are matched byembedding similarityand hit quality is continuous rather than binary. Through experiments on two datasets from MemoryBench-Full (LoCoMo, DialSim) with 8 replacement policies, we reveal a surprising finding: classic heuristics (LRU, LFU) consistently underperform the naive FIFO baseline on semantic workloads, due to the absence of temporal locality and frequency concentration. We propose SOLAR, a learning-augmented framework that derives modification timing fromregret accumulation(achieving sim17\% modification rate) and content selection fromBayesian online learningoverimplicit retrieval feedback. We prove SOLAR achieves a constantcompetitive ratioleq 3, independent of cache size and horizon (vs.\ Ω(K) for FIFO), andeviction regretO(KTlog T), matching the Ω(KT) lower bound up to logarithmic factors. Experiments demonstrate 5--75\% relative improvement over FIFO at tight cache sizes, with a clearly characterized phase transition at theworking set boundary. Synthetic experiments with 5000-item pools further reveal an inverted-U relationship between pool size andretrieval quality, justifying capacity constraints as a retrieval noise phenomenon rather than a storage limitation.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2607\.00394

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.00394 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.00394 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.00394 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Selective Memory Retention for Long-Horizon LLM Agents

arXiv cs.AI

This paper presents TraceRetain, a lightweight framework for bounded external memory in frozen LLM agents, demonstrating that selective retention differentiates from cache heuristics primarily when memory streams contain noise, offering task-success and efficiency benefits.