Tag
The paper investigates budgeted repair methods for stale KV caches in LLM systems after document edits, demonstrating that contiguous edit-local windows efficiently recover performance and are faster than full re-prefill.
This paper proposes residual dominance as a structural explanation for last-item reliance in causal self-attention based sequential recommenders, using prediction-time diagnostics and norm-based analysis to link this behavior to residual addition in transformer models.