How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Hugging Face Daily Papers Papers

Summary

This paper investigates the quantitative limits of parametric memory in LLMs using LoRA as a probe, establishing a power law relationship and introducing a threshold-guided optimization method called MemFT for improved memory performance.

Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adaptation (LoRA) is widely used for such memory updates, existing studies mainly rely on qualitative downstream evaluations, leaving the quantitative capacity limits and underlying dynamics of exact parametric memory largely unexplored. To bridge this gap, we employ LoRA as a controlled memory capacity probe within the latent space to systematically quantify exact parametric memory. We introduce the Parametric Memory Law, a robust power law linking loss reduction Delta L to effective parameters and sequence length. At the token level, fine-grained analysis reveals a deterministic phase transition, demonstrating that a prediction probability of p > 0.5 constitutes a sufficient condition for verbatim recall under greedy decoding. Driven by these insights, we introduce MemFT, a threshold-guided optimization strategy that dynamically redistributes the training budget toward sub-threshold tokens. Empirical evaluations demonstrate that MemFT can enhance memory fidelity and efficiency. Code will be released at https://github.com/zjunlp/ParametricMemoryLaw.
Original Article
View Cached Full Text

Cached at: 05/29/26, 02:59 AM

Paper page - How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Source: https://huggingface.co/papers/2605.30260

Abstract

Research investigates the quantitative limits of parametric memory in large language models using LoRA as a probe, establishing a power law relationship and developing a threshold-guided optimization method for improved memory performance.

Large Language Models(LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. WhileLow-Rank Adaptation(LoRA) is widely used for such memory updates, existing studies mainly rely on qualitative downstream evaluations, leaving the quantitative capacity limits and underlying dynamics of exactparametric memorylargely unexplored. To bridge this gap, we employ LoRA as a controlled memory capacity probe within thelatent spaceto systematically quantify exactparametric memory. We introduce theParametric MemoryLaw, a robustpower lawlinking loss reduction Delta L to effective parameters and sequence length. At the token level, fine-grained analysis reveals a deterministicphase transition, demonstrating that a prediction probability of p > 0.5 constitutes a sufficient condition forverbatim recallundergreedy decoding. Driven by these insights, we introduceMemFT, athreshold-guided optimizationstrategy that dynamically redistributes the training budget toward sub-threshold tokens. Empirical evaluations demonstrate thatMemFTcan enhance memory fidelity and efficiency. Code will be released at https://github.com/zjunlp/ParametricMemoryLaw.

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2605\.30260

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.30260 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.30260 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.30260 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Hugging Face Daily Papers

MemSFT is a research paper proposing to mitigate the alignment tax in LLM fine-tuning by using an external parametric memory that decouples domain specialization from backbone parameter updates, enabling reuse across different LLM sizes while preserving general performance.

Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

arXiv cs.CL

This paper proposes a Mixture of LoRA and Full (MoLF) fine-tuning framework that uses gradient-guided optimizer routing to adaptively switch between LoRA and full fine-tuning. It aims to overcome the structural limitations of relying solely on static adaptation methods by combining the plasticity of full tuning with the regularization of LoRA.