Learning New Facts with QLoRA: An Acquisition-Retention Frontier

arXiv cs.CL Papers

Summary

This paper explores the trade-off between factual knowledge acquisition and capability retention in language models when using QLoRA, showing that higher-rank QLoRA improves fact learning but may degrade out-of-domain performance.

arXiv:2608.25677v1 Announce Type: new Abstract: Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. We show that this assumption depends strongly on adapter capacity. We study factual acquisition in a controlled OpenStreetMap-derived benchmark where Qwen3-4B must acquire anonymized geographic associations while retaining unrelated capabilities. Comparing full fine-tuning (FFT) with quantized low-rank adaptation (QLoRA) at ranks 8, 16, 32, and 64, we find that rank induces a clear acquisition--retention frontier. Low-rank QLoRA preserves out-of-domain (OOD) performance but acquires fewer facts, whereas higher ranks improve same-fact paraphrase generalization at an increasing cost in performance on unrelated benchmarks. FFT behaves as a conservative baseline: it retains general capabilities well, but does not reach the highest factual-acquisition regime. Distributional, weight-space, and spectral diagnostics mirror this behavioral trade-off, with higher-rank QLoRA moving farther from the pretrained model. A separate math adaptation experiment shows a weaker frontier, suggesting that the effect is most pronounced when adaptation must install new factual associations rather than reinforce skills already supported by pretraining. Code and data are available at https://github.com/zhngstl/new_facts_forgetting.
Original Article
View Cached Full Text

Cached at: 08/27/26, 09:25 AM

# Learning New Facts with QLoRA: An Acquisition-Retention Frontier
Source: [https://arxiv.org/html/2608.25677](https://arxiv.org/html/2608.25677)
Estelle ZhengAffiliation:LORIA, CNRS, FranceAffiliation:Alcatel\-Lucent Enterprise, FranceEmail:[estelle\.zheng@loria\.fr](mailto:)Emmanuel HelbertAffiliation:Alcatel\-Lucent Enterprise, FranceChristophe CerisaraAffiliation:LORIA, CNRS, France

###### Abstract

Parameter\-efficient fine\-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters\. We show that this assumption depends strongly on adapter capacity\. We study factual acquisition in a controlled OpenStreetMap\-derived benchmark where Qwen3\-4B must acquire anonymized geographic associations while retaining unrelated capabilities\. Comparing full fine\-tuning \(FFT\) with quantized low\-rank adaptation \(QLoRA\) at ranks 8, 16, 32, and 64, we find that rank induces a clear acquisition\-\-retention frontier\. Low\-rank QLoRA preserves out\-of\-domain \(OOD\) performance but acquires fewer facts, whereas higher ranks improve same\-fact paraphrase generalization at an increasing cost in performance on unrelated benchmarks\. FFT behaves as a conservative baseline: it retains general capabilities well, but does not reach the highest factual\-acquisition regime\. Distributional, weight\-space, and spectral diagnostics mirror this behavioral trade\-off, with higher\-rank QLoRA moving farther from the pretrained model\. A separate math adaptation experiment shows a weaker frontier, suggesting that the effect is most pronounced when adaptation must install new factual associations rather than reinforce skills already supported by pretraining\.111Code and data are available at[https://github\.com/zhngstl/new\_facts\_forgetting](https://github.com/zhngstl/new_facts_forgetting)\.

## 1Introduction

Pretrained language models \(PLMs\) are often fine\-tuned for new domains, task\-specific skills, or factual knowledge\. Full fine\-tuning \(FFT\) updates all model parameters and can be effective, but it is costly and may degrade performance outside the adaptation distribution\. Parameter\-efficient fine\-tuning \(PEFT\) methods such as low\-rank adaptation \(LoRA\) and its quantized variant QLoRA reduce this cost by freezing pretrained weights and learning low\-rank updates\([Hu et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib17);[Dettmers et al\., 2023](https://arxiv.org/html/2608.25677#bib.bib18)\)\. This restriction is often assumed to improve retention of previous capabilities, but it may also limit what the model can acquire\.

Recent work shows that the relationship between LoRA and FFT is not simply one of efficiency\.[Biderman et al\. \(2024\)](https://arxiv.org/html/2608.25677#bib.bib15)find that LoRA can preserve more out\-of\-domain \(OOD\) behavior partly because it learns less from the target distribution\.[Shuttleworth et al\. \(2025\)](https://arxiv.org/html/2608.25677#bib.bib16)show that LoRA and FFT can reach similar accuracy from different regions of weight space\. Other comparisons study broad adaptation settings such as coding\([Männistö et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib28)\), mathematics\([Biderman et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib15)\), question answering\([Sun et al\., 2023](https://arxiv.org/html/2608.25677#bib.bib27)\), or instruction tuning\([Xin et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib26)\)\. These settings mix several gains: format adaptation, skill reinforcement, domain shift, or new information\. We focus next on factual knowledge acquisition\. In this context, retention is ambiguous unless acquisition is measured at the same time: a low\-capacity adapter may appear safer simply because it has not strongly incorporated the target facts\.

Our setting relates to factual knowledge editing, which modifies specific associations while preserving unrelated behavior\([Meng et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib19);[Mitchell et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib20);[Meng et al\., 2023](https://arxiv.org/html/2608.25677#bib.bib22);[Yang et al\., 2025b](https://arxiv.org/html/2608.25677#bib.bib21)\), and continual learning, which studies the stability–plasticity trade\-off under sequential updates\([Jang et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib32);[Shi et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib33)\)\. However, model\-editing benchmarks often focus on localized modifications of previously known facts, sometimes replacing existing associations, while continual\-learning approaches commonly evaluate sequences of tasks or updates, including PEFT\-based methods that regularize, initialize, or merge adapter subspaces\([Lu et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib31);[Qiao and Mahdavi, 2026](https://arxiv.org/html/2608.25677#bib.bib1)\)\. Our goal is different: we study a single controlled batch\-adaptation stage in which models acquire novel factual associations using standard FFT and QLoRA\. We measure how adaptation capacity affects acquisition, paraphrase generalization, and retention of general LLM capabilities\.

We introduce a factual acquisition benchmark derived from OpenStreetMap \(OSM\) in which models are trained to acquire anonymized geographic associations\. The anonymized entities reduce direct reliance on pretrained world knowledge, while the OSM structure preserves realistic relational dependencies\. We compare FFT with QLoRA ranksr∈\{8,16,32,64\}r\\in\\\{8,16,32,64\\\}and evaluate supervised fact memorization, same\-fact paraphrase generalization, and OOD retention\. We further connect behavioral performance to model\-drift diagnostics, including KL divergence from the base model, RMS\-normalized dense update norms, and SVD\-based spectral changes\. A standard\-LoRA rank sweep on Qwen3\-1\.7B separately tests whether the rank trend persists without quantization\.

Our results show that LoRA rank induces a clear acquisition–retention frontier\. Low\-rank LoRA preserves OOD performance but acquires fewer facts, whereas higher\-rank LoRA improves factual acquisition at the cost of larger OOD degradation\. FFT behaves as a baseline: it retains general capabilities relatively well, but does not reach the highest factual\-acquisition regime observed with higher\-rank LoRA\. The same trend appears in model\-drift diagnostics, where higher\-acquisition LoRA runs move farther away from the pretrained model\.

This paper makes three contributions: \(i\) we introduce a controlled OpenStreetMap\-derived benchmark for factual acquisition, using anonymized entities to reduce direct reliance on pretrained world knowledge; \(ii\) we show that QLoRA rank controls an acquisition–retention frontier for new factual associations; and \(iii\) we connect this behavioral trade\-off to model\-drift diagnostics, showing that stronger factual acquisition is associated with larger distributional and weight\-space shifts from the pretrained model\.

## 2Methodology

Standard fine\-tuning datasets often evaluate broad task adaptation rather than the acquisition of genuinely new facts\. Benchmarks commonly used to evaluate knowledge editing, such as ZsRE[Levy et al\. \(2017\)](https://arxiv.org/html/2608.25677#bib.bib30), CounterFact[Meng et al\. \(2022\)](https://arxiv.org/html/2608.25677#bib.bib19), MQuAKE[Zhong et al\. \(2023\)](https://arxiv.org/html/2608.25677#bib.bib23), and RippleEdits[Cohen et al\. \(2024\)](https://arxiv.org/html/2608.25677#bib.bib24)typically evaluate localized updates to known facts, including counterfactual or outdated associations\. They are complementary to our goal of studying standard adaptation on a batch of novel anonymized associations\. A fully synthetic benchmark could also provide novel facts, but its topology, relation frequencies, and cross\-relation dependencies would have to be chosen by the researcher\. We instead use OpenStreetMap \(OSM\) because it supplies a naturally occurring, internally coherent graph whose structure was generated independently of our experimental hypotheses\. Anonymization then reduces reliance on pretrained lexical knowledge while retaining this non\-uniform relational structure\.

### 2\.1OSM Factual Acquisition Dataset

We derive atomic facts from 14 city\-level OSM extracts, linking entities \(POIs, roads, and cities\) to five relation types: POI category, containing city, nearest road, nearest POI, and road\-length bucket\. The training split contains 1,938 instruction\-style question\-answer examples covering direct queries, paraphrases, locality\-preservation probes, spatial\-compositional questions, and inverse city\-signature examples\. Evaluation uses 900 held\-out examples derived from the same facts but expressed with disjoint surface templates, so performance measures the acquisition of factual associations and their generalization across surface forms rather than prompt memorization\. Because some relations have distinct answer types, the benchmark does not by itself establish that models learn abstract relation semantics\.

To reduce contamination from pretrained world knowledge, we restrict source cities to small cities and replace all entity names with synthetic identifiers \(e\.g\.,C\-TRAIN\-001,POI\-TRAIN\-000001\)\. Dataset details are in Appendix[A](https://arxiv.org/html/2608.25677#A1), and example prompts are in Appendix[D](https://arxiv.org/html/2608.25677#A4)\.

### 2\.2Base\-model prior knowledge diagnostic

Before fine\-tuning, we test whether the base model can already solve the task from prior knowledge or answer\-type biases\. We evaluate both anonymized and non\-anonymized versions of the data as a question\-answering task, where the model is prompted to generate the gold answer\. We report exact\-match \(EM\) generation accuracy and a teacher\-forced gold\-vs\-distractor preference score\. For each example, we sample five distractors from other gold answers in the same split, matching both answer type and relation whenever possible\. The model prefers the gold answer when its average per\-token log\-probability exceeds that of the distractor\. We report the percentage of gold\-preferred pairs and the mean log\-probability margin \(Δ\\Deltalp\)\. More details are in Appendix[C](https://arxiv.org/html/2608.25677#A3)\.

Table 1:Base\-model prior diagnostic\.Anonymization reduces EM accuracy, gold\-answer preference, and log\-probability margins, suggesting that real names activate relevant pretrained information\. The higher paraphrase EM partly reflects its larger share of constrained\-response questions; see Appendix[A\.1](https://arxiv.org/html/2608.25677#A1.SS1)\.Table[1](https://arxiv.org/html/2608.25677#S2.T1)shows that real entity names provide useful semantic cues, while anonymization sharply reduces exact match and answer\-likelihood margins\. Preference scores remain slightly above chance, indicating weak structural or answer\-type biases, but the base model cannot solve the anonymized task directly\. The higher EM on the paraphrase split is partly a response\-format effect\. The paraphrase split contains roughly twice the proportion of yes/no questions, increasing its approximate chance EM from 6\.68% to 10\.36%\. Appendix[A\.1](https://arxiv.org/html/2608.25677#A1.SS1)gives the complete counts and ratios\.

## 3Experimental Setup

### 3\.1Models and adaptation methods

We use Qwen3\-4B\([Yang et al\., 2025a](https://arxiv.org/html/2608.25677#bib.bib2)\)as the base model and compare full fine\-tuning \(FFT\) with QLoRA adapters of rankr∈\{8,16,32,64\}r\\in\\\{8,16,32,64\\\}\. Each training example consists of a question and its gold answer, with the autoregressive loss applied only to answer tokens\. All runs are repeated over five random seeds\. Hyperparameters are reported in Appendix[E](https://arxiv.org/html/2608.25677#A5)\.

To test whether the within\-adapter rank trend persists without quantization, we additionally run standard LoRA on Qwen3\-1\.7B\([Yang et al\., 2025a](https://arxiv.org/html/2608.25677#bib.bib2)\)at ranksr∈\{8,16,32\}r\\in\\\{8,16,32\\\}, using the same OSM task and OOD evaluation suite\. This reduced control changes model scale; within its rank sweep\.

### 3\.2Evaluation axes

We evaluate each adapted model along three axes\.

#### Factual acquisition

We report EM accuracy on two OSM splits\. Training accuracy measures recovery of the supervised facts, while paraphrase accuracy measures same\-fact generalization under held\-out templates disjoint from training\. Because the paraphrase set is derived from training facts, it does not test unseen OSM knowledge; rather, it tests whether the learned association is robust to phrasing variation\.

#### OOD retention

We use LM Evaluation Harness\([Gao et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib3)\)on five benchmarks: HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2608.25677#bib.bib4)\), IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2608.25677#bib.bib5)\), TruthfulQA\([Lin et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib6)\), MMLU\-Redux\-2\.0\([Gema et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib7)\), and BBH\([Suzgun et al\., 2023](https://arxiv.org/html/2608.25677#bib.bib8)\)\. These cover code generation, instruction following, truthfulness, general knowledge, and reasoning\. We define forgetting as the drop in average OOD score relative to the base model:

ΔOOD=OODbase−OODadapted\.\\Delta\_\{\\mathrm\{OOD\}\}=\\mathrm\{OOD\}\_\{\\mathrm\{base\}\}\-\\mathrm\{OOD\}\_\{\\mathrm\{adapted\}\}\.

#### Model\-drift diagnostics

Behavioral accuracy alone does not reveal how acquisition is achieved: two models can reach similar OSM accuracy while differing substantially in how far they move from the pretrained model, with different implications for retention\. Following prior work on LoRA retention and LoRA–FFT weight\-space differences\([Biderman et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib15);[Shuttleworth et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib16)\), we therefore measure drift using KL divergence from the base model\([Shenfeld et al\., 2026](https://arxiv.org/html/2608.25677#bib.bib25)\), teacher\-forced negative log\-likelihood on gold OSM answers, RMS\-normalized dense weight drift, and SVD\-based intruder dimensions\([Glorot and Bengio, 2010](https://arxiv.org/html/2608.25677#bib.bib29);[Shuttleworth et al\., 2025](https://arxiv.org/html/2608.25677#bib.bib16)\)\. For comparability, FFT and QLoRA are analyzed in the same dense update space:Wft−W0W\_\{\\mathrm\{ft\}\}\-W\_\{0\}for FFT andΔ​W=αr​B​A\\Delta W=\\frac\{\\alpha\}\{r\}BAfor QLoRA\. RMS normalization controls for differences in module size\. Full metric definitions are given in Appendix[C](https://arxiv.org/html/2608.25677#A3)\.

Figure 1:OSM paraphrase accuracy against average OOD performance\.Points show final\-checkpoint means and error bars show standard deviations over five seeds\. Higher\-rank QLoRA reaches stronger acquisition but lower retention, while FFT and rank 8 remain closer to the pretrained model\.

## 4Results

Figure 2:Model\-drift diagnostics for different QLoRA ranks\.\(a\) Higher\-rank QLoRA adapters show larger KL divergence from the pretrained model, \(b\) larger effective weight updates, and \(c\) larger spectral shifts under the SVD intruder diagnostic\. Dashed lines show FFT for comparison\. Points show means and error bars show std over five seeds\.### 4\.1QLoRA rank controls the acquisition–retention trade\-off

Figure[1](https://arxiv.org/html/2608.25677#S3.F1)shows that QLoRA rank acts as a plasticity control\. Low rank keeps the model close to the pretrained solution and therefore preserves OOD behavior, but this retention coincides with weaker same\-fact generalization\. Increasing rank allows the model to install the OSM associations more reliably, but moves it onto a lower\-retention part of the frontier\. Rank 64 occupies a high\-plasticity, low\-retention regime: factual accuracy remains high, but unrelated capabilities collapse\. Thus, QLoRA is not uniformly safer than FFT; its behavior depends on where rank places the model on the acquisition–retention frontier\. The per\-benchmark results in Appendix[B](https://arxiv.org/html/2608.25677#A2)show that degradation is broad on HumanEval, IFEval, MMLU\-Redux, and BBH, while TruthfulQA remains comparatively stable\.

#### Standard\-LoRA control\.

The unquantized Qwen3\-1\.7B control shows the same qualitative monotonic trade\-off: paraphrase EM rises from 76% atr=8r=8to 79% atr=16r=16and 86% atr=32r=32, while average OOD performance falls from 57\.0% to 52\.0% and 40\.2%, respectively\. This suggests that quantization is not required for the qualitative rank trend, although this reduced control changes model scale\.

### 4\.2Higher acquisition requires greater adaptation capacity

Endpoint comparisons can conflate adaptation method with achieved task performance: a method may appear to retain more simply because it has acquired fewer target facts\. We therefore compare, for each method and seed, the evaluated checkpoint closest to three target paraphrase accuracies in Table[2](https://arxiv.org/html/2608.25677#S4.T2)\.

FFT and QLoRAr=8r=8retain OOD performance well but do not reach the highest paraphrase accuracy\. Higher\-rank QLoRA configurations achieve stronger paraphrase performance only with larger OOD losses\. This suggests that the apparent robustness of low\-rank adaptation to forgetting is actually partly due to limited plasticity\.

Table 2:Target\-acquisition checkpoint comparison\.For each target paraphrase accuracy, we select the nearest evaluated checkpoint per seed and method\. We report mean achieved paraphrase accuracy and OOD retention as a percentage of base\-model OOD performance\. The table abbreviates QLoRA as QL\.
### 4\.3Model drift is associated with forgetting

Figure[2](https://arxiv.org/html/2608.25677#S4.F2)shows that configurations with stronger OOD degradation also exhibit larger drift from the pretrained model\. Higher\-rank QLoRA checkpoints have larger symmetric KL divergence and larger effective dense update magnitudes\. The strongest forgetting regime, QLoRA r=64, also has the largest SVD intruder excess, indicating a larger change in the leading spectral structure of adapted weight matrices\.

These diagnostics are consistent with the behavioral results\. Stronger OSM acquisition is reflected not only in higher paraphrase accuracy but also in larger distributional and weight\-space shifts\. High\-rank QLoRA therefore appears to install the target facts through more disruptive updates, whereas FFT and low\-rank QLoRA remain closer to the pretrained model\. This association motivates train\-time controls and diagnostics for the trade\-off\.

#### Additional math adaptation comparison\.

Table 3:Math\-adaptation results\.Pass@1 scores and averages are in percent\. OOD averages cover HumanEval, IFEval, TruthfulQA, MMLU\-Redux, and BBH; OOD drop is relative to the base Qwen3\-4B\.We run a separate reasoning experiment on a 94k\-example subset ofOpenR1\-Math\-220k\([Hugging Face, 2025](https://arxiv.org/html/2608.25677#bib.bib9)\)to test whether the OSM trend also appears in a larger skill\-adaptation regime\. We evaluate Pass@1 on MATH\-500\([Hendrycks et al\., 2021](https://arxiv.org/html/2608.25677#bib.bib10)\), AIME’24 and AIME’25\([Mathematical Association of America, 2024](https://arxiv.org/html/2608.25677#bib.bib11)\), AMC’23\([American Mathematics Competitions, 2023](https://arxiv.org/html/2608.25677#bib.bib12)\), Minerva Math\([Lewkowycz et al\., 2022](https://arxiv.org/html/2608.25677#bib.bib13)\), and OlympiadBench\([He et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib14)\)\. OOD degradation uses the same five\-benchmark average as the main experiment, relative to the base Qwen3\-4B; hyperparameters are in Appendix[E\.2](https://arxiv.org/html/2608.25677#A5.SS2)\.

Table[3](https://arxiv.org/html/2608.25677#S4.T3)shows that the OSM frontier does not directly transfer to math adaptation\. FFT and QLoRA obtain nearly identical average math performance: 42\.50 for FFT, 42\.60 for QLoRAr=16r=16, and 42\.03 for QLoRAr=32r=32\. Their OOD drops are also small at 1\.71, 2\.23, and 1\.58 points, respectively\. Math fine\-tuning exposes the model to reasoning traces and solution strategies that may already be supported by pretraining, rather than binding anonymized entities to novel associations\. Consistent with prior task\-adaptation results\([Biderman et al\., 2024](https://arxiv.org/html/2608.25677#bib.bib15)\), the strong rank\-dependent frontier observed on OSM is not evident in this math setting\. In our experiments, it is therefore most pronounced when adaptation installs new factual associations while preserving OOD behavior\.

## 5Conclusion

We studied factual acquisition under FFT and QLoRA using an anonymized OpenStreetMap\-derived benchmark\. Our results show that QLoRA rank controls an acquisition–retention trade\-off: low\-rank adapters preserve general capabilities but acquire fewer facts, while higher ranks improve same\-fact paraphrase generalization at increasing OOD cost\. FFT provides a conservative baseline, retaining general capabilities well but not reaching the highest acquisition regime observed with mid\-rank QLoRA\. Model\-drift diagnostics mirror this pattern: higher\-rank QLoRA produces larger KL divergence, larger effective dense updates, and stronger SVD intruder effects\. Thus, PEFT should not be treated as inherently safe for knowledge injection: adapter rank controls a plasticity trade\-off, determining both how much new factual knowledge is installed and how much pretrained behavior is disturbed\. The unquantized LoRA control suggests that this rank effect does not require QLoRA quantization, while the much weaker math frontier limits our conclusion to the present novel\-association setting rather than fine\-tuning in general\.

## Limitations

#### Benchmark scope\.

Our OSM dataset comprises 1,938 training examples across 14 small cities, so it remains unclear whether the acquisition–retention frontier generalizes to larger or more diverse factual corpora\. Additionally, the use of anonymized synthetic identifiers, while useful for controlling pretrained knowledge, may not fully reflect real\-world knowledge injection scenarios where new facts interact with existing world knowledge in richer and less controlled ways\. Because some relations have distinct answer types, the benchmark establishes the acquisition of question\-conditioned factual associations but does not fully separate entity association from abstract relation learning\. A stronger test would use relations with overlapping answer spaces or deliberately conflicting examples\.

#### Model coverage\.

The main five\-seed experiments are conducted with Qwen3\-4B, while the standard\-LoRA control uses Qwen3\-1\.7B\. The shape of the acquisition–retention frontier may differ for larger models, models with different pretraining data mixtures, or architectures with different weight structures\. Whether the rank\-dependent effects we observe persist at scale remains an open question\.

#### OOD benchmark coverage\.

The five OOD benchmarks used to measure retention \(i\.e\., HumanEval, IFEval, TruthfulQA, MMLU\-Redux, and BBH\) provide a reasonable but not exhaustive proxy for general model capability\. Retention on other dimensions, such as long\-context reasoning or multilingual tasks, is not assessed\.

#### Adaptation\-method coverage\.

The standard\-LoRA control supports the within\-adapter rank effect without quantization, but is limited to one smaller model and ranks 8–32\. The main QLoRA–FFT comparison still differs in quantization and optimization, and the control lacks matched FFT and QLoRA baselines on Qwen3\-1\.7B\. It therefore does not isolate every method\-level difference\.

#### Math experiment scope\.

The math adaptation comparison is limited to two QLoRA ranks \(r∈\{16,32\}r\\in\\\{16,32\\\}\) and a single epoch of training\. The conclusion that FFT and QLoRA behave more similarly in skill\-reinforcement settings therefore rests on a relatively narrow hyperparameter sweep, and a fuller rank ablation analogous to the OSM experiments would strengthen this claim\.

## Ethical Considerations

The benchmark uses public OpenStreetMap records under the ODbL 1\.0 license; full usage and attribution details are provided in Appendix[A\.2](https://arxiv.org/html/2608.25677#A1.SS2)\. Anonymized task instances remove original entity names and coordinates and contain no user\-level traces\. Because the underlying database describes real places, however, anonymization should not be treated as a guarantee against geographic re\-identification\.

## Acknowledgments

This project was provided with computing HPC and storage resources by GENCI at IDRIS thanks to the grant 2025\-AD011011668R5 and 2025\-AD011017250 on the supercomputer Jean Zay\.

## References

- American Mathematics Competitions \(2023\)American Mathematics CompetitionsAmerican mathematics contest 12\.Note:[https://huggingface\.co/datasets/AI\-MO/aimo\-validation\-amc](https://huggingface.co/datasets/AI-MO/aimo-validation-amc)Accessed: 2025\-06\-25Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Bidermanet al\.\(2024\)D\. Biderman, J\. Portes, J\. J\. G\. Ortiz, M\. Paul, P\. Greengard, C\. Jennings, D\. King, S\. Havens, V\. Chiley, J\. Frankle, C\. Blakeney, and J\. P\. CunninghamLoRA learns less and forgets less\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=aloEru2qCG)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px3.p1.1),[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p2.1)\.
- Chenet al\.\(2021\)M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. ZarembaEvaluating large language models trained on code\.External Links:2107\.03374Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.
- Cohenet al\.\(2024\)R\. Cohen, E\. Biran, O\. Yoran, A\. Globerson, and M\. GevaEvaluating the ripple effects of knowledge editing in language models\.Transactions of the Association for Computational Linguistics12,pp\. 283–298\.External Links:[Link](https://aclanthology.org/2024.tacl-1.16/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00644)Cited by:[§2](https://arxiv.org/html/2608.25677#S2.p1.1)\.
- Dettmerset al\.\(2023\)T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. ZettlemoyerQLoRA: efficient finetuning of quantized llms\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 10088–10115\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p1.1)\.
- Gaoet al\.\(2024\)L\. Gao, J\. Tow, B\. Abbasi, S\. Biderman, S\. Black, A\. DiPofi, C\. Foster, L\. Golding, J\. Hsu, A\. Le Noac’h, H\. Li, K\. McDonell, N\. Muennighoff, C\. Ociepa, J\. Phang, L\. Reynolds, H\. Schoelkopf, A\. Skowron, L\. Sutawika, E\. Tang, A\. Thite, B\. Wang, K\. Wang, and A\. ZouThe language model evaluation harness\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.12608602),[Link](https://zenodo.org/records/12608602)Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.
- Gemaet al\.\(2025\)A\. P\. Gema, J\. O\. J\. Leang, G\. Hong, A\. Devoto, A\. C\. M\. Mancino, R\. Saxena, X\. He, Y\. Zhao, X\. Du, M\. R\. Ghasemi Madani, C\. Barale, R\. McHardy, J\. Harris, J\. Kaddour, E\. Van Krieken, and P\. MinerviniAre we done with MMLU?\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 5069–5096\.External Links:[Link](https://aclanthology.org/2025.naacl-long.262/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.262),ISBN 979\-8\-89176\-189\-6Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.
- Glorot and Bengio \(2010\)X\. Glorot and Y\. BengioUnderstanding the difficulty of training deep feedforward neural networks\.InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics,Y\. W\. Teh and M\. Titterington \(Eds\.\),Proceedings of Machine Learning Research, Vol\.9,Chia Laguna Resort, Sardinia, Italy,pp\. 249–256\.External Links:[Link](https://proceedings.mlr.press/v9/glorot10a.html)Cited by:[Appendix C](https://arxiv.org/html/2608.25677#A3.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px3.p1.1)\.
- Heet al\.\(2024\)C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang, J\. Liu, L\. Qi, Z\. Liu, and M\. SunOlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 3828–3850\.External Links:[Link](https://aclanthology.org/2024.acl-long.211/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.211)Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the math dataset\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,J\. Vanschoren and S\. Yeung \(Eds\.\),Vol\.1,pp\.\.External Links:[Link](https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/be83ab3ecd0db773eb2dc1b0a17836a1-Paper-round2.pdf)Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p1.1)\.
- Hugging Face \(2025\)Hugging FaceOpen r1: a fully open reproduction of deepseek\-r1\.External Links:[Link](https://github.com/huggingface/open-r1)Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Janget al\.\(2022\)J\. Jang, S\. Ye, S\. Yang, J\. Shin, J\. Han, G\. Kim, S\. J\. Choi, and M\. SeoTowards continual knowledge learning of language models\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Levyet al\.\(2017\)O\. Levy, M\. Seo, E\. Choi, and L\. ZettlemoyerZero\-shot relation extraction via reading comprehension\.InProceedings of the 21st Conference on Computational Natural Language Learning \(CoNLL 2017\),R\. Levy and L\. Specia \(Eds\.\),Vancouver, Canada,pp\. 333–342\.External Links:[Link](https://aclanthology.org/K17-1034/),[Document](https://dx.doi.org/10.18653/v1/K17-1034)Cited by:[§2](https://arxiv.org/html/2608.25677#S2.p1.1)\.
- Lewkowyczet al\.\(2022\)A\. Lewkowycz, A\. J\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, Y\. Wu, B\. Neyshabur, G\. Gur\-Ari, and V\. MisraSolving quantitative reasoning problems with language models\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/18abbeef8cfe9203fdf9053c9c4fe191-Abstract-Conference.html)Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 3214–3252\.External Links:[Link](https://aclanthology.org/2022.acl-long.229/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.229)Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.
- Luet al\.\(2025\)Y\. Lu, B\. Qian, C\. Yuan, H\. Jiang, and X\. WangControlled low\-rank adaptation with subspace regularization for continued training on large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 19165–19181\.External Links:[Link](https://aclanthology.org/2025.acl-long.940/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.940),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Männistöet al\.\(2025\)J\. Männistö, J\. Attieh, and J\. TiedemannA comparative study of PEFT methods for python code generation\.InProceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies \(NoDaLiDa/Baltic\-HLT 2025\),R\. Johansson and S\. Stymne \(Eds\.\),Tallinn, Estonia,pp\. 390–396\.External Links:[Link](https://aclanthology.org/2025.nodalida-1.42/),ISBN 978\-9908\-53\-109\-0Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p2.1)\.
- Mathematical Association of America \(2024\)Mathematical Association of AmericaAmerican invitational mathematics examination\.Note:[https://artofproblemsolving\.com/wiki/index\.php?title=AIME\_Problems\_and\_Solutions](https://artofproblemsolving.com/wiki/index.php?title=AIME_Problems_and_Solutions)Accessed: 2025\-06\-25Cited by:[§4\.3](https://arxiv.org/html/2608.25677#S4.SS3.SSS0.Px1.p1.1)\.
- Menget al\.\(2022\)K\. Meng, D\. Bau, A\. Andonian, and Y\. BelinkovLocating and editing factual associations in gpt\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 17359–17372\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/6f1d43d5a82a37e89b0665b33bf3a182-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1),[§2](https://arxiv.org/html/2608.25677#S2.p1.1)\.
- Menget al\.\(2023\)K\. Meng, A\. S\. Sharma, A\. J\. Andonian, Y\. Belinkov, and D\. BauMass\-editing memory in a transformer\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MkbcAHIYgyS)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Mitchellet al\.\(2022\)E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, and C\. D\. ManningFast model editing at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0DcZxeWfOPt)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Qiao and Mahdavi \(2026\)F\. Qiao and M\. MahdaviMerge before forget: a single loRA continual learning via continual merging\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=i1Rj7yU6eF)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Shenfeldet al\.\(2026\)I\. Shenfeld, J\. Pari, and P\. AgrawalRL’s razor: why online reinforcement learning forgets less\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=7HNRYT4V44)Cited by:[Appendix C](https://arxiv.org/html/2608.25677#A3.SS0.SSS0.Px1.p1.2),[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px3.p1.1)\.
- Shiet al\.\(2025\)H\. Shi, Z\. Xu, H\. Wang, W\. Qin, W\. Wang, Y\. Wang, Z\. Wang, S\. Ebrahimi, and H\. WangContinual learning of large language models: a comprehensive survey\.ACM Comput\. Surv\.58\(5\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3735633),[Document](https://dx.doi.org/10.1145/3735633)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Shuttleworthet al\.\(2025\)R\. Shuttleworth, J\. Andreas, A\. Torralba, and P\. SharmaLoRA vs full fine\-tuning: an illusion of equivalence\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 174627–174662\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/ff541950d1e885af90f523571564a401-Paper-Conference.pdf)Cited by:[Appendix C](https://arxiv.org/html/2608.25677#A3.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.25677#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px3.p1.1)\.
- Sunet al\.\(2023\)X\. Sun, Y\. Ji, B\. Ma, and X\. LiA comparative study between full\-parameter and lora\-based fine\-tuning on chinese instruction data for instruction following large language model\.External Links:2304\.08109,[Link](https://arxiv.org/abs/2304.08109)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p2.1)\.
- Suzgunet al\.\(2023\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. Le, E\. Chi, D\. Zhou, and J\. WeiChallenging BIG\-bench tasks and whether chain\-of\-thought can solve them\.InFindings of the Association for Computational Linguistics: ACL 2023,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 13003–13051\.External Links:[Link](https://aclanthology.org/2023.findings-acl.824/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.824)Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.
- Xinet al\.\(2024\)C\. Xin, Y\. Lu, H\. Lin, S\. Zhou, H\. Zhu, W\. Wang, Z\. Liu, X\. Han, and L\. SunBeyond full fine\-tuning: harnessing the power of LoRA for multi\-task instruction tuning\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 2307–2317\.External Links:[Link](https://aclanthology.org/2024.lrec-main.206/)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p2.1)\.
- Yanget al\.\(2025a\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.Technical reportQwen Team\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.1](https://arxiv.org/html/2608.25677#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.25677#S3.SS1.p2.1)\.
- Yanget al\.\(2025b\)W\. Yang, F\. Sun, R\. Tang, H\. Zang, D\. Su, Q\. Cao, J\. Wang, H\. Shen, and X\. ChengFine\-tuning done right in model editing\.InSocially Responsible and Trustworthy Foundation Models at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=lJFPwtobkG)Cited by:[§1](https://arxiv.org/html/2608.25677#S1.p3.1)\.
- Zhonget al\.\(2023\)Z\. Zhong, Z\. Wu, C\. Manning, C\. Potts, and D\. ChenMQuAKE: assessing knowledge editing in language models via multi\-hop questions\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 15686–15702\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.971/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.971)Cited by:[§2](https://arxiv.org/html/2608.25677#S2.p1.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.External Links:2311\.07911,[Link](https://arxiv.org/abs/2311.07911)Cited by:[§3\.2](https://arxiv.org/html/2608.25677#S3.SS2.SSS0.Px2.p1.1)\.

## Appendix ADataset Construction

We construct the dataset from 14 city\-level OpenStreetMap \(OSM\) extracts\. The training split contains 1,938 instruction examples\. We retain points of interest \(POIs\) and roads with valid, locally unique names, and derive five atomic relation types: POI category, containing city, nearest road, nearest POI, and road\-length bucket\.

The training data combine direct fact queries, paraphrases of the same facts, locality\-preservation probes, spatial\-compositional questions, and inverse city\- signature examples\. Spatial examples include four\-way nearest\-POI selection and balanced yes/no road\-intersection predicates\. Examples are sampled with fixed seeds and relation\-balanced quotas to reduce dominance by common POI categories\.

For evaluation, we use a held\-out paraphrase set of 900 examples constructed from facts represented in the training data\. These evaluation prompts use disjoint lookup, slot\-query, and predicate templates, so they test whether the model recalls the learned factual associations under different surface forms rather than memorizing exact training prompts\.

Since current LLMs might have some prior knowledge of popular global cities, we focus on smaller cities with populations between 5,000 and 80,000\. To further reduce the influence of prior knowledge, all names of cities, POIs, and roads are replaced by synthetic identifiers such asC\-TRAIN\-001,POI\-TRAIN\-000001, andROAD\-TRAIN\-000001\. The anonymized task instances contain no source coordinates or user\-level data\. Representative examples appear in Appendix[D](https://arxiv.org/html/2608.25677#A4)\.

### A\.1Response\-format composition

The train and paraphrase splits differ in their proportions of constrained responses\. In particular, yes/no questions make up 6\.2% of the training split but 13\.3% of the paraphrase split\. Treating open\-ended exact\-match chance as negligible, four\-choice chance as 25%, and yes/no chance as 50%, this raises approximate chance EM from 6\.68% to 10\.36% and partly explains the base\-model difference in Table[1](https://arxiv.org/html/2608.25677#S2.T1)\.

Table 4:Response\-format composition\.Counts and within\-split ratios for training and paraphrase splits, with approximate chance EM for each split\.
### A\.2OpenStreetMap usage and license

We use OSM database records and geometries—not rendered map tiles—to select named POIs and roads, determine city membership, compute nearest\-neighbor and intersection relations, and bucket road lengths before anonymization\. The source data are[© OpenStreetMap contributors](https://www.openstreetmap.org/copyright), available under the Open Data Commons Open Database License \(ODbL\) 1\.0\.

## Appendix BPer\-benchmark OOD Results at Final Checkpoints

Table 5:Per\-benchmark OOD scores at the final checkpoint \(mean±\\pmstandard deviation over five seeds\)\.Degradation is broad on HumanEval, IFEval, MMLU\-Redux, and BBH; TruthfulQA is comparatively stable\.The final\-checkpoint task\-level results complement Figure[1](https://arxiv.org/html/2608.25677#S3.F1)and show that the average OOD degradation is not driven by a single benchmark\. HumanEval, IFEval, MMLU\-Redux, and BBH decline with increasing QLoRA rank, whereas TruthfulQA remains comparatively stable\.

## Appendix CDetails on metrics

#### Symmetric KL\.

Letp0\(⋅∣x<t\)p\_\{0\}\(\\cdot\\mid x\_\{<t\}\)denote the next\-token distribution of the pretrained base model andpθ\(⋅∣x<t\)p\_\{\\theta\}\(\\cdot\\mid x\_\{<t\}\)the corresponding distribution of the adapted checkpoint\. We compute token\-level KL divergences under teacher forcing, excluding padding positions\. The reported symmetric KL is

Dsym\(p0,pθ\)=12\[DKL\(p0∥pθ\)\+DKL\(pθ∥p0\)\]D\_\{\\mathrm\{sym\}\}\(p\_\{0\},p\_\{\\theta\}\)=\\tfrac\{1\}\{2\}\\\!\\left\[D\_\{\\mathrm\{KL\}\}\(p\_\{0\}\\,\\\|\\,p\_\{\\theta\}\)\+D\_\{\\mathrm\{KL\}\}\(p\_\{\\theta\}\\,\\\|\\,p\_\{0\}\)\\right\]

averaged over all non\-padding tokens and then over batches\. Instead of using the standard KL that can be dominated by low\-probability tokens, the symmetric KL emphasizes differences in high\-probability regions of the distribution, which are more likely to reflect changes in model behavior\. Symmetrization treats each model in turn as the reference distribution and captures changes in both directions\. This metric is inspired by[Shenfeld et al\. \(2026\)](https://arxiv.org/html/2608.25677#bib.bib25)on distribution shifts\.

#### Dense RMS drift\.

To compare weight\-space drift between FFT and QLoRA, we use the root mean square \(RMS\) of the effective dense update, following the scale normalization used in weight\-initialization analyses\([Glorot and Bengio, 2010](https://arxiv.org/html/2608.25677#bib.bib29)\)\. For FFT, the update of a selected linear module isΔ​W=Wθ−W0\\Delta W=W\_\{\\theta\}\-W\_\{0\}\. For QLoRA, the effective merged update is

Δ​W=αr​B​A,\\Delta W=\\frac\{\\alpha\}\{r\}BA,whereAAandBBare the LoRA factors,rris the adapter rank, andα\\alphais the LoRA scaling parameter\. For a module withdout×dind\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}dense shape, the module RMS drift is

RMS⁡\(Δ​W\)=‖Δ​W‖F2dout​din\.\\mathrm\{RMS\}\(\\Delta W\)=\\sqrt\{\\frac\{\\\|\\Delta W\\\|\_\{F\}^\{2\}\}\{d\_\{\\mathrm\{out\}\}d\_\{\\mathrm\{in\}\}\}\}\.The global dense RMS drift reported in the figures is the same quantity after summing‖Δ​W‖F2\\\|\\Delta W\\\|\_\{F\}^\{2\}and the dense parameter counts over all selected linear modules:

DRMS=∑m‖Δ​Wm‖F2∑mdout,m​din,m\.D\_\{\\mathrm\{RMS\}\}=\\sqrt\{\\frac\{\\sum\_\{m\}\\\|\\Delta W\_\{m\}\\\|\_\{F\}^\{2\}\}\{\\sum\_\{m\}d\_\{\\mathrm\{out\},m\}d\_\{\\mathrm\{in\},m\}\}\}\.

#### SVD intruder dimensions\.

The SVD diagnostic follows the intruder\-dimension construction of[Shuttleworth et al\. \(2025\)](https://arxiv.org/html/2608.25677#bib.bib16)\. For each selected linear module, we compute the topkkleft singular vectors of the adapted weight matrix and compare each of them to the topKKleft singular vectors of the corresponding pretrained base weight\. In our implementation, the defaults arek=10k=10andK=64K=64\. For an adapted singular vectoruiθu\_\{i\}^\{\\theta\}, define its best alignment with the selected base singular vectors as

ci=max1≤j≤K⁡\|⟨uiθ,uj0⟩\|\.c\_\{i\}=\\max\_\{1\\leq j\\leq K\}\|\\langle u\_\{i\}^\{\\theta\},u\_\{j\}^\{0\}\\rangle\|\.For a thresholdϵ\\epsilon, the vector is counted as an intruder whenci<ϵc\_\{i\}<\\epsilon\. The diagnostic summary reports the intruder rate,

IntruderRateϵ=\#⁡\{\(m,i\):cm,i<ϵ\}\#​\{\(m,i\)\},\\mathrm\{IntruderRate\}\_\{\\epsilon\}=\\frac\{\\\#\\\{\(m,i\):c\_\{m,i\}<\\epsilon\\\}\}\{\\\#\\\{\(m,i\)\\\}\},over all selected modules and top adapted singular vectors\. We use the intruder rate atϵ=0\.8\\epsilon=0\.8as the main SVD diagnostic\. To emphasize rank\-dependent excess beyond the FFT baseline, the plotted SVD quantity is

IntruderExcess\\displaystyle\\mathrm\{IntruderExcess\}=IntruderRateϵ=0\.8method\\displaystyle=\\mathrm\{IntruderRate\}^\{\\mathrm\{method\}\}\_\{\\epsilon=0\.8\}−IntruderRateϵ=0\.8FFT,\\displaystyle\-\\mathrm\{IntruderRate\}^\{\\mathrm\{FFT\}\}\_\{\\epsilon=0\.8\},matched by seed and closest checkpoint step\.

#### Answer log\-probability and distractor margin\.

For OSM answer\-likelihood diagnostics, we score only the answer continuation tokens under teacher forcing\. Given a promptqqand answeraa, the script forms the concatenated sequence\[q,a\]\[q,a\], masks out prompt tokens, and reports the average answer log\-probability

log⁡pθ​\(a∣q\)¯=1\|a\|​∑t∈alog⁡pθ​\(at∣q,a<t\)\.\\overline\{\\log p\_\{\\theta\}\(a\\mid q\)\}=\\frac\{1\}\{\|a\|\}\\sum\_\{t\\in a\}\\log p\_\{\\theta\}\(a\_\{t\}\\mid q,a\_\{<t\}\)\.The negative log\-likelihood is the negative of this average\. For the gold\-vs\-distractor diagnostic, distractor answers are sampled from examples in the same split, matching both relation and answer type whenever possible\. The reported margin is the difference between the average log\-probability of the gold answer and that of the sampled distractor; a positive margin means the model assigns higher teacher\-forced likelihood to the gold answer\.

## Appendix DAdditional dataset examples

Below are representative anonymized examples from the training and held\-out paraphrase validation splits\.

#### Training examples\.

1. 1\.Atomic fact\.Question:InC\-TRAIN\-001, what type of place isPOI\- TRAIN\-002699? Answer:AMENITY\-restaurant
2. 2\.Nearest POI\.Question:InC\-TRAIN\-001, which POI is nearest toPOI\- TRAIN\-001802? Answer:POI\-TRAIN\-001425
3. 3\.Road length bucket\.Question:InC\-TRAIN\-002, which length bucket applies toROAD\-TRAIN\-027122? Answer:LENGTH\-100\-200M
4. 4\.Spatial multiple choice\.Question:InC\-TRAIN\-001, which POI is closest toPOI\-TRAIN\-002343:POI\-TRAIN\-000318,POI\-TRAIN\-002124,POI\-TRAIN\-000864,POI\-TRAIN\-002699? Answer:POI\-TRAIN\-002699
5. 5\.Inverse city signature\.Question:Which city alias matches this local OSM signature? > POI\-TRAIN\-002699is aAMENITY\-restaurant\. POI\-TRAIN\-000340is closest toPOI\-TRAIN\-000682\. POI\-TRAIN\-001425appears in the same city asPOI\-TRAIN\-001802\. Answer:C\-TRAIN\-001

#### Held\-out paraphrase validation examples\.

1. 1\.Slot\-style category query\.Question:Snapshot slot query→\\rightarrowcity:C\-TRAIN\-001; key:POI\-TRAIN\-001463; slot: place\_type\. Answer:AMENITY\-school
2. 2\.Nearest\-road lookup\.Question:Map the pair \(C\-TRAIN\-002,POI\-TRAIN\-001715\) to its nearest road\. Answer:ROAD\-TRAIN\-019865
3. 3\.Road graph predicate\.Question:Evaluate this OSM road\-graph predicate for city=C\- TRAIN\-002: intersects\(ROAD\-TRAIN\-030210,ROAD\-TRAIN\-003453\)\. Return yes or no\. Answer:yes
4. 4\.Paraphrased road\-length query\.Question:Complete this fact: road\_length\_bucket\[C\-TRAIN\-002\] \[ROAD\-TRAIN\-017282\] = Answer:LENGTH\-050\-100M
5. 5\.Validation multiple choice\.Question:OSM relation lookup; city=C\-TRAIN\-009; relation=nearest\_poi; query=POI\-TRAIN\-001471; choices=\[POI\-TRAIN\-000156,POI\-TRAIN\-000138,POI\-TRAIN\-002766,POI\-TRAIN\-002758\]\. Return the matching choice only\. Answer:POI\-TRAIN\-000156

## Appendix EHyperparameters

We report the main hyperparameters for the OSM and math fine\-tuning experiments\.

### E\.1OpenStreetMap task

We run a small sweep over\{2×10−5,5×10−5,2×10−4\}\\\{2\\times 10^\{\-5\},5\\times 10^\{\-5\},2\\times 10^\{\-4\}\\\}for QLoRA and\{2×10−5,2×10−4\}\\\{2\\times 10^\{\-5\},2\\times 10^\{\-4\}\\\}for FFT\. We select the best learning rate for each method based on the lowest training loss\.

Table 6:Hyperparameters for the main OpenStreetMap fine\-tuning experiments\.
### E\.2Math task

We first fine\-tune the full model with the same learning rate as in the OSM experiment\. We then run a small sweep over\{1×10−5,2×10−5\}\\\{1\\times 10^\{\-5\},2\\times 10^\{\-5\}\\\}for QLoRA and select the learning rate with the lowest training loss after one epoch\.

Table 7:Hyperparameters for the additional math adaptation experiments\.

Similar Articles

Can a Language Model Learn Facts Continually in Its Weights?

arXiv cs.CL

This paper investigates whether language models can learn new facts in their weights through continual learning. Using invented facts and sequential writes into Qwen3 models, it finds that training data breadth determines knowledge type and retention: bare-statement facts are quickly forgotten (1% accuracy after 20 writes), while facts learned from diverse restatements retain 46% accuracy. Forgotten facts are not erased but become behaviorally inaccessible due to later writes redirecting questions, and context remains the reliable channel for fact composition and survival.

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs

arXiv cs.CL

This paper investigates whether reinforcement learning can improve the direct recall of parametric knowledge in LLMs beyond reasoning tasks. It demonstrates that RL with binary rewards yields significant gains in factual QA benchmarks by redistributing probability mass to unlock latent knowledge rather than acquiring new facts.