ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

arXiv cs.AI Papers

Summary

ALOE introduces a semantically addressed low-rank operator for knowledge editing in language models, improving edit scope and precision with high efficacy and locality on standard benchmarks.

arXiv:2609.29269v1 Announce Type: new Abstract: Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates too narrowly memorizes one prompt, while one that activates too broadly disrupts neighboring knowledge. Parametric editors encode this scope implicitly, whereas memory-based editors make the selection explicit but keep it outside the edited model. We propose ALOE (Addressed Low-rank Operator for Editing), which learns semantic addresses from paraphrases and hard same-subject negatives, aligns them with autoregressive hidden states through rollout refinement and gate calibration, and embeds the resulting gated low-rank operator within one MLP layer, so that the deployed model runs in a single forward pass with no external retriever or auxiliary router. Evaluated on CounterFact, ZSRE, and KnowEdit across three 7--8B model families, ALOE achieves efficacy between 0.955 and 0.999 and locality between 0.981 and 1.000; mechanistic analyses confirm that the learned geometry separates competing edits and that calibration suppresses out-of-scope activation. The remaining errors concentrate in paraphrase coverage and write fitting.
Original Article
View Cached Full Text

Cached at: 09/25/26, 09:45 AM

# ALOE: SEMANTICALLY ADDRESSED LOW-RANK OPERATORS FOR KNOWLEDGE EDITING
Source: [https://arxiv.org/html/2609.29269](https://arxiv.org/html/2609.29269)
###### Abstract

Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved\. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual\. An update that activates too narrowly memorizes one prompt, while one that activates too broadly disrupts neighboring knowledge\. Parametric editors encode this scope implicitly, whereas memory\-based editors make the selection explicit but keep it outside the edited model\. We propose ALOE \(Addressed Low\-rank Operator for Editing\), which learns semantic addresses from paraphrases and hard same\-subject negatives, aligns them with autoregressive hidden states through rollout refinement and gate calibration, and embeds the resulting gated low\-rank operator within one MLP layer, so that the deployed model runs in a single forward pass with no external retriever or auxiliary router\. Evaluated on CounterFact, ZSRE, and KnowEdit across three 7–8B model families, ALOE achieves efficacy between 0\.955 and 0\.999 and locality between 0\.981 and 1\.000; mechanistic analyses confirm that the learned geometry separates competing edits and that calibration suppresses out\-of\-scope activation\. The remaining errors concentrate in paraphrase coverage and write fitting\.

###### Index Terms:

knowledge editing, low\-rank operators, language models, model adaptation

††address:1Shanghai Jiao Tong University## 1Introduction

Language models store factual associations in their parameters, including feed\-forward layers and individual neurons\. Knowledge editing aims to revise such associations on request: the edited model should produce the new fact on the original request \(efficacy\), transfer the change to equivalent wordings of that request \(generalization\), and leave unrelated behavior untouched \(locality\)\. All three criteria depend on one choice that is rarely made explicit: where in activation space the update is allowed to act\.

Editing therefore involves two distinct problems\. The first is a*write problem*: specifying the residual that encodes the new fact\. The second is an*address problem*: deciding which hidden states should receive that residual\. A birthplace edit, for example, should cover rewordings of the same question without changing the model’s answer about the person’s employer, and at scale, different subjects that share a relation must receive distinct writes\. An update whose address is too narrow memorizes the training prompt and misses its rewordings; one whose address is too broad activates on inputs it should leave alone \(Fig\.[1](https://arxiv.org/html/2609.29269#S1.F1)\)\. These boundary failures are documented in current editors: an edit perturbs other facts about the same subject\[[12](https://arxiv.org/html/2609.29269#bib.bib15),[3](https://arxiv.org/html/2609.29269#bib.bib23)\], damages knowledge beyond its intended scope\[[19](https://arxiv.org/html/2609.29269#bib.bib25)\], and fails to propagate to facts implied by the edited one\[[1](https://arxiv.org/html/2609.29269#bib.bib14)\]\.

Existing editors handle the address problem in one of two ways\. ROME and MEMIT encode facts through low\-rank MLP updates, while PMET refines the participating states\[[13](https://arxiv.org/html/2609.29269#bib.bib3),[14](https://arxiv.org/html/2609.29269#bib.bib5),[11](https://arxiv.org/html/2609.29269#bib.bib11)\]\. In such an updateΔ​W=U​V⊤\\Delta W=UV^\{\\top\}, which acts on the layer inputh⁡\(x\)h\(x\)of a promptxx, the termV⊤​h​\(x\)V^\{\\top\}h\(x\)already serves as an address for each write column, but its factors are never trained to separate paraphrases from semantic neighbors, and causal localization need not identify the most editable layer\[[7](https://arxiv.org/html/2609.29269#bib.bib9)\]\. SERAC, GRACE, and WISE instead make the selection explicit through a classifier, a codebook, or a side memory\[[16](https://arxiv.org/html/2609.29269#bib.bib4),[6](https://arxiv.org/html/2609.29269#bib.bib6),[20](https://arxiv.org/html/2609.29269#bib.bib17)\], but their deployed systems perform this selection outside the edited weights\.

Figure 1:The address problem\. The same write can miss paraphrases or affect neighboring facts\. Filled cells denote an active write; empty cells denote no update \(schematic\)\.To obtain selection that is both explicitly learned and executed inside the edited model, we proposeALOE\(Addressed Low\-rank Operator for Editing\)\. During construction, asymmetric query/key maps learn a scope representation for each edit from three kinds of examples: paraphrases of the request, locality inputs that should remain unchanged, and same\-subject prompts whose relation differs from the edited one, which serve as hard negatives\. The resulting codes initialize address vectors inside one MLP layer, and these addresses are refined and calibrated on autoregressive rollouts so that they respond to the states the model produces at generation time\. A joint solve then fits the write directions under the calibrated gates\. At inference, the edited model needs no external component, because the address is a trained part of the MLP itself\. Our contributions are as follows\. First, we formulate knowledge editing as a coupled address–write problem and show empirically that the two components fail independently, so each can be measured and diagnosed on its own\. Second, we realize the address as an intrinsic gated low\-rank MLP operator: it learns semantic edit boundaries at construction time and executes them inside the edited layer at inference\. Third, on CounterFact, ZSRE, and KnowEdit across three 7–8B model families,ALOEattains efficacy between 0\.955 and 0\.999 and locality between 0\.981 and 1\.000 on edit streams of 839 to 1,301 facts\.

Figure 2:ALOE learns semantic addresses, refines and calibrates continuous edit gates, and inserts the resulting low\-dimensional residual into one MLP projection\. The edited LLM executes the operator in a standard forward pass\.
## 2Related work

### 2\.1Parametric knowledge editing

Parametric editors differ mainly in how they construct and constrain weight updates\. KnowledgeEditor and MEND learn to transform edit gradients into parameter changes, and MALMEN extends learned editing to large batches\[[2](https://arxiv.org/html/2609.29269#bib.bib1),[15](https://arxiv.org/html/2609.29269#bib.bib2),[17](https://arxiv.org/html/2609.29269#bib.bib10)\]\. ROME instead writes a factual association directly into an MLP through a rank\-one update, and MEMIT extends this construction to many edits at once\[[13](https://arxiv.org/html/2609.29269#bib.bib3),[14](https://arxiv.org/html/2609.29269#bib.bib5)\]; carefully configured fine\-tuning remains a useful reference\. More recent methods such as AlphaEdit constrain updates to preservation or orthogonal subspaces to reduce interference with existing knowledge\[[4](https://arxiv.org/html/2609.29269#bib.bib22),[22](https://arxiv.org/html/2609.29269#bib.bib24)\]\.

These methods decide*how*and*where*new knowledge is written into shared parameters\. The scope of the resulting edit—which inputs should activate it—is induced implicitly by the interaction between the update and the model’s hidden states\.

### 2\.2Selective access and edit scope

A second line of work makes access to edited knowledge explicit\. SERAC retrieves a counterfactual model, GRACE stores edits in discrete key–value adaptors, and WISE routes inputs among memories\[[16](https://arxiv.org/html/2609.29269#bib.bib4),[6](https://arxiv.org/html/2609.29269#bib.bib6),[20](https://arxiv.org/html/2609.29269#bib.bib17)\]\. In\-context editing avoids parameter modification altogether by placing updated facts in demonstrations\[[26](https://arxiv.org/html/2609.29269#bib.bib7)\]\. Locate\-then\-edit and associated\-knowledge methods improve the placement or propagation of edits\. In all of these systems, the selection mechanism lives outside the edited weights—a retriever, a codebook, a router, or the prompt—or targets where an edit should be placed rather than when it should fire\. ALOE learns semantic addresses that gate low\-rank writes directly inside an edited MLP, so selective access becomes part of the deployed model itself\.

The scope of an edit is also central to how editing is evaluated\. CounterFact\+ strengthens tests of specificity\[[8](https://arxiv.org/html/2609.29269#bib.bib8)\], and RippleEdits and ReCoE measure whether an edit propagates to related knowledge\[[1](https://arxiv.org/html/2609.29269#bib.bib14),[9](https://arxiv.org/html/2609.29269#bib.bib16)\]\. Sequential\-edit studies show that poorly controlled updates cause forgetting and degrade general abilities\. These evaluation results point to a shared requirement: an edit should activate for equivalent or relevant queries and remain inactive for nearby but out\-of\-scope knowledge\. We formulate this boundary explicitly as the*address*of an edit and study it separately from the write itself\.

## 3Method

### 3\.1An addressed residual operator

Figure[2](https://arxiv.org/html/2609.29269#S1.F2)illustrates the construction pipeline\. Letfθf\_\{\\theta\}be a frozen transformer, and letℰ=\{\(ti,yi\)\}i=1n\\mathcal\{E\}=\\\{\(t\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{n\}denote the requested edits, wheretit\_\{i\}is the request text andyiy\_\{i\}is its target\. We augment a single MLP down\-projectionW∈ℝdout×dW\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\}, whose input state ish∈ℝdh\\in\\mathbb\{R\}^\{d\}\. ALOE gives each edit two learned vectors: an*address*that decides which hidden states the edit applies to, and a*write direction*that carries the new fact\. The address gates the write direction, so the edit only enters the output where its address activates:

h¯\\displaystyle\\bar\{h\}=h/max⁡\(‖h‖2,ϵ\),\\displaystyle=h/\\max\(\\\|h\\\|\_\{2\},\\epsilon\),g⁡\(h\)\\displaystyle g\(h\)=ϕ⁡\(𝜶⊙\(V⊤​h¯−𝝉\)\),\\displaystyle=\\phi\\\!\\left\(\\boldsymbol\{\\alpha\}\\odot\(V^\{\\top\}\\bar\{h\}\-\\boldsymbol\{\\tau\}\)\\right\),\(1\)FALOE​\(h\)\\displaystyle F\_\{\\mathrm\{ALOE\}\}\(h\)=W​h\+U​g​\(h\)\.\\displaystyle=Wh\+Ug\(h\)\.\(2\)HereV=\[v1,…,vn\]∈ℝd×nV=\[v\_\{1\},\\ldots,v\_\{n\}\]\\in\\mathbb\{R\}^\{d\\times n\}contains the semantic addresses,U=\[u1,…,un\]∈ℝdout×nU=\[u\_\{1\},\\ldots,u\_\{n\}\]\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times n\}contains the write directions, and𝝉,𝜶∈ℝn\\boldsymbol\{\\tau\},\\boldsymbol\{\\alpha\}\\in\\mathbb\{R\}^\{n\}are per\-edit thresholds and temperatures\. Normalizing the input state inh¯\\bar\{h\}makes each gate depend on the direction of the hidden state rather than its magnitude\. The dead\-zone sigmoidϕ⁡\(z\)=max⁡\(σ⁡\(z\)−ϵg,0\)/\(1−ϵg\)\\phi\(z\)=\\max\(\\sigma\(z\)\-\\epsilon\_\{g\},0\)/\(1\-\\epsilon\_\{g\}\)withϵg=10−3\\epsilon\_\{g\}=10^\{\-3\}drives the gates of inactive addresses to exactly zero while retaining continuous activation above threshold, so an edit whose address does not match has no effect on the output\. The operator is therefore state\-dependent: the hidden state at each token determines which write directions enter the output, in contrast to the fixed linear updateW\+U​V⊤W\+UV^\{\\top\}of standard low\-rank editing\.

### 3\.2Learning the semantic boundary

The first construction stage learns where each residual should act\. For editii, the full request serves as the canonical key, so that the address retains the identity of both the subject and the relation\. Each training item contains a paraphrasexi\+x\_\{i\}^\{\+\}of the request, a locality inputxi−x\_\{i\}^\{\-\}that should not activate the edit, and same\-subject promptsri​k−r\_\{ik\}^\{\-\}whose predicates differ fromtit\_\{i\}; because these prompts share the subject with the edit, they are harder negatives than random prompts\.

We learn two distinct linear mapsQ,K∈ℝp×dQ,K\\in\\mathbb\{R\}^\{p\\times d\}and a positive scalecc\. Withki=norm⁡\(K​h​\(ti\)\)k\_\{i\}=\\operatorname\{norm\}\(Kh\(t\_\{i\}\)\)as the key of editii, the construction\-time score between a queryxxand editiiis

s⁡\(x,ti\)=c⁡⟨Q​h​\(x\),ki⟩\.s\(x,t\_\{i\}\)=c\\langle Qh\(x\),k\_\{i\}\\rangle\.\(3\)A high score means the query falls inside the edit’s scope\. The training loss rankss⁡\(xi\+,ti\)s\(x\_\{i\}^\{\+\},t\_\{i\}\)above the scores of wrong\-relation queries and keys, separates different requests within a minibatch, and enforces the locality margins⁡\(xi−,ti\)\+m<s⁡\(xi\+,ti\)s\(x\_\{i\}^\{\-\},t\_\{i\}\)\+m<s\(x\_\{i\}^\{\+\},t\_\{i\}\)\. Orthogonality regularization discourages the keys of different edits from collapsing onto each other, and relation\-balanced minibatches prevent frequent predicates from dominating the loss\. Labels and contrast sets are used only during construction; once the addresses are built, inference relies on the resulting address parameters alone\.

### 3\.3From construction to generation

The metric learned above separates construction prompts, but the deployed address must fire on hidden states encountered during generation, which need not coincide with the states seen at construction time: once the model begins producing its own tokens, its hidden states drift away from those measured on pre\-written prompts\. To bridge this shift, we initialize the address of editiias

vi\(0\)=c​Q⊤​norm⁡\(K​h​\(ti\)\),c\>0,v\_\{i\}^\{\(0\)\}=cQ^\{\\top\}\\operatorname\{norm\}\\\!\\left\(Kh\(t\_\{i\}\)\\right\),\\qquad c\>0,\(4\)wherenorm⁡\(z\)=z/‖z‖2\\operatorname\{norm\}\(z\)=z/\\\|z\\\|\_\{2\}\. Transposition preserves the ordering induced by the learned metric, and the normalization in Eq\. \([1](https://arxiv.org/html/2609.29269#S3.E1)\) removes the magnitude of the hidden state\. We then collect normalized anchor statesH\+H^\{\+\}and negative statesH−H^\{\-\}from left\-padded autoregressive rollouts, in which the model generates continuations of each request and we record the states it passes through, and refineV\(0\)V^\{\(0\)\}on these generation\-time states by

ℒaddr=1n​∑i\[mi−minh∈𝒫i⁡vi⊤​h\+maxh∈𝒩i⁡vi⊤​h\]\+2,\\mathcal\{L\}\_\{\\mathrm\{addr\}\}=\\frac\{1\}\{n\}\\sum\_\{i\}\\left\[m\_\{i\}\-\\min\_\{h\\in\\mathcal\{P\}\_\{i\}\}v\_\{i\}^\{\\top\}h\+\\max\_\{h\\in\\mathcal\{N\}\_\{i\}\}v\_\{i\}^\{\\top\}h\\right\]\_\{\+\}^\{2\},\(5\)for up to 3,000 AdamW steps at learning rate 0\.01, preserving the norms of the addresses\. This loss pushes each address toward its worst\-scoring anchor state and away from its worst\-scoring negative state, so that even the least typical valid phrasing still opens the gate\. Edits with identical input–target pairs share one address support, while conflicting targets remain separate constraints\. Calibration then converts the refined scores into gates\. For a separable slotii, it setsτi\\tau\_\{i\}andαi\\alpha\_\{i\}so that the worst negative falls below the dead zone and the worst positive maps to 0\.9; for a slot whose states cannot be separated, a finite\-temperature midpoint initialized at 8 minimizes a balanced worst\-case classification loss\. After refinement, the metric\-learning components are discarded:QQ,KK, the labels, and the prompts are no longer needed, and onlyVV,𝝉\\boldsymbol\{\\tau\}, and𝜶\\boldsymbol\{\\alpha\}remain in the edited model\.

### 3\.4Fitting writes jointly

Once the address support is fixed, the remaining task is to fit write directions that are compatible with the gates\. LetYYcontain target residuals obtained by independently optimizing each edit for 25 steps, compressing its down\-projection change to rank 16, and evaluating the compressed change on its anchor; columniiofYYis thus the output change that editiishould produce when its gate is fully open\. WithG\+=g⁡\(H\+\)G^\{\+\}=g\(H^\{\+\}\)andG−=g⁡\(H−\)G^\{\-\}=g\(H^\{\-\}\)collecting the gate activations of anchor and negative states, a preservation\-regularized ridge solve initializes all writes jointly:

U0=Y​G\+⁣⊤​\(G\+​G\+⁣⊤\+λ​G−​G−⁣⊤\+μ​I\)−1\.U\_\{0\}=YG^\{\+\\top\}\\\!\\left\(G^\{\+\}G^\{\+\\top\}\+\\lambda G^\{\-\}G^\{\-\\top\}\+\\mu I\\right\)^\{\-1\}\.\(6\)The gates of different edits can overlap, so one edit’s write direction can leak through another edit’s gate; solving for all writes in one system accounts for this cross\-activation, and theG−​G−⁣⊤G^\{\-\}G^\{\-\\top\}term keeps the writes small on states where no edit should fire\. WithVV,𝝉\\boldsymbol\{\\tau\}, and𝜶\\boldsymbol\{\\alpha\}frozen, we then optimizeUUfromU0U\_\{0\}for up to 100 AdamW steps using the negative log\-likelihood of the autoregressive targets\. At deployment, the model computes all gates densely and addsU​g​\(h\)Ug\(h\)toW​hWh, which requiresO⁡\(n⁡\(d\+dout\)\)O\(n\(d\+d\_\{\\mathrm\{out\}\}\)\)storage and compute linear in the number of edits\.

Table 1:Efficacy \(Eff\.\), generalization \(Gen\.\), and locality \(Loc\.\) across base models and benchmarks\. ALOE fits one operator to the complete benchmark stream \(839 CounterFact, 1,266 KnowEdit, and 1,301 ZSRE edits\); all results are means over seeds 0, 42, and 99\. Best and second\-best values areboldedandunderlined, respectively\.

## 4Experiments

### 4\.1Setup

We evaluate end\-to\-end editing on CounterFact\[[13](https://arxiv.org/html/2609.29269#bib.bib3)\], ZSRE\[[10](https://arxiv.org/html/2609.29269#bib.bib26)\], and KnowEdit\[[25](https://arxiv.org/html/2609.29269#bib.bib13)\], covering factual replacement, question\-answer rephrasing, and diverse knowledge domains\. The main comparison uses Llama\-2\-7B, Llama\-3\.1\-8B\-Instruct, and Qwen3\-8B\[[18](https://arxiv.org/html/2609.29269#bib.bib18),[5](https://arxiv.org/html/2609.29269#bib.bib19),[23](https://arxiv.org/html/2609.29269#bib.bib21)\], and the address\-transfer study adds Qwen2\.5\-7B\-Instruct\[[24](https://arxiv.org/html/2609.29269#bib.bib20)\]\. All models are run with seeds 0, 42, and 99, and every result reported in this section is the mean over the three runs\. The comparison includes five established knowledge editors: ROME\[[13](https://arxiv.org/html/2609.29269#bib.bib3)\], MEMIT\[[14](https://arxiv.org/html/2609.29269#bib.bib5)\], MEND\[[15](https://arxiv.org/html/2609.29269#bib.bib2)\], GRACE\[[6](https://arxiv.org/html/2609.29269#bib.bib6)\], and AlphaEdit\[[4](https://arxiv.org/html/2609.29269#bib.bib22)\]\. We report the standard dimensions separately:*efficacy*is exact match on the edited request,*generalization*is exact match on rephrases, and*locality*is normalized exact\-generation stability on out\-of\-scope prompts relative to the unedited model\. Baselines use their published configurations through EasyEdit where supported\[[21](https://arxiv.org/html/2609.29269#bib.bib12)\]\.

### 4\.2End\-to\-end editing

Table[1](https://arxiv.org/html/2609.29269#S3.T1)reports the three metrics separately for each model–benchmark pair\. The baselines split into two failure patterns\. On Llama\-2\-7B and Llama\-3\.1\-8B\-Instruct, the locate\-and\-edit and meta\-learning methods collapse under edit streams of this length: ROME, MEMIT, and MEND lose almost all efficacy on CounterFact, and GRACE retains efficacy only on Llama\-2 while activating on almost no rephrase anywhere\. On Qwen3\-8B these methods survive, but the survivors pay for generalization with locality: ROME’s rephrase accuracy comes with locality at or below 0\.20, and MEMIT and AlphaEdit show the same exchange at milder levels\. AlphaEdit is the strongest baseline overall, yet its profile is uneven across model families—competitive on Llama\-2, it keeps high generalization on Llama\-3 only while locality falls to 0\.33–0\.58, and on Qwen3 both its efficacy and its locality drop well below ALOE’s\. ALOE is the only method whose efficacy stays between 0\.955 and 0\.999 and whose locality stays between 0\.981 and 1\.000 in all nine model–benchmark cells, so its advantage is consistency: no failure cell, on any model family, under streams of up to 1,301 edits\. The cost is equally visible\. Generalization ranges from 0\.217 to 0\.472, below the best baseline cells on each model\.

![Refer to caption](https://arxiv.org/html/2609.29269v1/aloe-internal-mechanism.png)Figure 3:Runtime behavior after 839 CounterFact edits on Llama\-3\.1\-8B\-Instruct\. \(a\) Matched\-slot gates for edits and rephrases, and the maximum gate for locality queries\. \(b\) Responses of the first 32 sampled queries to their associated slots\. \(c\) Paired final\-block locality states under shared PCA, with empirical marginals and a full\-range inset\. PC1/PC2 explain 45\.1%/10\.7%; mean relative state drift is 11\.54% and next\-token agreement is 96\.5%\.
### 4\.3Mechanism analysis

We test the address mechanism on its own, using disjoint CounterFact train, calibration, and development partitions organized by predicate, with development subjects and exact surfaces excluded from training; the development set contains 632 examples from 34 predicates, each query paired with 16 fixed same\-subject wrong\-relation keys\. We measure scope AUC, calibration\-threshold recall, 16\-way same\-subject top\-1, and transferred\-threshold negative FPR against equal\-dimensional raw activation controls, with all runs at seeds 0, 42, and 99\. The learned address beats the raw controls by a wide margin: AUC 0\.9976 against 0\.6365, recall 0\.9910 against 0\.1804, same\-subject top\-1 0\.8576 against 0\.3117, and FPR 0\.0418\. Standard deviations across seeds stay below 0\.006 on every metric, so the selectivity comes from the learned geometry rather than from the hidden states themselves\. The choice of key matters just as much: an entity\-neutral key collapses 632 edits onto 179 unique addresses, while the full request keeps all 632 distinct and raises synthetic\-write top\-1 from 0\.698 to 0\.998\. The same construction objective transfers without retuning to the other three model families, which reach AUC 0\.997–0\.998 and FPR 0\.036–0\.042; their slightly lower same\-subject top\-1 \(0\.830–0\.848\) points to architecture\-dependent separation margins\.

The deployed checkpoint tells the same story from the runtime side and locates the remaining failures\. On 256 fixed CounterFact cases drawn from the 839\-edit model \(Fig\.[3](https://arxiv.org/html/2609.29269#S4.F3)\), the matched slot dominates 98\.4% of original requests, and locality states drift by 11\.54% on average with a median of zero, consistent with the near\-perfect locality scores; the visible weakness is on rephrases, whose mean gate activation is 0\.057 against 0\.860 for original requests, so most rephrasings never open the gate\. Controlled interventions on 256 CounterFact edits, removing one construction stage at a time with the data, layer, and write budget fixed, assign this gap to specific stages: rollout refinement contributes 55\.0 points of efficacy, confirming that the construction metric must be aligned with generation\-time states; calibration contributes 24\.8 points of efficacy and 96\.5 points of locality while slightly reducing generalization, because thresholds suppress false activation but cannot create paraphrase support that the address lacks; the preservation term changes every measure by at most half a point\. Paraphrase coverage is therefore an address problem, visible in the gate traces, while the residual gap between an open gate and a correct generation is a write\-fitting problem, visible in the transfer study where address quality holds but end\-to\-end accuracy does not\.

## 5Conclusion

We formulated knowledge editing as a coupled address–write problem and proposed ALOE, which makes the address intrinsic to an edited MLP\. An asymmetric metric learns the scope of each edit from paraphrases and same\-subject hard negatives, rollout refinement and gate calibration align this scope with generation\-time hidden states, and a joint solve fits the write directions under the calibrated gates\. On CounterFact, ZSRE, and KnowEdit across three 7–8B model families, ALOE attains efficacy between 0\.955 and 0\.999 and locality between 0\.981 and 1\.000 on streams of up to 1,301 edits\. Mechanism analyses show that the learned addresses separate in\-scope states from competing ones well beyond what raw hidden states provide, and that the same geometry transfers to an untuned model family\. Controlled ablations further show that rollout refinement drives most of the efficacy improvement, while gate calibration drives most of the locality improvement\. The remaining errors follow the same decomposition: missed paraphrases are address failures, while an open gate followed by a wrong generation is a write failure\. Generalization remains between 0\.217 and 0\.472 because calibrated gates rarely open for rephrasings far from the training paraphrases, which is the clearest limitation of the current system\. Our evaluation is limited to three English factual benchmarks, 7–8B models, and a single edited layer, with storage growing linearly in the number of edits\.

## References

- \[1\]R\. Cohen, E\. Biran, O\. Yoran, A\. Globerson, and M\. Geva\(2024\)Evaluating the ripple effects of knowledge editing in language models\.Transactions of the Association for Computational Linguistics12,pp\. 283–298\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00644),[Link](https://aclanthology.org/2024.tacl-1.16/)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p2.1)\.
- \[2\]N\. De Cao, W\. Aziz, and I\. Titov\(2021\)Editing factual knowledge in language models\.InProceedings of EMNLP,pp\. 6491–6506\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.522),[Link](https://aclanthology.org/2021.emnlp-main.522/)Cited by:[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1)\.
- \[3\]Z\. Duan, W\. Duan, Z\. Yin, Y\. Shen, S\. Jing, J\. Zhang,et al\.\(2025\)Related knowledge perturbation matters: rethinking multiple pieces of knowledge editing in same\-subject\.InProceedings of NAACL\-HLT: Short Papers,pp\. 363–373\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-short.31),[Link](https://aclanthology.org/2025.naacl-short.31/)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p2.1)\.
- \[4\]J\. Fang, H\. Jiang, K\. Wang, Y\. Ma, J\. Shi, X\. Wang,et al\.\(2025\)AlphaEdit: null\-space constrained knowledge editing for language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HvSytvg3Jh)Cited by:[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[5\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle,et al\.\(2024\)The Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.21783),[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[6\]T\. Hartvigsen, S\. Sankaranarayanan, H\. Palangi, Y\. Kim, and M\. Ghassemi\(2023\)Aging with GRACE: lifelong model editing with discrete key\-value adaptors\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/95b6e2ff961580e03c0a662a63a71812-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[7\]P\. Hase, M\. Bansal, B\. Kim, and A\. Ghandeharioun\(2023\)Does localization inform editing? surprising differences in causality\-based localization vs\. knowledge editing in language models\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Document](https://dx.doi.org/10.52202/075280-0774),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/3927bbdcf0e8d1fa8aa23c26f358a281-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1)\.
- \[8\]J\. Hoelscher\-Obermaier, J\. Persson, E\. Kran, I\. Konstas, and F\. Barez\(2023\)Detecting edit failures in large language models: an improved specificity benchmark\.InFindings of ACL,pp\. 11548–11559\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.733),[Link](https://aclanthology.org/2023.findings-acl.733/)Cited by:[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p2.1)\.
- \[9\]W\. Hua, J\. Guo, M\. Dong, H\. Zhu, P\. Ng, and Z\. Wang\(2024\)Propagation and pitfalls: reasoning\-based assessment of knowledge editing through counterfactual tasks\.InFindings of ACL,pp\. 12503–12525\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.743),[Link](https://aclanthology.org/2024.findings-acl.743/)Cited by:[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p2.1)\.
- \[10\]O\. Levy, M\. Seo, E\. Choi, and L\. Zettlemoyer\(2017\)Zero\-shot relation extraction via reading comprehension\.InProceedings of the 21st Conference on Computational Natural Language Learning,pp\. 333–342\.External Links:[Document](https://dx.doi.org/10.18653/v1/K17-1034),[Link](https://aclanthology.org/K17-1034/)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[11\]X\. Li, S\. Li, S\. Song, J\. Yang, J\. Ma, and J\. Yu\(2024\)PMET: precise model editing in a transformer\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 18564–18572\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v38i17.29818),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29818)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1)\.
- \[12\]J\. Ma, Z\. Ling, N\. Zhang, and J\. Gu\(2024\)Neighboring perturbations of knowledge editing on large language models\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 33839–33854\.External Links:[Link](https://proceedings.mlr.press/v235/ma24h.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p2.1)\.
- \[13\]K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov\(2022\)Locating and editing factual associations in GPT\.InAdvances in Neural Information Processing Systems,Vol\.35\.External Links:[Document](https://dx.doi.org/10.52202/068431-1262),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[14\]K\. Meng, A\. S\. Sharma, A\. Andonian, Y\. Belinkov, and D\. Bau\(2023\)Mass\-editing memory in a transformer\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MkbcAHIYgyS)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[15\]E\. Mitchell, C\. Lin, A\. Bosselut, C\. Finn, and C\. D\. Manning\(2022\)Fast model editing at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0DcZxeWfOPt)Cited by:[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[16\]E\. Mitchell, C\. Lin, A\. Bosselut, C\. D\. Manning, and C\. Finn\(2022\)Memory\-based model editing at scale\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 15817–15831\.External Links:[Link](https://proceedings.mlr.press/v162/mitchell22a.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p1.1)\.
- \[17\]C\. Tan, G\. Zhang, and J\. Fu\(2024\)Massive editing for large language models via meta learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=L6L1CJQ2PE)Cited by:[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1)\.
- \[18\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei,et al\.\(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2307.09288),[Link](https://arxiv.org/abs/2307.09288)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[19\]J\. Wang, Z\. Gu, X\. Zhu, L\. Zhang, H\. Ye, Z\. Xiong,et al\.\(2025\)The missing piece in model editing: a deep dive into the hidden damage brought by model editing\.In2025 IEEE International Conference on Acoustics, Speech, and Signal Processing \(ICASSP\),pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP49660.2025.10890406),[Link](https://doi.org/10.1109/ICASSP49660.2025.10890406)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p2.1)\.
- \[20\]P\. Wang, Z\. Li, N\. Zhang, Z\. Xu, Y\. Yao, Y\. Jiang,et al\.\(2024\)WISE: rethinking the knowledge memory for lifelong model editing of large language models\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Document](https://dx.doi.org/10.52202/079017-1703),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/60960ad78868fce5c165295fbd895060-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2609.29269#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p1.1)\.
- \[21\]P\. Wang, N\. Zhang, B\. Tian, Z\. Xi, Y\. Yao, Z\. Xu,et al\.\(2024\)EasyEdit: an easy\-to\-use knowledge editing framework for large language models\.InProceedings of ACL: System Demonstrations,pp\. 82–93\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-demos.9),[Link](https://aclanthology.org/2024.acl-demos.9/)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[22\]H\. Xu, P\. Lan, E\. Yang, G\. Guo, J\. Zhao, L\. Jiang,et al\.\(2025\)Knowledge decoupling via orthogonal projection for lifelong editing of large language models\.InProceedings of ACL,pp\. 13194–13213\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.646),[Link](https://aclanthology.org/2025.acl-long.646/)Cited by:[§2\.1](https://arxiv.org/html/2609.29269#S2.SS1.p1.1)\.
- \[23\]A\. Yang, A\. Li, B\. Yang, B\. Zhang,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.09388),[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[24\]A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu,et al\.\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2412.15115),[Link](https://arxiv.org/abs/2412.15115)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[25\]N\. Zhang, Y\. Yao, B\. Tian, P\. Wang, S\. Deng, M\. Wang,et al\.\(2024\)A comprehensive study of knowledge editing for large language models\.arXiv preprint arXiv:2401\.01286\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2401.01286),[Link](https://arxiv.org/abs/2401.01286)Cited by:[§4\.1](https://arxiv.org/html/2609.29269#S4.SS1.p1.1)\.
- \[26\]C\. Zheng, L\. Li, Q\. Dong, Y\. Fan, Z\. Wu, J\. Xu, and B\. Chang\(2023\)Can we edit factual knowledge by in\-context learning?\.InProceedings of EMNLP,pp\. 4862–4876\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.296),[Link](https://aclanthology.org/2023.emnlp-main.296/)Cited by:[§2\.2](https://arxiv.org/html/2609.29269#S2.SS2.p1.1)\.

Similar Articles