Erased, but Not Gone: Output Forgetting Is Not True Forgetting

arXiv cs.LG Papers

Summary

This paper argues that standard output-level evaluations of machine unlearning overestimate success, showing that methods can appear successful at the output layer while retaining structured representation-level discrepancies relative to retrained models. The authors propose retraining-consistent representation forgetting as a stronger evaluative lens.

arXiv:2606.25001v1 Announce Type: new Abstract: Machine unlearning (MU) is commonly judged by output forgetting, such as low forget-set accuracy or reduced logit-level membership inference. But if output-level success can coexist with retraining-inconsistent residuals in representation space, what kind of forgetting are current evaluations actually certifying? We study this question through retraining-consistent representation forgetting, using the retrained model (i.e., trained from scratch without the forget data) as an operational reference for correct forgetting. Across multiple unlearning methods, datasets, and models, our theoretical analysis and empirical results show that standard output-level evaluation can systematically overestimate the success of unlearning. Under this stronger lens, current methods often appear forgotten at the output layer while exhibiting a structured mismatch relative to retraining. They partially align with retraining on forget samples, remain more inconsistent on retain samples, and leave residual discrepancy concentrated along retraining-related directions rather than diffuse in representation space. This structured mismatch is characterized by forget/retain asymmetry, directional mismatch, and concentrated residuals along retraining-related directions. These results suggest that current MU is often evaluated for apparent forgetting rather than retraining-consistent forgetting. More broadly, retraining reveals what output forgetting hides.
Original Article
View Cached Full Text

Cached at: 06/25/26, 05:11 AM

# Output Forgetting Is Not True Forgetting
Source: [https://arxiv.org/html/2606.25001](https://arxiv.org/html/2606.25001)
## Erased, but Not Gone: Output Forgetting Is Not True Forgetting

Teresa Pui Yee Yong1, Win Kent Ong1, Chee Seng Chan1,2 1Universiti Malaya, Kuala Lumpur, Malaysia 2VinUniversity, Hanoi, Vietnam

\(April 2026\)

###### Abstract

Machine unlearning \(MU\) is commonly judged by*output forgetting*, such as low forget\-set accuracy or reduced logit\-level membership inference\. But if output\-level success can coexist with retraining\-inconsistent residuals in representation space, what kind of forgetting are current evaluations actually certifying? We study this question through*retraining\-consistent representation forgetting*, using the retrained model \(i\.e\.,trained from scratch without the forget data\) as an operational reference for correct forgetting\. Across multiple unlearning methods, datasets, and models, our theoretical analysis and empirical results show that standard output\-level evaluation can systematically overestimate the success of unlearning\. Under this stronger lens, current methods often appear forgotten at the output layer while exhibiting a structured mismatch relative to retraining\. They partially align with retraining on forget samples, remain more inconsistent on retain samples, and leave residual discrepancy concentrated along retraining\-related directions rather than diffuse in representation space\. This structured mismatch is characterized by forget/retain asymmetry, directional mismatch, and concentrated residuals along retraining\-related directions\. These results suggest that current MU is often evaluated for*apparent forgetting*rather than retraining\-consistent forgetting\. More broadly, retraining reveals what output forgetting hides\.

## 1Introduction

Machine unlearning \(MU\) is commonly evaluated through*output forgetting*\[[21](https://arxiv.org/html/2606.25001#bib.bib39),[5](https://arxiv.org/html/2606.25001#bib.bib12),[31](https://arxiv.org/html/2606.25001#bib.bib20),[13](https://arxiv.org/html/2606.25001#bib.bib11),[9](https://arxiv.org/html/2606.25001#bib.bib15)\]\. That is, if a model attains low forget\-set accuracy, low output\-level membership inference\[[4](https://arxiv.org/html/2606.25001#bib.bib40)\], and high retain accuracy, it is typically regarded as having forgotten successfully\. This has become the default success signal in much of the MU literature\[[5](https://arxiv.org/html/2606.25001#bib.bib12),[13](https://arxiv.org/html/2606.25001#bib.bib11),[9](https://arxiv.org/html/2606.25001#bib.bib15),[3](https://arxiv.org/html/2606.25001#bib.bib1),[2](https://arxiv.org/html/2606.25001#bib.bib2),[35](https://arxiv.org/html/2606.25001#bib.bib3),[10](https://arxiv.org/html/2606.25001#bib.bib4),[6](https://arxiv.org/html/2606.25001#bib.bib13),[32](https://arxiv.org/html/2606.25001#bib.bib14),[26](https://arxiv.org/html/2606.25001#bib.bib37)\]\. In this paper, we show through both theoretical analysis and extensive experimental results that this signal is too weak\. Across multiple unlearning methods, datasets, and models, methods that appear successful at the output layer often remain inconsistent with retraining in representation space\. In other words, current evaluation can systematically mistake*apparent forgetting*for successful forgetting\.

To expose this gap, we useretraining\-consistent representation forgettingas a stronger evaluative lens\. Concretely, we use the retrained model,i\.e\.,the model trained from scratch without the forget data, as the operational reference for correct forgetting in representation space\. This reference is important because comparing an unlearned model to the original reveals how much it has changed, whereas comparing it to retraining reveals whether it has changed in a way consistent with forgetting the designated data\. Under this lens, we find that many existing unlearning methods\[[21](https://arxiv.org/html/2606.25001#bib.bib39),[5](https://arxiv.org/html/2606.25001#bib.bib12),[31](https://arxiv.org/html/2606.25001#bib.bib20),[13](https://arxiv.org/html/2606.25001#bib.bib11),[9](https://arxiv.org/html/2606.25001#bib.bib15),[22](https://arxiv.org/html/2606.25001#bib.bib16)\]achieve apparent output\-level forgetting while remaining retraining\-inconsistent in representation space \(see Fig\.[1](https://arxiv.org/html/2606.25001#S1.F1)\)\. This shows that output\-level success is not sufficient to certify retraining\-consistent forgetting\. We emphasize that retraining\-consistent representation forgetting is not introduced as a universal axiom for all MU settings\. Rather, it serves here as a stronger evaluative lens for diagnosing failure modes that weaker endpoint\-based criteria can miss\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x1.png)Figure 1:Looks forgotten, but not close to retraining\.A schematic summary of the evaluation gap studied in this paper\. Output\-level metrics may suggest successful forgetting, yet the same methods can remain far from the exact retraining reference in representation space\. This discrepancy motivates our retraining\-consistent analysis of forget/retain asymmetry, directional mismatch, and concentrated residuals\.This raises the central question of the paper,i\.e\.,*if output\-level success can coexist with retraining\-inconsistent residuals in representation space, then what kind of forgetting is current evaluation actually certifying?*We further ask whether such residual mismatch is merely incidental, or whether it follows a stable structure relative to exact retraining\.

As illustrated in Fig\.[1](https://arxiv.org/html/2606.25001#S1.F1), our claim is not simply that representation\-level residuals can remain after unlearning\. Rather, under a retraining\-consistent lens, these residuals reveal a systematic gap between what current evaluations certify and what exact retraining would produce\. Across the evaluated methods, this gap is not random\. Unlearning often moves in partial agreement with retraining on forget samples, deviates more substantially on retain samples, and leaves residual discrepancy concentrated along retraining\-related directions\. In particular, the mismatch manifests through forget/retain asymmetry, directional mismatch, and concentrated residuals in representation space\. These findings suggest that many current methods\[[9](https://arxiv.org/html/2606.25001#bib.bib15),[2](https://arxiv.org/html/2606.25001#bib.bib2),[35](https://arxiv.org/html/2606.25001#bib.bib3),[10](https://arxiv.org/html/2606.25001#bib.bib4),[6](https://arxiv.org/html/2606.25001#bib.bib13),[32](https://arxiv.org/html/2606.25001#bib.bib14),[11](https://arxiv.org/html/2606.25001#bib.bib25),[12](https://arxiv.org/html/2606.25001#bib.bib38)\]are better understood as achieving*apparent forgetting*at the output level while leaving behind structured representation\-level residuals\.

We instantiate this analysis in controlled class\-unlearning settings, starting with CIFAR\-10 with ResNet\-18 and extending across dataset complexity, model size, architecture, the forget class, and seed\. Across these settings, the same structured mismatch persists beyond a single benchmark configuration\. Taken together, these results identify a hidden failure mode in current MU evaluation and algorithms\. The field often treats output\-level forgetting as evidence of successful unlearning, even when retraining\-consistent forgetting has not been achieved in representation space\.

Our contributions are threefold:

- •We show that*output forgetting*is a misleading success signal for MU, where standard output\-level evaluations can systematically overestimate successful forgetting\.
- •We introduce*retraining\-consistent representation forgetting*as a stronger evaluative lens, using the retrained model as the operational reference for what correct forgetting should look like in representation space\.
- •Under this lens, we uncover a*structured failure mode*of current unlearning methods, characterized by forget/retain asymmetry, directional mismatch, concentrated residual leakage, and persistence across scales\.

In this sense, the paper’s central message is simple\. Retraining reveals what output forgetting hides\.

## 2Problem Formulation and Retraining\-Consistent Representation Lens

The empirical contradiction motivating this paper is simple\. Output\-level success does not guarantee consistency with retraining in representation space\. The natural next question is therefore not only whether mismatch remains, but how to formalize that mismatch relative to retraining\. This section introduces the minimal language needed to pose that question precisely\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x2.png)Figure 2:Output forgetting can appear successful while forget set information remains recoverable in representation space\. We study this gap through a retraining\-consistent representation lens\.### 2\.1Preliminaries

Letθ\\thetadenote the initial model parameters, and let𝒟=\{\(xi,yi\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}be the full training dataset\. A learning algorithm𝒜\\mathcal\{A\}produces the original modelθo=𝒜​\(θ,𝒟\)\\theta\_\{o\}=\\mathcal\{A\}\(\\theta,\\mathcal\{D\}\)\. Given a designated forget set𝒟u⊂𝒟\\mathcal\{D\}\_\{u\}\\subset\\mathcal\{D\}, the remaining data form the retain set𝒟r=𝒟∖𝒟u\\mathcal\{D\}\_\{r\}=\\mathcal\{D\}\\setminus\\mathcal\{D\}\_\{u\}\. An unlearning algorithm𝒰\\mathcal\{U\}produces an unlearned modelθu=𝒰​\(θo,𝒟u\)\\theta\_\{u\}=\\mathcal\{U\}\(\\theta\_\{o\},\\mathcal\{D\}\_\{u\}\)with the goal of removing the influence of𝒟u\\mathcal\{D\}\_\{u\}fromθo\\theta\_\{o\}\. The ideal reference is the retrained modelθr=𝒜​\(θ,𝒟r\)\\theta\_\{r\}=\\mathcal\{A\}\(\\theta,\\mathcal\{D\}\_\{r\}\), trained from scratch on the𝒟r\\mathcal\{D\}\_\{r\}only\.

We write the model asfθ​\(x\)=gθ​\(hθ​\(x\)\)f\_\{\\theta\}\(x\)=g\_\{\\theta\}\(h\_\{\\theta\}\(x\)\), wherehθ:𝒳→ℝdh\_\{\\theta\}:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\}denotes the penultimate\-layer representation andgθg\_\{\\theta\}the prediction head\. Throughout the paper,θr\\theta\_\{r\}serves as the*operational reference*for correct forgetting\. Comparingθu\\theta\_\{u\}toθo\\theta\_\{o\}measures the extent of change, whereas comparingθu\\theta\_\{u\}toθr\\theta\_\{r\}assesses whether this change is consistent with retraining without𝒟u\\mathcal\{D\}\_\{u\}\.

### 2\.2Two Lenses on Machine Unlearning

Most MU evaluations\[[5](https://arxiv.org/html/2606.25001#bib.bib12),[13](https://arxiv.org/html/2606.25001#bib.bib11),[9](https://arxiv.org/html/2606.25001#bib.bib15),[3](https://arxiv.org/html/2606.25001#bib.bib1),[2](https://arxiv.org/html/2606.25001#bib.bib2),[35](https://arxiv.org/html/2606.25001#bib.bib3),[10](https://arxiv.org/html/2606.25001#bib.bib4),[6](https://arxiv.org/html/2606.25001#bib.bib13),[32](https://arxiv.org/html/2606.25001#bib.bib14),[26](https://arxiv.org/html/2606.25001#bib.bib37)\]are still dominated by*output\-level success signals*, such as low forget\-set accuracy, low output\-level membership inference, and high retain accuracy\. These metrics assess whether a model*appears*to forget at the prediction layer\.

###### Definition 1\(Output Forgetting\)\.

For a setSS, letℱθS=\{fθ​\(x\):x∈S\}\\mathcal\{F\}\_\{\\theta\}^\{S\}=\\\{f\_\{\\theta\}\(x\):x\\in S\\\}denote the model outputs with empirical distributionℙθS​\(ℱ\)\\mathbb\{P\}\_\{\\theta\}^\{S\}\(\\mathcal\{F\}\)\. We say that an unlearned modelθu\\theta\_\{u\}achieves*output forgetting*if its output behavior matches that of the retrained modelθr\\theta\_\{r\}on both the forget set𝒟u\\mathcal\{D\}\_\{u\}and the retain set𝒟r\\mathcal\{D\}\_\{r\}, i\.e\., for eachS∈\{𝒟u,𝒟r\}S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}:

ℙθuS​\(ℱ\)≈ℙθrS​\(ℱ\),∀S∈\{𝒟u,𝒟r\}\.\\mathbb\{P\}\_\{\\theta\_\{u\}\}^\{S\}\(\\mathcal\{F\}\)\\approx\\mathbb\{P\}\_\{\\theta\_\{r\}\}^\{S\}\(\\mathcal\{F\}\),\\quad\\forall S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}\.\(1\)

However, output forgetting alone does not establish whether the model is also consistent with retraining in representation space\. To formalize this gap, we use*retraining\-consistent representation forgetting*as a stronger evaluative lens, illustrated in Fig\.[2](https://arxiv.org/html/2606.25001#S2.F2)\.

For a setSS, letℋθS=\{hθ​\(x\):x∈S\}\\mathcal\{H\}\_\{\\theta\}^\{S\}=\\\{h\_\{\\theta\}\(x\):x\\in S\\\}denote the representations with empirical distributionℙθS​\(ℋ\)\\mathbb\{P\}\_\{\\theta\}^\{S\}\(\\mathcal\{H\}\)\. We define the representation transformation fromθ\\thetatoθ′\\theta^\{\\prime\}onSSas𝒯θ→θ′S=\{hθ′​\(x\)−hθ​\(x\):x∈S\}\\mathcal\{T\}\_\{\\theta\\rightarrow\\theta^\{\\prime\}\}^\{S\}=\\\{h\_\{\\theta^\{\\prime\}\}\(x\)\-h\_\{\\theta\}\(x\):x\\in S\\\}, with empirical distributionℙθ→θ′S​\(𝒯\)\\mathbb\{P\}\_\{\\theta\\rightarrow\\theta^\{\\prime\}\}^\{S\}\(\\mathcal\{T\}\)\.

###### Definition 2\(Retraining\-Consistent Representation Forgetting\)\.

We define*retraining\-consistent representation forgetting*as follows\. Under this lens, an unlearned modelθu\\theta\_\{u\}is assessed not only by whether its representations resemble those of the retrained modelθr\\theta\_\{r\}, but also by whether its representation changes relative to the original modelθo\\theta\_\{o\}match those induced by retraining\. Concretely, for eachS∈\{𝒟u,𝒟r\}S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}, we assess:

ℙθuS​\(ℋ\)≈ℙθrS​\(ℋ\),ℙθo→θuS​\(𝒯\)≈ℙθo→θrS​\(𝒯\),∀S∈\{𝒟u,𝒟r\}\.\\begin\{aligned\} \\mathbb\{P\}\_\{\\theta\_\{u\}\}^\{S\}\(\\mathcal\{H\}\)&\\approx\\mathbb\{P\}\_\{\\theta\_\{r\}\}^\{S\}\(\\mathcal\{H\}\),\\\\ \\mathbb\{P\}\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{u\}\}^\{S\}\(\\mathcal\{T\}\)&\\approx\\mathbb\{P\}\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\(\\mathcal\{T\}\),\\end\{aligned\}\\quad\\forall S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}\.\(2\)

We use this notion as a*stronger evaluative lens*, not as a universal axiom for all MU settings\. Its role in this paper is diagnostic rather than doctrinal, as endpoint similarity is insufficient to rule out a hidden residual mismatch, since models can agree at the output layer while remaining substantially different in the representation space\. In our benchmark\-scale setting, retraining provides the strongest available reference for detecting such mismatch\.

### 2\.3Why Output Forgetting is Too Weak

The following result formalizes a limitation of output\-level metrics\. See Appen\.[A](https://arxiv.org/html/2606.25001#A1)for the full proof\.

###### Theorem 1\(Output Forgetting Does Not Imply Retraining\-Consistent Representation Forgetting\)\.

Letfθ​\(x\)=gθ​\(hθ​\(x\)\)f\_\{\\theta\}\(x\)=g\_\{\\theta\}\(h\_\{\\theta\}\(x\)\)be a classifier composed of a feature extractorhθ:𝒳→ℝdh\_\{\\theta\}:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\}and a prediction headgθg\_\{\\theta\}\. There exists an unlearned modelθu\\theta\_\{u\}such thatfθu​\(x\)=fθr​\(x\),∀x∈𝒟,f\_\{\\theta\_\{u\}\}\(x\)=f\_\{\\theta\_\{r\}\}\(x\),\\quad\\forall x\\in\\mathcal\{D\},whileℙθuS​\(ℋ\)≠ℙθrS​\(ℋ\)\\mathbb\{P\}\_\{\\theta\_\{u\}\}^\{S\}\(\\mathcal\{H\}\)\\neq\\mathbb\{P\}\_\{\\theta\_\{r\}\}^\{S\}\(\\mathcal\{H\}\)for at least oneS∈\{𝒟u,𝒟r\}S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}\. Hence, perfect output alignment with the retrained model does not guarantee retraining\-consistent representations, and representation\-level residuals may persist\.

### 2\.4Diagnostics for the Hidden Failure Mode

Our goal is not to introduce geometry for its own sake, but to use it as a diagnostic toolkit for exposing*how*current methods fail under a retraining\-consistent representation lens\.

#### Representation\-level leakage\.

Membership inference attack \(MIA\)\[[4](https://arxiv.org/html/2606.25001#bib.bib40)\]evaluates whether a sample belongs to the training set\. In the unlearning setting,MIAlogit\\mathrm\{MIA\}\_\{\\mathrm\{logit\}\}uses output statistics such as confidence or entropy, whileMIArep\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}operates on feature embeddingshθu​\(x\)h\_\{\\theta\_\{u\}\}\(x\)\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]and captures separability in representation space\. To compare residual leakage consistently across output and representation levels, we use the normalized leakage ratio:

ρ∗=\|MIA∗u−MIA∗r\|\|MIA∗o−MIA∗r\|,∗∈\{logit,rep\},\\rho\_\{\*\}=\\frac\{\|\\mathrm\{MIA\}\_\{\*\}^\{u\}\-\\mathrm\{MIA\}\_\{\*\}^\{r\}\|\}\{\|\\mathrm\{MIA\}\_\{\*\}^\{o\}\-\\mathrm\{MIA\}\_\{\*\}^\{r\}\|\},\\qquad\*\\in\\\{\\mathrm\{logit\},\\mathrm\{rep\}\\\},\(3\)whereMIA∗o\\mathrm\{MIA\}\_\{\*\}^\{o\},MIA∗u\\mathrm\{MIA\}\_\{\*\}^\{u\}, andMIA∗r\\mathrm\{MIA\}\_\{\*\}^\{r\}denote the attack success rates of the original, unlearned, and retrained models, respectively\. Lowerρ∗\\rho\_\{\*\}indicates less residual leakage relative to retraining\. A systematic analysis of MIA robustness across configurations and the selection rationale are provided in Appendix[B](https://arxiv.org/html/2606.25001#A2)\.

#### Representation similarity\.

To quantify how closelyθu\\theta\_\{u\}matchesθr\\theta\_\{r\}in representation geometry, we use centered kernel alignment \(CKA\)\[[19](https://arxiv.org/html/2606.25001#bib.bib22)\]\. Given representation matricesX,Y∈ℝn×dX,Y\\in\\mathbb\{R\}^\{n\\times d\}evaluated on the same data, linear CKA is defined as:

CKA​\(X,Y\)=⟨X​X⊤,Y​Y⊤⟩F‖X​X⊤‖F​‖Y​Y⊤‖F\.\\mathrm\{CKA\}\(X,Y\)=\\frac\{\\langle XX^\{\\top\},YY^\{\\top\}\\rangle\_\{F\}\}\{\\\|XX^\{\\top\}\\\|\_\{F\}\\\|YY^\{\\top\}\\\|\_\{F\}\}\.\(4\)Higher CKA indicates that the unlearned representation is closer to the retrained representation\.

#### Transformation mismatch\.

To characterize how unlearning moves representations relative to retraining, we define the mean representation shift fromθo\\theta\_\{o\}toθ′\\theta^\{\\prime\}on a setSSas:

Δθo→θ′S=1\|S\|​∑x∈Shθ′​\(x\)−1\|S\|​∑x∈Shθo​\(x\),θ′∈\{θu,θr\}\.\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta^\{\\prime\}\}^\{S\}=\\frac\{1\}\{\|S\|\}\\sum\_\{x\\in S\}h\_\{\\theta^\{\\prime\}\}\(x\)\-\\frac\{1\}\{\|S\|\}\\sum\_\{x\\in S\}h\_\{\\theta\_\{o\}\}\(x\),\\qquad\\theta^\{\\prime\}\\in\\\{\\theta\_\{u\},\\theta\_\{r\}\\\}\.\(5\)We then measure directional alignment via cosine similarity:

cos⁡\(Δθo→θuS,Δθo→θrS\)=Δθo→θuS⋅Δθo→θrS‖Δθo→θuS‖​‖Δθo→θrS‖\.\\cos\\\!\\left\(\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{u\}\}^\{S\},\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\\right\)=\\frac\{\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{u\}\}^\{S\}\\cdot\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\}\{\\\|\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{u\}\}^\{S\}\\\|\\\|\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\\\|\}\.\(6\)This quantity is not itself the paper’s main contribution; rather, it diagnoses whether unlearning follows a retraining\-consistent transformation on forget and retain samples\. To further localize the discrepancy, we decompose representations along the retraining direction\. LetvS=Δθo→θrS‖Δθo→θrS‖,v^\{S\}=\\frac\{\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\}\{\\\|\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\\\|\},and decomposehθu∥​\(x\)=\(hθu​\(x\)⋅vS\)​vS,hθu⟂​\(x\)=hθu​\(x\)−hθu∥​\(x\)\.h\_\{\\theta\_\{u\}\}^\{\\parallel\}\(x\)=\(h\_\{\\theta\_\{u\}\}\(x\)\\cdot v^\{S\}\)\\,v^\{S\},\\,h\_\{\\theta\_\{u\}\}^\{\\perp\}\(x\)=h\_\{\\theta\_\{u\}\}\(x\)\-h\_\{\\theta\_\{u\}\}^\{\\parallel\}\(x\)\.This subspace decomposition is used to test whether residual leakage and representation mismatch are diffuse or instead concentrated along retraining\-related directions\.

## 3Experimental Results

### 3\.1Experimental Setup

#### Datasets, models, and unlearning scenario\.

We study class unlearning, where the goal is to remove the influence of one designated class while preserving performance on the retain set𝒟r\\mathcal\{D\}\_\{r\}\. We evaluate on CIFAR\-10, CIFAR\-100\[[20](https://arxiv.org/html/2606.25001#bib.bib33)\], and TinyImageNet\[[23](https://arxiv.org/html/2606.25001#bib.bib35)\], which are standard image classification benchmarks for class\-level unlearning\. Our primary analysis is conducted on ResNet\-18\[[18](https://arxiv.org/html/2606.25001#bib.bib23)\], with ResNet\-50 for model\-capacity scaling and ViT\-Tiny\[[8](https://arxiv.org/html/2606.25001#bib.bib36)\]for cross\-backbone validation\.

#### Baselines and unlearning methods\.

We use the original modelθo\\theta\_\{o\}\(without unlearning\) and the retrained modelθr\\theta\_\{r\}\(trained from scratch on𝒟r\\mathcal\{D\}\_\{r\}only\) as the lower and upper references\.We evaluate five output\-oriented methods, including SCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\], Boundary Shrink\[[5](https://arxiv.org/html/2606.25001#bib.bib12)\], UNSIR\[[31](https://arxiv.org/html/2606.25001#bib.bib20)\], Amnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\], and SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\]\. We also include POUR\-P and POUR\-D\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\], which are representation\-aware unlearning methods, to test whether the same failure mode persists even when the methods explicitly operate in representation space\. Additional experimental details are provided in Appendix[C](https://arxiv.org/html/2606.25001#A3)\.

#### Evaluation metrics\.

Following Sec\.[2](https://arxiv.org/html/2606.25001#S2), we report both output\-level and representation\-level quantities\. Output\-level metrics are forget\-set accuracyAccu\\text\{Acc\}\_\{u\}, retain\-set accuracyAccr\\text\{Acc\}\_\{r\}, and logit\-level membership inferenceMIAlogit\\text\{MIA\}\_\{\\text\{logit\}\}\. Representation\-level diagnostics are forget\-set and retain\-set CKA againstθr\\theta\_\{r\}\(CKAu\\text\{CKA\}\_\{u\}andCKAr\\text\{CKA\}\_\{r\}\), representation\-level membership inferenceMIArep\\text\{MIA\}\_\{\\text\{rep\}\}, cosine alignment between unlearning and retraining shifts, and subspace\-localized residual analyses\. LowerAccu\\text\{Acc\}\_\{u\}and lower MIA indicate stronger forgetting under the corresponding metric, whereas higherAccr\\text\{Acc\}\_\{r\}, higher CKA, and higher directional alignment indicate closer approximation to retraining\.

MethodAppears forgotten understandard output evaluationActually close to retrainingin representation space?InterpretationAccu\(%\)↓\\downarrowAccr\(%\)↑\\uparrowMIAlogit\{\}\_\{\\text\{logit\}\}\(%\)↓\\downarrowCKAu↑\\uparrowCKAr↑\\uparrowMIArep\{\}\_\{\\text\{rep\}\}\(%\)↓\\downarrowOriginal98\.6598\.3581\.10\.5640\.91347\.7Original modelRetrain0\.0098\.7426\.31\.0001\.00044\.5Reference for correct forgettingSCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\]0\.0098\.9840\.20\.5130\.80747\.4Looks forgotten, still far from retrainBoundary Shrink\[[5](https://arxiv.org/html/2606.25001#bib.bib12)\]13\.5584\.2722\.10\.4380\.74347\.3Weak on both viewsUNSIR\[[31](https://arxiv.org/html/2606.25001#bib.bib20)\]0\.0027\.5520\.10\.0150\.04245\.0Forgets destructivelyAmnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\]0\.0097\.874\.20\.0860\.76647\.0Strong output forgetting, weak rep consistencySSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\]0\.0097\.394\.90\.5930\.89250\.9Strong output forgetting, highest rep leakagePOUR\-P\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.0098\.3610\.80\.6120\.91248\.3Better rep consistency, still not retrainPOUR\-D\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.0296\.9011\.00\.6180\.88347\.9Better rep consistency, still not retrain

Table 1:Standard output\-level evaluation can overestimate successful unlearning\. Several methods, notably SCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\], Amnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\], and SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\], appear strong under forget accuracy, retain accuracy, and logit\-level MIA, yet remain substantially inconsistent with the retrained reference in representation space\. This gap is the first empirical signal of the hidden failure mode studied in this paper\.

### 3\.2Standard Output\-level Evaluation overestimates Successful Unlearning

We begin with the paper’s central claim,i\.e\.,whether the field’s default success signal, namely output\-level forgetting, can overestimate successful unlearning\. The answer isYES\.

As shown in Tab\.[1](https://arxiv.org/html/2606.25001#S3.T1), several methods, including SCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\], Amnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\], and SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\], achieve near\-zeroAccu\\text\{Acc\}\_\{u\}together with reducedMIAlogit\\text\{MIA\}\_\{\\text\{logit\}\}, which would typically be interpreted as successful forgetting under standard output\-level evaluation\. However, these same models remain far from the retrained reference in representation space\. For instance, theirCKAu\\text\{CKA\}\_\{u\}is substantially below that ofθr\\theta\_\{r\}, and theirMIArep\\text\{MIA\}\_\{\\text\{rep\}\}remains elevated\. Under the stronger retraining\-consistent representation lens adopted in this paper, these methods therefore remain substantially inconsistent with retraining\.

The same pattern is visible in Fig\.[4](https://arxiv.org/html/2606.25001#S3.F4.5)\. Most methods lie above the diagonal, indicating that normalized leakage remains higher in the representation space than at the output\-level\. In other words, standard output\-level evaluation paints a more optimistic picture than the stronger representation\-level diagnostics support\. This is the first component of the hidden failure mode \-*the field’s dominant success signal is too weak*\. Linear probing and t\-SNE visualization further support this claim in Appendix[D](https://arxiv.org/html/2606.25001#A4)\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x3.png)Figure 3:ρlogit\\rho\_\{\\text\{logit\}\}againstρrep\\rho\_\{\\text\{rep\}\}on CIFAR\-10 with ResNet\-18\. Methods above the diagonal exhibit more residual leakage in representation space than at the output\-level\.MethodForget setcos⁡\(𝚫𝐮\)↑\\mathbf\{\\cos\(\\Delta\_\{u\}\)\}\\uparrowRetain setcos⁡\(𝚫𝐫\)↑\\mathbf\{\\cos\(\\Delta\_\{r\}\)\}\\uparrowAsymmetry gapcos⁡\(𝚫𝐮\)−cos⁡\(𝚫𝐫\)↓\\mathbf\{\\cos\(\\Delta\_\{u\}\)\-\\cos\(\\Delta\_\{r\}\)\}\\downarrowSCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\]0\.945−\-0\.3451\.290Boundary Shrink\[[5](https://arxiv.org/html/2606.25001#bib.bib12)\]0\.916−\-0\.3051\.221UNSIR\[[31](https://arxiv.org/html/2606.25001#bib.bib20)\]0\.7200\.2060\.514Amnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\]0\.7950\.4320\.364SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\]0\.911−\-0\.3891\.300POUR\-P\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.8540\.4680\.386POUR\-D\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.8840\.0740\.810

Figure 4:The mismatch with retraining is asymmetric across forget and retain samples\. Most methods align much more strongly with retraining on the forget set than on the retain set; negative values ofcos⁡\(Δr\)\\cos\(\\Delta\_\{r\}\)indicate especially poor retain\-side alignment, and the asymmetry gapcos⁡\(Δu\)−cos⁡\(Δr\)\\cos\(\\Delta\_\{u\}\)\-\\cos\(\\Delta\_\{r\}\)makes this imbalance explicit\.

### 3\.3Mismatch with Retraining is Asymmetric across Forget and Retain Samples

We next ask whether the mismatch with retraining is uniform across samples\. It isNOT\.

Instead, the discrepancy is asymmetric between the forget set𝒟u\\mathcal\{D\}\_\{u\}and the retain set𝒟r\\mathcal\{D\}\_\{r\}\. As shown in Tab\.[4](https://arxiv.org/html/2606.25001#S3.F4.5)and Fig\.[6](https://arxiv.org/html/2606.25001#S3.F6.fig2), unlearning methods often exhibit relatively high directional alignment with retraining on𝒟u\\mathcal\{D\}\_\{u\}, but much weaker alignment on𝒟r\\mathcal\{D\}\_\{r\}\. At the same time, this directional agreement on𝒟u\\mathcal\{D\}\_\{u\}does not translate into retraining\-consistent representation geometry asCKAu\\text\{CKA\}\_\{u\}remains substantially lower than expected under the retrained reference\. In contrast,𝒟r\\mathcal\{D\}\_\{r\}often preserves higher geometric similarity despite weaker directional alignment\. Shift magnitude diagnostics are provided as a descriptive complement in Appendix[E](https://arxiv.org/html/2606.25001#A5)\.

This asymmetry is important\. It shows that current methods are not simply “wrong everywhere\.” Instead, they exhibit a partial and uneven approximation of retraining\. They often move in the right coarse direction on forget samples, yet fail to reproduce retraining\-consistent behavior on retain samples and fail to recover retraining\-consistent geometry on the forget set\. This is the second component of the hidden failure mode \-*unlearning is asymmetrically misaligned with retraining across forget and retain samples*\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x4.png)\(a\)cos⁡\(Δ\)\\cos\(\\Delta\)
![Refer to caption](https://arxiv.org/html/2606.25001v1/x5.png)\(b\)CKA

Figure 5:Directional alignment and representation alignment on CIFAR\-10 with ResNet\-18\. Forget and retain samples exhibit different patterns of mismatch relative to retraining\.\\nextfloat

![Refer to caption](https://arxiv.org/html/2606.25001v1/x6.png)\(c\)Δ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}
![Refer to caption](https://arxiv.org/html/2606.25001v1/x7.png)\(d\)CKAu\\text\{CKA\}\_\{u\}

Figure 6:Residual leakage and representation mismatch by subspace on CIFAR\-10 with ResNet\-18\.![Refer to caption](https://arxiv.org/html/2606.25001v1/x8.png)\(a\)Accu\\text\{Acc\}\_\{u\}
![Refer to caption](https://arxiv.org/html/2606.25001v1/x9.png)\(b\)CKAu\\text\{CKA\}\_\{u\}

Figure 7:Output\-level and representation\-level forgetting across dataset complexity with ResNet\-18\.\\nextfloat

![Refer to caption](https://arxiv.org/html/2606.25001v1/x10.png)\(c\)Δ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}by subspace
![Refer to caption](https://arxiv.org/html/2606.25001v1/x11.png)\(d\)CKAu\\text\{CKA\}\_\{u\}by subspace

Figure 8:Residual leakage and representation mismatch across model size on CIFAR\-100\.![Refer to caption](https://arxiv.org/html/2606.25001v1/x12.png)\(a\)cos⁡\(Δ\)\\cos\(\\Delta\)
![Refer to caption](https://arxiv.org/html/2606.25001v1/x13.png)\(b\)CKA

Figure 9:Cross\-backbone validation on CIFAR\-100 with ViT\-Tiny\.\\nextfloat

![Refer to caption](https://arxiv.org/html/2606.25001v1/x14.png)Figure 10:Directional asymmetry across forget classes\.\\nextfloat

![Refer to caption](https://arxiv.org/html/2606.25001v1/x15.png)

Figure 11:Directional asymmetry across random seeds\.
### 3\.4Residual Discrepancy is Structured rather than Random

The next question is whether the remaining discrepancy is merely diffuse error, or whether it has structure\. Our results support the latter\.

Fig\.[6](https://arxiv.org/html/2606.25001#S3.F6.fig2)shows that both the residualMIArep\\text\{MIA\}\_\{\\text\{rep\}\}toθr\\theta\_\{r\}and the reduction inCKAu\\text\{CKA\}\_\{u\}are strongest along the retraining shift direction and substantially weaker in the orthogonal subspace\. This pattern also persists in the representation\-aware POUR methods\. The discrepancy is hence not well described as diffuse noise around retraining\. Instead, it is organized along retraining\-related directions\. This is the third component of the hidden failure mode\. The results suggest current methods often leave behind a*structured residual*rather than fully matching retraining in representation space\. Geometry is useful here not as the paper’s main point, rather it reveals*how*the hidden failure mode is organised\. Additional subspace presentations are provided in Appendix[F](https://arxiv.org/html/2606.25001#A6)for further illustration\.

### 3\.5Same Structured Mismatch Persists across Scale

We now test whether the diagnosed mismatch is a narrow artifact of one benchmark setting or a stable property of current unlearning behavior\. Specifically, we rule out four weaker explanations,i\.e\.,easy datasets, underpowered models, convolutional backbones, and favorable class or seed choices\.

#### Dataset complexity\.

A first weak explanation is that the diagnosed contradiction is merely a byproduct of an easy benchmark\. To test this, we increase dataset complexity from CIFAR\-10 to CIFAR\-100 and TinyImageNet\. If the gap appeared only in simple datasets, it would be much less convincing as a broader weakness of current unlearning methods\. Empirically, however, the same pattern remains\. TheAccu\\text\{Acc\}\_\{u\}stays near zero while representation similarity to retraining remains substantially below the retrained reference \(Fig\.[8](https://arxiv.org/html/2606.25001#S3.F8.fig2)\)\. This shows that the paper’s central contradiction persists as task complexity increases, where models can continue to look forgotten at the output layer while remaining retraining\-inconsistent in the representation space\.

#### Model size\.

A second weak explanation is that the hidden failure mode is simply a small\-model artifact\. To test this, we increase model capacity from ResNet\-18 to ResNet\-50 on CIFAR\-100\. If larger backbones could realize retraining\-consistent forgetting more easily, the mismatch might disappear with scale\. Instead, although overall representation quality improves, the mismatch remains, and the residual discrepancy becomes more concentrated along retraining\-related directions \(Fig\.[8](https://arxiv.org/html/2606.25001#S3.F8.fig2)\)\. Scale therefore does not wash out the failure mode; if anything, it sharpens its structure\.

#### Architecture\.

A third weak explanation is that the diagnosed mismatch is tied to one specific representation mechanism\. To test this, we replace the backbone with ViT\-Tiny\. This matters because convolutional backbones and vision transformers differ substantially in their inductive biases, feature formation, and representation geometries\. If the same gap only appeared in ResNets, the diagnosis could be dismissed as a backbone\-specific artifact rather than a broader issue in current unlearning methods\. Empirically, the same direction/geometry gap persists on ViT\-Tiny \(Fig\.[11](https://arxiv.org/html/2606.25001#S3.F11.2)\), showing that the failure mode is not merely a ResNet\- or convolution\-specific phenomenon, but survives a substantial change in architecture\.

#### Forget class and seed\.

A fourth weak explanation is that the observed pattern is contingent on one favorable class choice or one lucky training run\. To test this, we vary the forgotten class and the random seed\. If so, the diagnosed mismatch would be better understood as a fragile empirical coincidence than as a stable property of current unlearning behavior\. Instead, the same forget/retain asymmetry recurs across classes and seeds with low variance \(Fig\.[11](https://arxiv.org/html/2606.25001#S3.F11.2)and[11](https://arxiv.org/html/2606.25001#S3.F11.2)\), indicating that the diagnosed mismatch is neither class\-specific nor a stochastic accident\.

Taken together, these experiments rule out the most immediate narrow explanations of the observed gap\. What persists across scale is not merely a metric discrepancy, but the same structured mismatch relative to retraining\. In other words, the output\-level success remains too optimistic, while forget/retain asymmetry, directional inconsistency, and concentrated residual discrepancy repeatedly reappear\. This strengthens the claim that the hidden failure mode identified here reflects a broader weakness of current MU evaluation and algorithms, rather than a narrow artifact of one dataset, backbone, or training run\. Extended supporting scaling results are provided in Appendix[G](https://arxiv.org/html/2606.25001#A7)\.

## 4Related Work

#### Existing success signals for machine unlearning\.

Most MU methods are designed and evaluated through*output\-level forgetting*where successful unlearning is associated with low forget\-set accuracy, reduced output\-level membership inference, and preserved retain accuracy\[[5](https://arxiv.org/html/2606.25001#bib.bib12),[13](https://arxiv.org/html/2606.25001#bib.bib11),[9](https://arxiv.org/html/2606.25001#bib.bib15),[3](https://arxiv.org/html/2606.25001#bib.bib1),[2](https://arxiv.org/html/2606.25001#bib.bib2),[35](https://arxiv.org/html/2606.25001#bib.bib3),[10](https://arxiv.org/html/2606.25001#bib.bib4),[6](https://arxiv.org/html/2606.25001#bib.bib13),[32](https://arxiv.org/html/2606.25001#bib.bib14)\]\. This output\-centered view is natural because many methods manipulate decision boundaries, output distributions, or forget\-set predictions\. As a result, the dominant evaluation framework in the literature focuses on output behavior, commonly through forget accuracy and logit\-level MIA\[[5](https://arxiv.org/html/2606.25001#bib.bib12),[31](https://arxiv.org/html/2606.25001#bib.bib20),[9](https://arxiv.org/html/2606.25001#bib.bib15),[6](https://arxiv.org/html/2606.25001#bib.bib13),[28](https://arxiv.org/html/2606.25001#bib.bib21)\]\. However, these signals only indicate whether the model*appears*to forget at the prediction layer\. They do not establish whether forgotten information has actually been removed from the model’s internal representation\. Our work builds on this limitation, not by arguing that output\-level metrics are useless, but by showing that they can be systematically too weak as success signals\.

#### Representation\-level unlearning and representation\-aware evaluation\.

A growing body of work moves beyond output suppression and operates directly on the feature \(i\.e\.,representation\) space to suppress forget information while preserving retained knowledge\[[22](https://arxiv.org/html/2606.25001#bib.bib16),[18](https://arxiv.org/html/2606.25001#bib.bib23),[25](https://arxiv.org/html/2606.25001#bib.bib19),[15](https://arxiv.org/html/2606.25001#bib.bib24),[1](https://arxiv.org/html/2606.25001#bib.bib6),[37](https://arxiv.org/html/2606.25001#bib.bib8),[24](https://arxiv.org/html/2606.25001#bib.bib5),[7](https://arxiv.org/html/2606.25001#bib.bib7),[27](https://arxiv.org/html/2606.25001#bib.bib17)\]\. Common strategies include projection\-based removal, contrastive separation, adversarial representation erasure, and feature\-space transformation\. In parallel, recent evaluation protocols increasingly incorporate representation\-aware diagnostics such as CKA similarity, representation\-level MIA, linear probing, and mutual\-information\-based measures\[[22](https://arxiv.org/html/2606.25001#bib.bib16),[24](https://arxiv.org/html/2606.25001#bib.bib5),[27](https://arxiv.org/html/2606.25001#bib.bib17),[14](https://arxiv.org/html/2606.25001#bib.bib18),[17](https://arxiv.org/html/2606.25001#bib.bib9),[16](https://arxiv.org/html/2606.25001#bib.bib42)\]\. These works represent an important step forward by showing that output\-level success does not necessarily imply representation\-level forgetting, and that unlearned models can remain distinguishable from retrained ones in feature space\. Our paper is aligned with this broader shift toward representation\-aware analysis\.

#### The hidden failure mode behind apparent forgetting\.

Despite this progress, prior work mainly studies*whether*unlearned representations differ from retrained ones, rather than what kind of mismatch remains once output\-level success appears to be achieved\. This leaves four questions insufficiently characterized, namely \(i\) whether current methods follow the same transformation as retraining, \(ii\) whether mismatch differs across forget and retain samples, \(iii\) whether residual discrepancy is diffuse or directionally organized, and finally \(iv\) whether the same pattern persists across scale\. These questions motivate our paper\.

Our contribution is therefore not another representation metric, but a diagnostic analysis of a hidden failure mode in current MU evaluation and algorithms\. We show that standard output\-level evaluation can systematically overestimate successful forgetting, while a stronger retraining\-consistent representation lens reveals that many methods achieve only*apparent forgetting*, leaving behind a structured residual in representation space\.

## 5Discussion

### 5\.1A Hidden Failure Mode of Current Machine Unlearning

Retraining reveals what output forgetting hides\.Concretely, our results show that current MU exhibits a hidden failure mode that standard output\-level evaluation systematically fails to detect\. First, standard output\-level success signals are too weak as methods can attain low forget\-set accuracy, low logit\-level membership inference, and strong retain accuracy while remaining substantially inconsistent with retraining in representation space\. In this sense, current evaluation can mistake*apparent forgetting*for successful forgetting\. Second, this failure is structured rather than random\. Under a retraining\-consistent representation lens, current methods often partially align with retraining on forget samples, diverge on retain samples, and leave concentrated residual discrepancy rather than diffuse error\. Extended discussion is provided in Appendix[H](https://arxiv.org/html/2606.25001#A8)\.

### 5\.2Implications for Evaluation

The first implication is for evaluation practice\. Output\-level metrics remain useful, especially in black\-box settings, because they measure prediction\-layer behavior and output\-level leakage\. But our results show that they are too weak to serve as the sole success signal when the goal is to remove the influence of the forget set from the model\. A method can satisfy standard output\-level criteria while remaining substantially inconsistent with retraining in representation space\. This suggests a stricter evaluation principle\. Unlearning should be assessed not only by whether the model*looks*forgotten at the output layer, but also by whether it remains consistent with retraining under a stronger representation\-level lens\. Under this view, the key question is no longer just whether forget\-set accuracy is low or output\-level MIA is reduced, but whether the unlearned model changes in a way that resembles retraining without the forget data\.

### 5\.3Implications for Method Design

The second implication is for algorithm design\. Many current methods are optimized to suppress outputs, reduce forget\-set confidence, or weaken output\-level membership signals\. Our results suggest that this is not enough\. A method can satisfy these objectives while leaving behind asymmetric and structured retraining\-inconsistent residuals in representation space\. This does not imply that current methods are useless, nor that all of them fail in the same way\. Rather, it suggests that many current methods are better understood as optimizing for*apparent forgetting*rather than retraining\-consistent forgetting\. Future unlearning methods should therefore be designed against a stronger target that not only reduces output\-level traces of the forget set but also reduces structured residual discrepancy and better approximates retraining\-consistent transformations\.

### 5\.4Implications for Community Practice

The broader message of this paper is not that geometry, CKA, or representation\-level MIA are themselves the main contribution\. They are diagnostic tools whose value lies in revealing a hidden failure mode that standard evaluation leaves unseen\. More broadly, the most important structure in unlearning is not fully visible at the output layer\. Output\-level metrics remain useful, but they are insufficient to determine whether unlearning actually follows retraining in the representation space\. In this sense,*retraining reveals what output forgetting hides*,i\.e\.,not just that residual discrepancy remains, but how that discrepancy is organized\.

Seen this way, the central issue is not merely that some metrics are incomplete\. It is the field’s default practice that can certify successful unlearning on the basis of a signal that is too weak to measure the goal it claims to measure\. Addressing this gap will require progress not only in how MU is evaluated, but also in how its objectives are defined and how its methods are justified\.

## 6Conclusion

This paper presents a diagnostic analysis of a hidden failure mode in current MU\. Standard output\-level evaluations can systematically overestimate successful forgetting, where methods that appear successful under forget accuracy, output\-level membership inference, and retain accuracy often remain retraining\-inconsistent in representation space\. Using retraining\-consistent representation forgetting as a stronger evaluative lens, we show that this discrepancy is structured rather than random\. Current methods often partially align with retraining on forget samples, remain more inconsistent on retain samples, and leave concentrated residual discrepancy rather than closely matching retraining in representation space\. The implication is not that output\-level metrics are useless, but that they are too weak to serve as the sole success signal for unlearning\. If the goal is to remove the influence of the forget set from the model, then success must be judged not only by what the model suppresses at the output layer, but also by whether it remains consistent with retraining in representation space\. Retraining reveals what output forgetting hides\.

## References

- \[1\]A\. Almudévar and A\. Ortega\(2026\)Representation unlearning: forgetting through information compression\.arXiv preprint arXiv:2601\.21564\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[2\]L\. Bourtoule, V\. Chandrasekaran, C\. A\. Choquette\-Choo, H\. Jia, A\. Travers, B\. Zhang, D\. Lie, and N\. Papernot\(2021\)Machine unlearning\.In2021 IEEE symposium on security and privacy \(SP\),pp\. 141–159\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[3\]Y\. Cao and J\. Yang\(2015\)Towards making systems forget with machine unlearning\.In2015 IEEE symposium on security and privacy,pp\. 463–480\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[4\]N\. Carlini, S\. Chien, M\. Nasr, S\. Song, A\. Terzis, and F\. Tramer\(2022\)Membership inference attacks from first principles\.In2022 IEEE Symposium on Security and Privacy \(SP\),pp\. 1897–1914\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§2\.4](https://arxiv.org/html/2606.25001#S2.SS4.SSS0.Px1.p1.3)\.
- \[5\]M\. Chen, W\. Gao, G\. Liu, K\. Peng, and C\. Wang\(2023\)Boundary unlearning: rapid forgetting of deep networks via shifting the decision boundary\.InProceedings of the IEEE/CVF Conference on CVPR,pp\. 7766–7775\.Cited by:[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.8.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.5.5.5.5.5.5.5.5.2),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.17.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[6\]V\. S\. Chundawat, A\. K\. Tarun, M\. Mandal, and M\. Kankanhalli\(2023\)Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 7210–7217\.Cited by:[Table 2](https://arxiv.org/html/2606.25001#A2.T2.1.1.4.2.2.1.1.1),[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[7\]M\. Cotogni, J\. Bonato, L\. Sabetta, F\. Pelosin, and A\. Nicolosi\(2023\)Duck: distance\-based unlearning via centroid kinematics\.arXiv preprint arXiv:2312\.02052\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[8\]A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly,et al\.\(2020\)An image is worth 16x16 words: transformers for image recognition at scale\.arXiv preprint arXiv:2010\.11929\.Cited by:[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px1.p1.1)\.
- \[9\]J\. Foster, S\. Schoepf, and A\. Brintrup\(2024\)Fast machine unlearning without retraining through selective synaptic dampening\.InProceedings of the AAAI Conference,Vol\.38,pp\. 12043–12051\.Cited by:[Table 2](https://arxiv.org/html/2606.25001#A2.T2.1.1.4.2.2.1.2.1),[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.11.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.6.6.6.6.6.6.6.6.2),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[§3\.2](https://arxiv.org/html/2606.25001#S3.SS2.p2.5),[Table 1](https://arxiv.org/html/2606.25001#S3.T1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.20.1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.15.2),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[10\]A\. Ginart, M\. Guan, G\. Valiant, and J\. Y\. Zou\(2019\)Making ai forget you: data deletion in machine learning\.Advances in neural information processing systems32\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[11\]A\. Golatkar, A\. Achille, and S\. Soatto\(2020\)Eternal sunshine of the spotless net: selective forgetting in deep networks\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 9304–9312\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p4.1)\.
- \[12\]A\. Golatkar, A\. Achille, and S\. Soatto\(2020\)Forgetting outside the box: scrubbing deep networks of information accessible from input\-output observations\.InEuropean Conference on Computer Vision,pp\. 383–398\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p4.1)\.
- \[13\]L\. Graves, V\. Nagisetty, and V\. Ganesh\(2021\)Amnesiac machine learning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 11516–11524\.Cited by:[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.10.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.6.6.6.6.6.6.6.8.1),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[§3\.2](https://arxiv.org/html/2606.25001#S3.SS2.p2.5),[Table 1](https://arxiv.org/html/2606.25001#S3.T1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.19.1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.15.2),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[14\]T\. Guo, S\. Guo, J\. Zhang, W\. Xu, and J\. Wang\(2022\)Efficient attribute unlearning: towards selective removal of input attributes from feature representations\.arXiv preprint arXiv:2202\.13295\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[15\]T\. Hoang, S\. Rana, S\. Gupta, and S\. Venkatesh\(2024\)Learn to unlearn for deep neural networks: minimizing unlearning interference with gradient projection\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 4819–4828\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[16\]D\. Jeon, W\. Jeung, T\. Kim, A\. No, and J\. Choi\(2026\)An information theoretic evaluation metric for strong unlearning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 22173–22181\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[17\]Y\. Kim, S\. Cha, and D\. Kim\(2026\)Are we truly forgetting? a critical re\-examination of machine unlearning evaluation protocols\.Engineering Applications of Artificial Intelligence167,pp\. 113785\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[18\]S\. Kodge, G\. Saha, and K\. Roy\(2024\)Deep unlearning: fast and efficient gradient\-free class forgetting\.Transactions on Machine Learning Research\.Cited by:[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[19\]S\. Kornblith, M\. Norouzi, H\. Lee, and G\. Hinton\(2019\)Similarity of neural network representations revisited\.InInternational conference on machine learning,pp\. 3519–3529\.Cited by:[§2\.4](https://arxiv.org/html/2606.25001#S2.SS4.SSS0.Px2.p1.3)\.
- \[20\]A\. Krizhevsky, G\. Hinton,et al\.\(2009\)Learning multiple layers of features from tiny images\.Cited by:[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px1.p1.1)\.
- \[21\]M\. Kurmanji, P\. Triantafillou, J\. Hayes, and E\. Triantafillou\(2023\)Towards unbounded machine unlearning\.Advances in neural information processing systems36,pp\. 1957–1987\.Cited by:[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.7.1),[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.4.4.4.4.4.4.4.4.2),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[§3\.2](https://arxiv.org/html/2606.25001#S3.SS2.p2.5),[Table 1](https://arxiv.org/html/2606.25001#S3.T1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.16.1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.15.2)\.
- \[22\]A\. Le, C\. Peng, Y\. Liu, and J\. A\. Noble\(2025\)POUR: a provably optimal method for unlearning representations via neural collapse\.arXiv preprint arXiv:2511\.19339\.Cited by:[Table 2](https://arxiv.org/html/2606.25001#A2.T2.1.1.4.3),[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.12.1),[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.13.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[§2\.4](https://arxiv.org/html/2606.25001#S2.SS4.SSS0.Px1.p1.3),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.6.6.6.6.6.6.6.10.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.6.6.6.6.6.6.6.9.1),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.21.1),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.22.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[23\]Y\. Le, X\. Yang,et al\.\(2015\)Tiny imagenet visual recognition challenge\.CS 231N7\(7\),pp\. 3\.Cited by:[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px1.p1.1)\.
- \[24\]J\. Lee, Y\. Kim, and D\. Kim\(2026\)Erase at the core: representation unlearning for machine unlearning\.arXiv preprint arXiv:2602\.05375\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[25\]T\. Lee, S\. Park, M\. Jeon, H\. Hwang, and G\. Park\(2025\)ESC: erasing space concept for knowledge deletion\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 5010–5019\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[26\]W\. K\. Ong and C\. S\. Chan\(2025\)Maverick: collaboration\-free federated unlearning for medical privacy\.InInternational Conference on Medical Image Computing and Computer\-Assisted Intervention,pp\. 358–368\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1)\.
- \[27\]N\. M\. Sepahvand, E\. Triantafillou, H\. Larochelle, D\. Precup, J\. J\. Clark, D\. M\. Roy, and G\. K\. Dziugaite\(2025\)Selective unlearning via representation erasure using domain adversarial training\.InThe Thirteenth International Conference on Learning Representations,Cited by:[Table 2](https://arxiv.org/html/2606.25001#A2.T2.1.1.4.4),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.
- \[28\]R\. Shokri, M\. Stronati, C\. Song, and V\. Shmatikov\(2017\)Membership inference attacks against machine learning models\.In2017 IEEE symposium on security and privacy \(SP\),pp\. 3–18\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[29\]C\. N\. Spartalis, T\. Semertzidis, E\. Gavves, and P\. Daras\(2025\)Lotus: large\-scale machine unlearning with a taste of uncertainty\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 10046–10055\.Cited by:[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px2.p1.2)\.
- \[30\]J\. Sun, Z\. Wei, J\. Zou, J\. Gong, G\. Wang, C\. Dong, J\. Li, and B\. Liu\(2026\)Statistical mia: rethinking membership inference attack for reliable unlearning auditing\.arXiv preprint arXiv:2602\.01150\.Cited by:[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px1.p1.1)\.
- \[31\]A\. K\. Tarun, V\. S\. Chundawat, M\. Mandal, and M\. Kankanhalli\(2023\)Fast yet effective machine unlearning\.IEEE transactions on neural networks and learning systems35\(9\),pp\. 13046–13055\.Cited by:[Table 12](https://arxiv.org/html/2606.25001#A5.T12.6.9.1),[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p2.1),[Figure 4](https://arxiv.org/html/2606.25001#S3.F4.6.p1.6.6.6.6.6.6.6.7.1),[§3\.1](https://arxiv.org/html/2606.25001#S3.SS1.SSS0.Px2.p1.3),[Table 1](https://arxiv.org/html/2606.25001#S3.T1.12.12.12.12.12.12.12.18.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[32\]A\. Thudi, G\. Deza, V\. Chandrasekaran, and N\. Papernot\(2022\)Unrolling sgd: understanding factors influencing machine unlearning\.In2022 IEEE 7th EuroS&P,pp\. 303–319\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[33\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.\(2023\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px1.p1.1)\.
- \[34\]L\. Van der Maaten and G\. Hinton\(2008\)Visualizing data using t\-sne\.\.Journal of machine learning research9\(11\)\.Cited by:[§D\.2](https://arxiv.org/html/2606.25001#A4.SS2.p1.2)\.
- \[35\]J\. Xu, Z\. Wu, C\. Wang, and X\. Jia\(2024\)Machine unlearning: solutions and challenges\.IEEE Transactions on Emerging Topics in Computational Intelligence\.Cited by:[§1](https://arxiv.org/html/2606.25001#S1.p1.1),[§1](https://arxiv.org/html/2606.25001#S1.p4.1),[§2\.2](https://arxiv.org/html/2606.25001#S2.SS2.p1.1),[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px1.p1.1)\.
- \[36\]J\. Yu, Y\. He, A\. Goyal, and S\. Arora\(2025\)On the impossibility of retrain equivalence in machine unlearning\.arXiv preprint arXiv:2510\.16629\.Cited by:[Appendix I](https://arxiv.org/html/2606.25001#A9.SS0.SSS0.Px2.p1.2)\.
- \[37\]Q\. Zhang, C\. Yang, J\. Lou, L\. Xiong,et al\.\(2024\)Contrastive unlearning: a contrastive approach to machine unlearning\.arXiv preprint arXiv:2401\.10458\.Cited by:[§4](https://arxiv.org/html/2606.25001#S4.SS0.SSS0.Px2.p1.1)\.

## Appendix AProof of Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)

Definition 2 motivates a retraining\-consistent representation lens because output agreement alone does not determine the internal state of a model\. A model may match the retrained model’s predictions while retaining a different representation geometry or following a different representation transformation from the original model\. Thus, output forgetting can certify apparent prediction\-level behavior, but it cannot by itself certify that the model has moved toward the retrained solution in representation space\. The following theorem formalizes this limitation\. Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)is modest and only serves to justify why output\-level alignment alone cannot certify retraining\-consistent forgetting\.

###### Theorem 1\(Output Forgetting Does Not Imply Retraining\-Consistent Representation Forgetting\)\.

Letfθ​\(x\)=gθ​\(hθ​\(x\)\)f\_\{\\theta\}\(x\)=g\_\{\\theta\}\(h\_\{\\theta\}\(x\)\)be a classifier composed of a feature extractorhθ:𝒳→ℝdh\_\{\\theta\}:\\mathcal\{X\}\\to\\mathbb\{R\}^\{d\}and a prediction headgθg\_\{\\theta\}\. There exists an unlearned modelθu\\theta\_\{u\}such that

fθu​\(x\)=fθr​\(x\),∀x∈𝒟,f\_\{\\theta\_\{u\}\}\(x\)=f\_\{\\theta\_\{r\}\}\(x\),\\qquad\\forall x\\in\\mathcal\{D\},\(7\)while

ℙθuS​\(ℋ\)≠ℙθrS​\(ℋ\)\\mathbb\{P\}\_\{\\theta\_\{u\}\}^\{S\}\(\\mathcal\{H\}\)\\neq\\mathbb\{P\}\_\{\\theta\_\{r\}\}^\{S\}\(\\mathcal\{H\}\)\(8\)for at least oneS∈\{𝒟u,𝒟r\}S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\}\. Hence, perfect output alignment with retraining does not imply representation\-level consistency with retraining\.

###### Proof\.

Consider a linear prediction headgθ​\(h\)=W​hg\_\{\\theta\}\(h\)=Wh, whereW∈ℝk×dW\\in\\mathbb\{R\}^\{k\\times d\}, and assumeker⁡\(W\)\\ker\(W\)is non\-trivial\. Letz0∈ker⁡\(W\)z\_\{0\}\\in\\ker\(W\)withz0≠0z\_\{0\}\\neq 0, and define

hθu​\(x\)=\{hθr​\(x\)\+z0,x∈𝒟u,hθr​\(x\),x∈𝒟r\.h\_\{\\theta\_\{u\}\}\(x\)=\\begin\{cases\}h\_\{\\theta\_\{r\}\}\(x\)\+z\_\{0\},&x\\in\\mathcal\{D\}\_\{u\},\\\\ h\_\{\\theta\_\{r\}\}\(x\),&x\\in\\mathcal\{D\}\_\{r\}\.\\end\{cases\}\(9\)Then for anyx∈𝒟x\\in\\mathcal\{D\},

fθu​\(x\)=W​hθu​\(x\)=W​hθr​\(x\)\+W​z0=W​hθr​\(x\)=fθr​\(x\),f\_\{\\theta\_\{u\}\}\(x\)=Wh\_\{\\theta\_\{u\}\}\(x\)=Wh\_\{\\theta\_\{r\}\}\(x\)\+Wz\_\{0\}=Wh\_\{\\theta\_\{r\}\}\(x\)=f\_\{\\theta\_\{r\}\}\(x\),\(10\)so the two models are identical in output space on𝒟\\mathcal\{D\}\. However, becausez0≠0z\_\{0\}\\neq 0, the representation distribution on𝒟u\\mathcal\{D\}\_\{u\}is shifted, and thus

ℙθu𝒟u​\(ℋ\)≠ℙθr𝒟u​\(ℋ\)\.\\mathbb\{P\}\_\{\\theta\_\{u\}\}^\{\\mathcal\{D\}\_\{u\}\}\(\\mathcal\{H\}\)\\neq\\mathbb\{P\}\_\{\\theta\_\{r\}\}^\{\\mathcal\{D\}\_\{u\}\}\(\\mathcal\{H\}\)\.\(11\)Therefore, output\-level alignment does not guarantee retraining\-consistent representation forgetting\. ∎

Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)shows that a model can appear to forget at the output layer while still remaining inconsistent with retraining in representation space\. This is precisely why standard output\-level evaluation can overestimate successful unlearning\.

#### A stronger operational variant\.

Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)shows that output agreement does not imply representation agreement\. The following proposition gives an operational form of this argument, showing that residual discrepancy may remain visible after a linear projection in representation space even when outputs match\.

###### Proposition 1\(Output Agreement Does Not Rule Out Linearly Visible Representation\-Level Residuals\)\.

There exist an unlearned modelθu\\theta\_\{u\}and a retrained modelθr\\theta\_\{r\}such that

fθu​\(x\)=fθr​\(x\),∀x∈𝒟,f\_\{\\theta\_\{u\}\}\(x\)=f\_\{\\theta\_\{r\}\}\(x\),\\qquad\\forall x\\in\\mathcal\{D\},\(12\)while there exists a linear functionalℓ​\(h\)=w⊤​h\+b\\ell\(h\)=w^\{\\top\}h\+bfor which the induced one\-dimensional distributions ofℓ​\(hθu​\(x\)\)\\ell\(h\_\{\\theta\_\{u\}\}\(x\)\)andℓ​\(hθr​\(x\)\)\\ell\(h\_\{\\theta\_\{r\}\}\(x\)\)differ on𝒟u\\mathcal\{D\}\_\{u\}\. Hence, exact output agreement does not preclude linearly visible residual discrepancy in representation space\.

###### Proof\.

Use the same construction as in Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)\. Let

hθu​\(x\)=\{hθr​\(x\)\+z0,x∈𝒟u,hθr​\(x\),x∈𝒟r,h\_\{\\theta\_\{u\}\}\(x\)=\\begin\{cases\}h\_\{\\theta\_\{r\}\}\(x\)\+z\_\{0\},&x\\in\\mathcal\{D\}\_\{u\},\\\\ h\_\{\\theta\_\{r\}\}\(x\),&x\\in\\mathcal\{D\}\_\{r\},\\end\{cases\}wherez0≠0z\_\{0\}\\neq 0lies inker⁡\(W\)\\ker\(W\)for the linear prediction headgθ​\(h\)=W​hg\_\{\\theta\}\(h\)=Wh\. Then

fθu​\(x\)=fθr​\(x\),∀x∈𝒟,f\_\{\\theta\_\{u\}\}\(x\)=f\_\{\\theta\_\{r\}\}\(x\),\\qquad\\forall x\\in\\mathcal\{D\},exactly as in Theorem[1](https://arxiv.org/html/2606.25001#Thmtheorem1)\.

Now choose anyw∈ℝdw\\in\\mathbb\{R\}^\{d\}such thatw⊤​z0≠0w^\{\\top\}z\_\{0\}\\neq 0; such a vector exists becausez0≠0z\_\{0\}\\neq 0\. Define the linear functional

ℓ​\(h\)=w⊤​h\+b\\ell\(h\)=w^\{\\top\}h\+bfor any scalarbb\. Then for everyx∈𝒟ux\\in\\mathcal\{D\}\_\{u\},

ℓ​\(hθu​\(x\)\)=w⊤​\(hθr​\(x\)\+z0\)\+b=ℓ​\(hθr​\(x\)\)\+w⊤​z0\.\\ell\(h\_\{\\theta\_\{u\}\}\(x\)\)=w^\{\\top\}\(h\_\{\\theta\_\{r\}\}\(x\)\+z\_\{0\}\)\+b=\\ell\(h\_\{\\theta\_\{r\}\}\(x\)\)\+w^\{\\top\}z\_\{0\}\.Sincew⊤​z0≠0w^\{\\top\}z\_\{0\}\\neq 0, the projected representations on𝒟u\\mathcal\{D\}\_\{u\}differ by a nonzero constant shift, and therefore their induced one\-dimensional distributions are not equal\. Hence, exact output agreement does not preclude residual discrepancy that remains visible after a linear projection\. ∎

## Appendix BMetric Robustness of the Leakage Diagnosis

A natural concern is that the observed gap between output\-level and representation\-level leakage may depend on the particular MIA configuration used\. Since our main analysis compares forgetting across both levels, the diagnosis is only meaningful if the chosen attack is valid and discriminative at each level\. This section, therefore, examines representative MIA configurations from prior work to test whether the main leakage pattern is stable across attack designs and to justify the configuration used in the paper\.

### B\.1Alternative MIA Configurations

MU evaluations use MIA implementations that differ in both the membership signal and the attack model\. To assess whether our leakage diagnosis depends on these design choices, we compare three representative configurations from prior work at both logit and representation levels\. The configurations are summarized in Tab\.[2](https://arxiv.org/html/2606.25001#A2.T2), named after the unlearning methods that introduced them\.

DifferencesBadT MIAPOUR MIASURE MIAMembership SignalRetain set \(Member\)Test set \(Non\-member\)Train set \(Member\)Test set \(Non\-member\)Train set \(Member\)Test set \(Non\-member\)Attack ModelLogistic RegressionLogistic Regressionkk\-Nearest NeighbourLiteratureBad Teacher\[[6](https://arxiv.org/html/2606.25001#bib.bib13)\]SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\]POUR\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]SURE\[[27](https://arxiv.org/html/2606.25001#bib.bib17)\]

Table 2:Comparison of alternative MIA configurations\.A logistic regression\-based MIA assumes that member and non\-member samples can be separated by a single hyperplane\. It is therefore sensitive to global linear separability in the feature space\. In contrast, akk\-nearest neighbors \(kkNN\) MIA is sensitive to local neighborhood structure and captures cluster\-level memorization\. Across both attack models, membership signals can also be defined differently\. A Retain \(Member\) vs\. Test \(Non\-member\) setting evaluates whether forget samples behave similarly to retained training data, while a Train \(Member\) vs\. Test \(Non\-member\) setting evaluates whether forget samples still resemble the overall training distribution\. These alternatives provide complementary views of how membership information may persist after unlearning\.

For each configuration, we implement logit\-level and representation\-level membership inference using the classifier output entropy and the post\-average\-pooling representation, respectively\. We then assess whether the resulting attacks provide a stable and meaningful basis for comparing leakage across the two levels\.

### B\.2Robustness of the Leakage Pattern Across MIA Configurations

![Refer to caption](https://arxiv.org/html/2606.25001v1/x16.png)Figure 12:ASR across MIA configurations on CIFAR\-10 with ResNet\-18\. Lower ASR indicates less membership leakage\. At the logit level, SURE MIA compresses the large method\-dependent differences seen under BadT MIA and POUR MIA\. At the representation level, BadT MIA yields uniformly low ASR, including for the original model\.As to Fig\.[12](https://arxiv.org/html/2606.25001#A2.F12), at the logit level, BadT MIA and POUR MIA produce similar trends across unlearning methods, with values spread over a larger range, whereas SURE MIA yields values that are more compressed across methods\. Since POUR MIA and SURE MIA share the same membership signal but differ only in attack model, this compression suggests that thekkNN classifier is less discriminative than logistic regression on output entropy\. This makes BadT MIA and POUR MIA better suited for capturing leakage at the logit level\.

Meanwhile, at the representation level, Fig\.[12](https://arxiv.org/html/2606.25001#A2.F12)shows POUR MIA and SURE MIA produce much closer values to each other across all unlearning methods, whereas BadT MIA yields uniformly low values, including for the original model\. This behaviour is plausibly explained by its membership signal, where the attack model is trained on the retain set as members and the test set as non\-members\. Under class\-level unlearning, the forget set is absent from the member set, so the attack can become sensitive to class\-distribution similarity rather than true training membership\. This interpretation is supported by the observation that the retrained model achieves a higher BadT MIA than the original model, contrary to the expected trend\. In the retrained model, the representation of the forget class collapses toward other classes and becomes harder to distinguish from the retain\-set distribution, causing the attack model to identify it as a member\. By contrast, in the original model, the forget\-class representation forms a more distinct cluster from the retain classes, so the attack model is more likely to identify it as a non\-member\.

Taken together, these observations \(Fig\.[12](https://arxiv.org/html/2606.25001#A2.F12)\) support the use of a train\-vs\.\-test membership signal with logistic regression, such as POUR MIA, as the most suitable diagnostic for comparing leakage across logit and representation levels in our setting\.More importantly, they show that the main qualitative picture is not an artifact of a single attack design\. While the MIA configuration affects sensitivity and interpretation, disciplined representation\-level leakage diagnosis remains necessary\.

### B\.3Implementation Verification Details

We implement the alternative MIA configurations according to the membership signal and attack model specified in Tab\.[2](https://arxiv.org/html/2606.25001#A2.T2)\. The purpose of this subsection is to ensure that differences across configurations are not artifacts of preprocessing or attack implementation\.

In all cases, we apply balanced class weights, standard feature normalization fitted on the training split, and stratified subsampling of the training set to match the test\-set size\. The subsampling is stratified over the joint class and membership label so that the forget class is preserved during resampling\. For logit\-level MIA, the membership signal is the prediction entropy of the model’s output distribution\. For representation\-level MIA, features are extracted after average pooling and then flattened\. Both entropy and representations are extracted without augmentation at inference time, ensuring that cross\-model comparisons are based on deterministic outputs\.

## Appendix CExperimental Details

This appendix provides full experimental details supporting the results reported in the main paper\. We describe the datasets used across all settings \(Tab\.[3](https://arxiv.org/html/2606.25001#A3.T3)\), the hyperparameters used to train the original and retrained baseline models \(Tab\.[4](https://arxiv.org/html/2606.25001#A3.T4)\), and the per\-method unlearning hyperparameters for each of the five dataset / architecture combinations studied: CIFAR\-10 / ResNet\-18, CIFAR\-100 / ResNet\-18, CIFAR\-100 / ResNet\-50, TinyImageNet / ResNet\-18, and CIFAR\-100 / ViT\-Tiny \(Tab\.[5](https://arxiv.org/html/2606.25001#A3.T5),[6](https://arxiv.org/html/2606.25001#A3.T6),[7](https://arxiv.org/html/2606.25001#A3.T7),[8](https://arxiv.org/html/2606.25001#A3.T8),[9](https://arxiv.org/html/2606.25001#A3.T9),[10](https://arxiv.org/html/2606.25001#A3.T10)\)\. Dashes \(—\) indicate the entry is not applicable for that setting\. All experiments were conducted on NVIDIA A100 GPUs\.

Dataset StatisticsCIFAR\-10CIFAR\-100TinyImageNetNumber of classes10100200Training images50,00050,000100,000Validation images10,00010,00010,000Forget\-class images5,000500500Image size32×3232\\times 3232×3232\\times 3264×6464\\times 64Table 3:Dataset statistics for CIFAR\-10, CIFAR\-100 and TinyImageNet\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyEpochs50200200200300OptimizerSGDSGDSGDSGDAdamWBatch size128128128256256Learning rate0\.010\.10\.10\.013×10−43\\times 10^\{\-4\}Momentum0\.90\.90\.90\.9—Weight decay5×10−45\\times 10^\{\-4\}2\.5×10−32\.5\\times 10^\{\-3\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}0\.07Label smoothing00000\.1Scheduler—ReduceLROnPlateauReduceLROnPlateauReduceLROnPlateauCosineAnnealingLRScheduler LR factor—0\.10\.10\.1—Scheduler patience—555—Min LR————10−610^\{\-6\}Warmup epochs————10Early stopping✓✓✓✓—Pretrained weight†——✓—

Table 4:Hyperparameters for original and retrained baseline training\.†CIFAR\-10 / ResNet\-18 was initialised from weights pretrained on CIFAR\-100 for 30 epochs at learning rate 0\.1\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyGamma0\.991111Alpha0\.0010\.50\.50\.50\.5Max steps25555Min steps35545OptimizerSGDAdamAdamAdamAdamLearning rate5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}LR decay rate0\.10\.10\.10\.10\.1LR decay epochs\[3, 5, 9\]\[2\]\[2\]\[2\]\[2\]Momentum0\.9————Weight decay5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}

Table 5:Hyperparameters for SCRUB\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyOptimizerAdamAdamAdamAdamAdamBatch size6464646464Epochs55535Learning rate10−310^\{\-3\}10−310^\{\-3\}10−310^\{\-3\}10−410^\{\-4\}10−310^\{\-3\}

Table 6:Hyperparameters for Amnesiac\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyOptimizerSGDSGDSGDSGDSGDEps0\.10\.050\.050\.10\.3Poison epochs101010510Fine\-tune LR10−510^\{\-5\}10−510^\{\-5\}10−510^\{\-5\}10−510^\{\-5\}10−410^\{\-4\}Momentum0\.90\.90\.90\.90\.9

Table 7:Hyperparameters for Boundary Shrink\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyNoise batch size25625625616256Samples per retain class1000100010004501000Noise epochs4040402040Noise LR0\.10\.10\.10\.010\.1Noiseℓ2\\ell\_\{2\}lambda0\.10\.10\.10\.20\.1Noise copies2020201520Impair epochs11161Impair LR0\.020\.020\.020\.00310−510^\{\-5\}Repair epochs11181Repair LR0\.010\.010\.010\.0083×10−43\\times 10^\{\-4\}Repair weight decay00000\.05OptimizerAdamAdamAdamAdamAdam

Table 8:Hyperparameters for UNSIR\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyLambda11111Alpha1010105050

Table 9:Hyperparameters for SSD\.HyperparameterCIFAR\-10ResNet\-18CIFAR\-100ResNet\-18CIFAR\-100ResNet\-50TinyImageNetResNet\-18CIFAR\-100ViT\-TinyEpochs1010010010010OptimizerAdamAdamAdamAdamAdamLearning rate10−410^\{\-4\}10−410^\{\-4\}10−410^\{\-4\}10−410^\{\-4\}10−410^\{\-4\}

Table 10:Hyperparameters for POUR\-D\.
## Appendix DAdditional Representation\-Space Sanity Checks

The main paper diagnoses representation\-level forgetting through directional alignment, subspace decomposition, representation\-level membership inference, and centered kernel alignment\. This appendix adds two complementary sanity checks, namely linear probing and t\-SNE visualization, to test whether the same gap remains visible under simpler readouts of the learned representation\. The goal is not to introduce new primary metrics, but to verify that the main diagnosis is not specific to the particular representation\-level diagnostics used in the paper\.

### D\.1Linear Probing

We first test whether forget\-class information remains linearly recoverable from the learned representations after unlearning\. For each model, we freeze the backbone and train a single linear classification layer on the full dataset𝒟\\mathcal\{D\}, i\.e\.,𝒟u∪𝒟r\\mathcal\{D\}\_\{u\}\\cup\\mathcal\{D\}\_\{r\}, using the original class labels\. We then report forget\-set probe accuracyAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}and retain\-set probe accuracyAccrprobe\\mathrm\{Acc\}\_\{r\}^\{\\mathrm\{probe\}\}\. These quantities measure, respectively, how much class\-identifying structure for the forget class remains linearly accessible in the backbone representation and how well retain\-class structure is preserved\.

The intuition is simple\. If a backbone no longer encodes forget\-class features in a linearly separable way, then even a linear head trained on the original labels should fail to recover that class from𝒟u\\mathcal\{D\}\_\{u\}\. Conversely, highAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}indicates that forget\-class identity remains recoverable from the representation, even if output\-level forgetting appears successful\. All experiments are conducted on CIFAR\-10 with ResNet\-18\. The linear probe is trained with SGD for 10 epochs using a learning rate of1×10−31\\times 10^\{\-3\}across all seven unlearning methods, together with the original and retrained baselines\.

MethodLinear ProbeOutput\-levelReadingAccuprobe↓\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}\\\!\\downarrowAccrprobe↑\\mathrm\{Acc\}\_\{r\}^\{\\mathrm\{probe\}\}\\\!\\uparrowAccu↓\\mathrm\{Acc\}\_\{u\}\\\!\\downarrowAccr↑\\mathrm\{Acc\}\_\{r\}\\\!\\uparrowOriginal98\.2098\.1198\.6598\.35Original representationRetrain45\.1498\.260\.0098\.74Reference for forgettingSCRUB46\.0298\.080\.0098\.98Closest to retrain under probeBoundary Shrink49\.3885\.1913\.5584\.27Reduced recoverability, damaged retainUNSIR54\.1823\.910\.0027\.55Destructive forgettingAmnesiac74\.6897\.450\.0097\.87Output forgotten, probe recoverableSSD70\.2497\.370\.0097\.39Output forgotten, probe recoverablePOUR\-P83\.3898\.100\.0098\.36Output forgotten, probe recoverablePOUR\-D70\.2496\.500\.0296\.90Output forgotten, probe recoverableTable 11:Linear probe class\-recovery accuracy and output\-level accuracy on CIFAR\-10 with ResNet\-18\. Low output forget accuracy alone does not imply representation behavior close to retraining\.Table[11](https://arxiv.org/html/2606.25001#A4.T11)supports the main paper’s diagnosis\. The original model attains near\-perfectAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}\(98\.2%98\.2\\%\), confirming that forget\-class identity is fully encoded in its representation\. In contrast, the retrained model achieves only45\.1%45\.1\\%, despite the probe being trained on the original class labels\. This indicates that retraining no longer maintains the forget class as linearly separable in the backbone representation\.

Among methods that achieve complete output\-level forgetting, i\.e\.,Accu=0\\mathrm\{Acc\}\_\{u\}=0, several still retain strong linear recoverability of the forget class\. In particular, POUR\-P \(83\.4%83\.4\\%\), Amnesiac \(74\.7%74\.7\\%\), SSD \(70\.2%70\.2\\%\), and POUR\-D \(70\.2%70\.2\\%\) all remain far above retraining inAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}\. Thus, although these methods appear successful under output\-level evaluation, forget\-class identity remains linearly accessible in their representations\.

Boundary Shrink and UNSIR produceAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}values closer to retraining \(49\.4%49\.4\\%and54\.2%54\.2\\%, respectively\), but this comes with severe degradation inAccrprobe\\mathrm\{Acc\}\_\{r\}^\{\\mathrm\{probe\}\}\(85\.2%85\.2\\%and23\.9%23\.9\\%\), mirrored by their poorAccr\\mathrm\{Acc\}\_\{r\}\. This pattern is more consistent with broad representational disruption than with selective forgetting\.

SCRUB is the strongest case under linear probing\. ItsAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}\(46\.0%46\.0\\%\) is close to retraining, whileAccrprobe\\mathrm\{Acc\}\_\{r\}^\{\\mathrm\{probe\}\}\(98\.1%98\.1\\%\) remains essentially fully preserved\. Under this probe\-based view, SCRUB appears close to retraining\. However, the main paper shows that this agreement is incomplete, with SCRUB remaining directionally misaligned with retraining, geometrically inconsistent under CKA, and elevated underMIArep\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}\. Moreover, its output\-level forgetting degrades under larger datasets and models \(Tab\.[18](https://arxiv.org/html/2606.25001#A7.F18)and Tab\.[20](https://arxiv.org/html/2606.25001#A7.F20)\)\. Thus, linear probing provides a useful sanity check, but not a complete account of retraining consistency\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x17.png)Figure 13:t\-SNE visualization of feature representations on CIFAR\-10 with ResNet\-18\. The forget class is Class 0 \(airplane\), shown in red\.
### D\.2t\-SNE Visualization

We next provide a qualitative sanity check through t\-SNE visualization\[[34](https://arxiv.org/html/2606.25001#bib.bib41)\]\. Fig\.[13](https://arxiv.org/html/2606.25001#A4.F13)visualizes the learned representations of𝒟u\\mathcal\{D\}\_\{u\}and𝒟r\\mathcal\{D\}\_\{r\}for each model\. We stress that t\-SNE is used here only as an illustrative complement to the quantitative diagnostics above; it is not a formal basis for the paper’s claims\.

The visual patterns broadly agree with the linear\-probe results\. Amnesiac, SSD, POUR\-P, and POUR\-D retain a visually distinct forget\-class cluster that remains separated from the retain\-class representations\. This resembles the original model more than the retrained model and is consistent with their highAccuprobe\\mathrm\{Acc\}\_\{u\}^\{\\mathrm\{probe\}\}, indicating that forget\-class geometry remains substantially intact despite complete output\-level forgetting\.

Boundary Shrink and UNSIR show a different behavior\. Their forget\-class cluster becomes diffuse, but retain\-class clusters are also visibly distorted relative to the original model\. This aligns with their degradedAccrprobe\\mathrm\{Acc\}\_\{r\}^\{\\mathrm\{probe\}\}andAccr\\mathrm\{Acc\}\_\{r\}, suggesting broad representational disruption rather than selective forgetting\.

SCRUB appears visually closest to retraining, as its forget cluster is diffuse while its retain clusters remain relatively well separated\. This agrees with its near\-retraining linear\-probe behavior\. However, as shown in the main paper, this visual similarity does not extend to directional alignment, CKA geometry, or representation\-level leakage, and it does not persist under increased dataset complexity or model size\.

In summary, linear probing and t\-SNE support the same qualitative conclusion as the main paper\. Even under simpler and more intuitive representation\-space diagnostics, existing unlearning methods do not faithfully reproduce the representation\-level behavior of retraining, and output\-level forgetting alone remains too weak to reveal this gap\.

## Appendix EAdditional Directional Diagnostics

The main paper characterizes forget/retain asymmetry through the directional alignment between unlearning and retraining shifts\. This appendix adds a complementary diagnostic for the CIFAR\-10 / ResNet\-18 baseline by comparing not only the direction, but also the magnitude of those shifts relative to retraining\. The goal is to test whether apparent directional agreement is accompanied by retraining\-like adjustment strength, or whether unlearning under\- or over\-adjusts even when the direction appears partially aligned\.

### E\.1Shift Magnitude Ratios

For each unlearning method, we compute the relative shift magnitudeRS=‖Δθo→θuS‖‖Δθo→θrS‖,S∈\{𝒟u,𝒟r\},R^\{S\}=\\frac\{\\\|\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{u\}\}^\{S\}\\\|\}\{\\\|\\Delta\_\{\\theta\_\{o\}\\rightarrow\\theta\_\{r\}\}^\{S\}\\\|\},\\qquad S\\in\\\{\\mathcal\{D\}\_\{u\},\\mathcal\{D\}\_\{r\}\\\},separately for forget samples𝒟u\\mathcal\{D\}\_\{u\}and retain samples𝒟r\\mathcal\{D\}\_\{r\}\. A value ofRS=1R^\{S\}=1indicates that unlearning matches retraining in shift magnitude,RS\>1R^\{S\}\>1indicates over\-adjustment, andRS<1R^\{S\}<1indicates under\-adjustment\.

MethodForget Samples\(𝒟u\\mathcal\{D\}\_\{u\}\)Retain Samples\(𝒟r\\mathcal\{D\}\_\{r\}\)cos⁡\(Δu\)↑\\cos\(\\Delta\_\{u\}\)\\uparrowRuR^\{u\}cos⁡\(Δr\)↑\\cos\(\\Delta\_\{r\}\)\\uparrowRrR^\{r\}SCRUB\[[21](https://arxiv.org/html/2606.25001#bib.bib39)\]0\.9451\.074\-0\.3454\.578Boundary Shrink\[[5](https://arxiv.org/html/2606.25001#bib.bib12)\]0\.9161\.170\-0\.3054\.606UNSIR\[[31](https://arxiv.org/html/2606.25001#bib.bib20)\]0\.7201\.7410\.2067\.282Amnesiac\[[13](https://arxiv.org/html/2606.25001#bib.bib11)\]0\.7951\.6120\.4321\.776SSD\[[9](https://arxiv.org/html/2606.25001#bib.bib15)\]0\.9111\.312\-0\.3892\.608POUR\-P\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.8541\.0920\.4680\.761POUR\-D\[[22](https://arxiv.org/html/2606.25001#bib.bib16)\]0\.8841\.0970\.0741\.191Table 12:Directional alignment and relative shift magnitude of unlearning with respect to retraining on CIFAR\-10 with ResNet\-18\. HereRuR^\{u\}andRrR^\{r\}denote the ratio of unlearning shift magnitude to retraining shift magnitude on forget and retain samples, respectively\.Table[12](https://arxiv.org/html/2606.25001#A5.T12)refines the main paper’s directional diagnosis\. On forget samples𝒟u\\mathcal\{D\}\_\{u\}, methods such as SCRUB, Boundary Shrink, SSD, and the POUR variants show high cosine similarity together with magnitude ratios close to11, indicating that their forget\-side shifts are at least roughly similar to retraining in both direction and scale\. In contrast, UNSIR and Amnesiac show weaker directional agreement together with substantially larger magnitude ratios, suggesting that they over\-adjust even on the forget set\.

On retain samples𝒟r\\mathcal\{D\}\_\{r\}, the picture changes sharply\. Most methods have magnitude ratios above11, often far above11, while cosine similarity is near zero or even negative\. This indicates that retain\-side behavior is not merely misdirected relative to retraining, but frequently over\-displaced as well\. In other words, the retain\-side mismatch identified in the main paper reflects both directional inconsistency and excessive adjustment strength\.

As a summary, these magnitude comparisons support the same qualitative conclusion as the main paper\. The asymmetry between forget and retain samples is not only directional\. It also appears in the scale of the representation shift, reinforcing the view that current unlearning methods only partially approximate retraining and often do so unevenly across the forget and retain sets\.

## Appendix FExtended Projection / Subspace Analysis

The main paper shows that residual leakage and representation mismatch are concentrated along the retraining shift direction rather than being diffusely distributed across representation space\. This appendix provides additional views of the same projection decomposition\. The goal is not to introduce a new analysis, but to verify that the main structured\-residual diagnosis remains visible under complementary summaries of the parallel and orthogonal components\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x18.png)\(a\)Δ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}
![Refer to caption](https://arxiv.org/html/2606.25001v1/x19.png)\(b\)CKAu\\mathrm\{CKA\}\_\{u\}

Figure 14:Residual discrepancy of the raw representation together with its parallel and orthogonal components\.\\nextfloat

![Refer to caption](https://arxiv.org/html/2606.25001v1/x20.png)\(c\)Δ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}
![Refer to caption](https://arxiv.org/html/2606.25001v1/x21.png)\(d\)CKAu\\mathrm\{CKA\}\_\{u\}

Figure 15:Residual discrepancy ratio relative to the raw representation, showing how discrepancy is distributed across the parallel and orthogonal components\.### F\.1Raw, Parallel, and Orthogonal Comparisons

Fig\.[15](https://arxiv.org/html/2606.25001#A6.F15.fig2)compares the raw representation with its parallel and orthogonal components for bothΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}andCKAu\\mathrm\{CKA\}\_\{u\}\. HigherΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}and lowerCKAu\\mathrm\{CKA\}\_\{u\}indicate larger discrepancy relative to retraining\.

The same qualitative pattern appears across methods\. The orthogonal component typically remains close to the raw representation, whereas the parallel component shows stronger deviation\. In the original model, the parallel component already carries a substantially larger discrepancy than the orthogonal component\. After unlearning, this imbalance persists\. Although the absolute discrepancy often decreases relative to the original model, the parallel component remains the more informative direction of residual mismatch\. This supports the main paper’s claim that the deviation from retraining is not uniformly distributed across representation space, but organized along the retraining\-related direction\.

### F\.2Parallel\-to\-Raw and Orthogonal\-to\-Raw Ratios

To quantify how much of the residual discrepancy is concentrated in each component, we normalize each subspace metric by its raw counterpart\. For bothΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}andCKAu\\mathrm\{CKA\}\_\{u\}, we compute

r∥=parallelraw,r⟂=orthogonalraw\.r\_\{\\parallel\}=\\frac\{\\mathrm\{parallel\}\}\{\\mathrm\{raw\}\},\\qquad r\_\{\\perp\}=\\frac\{\\mathrm\{orthogonal\}\}\{\\mathrm\{raw\}\}\.A ratio near11indicates that the subspace carries discrepancy comparable to the raw representation\. ForΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}, values above11indicate excess leakage relative to the raw representation\. ForCKAu\\mathrm\{CKA\}\_\{u\}, values below11indicate stronger mismatch than the raw representation\.

Fig\.[15](https://arxiv.org/html/2606.25001#A6.F15.fig2)shows a stable asymmetry across methods\. The orthogonal ratios remain close to11, indicating that the orthogonal component largely tracks the raw discrepancy\. In contrast, the parallel ratios consistently point to stronger discrepancy, withr∥r\_\{\\parallel\}typically exceeding11forΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}and falling below11forCKAu\\mathrm\{CKA\}\_\{u\}\. This again indicates that the retraining direction carries disproportionately large residual discrepancy\.

### F\.3Additional Scatter Plots

Fig\.[16](https://arxiv.org/html/2606.25001#A6.F16.fig1)provides a final visual summary of the same decomposition\. ForΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}, we plotΔ​MIArep∥\{\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}\}\_\{\\parallel\}againstΔ​MIArep⟂\{\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}\}\_\{\\perp\}for each method\. ForCKAu\\mathrm\{CKA\}\_\{u\}, we analogously plotCKAu∥\{\\mathrm\{CKA\}\_\{u\}\}\_\{\\parallel\}againstCKAu⟂\{\\mathrm\{CKA\}\_\{u\}\}\_\{\\perp\}\. The diagonal corresponds to equal discrepancy in the two components\.

If the residual were diffuse, points would concentrate near the diagonal\. Instead, the points systematically fall away from it\. ForΔ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}, most methods lie in the region where the parallel component shows stronger leakage than the orthogonal component\. ForCKAu\\mathrm\{CKA\}\_\{u\}, most methods occupy the complementary regime in which the parallel component exhibits lower similarity to retraining than the orthogonal component\. Across diverse methods, this off\-diagonal structure reinforces the main paper’s conclusion that residual discrepancy is organized, not diffuse, and remains concentrated along retraining\-related directions\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x22.png)\(a\)Δ​MIArep\\Delta\\mathrm\{MIA\}\_\{\\mathrm\{rep\}\}
![Refer to caption](https://arxiv.org/html/2606.25001v1/x23.png)\(b\)CKAu\\mathrm\{CKA\}\_\{u\}

Figure 16:Scatter plots of residual discrepancy across the parallel and orthogonal components\.

## Appendix GExtended Scaling Summaries

The main paper argues that the diagnosed mismatch is not a narrow artifact of one benchmark configuration, but a stable property of current unlearning behavior\. This appendix provides additional output\-level and directional summaries across dataset complexity, model size, and architecture\. The purpose is not to expand the paper into a benchmark comparison, but to show that the main persistence claim continues to hold across a broader range of experiments\.

### G\.1Additional Output\-level Accuracy Tables

We first report output\-level accuracy across four additional settings to rule out the weaker explanation that the representation\-level failures identified in the main paper are simply caused by poor output\-level forgetting\. For each setting, we report forget\-set accuracyAccu\\mathrm\{Acc\}\_\{u\}, retain\-set accuracyAccr\\mathrm\{Acc\}\_\{r\}, and the harmonic meanHM=2⋅Accr⋅\(100−Accu\)Accr\+100−Accu,\\mathrm\{HM\}=\\frac\{2\\cdot\\mathrm\{Acc\}\_\{r\}\\cdot\(100\-\\mathrm\{Acc\}\_\{u\}\)\}\{\\mathrm\{Acc\}\_\{r\}\+100\-\\mathrm\{Acc\}\_\{u\}\},with the original and retrained models serving as lower and upper references, respectively\. Bold indicates the top threeHM\\mathrm\{HM\}values among unlearning methods\. All metrics are reported in%\\%\.

MethodAccu\\mathrm\{Acc\}\_\{u\}↓\\downarrowAccr\\mathrm\{Acc\}\_\{r\}↑\\uparrowHM\\mathrm\{HM\}↑\\uparrowOriginal99\.4196\.871\.17Retrain0\.0095\.9197\.91SCRUB0\.3999\.2099\.40Boundary Shrink1\.4149\.7766\.15UNSIR0\.0021\.3335\.16Amnesiac0\.0091\.1895\.39SSD0\.0064\.1178\.13POUR\-P0\.0096\.8798\.41POUR\-D0\.0065\.5779\.21Figure 17:Output\-level accuracy for CIFAR\-100 / ResNet\-18\.

MethodAccu\\mathrm\{Acc\}\_\{u\}↓\\downarrowAccr\\mathrm\{Acc\}\_\{r\}↑\\uparrowHM\\mathrm\{HM\}↑\\uparrowOriginal96\.7991\.996\.20Retrain0\.0093\.5796\.68SCRUB67\.7798\.2848\.54Boundary Shrink0\.2244\.5561\.60UNSIR0\.0019\.6932\.90Amnesiac0\.0090\.8595\.21SSD0\.0068\.6281\.39POUR\-P0\.0091\.6595\.64POUR\-D0\.0078\.8888\.19Figure 18:Output\-level accuracy for CIFAR\-100 / ResNet\-50\. SCRUB fails to complete forgetting, withAccu=67\.77%\\mathrm\{Acc\}\_\{u\}=67\.77\\%andHM=48\.54%\\mathrm\{HM\}=48\.54\\%\.

MethodAccu\\mathrm\{Acc\}\_\{u\}↓\\downarrowAccr\\mathrm\{Acc\}\_\{r\}↑\\uparrowHM\\mathrm\{HM\}↑\\uparrowOriginal92\.2274\.6314\.09Retrain0\.0078\.0787\.68SCRUB26\.5079\.0776\.18Boundary Shrink10\.0458\.2370\.70UNSIR0\.0053\.4469\.66Amnesiac1\.0076\.4886\.29SSD0\.0022\.9137\.28POUR\-P0\.0074\.6085\.45POUR\-D0\.4063\.0777\.23Figure 19:Output\-level accuracy for TinyImageNet / ResNet\-18\.
MethodAccu\\mathrm\{Acc\}\_\{u\}↓\\downarrowAccr\\mathrm\{Acc\}\_\{r\}↑\\uparrowHM\\mathrm\{HM\}↑\\uparrowOriginal99\.8099\.950\.39Retrain0\.0099\.9599\.97SCRUB0\.0099\.9499\.97Boundary Shrink9\.8187\.4488\.79UNSIR0\.0039\.9257\.06Amnesiac0\.0084\.1591\.39SSD0\.0080\.5489\.22POUR\-P0\.0099\.9599\.98POUR\-D0\.2198\.7799\.28Figure 20:Output\-level accuracy for CIFAR\-100 / ViT\-Tiny\.

Taken together, these tables support the main paper’s interpretation rather than weaken it\. In most settings, several methods still achieve strong output\-level forgetting and high harmonic mean, so the representation\-level failures discussed in the paper cannot be dismissed as a trivial byproduct of poor output behavior\. This is especially true for methods such as Amnesiac and POUR\-P, which often remain strong under output\-level evaluation while still exhibiting directional and geometric mismatch relative to retraining\. At the same time, scaling makes some failures more visible, with SCRUB no longer forgetting effectively on CIFAR\-100 with ResNet\-50 \(Tab\.[18](https://arxiv.org/html/2606.25001#A7.F18)\) and on TinyImageNet with ResNet\-18 \(Tab\.[20](https://arxiv.org/html/2606.25001#A7.F20)\), while Boundary Shrink and UNSIR continue to suffer severe retain\-side degradation \(Tab\.[18](https://arxiv.org/html/2606.25001#A7.F18),[18](https://arxiv.org/html/2606.25001#A7.F18),[20](https://arxiv.org/html/2606.25001#A7.F20)\) \. Thus, broader output\-level summaries do not overturn the main diagnosis\. They reinforce the central claim that strong output\-level behavior can coexist with retraining\-inconsistent representation\-level structure\.

### G\.2Additional Directional Alignment Plots

We next report directional alignment across the same four settings to test whether the forget/retain asymmetry identified in the main paper persists beyond the CIFAR\-10 / ResNet\-18 baseline\. For each method, we plot the cosine similarity between the unlearning and retraining shifts, computed separately for forget and retain samples\. Values closer to11indicate stronger directional agreement with retraining, values near0indicate orthogonality, and negative values indicate opposing directions\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x24.png)Figure 21:Directional alignment for CIFAR\-100 / ResNet\-18\.
![Refer to caption](https://arxiv.org/html/2606.25001v1/x25.png)Figure 22:Directional alignment for CIFAR\-100 / ResNet\-50\. SCRUB shows negative alignment on both forget and retain samples, indicating that its representation shift moves in the opposite direction from retraining\.

![Refer to caption](https://arxiv.org/html/2606.25001v1/x26.png)Figure 23:Directional alignment for TinyImageNet / ResNet\-18\. Amnesiac shows slightly higher alignment on retain samples than on forget samples in this setting, deviating from the more common pattern observed elsewhere\.
![Refer to caption](https://arxiv.org/html/2606.25001v1/x27.png)Figure 24:Directional alignment for CIFAR\-100 / ViT\-Tiny\.

Across these settings, the same qualitative asymmetry persists\. Forget\-side alignment is typically moderate to high, whereas retain\-side alignment often remains near zero or below, especially in Fig\.[22](https://arxiv.org/html/2606.25001#A7.F22)and[24](https://arxiv.org/html/2606.25001#A7.F24)\. This shows that the main paper’s directional diagnosis is not tied to a single benchmark or backbone\. The two clearest exceptions, namely Amnesiac on TinyImageNet \(Fig\.[24](https://arxiv.org/html/2606.25001#A7.F24)\) and SCRUB on CIFAR\-100 with ResNet\-50 \(Fig\.[22](https://arxiv.org/html/2606.25001#A7.F22)\), do not weaken the argument; they represent stricter failures in which the method does not align with retraining even on forget samples\. Overall, these additional plots reinforce the main paper’s conclusion that current unlearning methods may partially align forget representations with retraining, yet remain persistently inconsistent with retain representations across increased dataset complexity, model size, and architectural change\.

## Appendix HExtended Discussion

This appendix expands the discussion in Section[5](https://arxiv.org/html/2606.25001#S5)\. It provides a fuller interpretation of the diagnosed failure mode, together with its implications for evaluation and method design\.

### H\.1A Hidden Failure Mode of Current Machine Unlearning

The central result of this paper is that current machine unlearning suffers from a hidden failure mode that standard output\-level evaluation systematically fails to detect\.

First, standard output\-level success signals are too weak\. Low forget\-set accuracy, low logit\-level membership inference, and strong retain accuracy can make an unlearned model appear successful even when substantial forget information remains recoverable in representation space\. In this sense, current evaluation can systematically mistake*apparent forgetting*for successful forgetting\.

Second, this failure is*structured rather than random*\. Current methods are not simply wrong in arbitrary ways\. Instead, a retraining\-consistent representation lens reveals a repeated pattern in which these methods often partially align with retraining on forget samples, diverge on retain samples, and leave residual leakage that is concentrated rather than diffuse\. The residual mismatch is therefore not well described as generic optimization error; it reflects a systematic deviation from retraining\-consistent forgetting\.

Taken together, these findings suggest that many current unlearning methods remain substantially inconsistent with retraining in representation space, even when they appear successful under output\-level evaluation\. In this sense, output\-level forgetting alone can be a misleading indicator of successful unlearning under a stronger representation\-level lens\.

### H\.2Implications for Evaluation

The most immediate implication of this failure mode is that current evaluation practice is incomplete\. If output\-level metrics alone can be satisfied while representation\-level residuals remain recoverable, then prevailing evaluation protocols can systematically overestimate successful unlearning\. This does not mean that output\-level metrics are uninformative\. They remain useful for measuring prediction\-layer behavior and black\-box leakage\. However, they are not strong enough to serve as the sole success signal when the goal is to remove the influence of the forget set from the model\.

Our results therefore support a stricter evaluation principle in which MU is assessed relative to a retrained reference rather than by output behavior alone\. In particular, evaluation should ask whether the unlearned model is*retraining\-consistent*in representation space, and whether the residual discrepancy is random or structured\. Under this lens, the key question is no longer just whether the model looks forgotten at the output layer, but whether it changes in a way that resembles retraining without the forget data\.

### H\.3Implications for Method Design

The same failure mode also has implications for algorithm design\. Many current methods are optimized around output\-level objectives, such as reducing forget\-set confidence or weakening output\-level membership signals\. Our results suggest that this is not sufficient to ensure consistency with retraining in representation space\. A method can satisfy these output\-level objectives while still exhibiting asymmetric and structured retraining\-inconsistent residuals\.

This does not imply that current methods are useless, nor that all of them fail in the same way\. Rather, it suggests that many methods are better understood as achieving*apparent forgetting*rather than retraining\-consistent forgetting\. Future unlearning methods may therefore benefit from stronger representation\-level principles, including reducing structured residuals in representation space and better approximating retraining\-consistent transformations\.

### H\.4Broader Perspective

The broader message of this paper is that the most informative structure in unlearning is not fully visible at the output layer\. Output\-level metrics remain useful, but they can be too weak to reveal whether unlearning actually follows retraining in representation space\. In this sense, retraining reveals what output forgetting hides by showing not only that residual discrepancy remains, but also how that discrepancy is organized\.

Seen this way, the central issue is not merely that some metrics are incomplete, but that current evaluation practice can certify successful\-looking unlearning on the basis of a signal that is too weak to rule out structured retraining\-inconsistent residuals\.

## Appendix ILimitations and Scope

Our goal is not to argue that output\-level metrics or retraining\-based references are useless in all settings\. Rather, the claim of this paper is narrower and more specific\. By themselves, they are*insufficient*to rule out a hidden failure mode in which apparent output\-level forgetting coexists with retraining\-inconsistent residuals in representation space\.

#### Output\-level metrics may be sufficient in practice\.

A natural counterargument is that many realistic adversaries only have black\-box access to model outputs, so evaluating MU through output\-level quantities such asMIAlogit\\mathrm\{MIA\}\_\{\\mathrm\{logit\}\}may be practically sufficient\[[21](https://arxiv.org/html/2606.25001#bib.bib39),[6](https://arxiv.org/html/2606.25001#bib.bib13)\]\. We agree that output\-level metrics remain useful in such settings\. However, this does not alter the paper’s central diagnosis\. A growing number of models are released with open weights or otherwise permit access to internal representations\[[33](https://arxiv.org/html/2606.25001#bib.bib30)\], and white\-box access is also natural in auditing and compliance settings for data deletion\[[30](https://arxiv.org/html/2606.25001#bib.bib29)\]\. In these cases, output\-level success alone is insufficient to rule out structured residual discrepancies in the representation space\. Our results show that a model can appear forgotten at the output layer while still remaining substantially inconsistent with retraining internally\. Thus, output\-level evaluation remains informative, but it is not sufficient as a general success signal for the stronger notion of forgetting studied here\.

#### Retraining may be too expensive to serve as a practical reference\.

A second objection is that retraining from scratch is often computationally infeasible in large\-scale settings\[[29](https://arxiv.org/html/2606.25001#bib.bib31),[36](https://arxiv.org/html/2606.25001#bib.bib32)\], makingθr\\theta\_\{r\}an impractical operational target\. We agree that retraining is often too costly to use as a deployment\-time procedure\. But this is different from the role it plays in our paper\. Here,θr\\theta\_\{r\}is used as a*benchmark\-scale reference*for what stronger forgetting should look like, not as a claim that every practical system must retrain in deployment\. If the goal of MU is to approximate the model that would have been obtained without the forget set, then retraining remains a natural comparative reference against which to diagnose failure\. Under this interpretation, retraining is useful not because it must always be deployed, but because it reveals mismatches that weaker endpoint\-based criteria can miss\.

More broadly, our argument does not require retraining\-consistent geometry to be the only valid notion of forgetting\. We use retraining as a stronger reference because, in our controlled setting, it captures the model that would have been obtained without the forget data\. The key point is therefore comparative, not absolute, because weaker output\-level criteria may certify apparent forgetting even when substantial retraining\-inconsistent residuals remain detectable in representation space\.

As a whole, these considerations clarify the scope of our claim rather than weaken it\. Output\-level evaluation remains useful, and retraining is not always operationally feasible\. Our point is that current machine unlearning is often evaluated using a signal that is too weak to detect the structured hidden failure mode studied in this paper\.

Similar Articles

Forgetting is Not Erasure: Recovering Latent Knowledge via Transport Keys

arXiv cs.LG

This paper argues that catastrophic forgetting in neural networks is not erasure but an interface alignment problem. It introduces 'transport keys' to recover latent task-specific features from sequentially trained models, demonstrating significant performance recovery on split CIFAR-100.

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

arXiv cs.CL

This paper shows that a language model with a lossy memory that retains a wrong conclusion but drops the evidence produces confident incorrect answers, whereas an empty memory leads to abstention. The authors propose a source-first compression policy that preserves recomputable sources instead of conclusions to maintain correctability, and demonstrate the mechanism across multiple models and dialogue systems.