Data Unlearning via Inverse Distillation
Summary
The paper introduces Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a multi-step flow-matching or diffusion teacher into an efficient one-step student while suppressing generation of forgotten training data, requiring only the teacher and forget-set samples without access to retained data or auxiliary classifiers.
View Cached Full Text
Cached at: 09/30/26, 09:45 AM
# Data Unlearning via Inverse Distillation
Source: [https://arxiv.org/html/2609.36099](https://arxiv.org/html/2609.36099)
Nikita KornilovApplied AI Institute, Moscow, RussiaMIRAI, Moscow, RussiaBRAIn Lab, Moscow, Russiajhomanik14@gmail\.comZhang ZhenheAI Foundation lab, Moscow, RussiaEvgeny BurnaevApplied AI Institute, Moscow, RussiaAXXX, Moscow, RussiaIaroslav KoshelevAI Foundation lab, Moscow, RussiaAlexander KorotinApplied AI Institute, Moscow, RussiaAXXX, Moscow, Russiaiamalexkorotin@gmail\.com
###### Abstract
Multi\-step matching models, including flow and diffusion models, produce high\-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets\. We introduce Inverse Distillation Unlearning \(IDU\), a unified framework that simultaneously distills a teacher multi\-step matching model into an efficient one\-step student generator and suppresses outputs corresponding to a designated training subset\. We first formulate distillation as a min\-max objective over a data distribution and then represent this distribution as a mixture of the forget\-set and the generated distributions\. This allows us to compare this mixture with the teacher’s training distribution and recover only the retained data at the optimum\. Our method requires only a pretrained full\-data teacher and data from the forget set, without access to retained training examples, extra feature extractors or classifiers\. Extensive experiments on MNIST and CIFAR\-10 datasets under flow\-matching and score\-based diffusion settings demonstrate that IDU substantially reduces the generation frequency of forgotten classes while preserving generation quality on the retained classes\. To the best of our knowledge, IDU is the first unified framework for simultaneous unlearning and distillation in unconditional flow\-matching and score\-based models\.
## 1Introduction
Multi\-step diffusion\([Sohl\-Dickstein et al\., 2015](https://arxiv.org/html/2609.36099#bib.bib19);[Ho et al\., 2020](https://arxiv.org/html/2609.36099#bib.bib20);[Song et al\., 2021](https://arxiv.org/html/2609.36099#bib.bib21)\), flow\([Lipman et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib25);[Liu et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib26)\), and other matching models\([Holderrieth et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib17);[Gao et al\., 2025b](https://arxiv.org/html/2609.36099#bib.bib18)\), followed by one\-step generators\([Kim et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib41);[Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14);[Yin et al\., 2024a](https://arxiv.org/html/2609.36099#bib.bib13);[Frans et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib30)\), now produce high\-quality and diverse samples\. Yet their reliance on large, often imperfectly moderated web datasets\([Schuhmann et al\., 2022](https://arxiv.org/html/2609.36099#bib.bib1)\)raises privacy, copyright, legal, and ethical concerns\([Voigt and Von dem Bussche, 2017](https://arxiv.org/html/2609.36099#bib.bib6);[Goldman, 2020](https://arxiv.org/html/2609.36099#bib.bib12)\): models may memorize and reproduce unwanted training content\([Carlini et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib2)\)\. To avoid such incidents, a critical research direction known asmachine unlearning \(MU\)has emerged within the field of trustworthy machine learning\([Bourtoule et al\., 2021](https://arxiv.org/html/2609.36099#bib.bib15);[Nguyen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib16)\)\. At its core, MU aims to remove the influence of specific training data, classes, or semantic concepts from trained generative models, while preserving the same generation quality for the remaining data\. The MU literature splits into two main paradigms:data unlearning and class unlearning\.
Data unlearning\([Alberti et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib43)\)addresses scenarios where the model is fine\-tuned to completely erase the influence of*particular data*for which no class or prompt anchor is available — such as individual faces or Not\-Safe\-For\-Work \(NSFW\) images\. In this setup, the forget data is typically provided as selected samples\. Ideally, the goal is to obtain a model as if it were trained only on the remaining data, yet without retraining from scratch\. Data unlearning naturally appears in unconditional models, but can also be applied to conditional variants when some data should be erased within a given class\.
The key challenge in this setup is that the forget and remaining data are deeply entangled within the model parameters — the forget data cannot be directly prompted during generation\. Furthermore, the quality of the remaining data must be preserved, even though its samples may be unavailable\. As a result, typical data unlearning algorithms distinguish between remaining and forget data by employing different losses on them\([Alberti et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib43);[Wu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib40);[Jiang et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib46)\), constrained optimization\([Khalafi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib48)\), variational framework\([Panda et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib5)\), importance sampling\([Shi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib47)\), energy functions\([Simone et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib49)\), or transport costs\([Choi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib45)\)\.
Data unlearning algorithms are well\-explored for matching models, with both general frameworks and model\-specific solutions \(e\.g\., flow matching\([Simone et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib49)\)or score\-based models\([Jiang et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib46)\)\)\.*Nevertheless, little progress\([Choi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib45)\)has been made toward effective, easy\-to\-tune, and model\-agnostic approaches for one\-step generators\.*
Class unlearning\([Gandikota et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib38)\)aims to removeall knowledgepertaining to unwanted classes in the conditional models\. More specifically, unlearned models are fine\-tuned to output noise or irrelevant data when conditioned on a particular class, such as an unsafe object category, while preserving the generative quality for all other classes\. In contrast to data unlearning, the forget data in this setup is typically provided as class labels, without any data samples\. Concept unlearning extends this idea to text\-to\-image models, where the goal is to remove entire concepts or styles \(e\.g\., artistic style, nudity or celebrity’s likeness\) that may be triggered by textual prompts\.
Most class unlearning approaches optimize a combination of two losses: a forget loss that swaps the undesirable class with another one, and a remaining loss that preserves quality for other classes\. Various techniques have been developed for designing and combining these losses for diffusion and flow models, including steer\-away guidance\([Gandikota et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib38)\), Bayesian continual learning\([Heng and Soh, 2023](https://arxiv.org/html/2609.36099#bib.bib36)\), saliency\-based weight updates\([Fan et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib37)\), cross\-attention editing\([Gandikota et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib31);[Lu et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib33)\), attention re\-steering\([Zhang et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib27)\), pruning\([Chavhan et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib34)\), adversarial strategies\([Bui et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib29)\), multi\-objective optimization to avoid conflicts between loss gradients\([Wu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib40)\), class swapping during distillation\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\), attention regularization\([Gao et al\., 2025a](https://arxiv.org/html/2609.36099#bib.bib50);[Fan et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib51)\), and unlearning irreversibility techniques\([Sharma et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib32);[Liu and Zhang, 2025](https://arxiv.org/html/2609.36099#bib.bib35);[Lu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib28)\)\.However, only a few papers\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\)tackle class unlearning in one\-step models, motivating further research\.
### 1\.1Contributions
We fill the gap in data unlearning for one\-step models and propose our novelInverse Distillation Unlearning \(IDU\)approach\. Our IDU can erase undesired data samples from a pretrained one\-step generator or build a new one without them, using only the teacher matching model trained on the original dataset\. Our method follows an inverse distillation pipeline and compares the current mixture of the generated and forget data with the teacher’s correct one, penalizing the generator for reproducing unwanted samples\. Moreover, IDU can work with different matching teachers, such as diffusion or flow models, and requires neither extra training\-time feature extractors nor classifiers; its only method\-specific trade\-off parameter isρ∈\[0,1\)\\rho\\in\[0,1\)\.
## 2Related Work
### 2\.1Diffusion, flow and matching models
Denoising diffusion probabilistic model\([Ho et al\., 2020](https://arxiv.org/html/2609.36099#bib.bib20),DDPM\)defines a multi\-step forward process, mapping data to Gaussian noise, and learns to reverse it via a trained denoiser\. Score\-based generative model\([Song et al\., 2021](https://arxiv.org/html/2609.36099#bib.bib21),SGM\)extends this idea to continuous time and approximates a score function to simulate the reverse SDE\. Flow matching\([Lipman et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib25),FM\)instead learns an ODE drift that interpolates between data and noise, enabling fewer sampling steps via advanced ODE solvers and flexible interpolations\. In all of the above cases, a model must approximate an intractable reverse\-process function \(denoiser, score, drift, etc\.\)\. It is done via available unbiased estimates of this function, conditioned on initial data samples\. Since a model is usually matched with these conditional function estimates, we refer to such models collectively asmatching models\.
Formally, a matching model constructs a probability pathptp\_\{t\}on the time interval\[0,T\]\[0,T\], transforming the selected datap0p\_\{0\}to noisepTp\_\{T\}\. This pathpt\(xt\)=∫ℝDpt\(xt\|x0\)p0\(x0\)dx0p\_\{t\}\(x\_\{t\}\)=\\int\_\{\\mathbb\{R\}^\{D\}\}p\_\{t\}\(x\_\{t\}\|x\_\{0\}\)p\_\{0\}\(x\_\{0\}\)dx\_\{0\}is built as a mixture of simple conditional pathspt\(⋅\|x0\)p\_\{t\}\(\\cdot\|x\_\{0\}\)conditioned on samplesx0∼p0x\_\{0\}\\sim p\_\{0\}\. Then, the standard universal matching \(UM\) lossℒUM\(f,p0\)\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}\)matches a modelf:\[0,T\]×ℝD→ℝDf:\[0,T\]\\times\\mathbb\{R\}^\{D\}\\to\\mathbb\{R\}^\{D\}with conditional estimatesftp0\(⋅\|x0\)f^\{p\_\{0\}\}\_\{t\}\(\\cdot\|x\_\{0\}\)at each timettand pointxt∼ptx\_\{t\}\\sim p\_\{t\}:
ℒUM\(f,p0\):=𝔼t,x0∼p0,xt∼pt\(⋅\|x0\)\[∥ft\(xt\)−ftp0\(xt\|x0\)∥2\]\.\\\!\\\!\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}\)\\\!\\\!:=\\mathbb\{E\}\_\{t,x\_\{0\}\\sim p\_\{0\},x\_\{t\}\\sim p\_\{t\}\(\\cdot\|x\_\{0\}\)\}\[\\\|f\_\{t\}\(x\_\{t\}\)\-f^\{p\_\{0\}\}\_\{t\}\(x\_\{t\}\|x\_\{0\}\)\\\|^\{2\}\]\.\(1\)Here, the notation𝔼t\\mathbb\{E\}\_\{t\}hides the time sampling and loss weighting inherent to the given matching model\. For example, denoising models recover unnoised samplesftp0\(xt\|x0\)=x0f\_\{t\}^\{p\_\{0\}\}\(x\_\{t\}\|x\_\{0\}\)=x\_\{0\}, score\-based models match the conditional scoreftp0\(xt\|x0\)=∇xtlnpt\(xt\|x0\)f\_\{t\}^\{p\_\{0\}\}\(x\_\{t\}\|x\_\{0\}\)=\\nabla\_\{x\_\{t\}\}\\ln p\_\{t\}\(x\_\{t\}\|x\_\{0\}\), and flow models with the linear interpolationxt=\(1−t/T\)x0\+\(t/T\)xTx\_\{t\}=\(1\-t/T\)x\_\{0\}\+\(t/T\)x\_\{T\}match the conditional driftftp0\(xt\|x0\)=\(xT−x0\)/T=\(xt−x0\)/tf\_\{t\}^\{p\_\{0\}\}\(x\_\{t\}\|x\_\{0\}\)=\(x\_\{T\}\-x\_\{0\}\)/T=\(x\_\{t\}\-x\_\{0\}\)/tfort\>0t\>0\.
### 2\.2Distillation and one\-step models
Multi\-step sampling makes matching models slower than one\-step generators such as VAEs\([Kingma and Welling, 2013](https://arxiv.org/html/2609.36099#bib.bib22)\)and GANs\([Goodfellow et al\., 2014](https://arxiv.org/html/2609.36099#bib.bib24)\)\. Beyond sampling acceleration\([Lu et al\., 2022](https://arxiv.org/html/2609.36099#bib.bib7);[Karras et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib8)\), matching distillation methods\([Yin et al\., 2024b](https://arxiv.org/html/2609.36099#bib.bib10)\)train aone\-step generatorGθG\_\{\\theta\}under amulti\-step teacherf∗f^\{\*\}: a fake model fits the generated distribution, while the generator reduces its discrepancy from the teacher\. Despite differences in model type and discrepancy measure\([Yin et al\., 2024a](https://arxiv.org/html/2609.36099#bib.bib13);[Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14);[Gushchin et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib11)\), these methods admit a common inverse\-optimization view\([Kornilov et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib9)\)\. Givenf∗=argminfℒUM\(f,p0∗\)f^\{\*\}=\\argmin\_\{f\}\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\*\}\)trained via UM loss minimization, they recover its data distributionp0∗p\_\{0\}^\{\*\}\. To do this, they optimize the following min\-max inverse distillation scheme over trainable generated distributionp0θp\_\{0\}^\{\\theta\}:
minmaxθ\{ℒUM\(f∗,p0θ\)−ℒUM\(f,p0θ\)\}f=min\{ℒUM\(f∗,p0θ\)−min\{ℒUM\(f,p0θ\)\}f\}θ\.\\min\{\{\}\_\{\\theta\}\}\\max\{\{\}\_\{f\}\}\\left\\\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p^\{\\theta\}\_\{0\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p^\{\\theta\}\_\{0\}\)\\right\\\}=\\min\{\{\}\_\{\\theta\}\}\\bigl\\\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p\_\{0\}^\{\\theta\}\)\-\\min\{\{\}\_\{f\}\}\\\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\theta\}\)\\\}\\bigr\\\}\.\(2\)The non\-negative difference between losses in this scheme measures how well the teacher fits the current data compared to the best possible fake model\. When the teacher data is retrieved, i\.e\.,p0θ=p0∗p^\{\\theta\}\_\{0\}=p\_\{0\}^\{\*\}, the difference becomes 0 and the scheme attains optimum\.
### 2\.3Data unlearning
In data unlearning, we have an original data distributionp0∗p^\{\*\}\_\{0\}, a forget\-data distributionp0Fp^\{F\}\_\{0\}whose influence we would like to remove, and the remaining datap0Rp^\{R\}\_\{0\}\. The ideal solution would be a model trained on the remaining datapRp^\{R\}via the vanilla lossℒvanilla\\mathcal\{L\}\_\{\\text\{vanilla\}\}\. However, instead of training from scratch, the early works\([Golatkar et al\., 2020](https://arxiv.org/html/2609.36099#bib.bib39);[Thudi et al\., 2022](https://arxiv.org/html/2609.36099#bib.bib4);[Tang and Khanna, 2026](https://arxiv.org/html/2609.36099#bib.bib3)\)propose to fine\-tune a weighted sum of the forget and remaining losses with a trade\-off factorρ∈\[0,1\)\\rho\\in\[0,1\):
ℒbase=ρ⋅ℒforget\+\(1−ρ\)⋅ℒremain,\\mathcal\{L\}\_\{\\text\{base\}\}=\\rho\\cdot\\mathcal\{L\}\_\{\\text\{forget\}\}\+\(1\-\\rho\)\\cdot\\mathcal\{L\}\_\{\\text\{remain\}\},\(3\)where the forget lossℒforget\\mathcal\{L\}\_\{\\text\{forget\}\}keeps the model away from the forget data, whereas the remaining lossℒremain\\mathcal\{L\}\_\{\\text\{remain\}\}ensures that the quality on the remaining samples does not degrade\.
Forunlearning in matching models, the typical losses areℒvanilla\(f\)=ℒUM\(f,p0R\)\\mathcal\{L\}\_\{\\text\{vanilla\}\}\(f\)=\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{R\}\),ℒforget\(f\)=−ℒUM\(f,p0F\)\\mathcal\{L\}\_\{\\text\{forget\}\}\(f\)=\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{F\}\), andℒremain\(f\)=ℒUM\(f,p0R\)\\mathcal\{L\}\_\{\\text\{remain\}\}\(f\)=\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{R\}\)\. Many data unlearning methods modify the forget and remaining losses along with the optimization procedure between them\.NegGradminimizes only forget lossℒforget\(f\)\\mathcal\{L\}\_\{\\text\{forget\}\}\(f\)\. However, this approach often leads to instability and catastrophic forgetting, degrading the overall generation quality\.SA\([Heng and Soh, 2023](https://arxiv.org/html/2609.36099#bib.bib36)\)adds a computationally heavy Elastic Weight Consolidation penalty to the base loss \- the quadratic form of divergence from the initial weights, computed using the Fisher Information Matrix on the current generated data\.SalUn\([Fan et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib37)\)optimizes the base loss and applies a binary gradient mask, selecting only the most significant parameters for forgetting\. This mask is calculated by thresholding the gradient magnitudes of the forget loss\.SISS\([Alberti et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib43)\)employs importance sampling to call the model only once per base lossℒbase\(f\)\\mathcal\{L\}\_\{\\text\{base\}\}\(f\)calculation, but basically does not change the loss structure\.MGSM\([Jiang et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib46)\)incorporates more natural score\-function orthogonality instead ofℓ2\\ell\_\{2\}\-loss for the forget part\.EraseDiff\([Wu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib40)\)utilizes random\-noise matching loss on forget samples and employs constrained optimization to smoothly merge the gradient directions of the losses\.Retrack\([Shi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib47)\)equips the remaining lossℒremain\(f\)\\mathcal\{L\}\_\{\\text\{remain\}\}\(f\)with the nearest\-neighbor importance weighting from the forget subset\. The work\([Khalafi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib48)\)optimizes KL divergence between noising processes, also under the constrained optimization formulation\.
Other methods use various types of guidance to avoid the forget data\.VDU\([Panda et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib5)\)uses a variational inference framework with a plasticity inducer for reducing the likelihood of unwanted data and a stability regularizer for quality preservation\. For flow models,ContinualFlow\([Simone et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib49)\)applies energy\-function weighting to the original loss to suppress unwanted data\.
One\-step model unlearning\.UOT\-Unlearn\([Choi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib45)\)uses unbalanced optimal transport to shift a pretrained one\-step generator away from unwanted samples\. Its cost function requires a precomputed feature extractor and hyperparameter tuning rather than a trained teacher\. However, the OT framework has limited scalability and generation diversity, compared with matching models distillation\. The same work adapts VDU, SalUn, and SA to consistency models\([Kim et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib41);[Geng et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib42);[Frans et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib30)\); these adaptations use a classifier to identify forget samples and do not directly extend to non\-self\-consistent generators\.
### 2\.4Class unlearning
Class unlearning aims to erase entire unwanted classes from a conditional model while keeping other classes intact\. Unlike specific data samples, whose influence in an unconditional model is difficult to trace, classes in a conditional model can be efficiently prompted and distinguished from one another\. For the same reason, the majority of class\-forgetting methods are data\-free\. These features enable a variety of methods formatching model unlearningthat bind the attributes of the unwanted classes to completely different entities, especially within the attention or inference structure\([Gandikota et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib38);[Gandikota et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib31);[Lu et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib33);[Zhang et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib27);[Gao et al\., 2025a](https://arxiv.org/html/2609.36099#bib.bib50);[Fan et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib51)\)\. Such methods are not always adaptable to data unlearning; nevertheless, they often use the same combination of the forget and remaining losses: the forget loss changes the model’s behavior on unwanted classes, while the remaining loss preserves it on the others\. For example, SA\([Heng and Soh, 2023](https://arxiv.org/html/2609.36099#bib.bib36)\), SalUn\([Fan et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib37)\), and EraseDiff\([Wu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib40)\)can be leveraged for both data and class unlearning tasks\.
One\-step model unlearning\.SFD\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\)performs class forgetting during conditional inverse distillation \([2](https://arxiv.org/html/2609.36099#S2.E2)\) by replacing the teacher score for forgotten classes with a safe\-class score in the generator loss; its full generator losses are given in Appendix[A\.3](https://arxiv.org/html/2609.36099#A1.SS3)\.
## 3Inverse Distillation Unlearning
### 3\.1Method Description
#### Setup and preliminaries\.
In data unlearning, we are given a forget\-data distributionp0Fp^\{F\}\_\{0\}that we would like to remove from the original distributionp0∗p^\{\*\}\_\{0\}, so that only the remaining \(or retained\) datap0Rp^\{R\}\_\{0\}is preserved:
p0∗=πp0F\+\(1−π\)p0R,p^\{\*\}\_\{0\}=\\pi\\,p^\{F\}\_\{0\}\+\(1\-\\pi\)\\,p^\{R\}\_\{0\},\(4\)whereπ∈\[0,1\)\\pi\\in\[0,1\)is the proportion of the forget data\. In our setup, neither the original nor the remaining data is available, but we do have access to a teacher modelf∗=argminfℒUM\(f,p0∗\)f^\{\*\}=\\argmin\_\{f\}\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f,p^\{\*\}\_\{0\}\)of an arbitrary matching type, trained on the original data\. We aim to train a one\-step generatorGθ:𝒵→ℝDG\_\{\\theta\}:\\mathcal\{Z\}\\to\\mathbb\{R\}^\{D\}with parametersθ\\thetathat will eventually reproduce only the remaining datap0Rp\_\{0\}^\{R\}\. The generator maps the latent distributionp𝒵p\_\{\\mathcal\{Z\}\}to a distributionp0θp\_\{0\}^\{\\theta\}and can be pretrained or initialized with a one\-step teacher inference scheme\.
The standard way to distill the datap0∗p^\{\*\}\_\{0\}, stored inside the teacher modelf∗f^\{\*\}, into the trainable distributionp0p\_\{0\}is to apply the inverse distillation scheme \([2](https://arxiv.org/html/2609.36099#S2.E2)\)\. This min\-max scheme reverses the forward minimization problem for obtaining the teacher from the fixed input data, i\.e\., it retrieves the data from which the fixed teacher was obtained:
minmaxp0\{ℒUM\(f∗,p0\)−ℒUM\(f,p0\)\}f∼min\{ℒUM\(f∗,p0\)−minf\{ℒUM\(f,p0\)⏟≥0\}p0\}\.\\min\{\{\}\_\{p\_\{0\}\}\}\\max\{\{\}\_\{f\}\}\\left\\\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p\_\{0\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}\)\\right\\\}\\sim\\min\{\{\}\_\{p\_\{0\}\}\}\\bigl\\\{\\underset\{\\geq 0\}\{\\underbrace\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p\_\{0\}\)\-\\min\{\{\}\_\{f\}\}\\\{\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}\)\}\}\\\}\\bigr\\\}\.\(5\)The non\-negative difference between losses in this scheme measures how well the teacher fits the current data compared to the best possible fake model\. For the teacher datap0=p0∗p\_\{0\}=p\_\{0\}^\{\*\}, the optimum is attained with the zero difference\.
#### Our approach\.
We build ourInverse Distillation Unlearning \(IDU\)method as follows: we run the inverse distillation scheme \([5](https://arxiv.org/html/2609.36099#S3.E5)\) but parametrize the optimized distributionp0p\_\{0\}as the mixed datap0=ρp0F\+\(1−ρ\)p0θ=:p0mixp\_\{0\}=\\rho\\,p^\{F\}\_\{0\}\+\(1\-\\rho\)\\,p^\{\\theta\}\_\{0\}=:p^\{\\mathrm\{mix\}\}\_\{0\}, similar to the data mix \([4](https://arxiv.org/html/2609.36099#S3.E4)\) with the proportionρ∈\[0,1\)\\rho\\in\[0,1\), where we substitute the desired remaining datap0Rp^\{R\}\_\{0\}with the generated onep0θp\_\{0\}^\{\\theta\}\. Thus, the optimal generator with the right proportionρ=π\\rho=\\pihas to learn only the remaining data in order to recover the full teacher distribution, sincethe forget component is already accounted for by the forget samples\. The following theorem formalizes the generator’s forgetting property\. The proof is given in Appendix[A\.1](https://arxiv.org/html/2609.36099#A1.SS1)\.
###### Theorem 1\(IDU’s forgetting property\)
Optimization of IDU loss \([6](https://arxiv.org/html/2609.36099#S3.E6)\) withρ=π\\rho=\\piretrieves only the remaining data, i\.e\., the optimal generator parametersθopt\\theta\_\{\\text\{opt\}\}yieldp0θopt=p0Rp^\{\\theta\_\{\\text\{opt\}\}\}\_\{0\}=p\_\{0\}^\{R\}\.
More specifically, we optimize the followingmin\-max IDU objectiveℒIDU\(f,p0θ\)\\mathcal\{L\}\_\{\\text\{IDU\}\}\(f,p\_\{0\}^\{\\theta\}\)over generator parametersθ\\thetaand fake modelff:
ℒIDU\(f,p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{IDU\}\}\(f,p\_\{0\}^\{\\theta\}\):=\\displaystyle:=ℒUM\(f∗,p0mix\)−ℒUM\(f,p0mix\)\\displaystyle\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p\_\{0\}^\{\\mathrm\{mix\}\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\)\(6\)=\\displaystyle=ℒUM\(f∗,ρp0F\+\(1−ρ\)p0θ\)−ℒUM\(f,ρp0F\+\(1−ρ\)p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},\\rho\\,p^\{F\}\_\{0\}\+\(1\-\\rho\)\\,p^\{\\theta\}\_\{0\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,\\rho\\,p^\{F\}\_\{0\}\+\(1\-\\rho\)\\,p^\{\\theta\}\_\{0\}\)=\\displaystyle=ρ⋅\[ℒUM\(f∗,p0F\)−ℒUM\(f,p0F\)\]\+\(1−ρ\)⋅\[ℒUM\(f∗,p0θ\)−ℒUM\(f,p0θ\)\],\\displaystyle\\rho\\cdot\[\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p^\{F\}\_\{0\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p^\{F\}\_\{0\}\)\]\+\(1\-\\rho\)\\cdot\[\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p^\{\\theta\}\_\{0\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p^\{\\theta\}\_\{0\}\)\],where the second equality holds because the UM losses \([1](https://arxiv.org/html/2609.36099#S2.E1)\) are linear in the input data, which appears only inside a mathematical expectation\.
### 3\.2Practical details
#### Optimization procedure\.
We alternate between two steps to optimize our min\-max IDU loss \([6](https://arxiv.org/html/2609.36099#S3.E6)\):
1\) First, we update the fake modelffvia the fake lossℒIDU\-fake\(f\)\\mathcal\{L\}\_\{\\text\{IDU\-fake\}\}\(f\):
ℒIDU\-fake\(f\)=ρ⋅ℒUM\(f,p0F\)\+\(1−ρ\)⋅ℒUM\(f,p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{IDU\-fake\}\}\(f\)=\\rho\\cdot\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p^\{F\}\_\{0\}\)\+\(1\-\\rho\)\\cdot\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p^\{\\theta\}\_\{0\}\)=𝔼t,x0θ∼pθ0,xtθ∼pθt\(⋅\|x0θ\)x0F∼pF0,xtF∼pFt\(⋅\|x0F\)\[ρ‖ft\(xtF\)−ftF\(xtF\|x0F\)‖2\+\(1−ρ\)‖ft\(xtθ\)−ftθ\(xtθ\|x0θ\)‖2\],\\displaystyle=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}t,x\_\{0\}^\{\\theta\}\\sim p^\{\\theta\}\_\{0\},x\_\{t\}^\{\\theta\}\\sim p^\{\\theta\}\_\{t\}\(\\cdot\|x\_\{0\}^\{\\theta\}\)\\\\ x\_\{0\}^\{F\}\\sim p^\{F\}\_\{0\},x\_\{t\}^\{F\}\\sim p^\{F\}\_\{t\}\(\\cdot\|x\_\{0\}^\{F\}\)\\end\{subarray\}\}\[\\rho\\\|f\_\{t\}\(x^\{F\}\_\{t\}\)\-f^\{F\}\_\{t\}\(x\_\{t\}^\{F\}\|x\_\{0\}^\{F\}\)\\\|^\{2\}\+\(1\-\\rho\)\\\|f\_\{t\}\(x\_\{t\}^\{\\theta\}\)\-f^\{\\theta\}\_\{t\}\(x\_\{t\}^\{\\theta\}\|x\_\{0\}^\{\\theta\}\)\\\|^\{2\}\],\(7\)whereptθ\(⋅\|x0θ\)p^\{\\theta\}\_\{t\}\(\\cdot\|x\_\{0\}^\{\\theta\}\)andptF\(⋅\|x0F\)p^\{F\}\_\{t\}\(\\cdot\|x\_\{0\}^\{F\}\)are the conditional forward noising processes built on the generated and forget data with the corresponding conditional estimatesftθ\(xtθ\|x0θ\)f^\{\\theta\}\_\{t\}\(x\_\{t\}^\{\\theta\}\|x\_\{0\}^\{\\theta\}\)andftF\(xtF\|x0F\)f^\{F\}\_\{t\}\(x\_\{t\}^\{F\}\|x\_\{0\}^\{F\}\)\.
2\) Next, we update the generator parametersθ\\thetawith a fixed fake modelffvia the generator lossℒIDU\-gen\(p0θ\)\\mathcal\{L\}\_\{\\text\{IDU\-gen\}\}\(p\_\{0\}^\{\\theta\}\); however, instead of the default lossℒIDU\-gen\(p0θ\)=ℒUM\(f∗,p0θ\)−ℒUM\(f,p0θ\),\\mathcal\{L\}\_\{\\text\{IDU\-gen\}\}\(p\_\{0\}^\{\\theta\}\)=\\mathcal\{L\}\_\{\\text\{UM\}\}\(f^\{\*\},p\_\{0\}^\{\\theta\}\)\-\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\theta\}\),we use its modified version, proposed in the Score identity Distillation \(SiD\) framework\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\):
ℒIDU\-gen\(p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{IDU\-gen\}\}\(p\_\{0\}^\{\\theta\}\)=2𝔼t,x0θ∼p0θ,xθt∼pθt\(⋅\|xθ0\)\[⟨ft∗\(xtθ\)−ft\(xtθ\),ft∗\(xtθ\)−ftθ\(xtθ\|x0θ\)⟩−αSiD‖ft∗\(xtθ\)−ft\(xtθ\)‖2\],\\displaystyle=2\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}t,x\_\{0\}^\{\\theta\}\\sim p^\{\\theta\}\_\{0\},\\\\ x^\{\\theta\}\_\{t\}\\sim p^\{\\theta\}\_\{t\}\(\\cdot\|x^\{\\theta\}\_\{0\}\)\\end\{subarray\}\}\\\!\\\!\[\\langle f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\),f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\)\-f\_\{t\}^\{\\theta\}\(x^\{\\theta\}\_\{t\}\|x^\{\\theta\}\_\{0\}\)\\rangle\-\\alpha\_\{\\text\{SiD\}\}\\\|f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\)\\\|^\{2\}\],where we heuristically scale the second term by exactlyαSiD\\alpha\_\{\\text\{SiD\}\}times\. This scale factorαSiD\\alpha\_\{\\text\{SiD\}\}is usually taken from the rangeαSiD∈\[0\.5,1\.2\]\\alpha\_\{\\text\{SiD\}\}\\in\[0\.5,1\.2\]\. For example, in caseαSiD=0\.5\\alpha\_\{\\text\{SiD\}\}=0\.5, we end up with the default theoretical loss, whereas greater factor values can yield better performance in practice\. Nevertheless, the optimal value depends strongly on the matching model type and neural network architecture\.
Figure 1:Pipeline of our IDU framework\. Forget and generated samples form two forward\-noising branches\. The framework alternates between two steps: first, update the fake model on both branches with weightsρ\\rhoand1−ρ1\-\\rhoto detect generations similar to the forget samples, then update the generator to suppress these generations using the frozen teacher and the fake model\.
#### Algorithm pseudocode\.
In Algorithm[1](https://arxiv.org/html/2609.36099#alg1), we provide IDU pseudocode for the flow\-matching setting\. The general IDU training pipeline is illustrated in Figure[1](https://arxiv.org/html/2609.36099#S3.F1)\.
Algorithm 1Inverse Distillation Unlearning0:teacher model
f∗f^\{\*\}, generator
GθG\_\{\\theta\}\(pretrained or initialized by the teacher\), fake model
fψf\_\{\\psi\}, forget data
p0Fp\_\{0\}^\{F\}, forgetting scale
ρ∈\[0,1\)\\rho\\in\[0,1\), SiD scale
αSiD∈\[0\.5,1\.2\]\\alpha\_\{\\text\{SiD\}\}\\in\[0\.5,1\.2\], number of iterations
NN, batch size
BB, latent distribution
p𝒵p\_\{\\mathcal\{Z\}\}, noise distribution
pTp\_\{T\}\.
1:for
n=0,…,N−1n=0,\\ldots,N\-1do
2:Sample generated and noise batches
\{x0,iθ=Gθ\(zi\)\}i=1B\\\{x^\{\\theta\}\_\{0,i\}=G\_\{\\theta\}\(z\_\{i\}\)\\\}\_\{i=1\}^\{B\},
zi∼p𝒵z\_\{i\}\\sim p\_\{\\mathcal\{Z\}\}and
\{xT,i\}i=1B∼pT\\\{x\_\{T,i\}\\\}\_\{i=1\}^\{B\}\\sim p\_\{T\};
3:Sample times
\{ti\}i=1B\\\{t\_\{i\}\\\}\_\{i=1\}^\{B\}and noised samples
\{xti,iθ\}i=1B\\\{x^\{\\theta\}\_\{t\_\{i\},i\}\\\}\_\{i=1\}^\{B\}according to the model type;For flow models:xti,iθ=\(1−ti/T\)⋅x0,iθ\+ti/T⋅xT,ix^\{\\theta\}\_\{t\_\{i\},i\}=\(1\-t\_\{i\}/T\)\\cdot x^\{\\theta\}\_\{0,i\}\+t\_\{i\}/T\\cdot x\_\{T,i\};
4:Sample forget data batch
\{x0,iF\}i=1B∼p0F\\\{x^\{F\}\_\{0,i\}\\\}\_\{i=1\}^\{B\}\\sim p\_\{0\}^\{F\}and noised forget samples
\{xti,iF\}i=1B\\\{x^\{F\}\_\{t\_\{i\},i\}\\\}\_\{i=1\}^\{B\};For flow models:xti,iF=\(1−ti/T\)⋅x0,iF\+ti/T⋅xT,ix^\{F\}\_\{t\_\{i\},i\}=\(1\-t\_\{i\}/T\)\\cdot x^\{F\}\_\{0,i\}\+t\_\{i\}/T\\cdot x\_\{T,i\};
5:Compute both matching targets
sti,iθ=ftiθ\(xti,iθ\|x0,iθ\)s^\{\\theta\}\_\{t\_\{i\},i\}=f\_\{t\_\{i\}\}^\{\\theta\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\|x^\{\\theta\}\_\{0,i\}\)and
sti,iF=ftiF\(xti,iF\|x0,iF\)s^\{F\}\_\{t\_\{i\},i\}=f\_\{t\_\{i\}\}^\{F\}\(x^\{F\}\_\{t\_\{i\},i\}\|x^\{F\}\_\{0,i\}\);For flow models:sti,iθ=\(xT,i−x0,iθ\)/Ts^\{\\theta\}\_\{t\_\{i\},i\}=\(x\_\{T,i\}\-x^\{\\theta\}\_\{0,i\}\)/Tandsti,iF=\(xT,i−x0,iF\)/Ts^\{F\}\_\{t\_\{i\},i\}=\(x\_\{T,i\}\-x^\{F\}\_\{0,i\}\)/T;
6:Update fake model parameters
ψ\\psivia loss:
1B∑i=1B\[ρ‖fψ,ti\(xti,iF\)−sti,iF‖2\+\(1−ρ\)‖fψ,ti\(xti,iθ\)−sti,iθ‖2\];\\frac\{1\}\{B\}\\sum\\limits^\{B\}\_\{i=1\}\\left\[\\rho\\\|f\_\{\\psi,t\_\{i\}\}\(x^\{F\}\_\{t\_\{i\},i\}\)\-s^\{F\}\_\{t\_\{i\},i\}\\\|^\{2\}\+\(1\-\\rho\)\\\|f\_\{\\psi,t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\)\-s^\{\\theta\}\_\{t\_\{i\},i\}\\\|^\{2\}\\right\];
7:Update generator parameters
θ\\thetavia loss:
1B∑i=1B\[2⟨fti∗\(xti,iθ\)−fψ,ti\(xti,iθ\),fti∗\(xti,iθ\)−sti,iθ⟩−2⋅αSiD‖fti∗\(xti,iθ\)−fψ,ti\(xti,iθ\)‖2\];\\frac\{1\}\{B\}\\sum\\limits^\{B\}\_\{i=1\}\[2\\langle\{f^\{\*\}\_\{t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\)\}\-f\_\{\\psi,t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\),f^\{\*\}\_\{t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\)\-s^\{\\theta\}\_\{t\_\{i\},i\}\\rangle\-2\\cdot\\alpha\_\{\\text\{SiD\}\}\\\|f^\{\*\}\_\{t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\)\-f\_\{\\psi,t\_\{i\}\}\(x^\{\\theta\}\_\{t\_\{i\},i\}\)\\\|^\{2\}\];
8:endfor
#### Hyperparameters\.
The only new hyperparameter introduced by our IDU method is the scale factorρ∈\[0,1\)\\rho\\in\[0,1\)which controls the strength of forgetting applied to the selected data\. The theoretically justified value ofρ=π\\rho=\\pican be approximately derived from the data split \([4](https://arxiv.org/html/2609.36099#S3.E4)\) as the proportion of the forget data within the overall dataset\. Nevertheless, we still recommend trying other values ofρ\\rhoin practice to find the best trade\-off between forgetting rate and retention quality\. We demonstrate this trade\-off for FM and SiD in Table[3](https://arxiv.org/html/2609.36099#S4.T3)\. For the backbone\-specific distillation scale, we useαSiD=0\.5\\alpha\_\{\\mathrm\{SiD\}\}=0\.5for FM andαSiD=1\.2\\alpha\_\{\\mathrm\{SiD\}\}=1\.2for SiD, following the corresponding original distillation recipes\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14);[Kornilov et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib9)\)\. Further optimization details are provided in Appendix[B\.5](https://arxiv.org/html/2609.36099#A2.SS5)\.
## 4Experiments
#### Experimental setup\.
We evaluate IDU with flow matching \(FM\)\([Tong et al\., 2023](https://arxiv.org/html/2609.36099#bib.bib23)\)and the EDM\-VP score\-based backbone\([Karras et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib8)\)used by Score identity Distillation \(SiD\)\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\), each on MNIST and CIFAR\-10\. The main experiments forget digits3,7\{3,7\}or CIFAR\-10 classes1,9\{1,9\}\(automobile and truck\); the latter pair tests jointly removing related modes\. Class membership makes forgetting measurable, but IDU receives only forget samples, not labels or prompts\. From each frozen full\-data teacher, we run ordinary distillation and IDU, with the former verifying that the same framework also recovers a standard one\-step generator when no forget data are mixed in\. IDU optimization accesses only the teacher and forget samples\. Retained images and evaluation classifiers are never supplied to the training objective\. Architectures, samplers, and hyperparameters are detailed in Appendices[B\.1](https://arxiv.org/html/2609.36099#A2.SS1),[B\.2](https://arxiv.org/html/2609.36099#A2.SS2), and[B\.5](https://arxiv.org/html/2609.36099#A2.SS5)\.
#### Evaluation protocol and metrics\.
We compute FID against the full training set for full\-data teachers and ordinary distillation, and against retained data for IDU and retraining \(*Retain FID*\)\. A teacher trained on retained data and its distilled student serve as retraining baselines\. We also report the*forgotten\-class generation rate \(FGR\)*: the percentage of 50,000 generated images assigned to a forgotten class by an off\-the\-shelf classifier\. Classifier and FID protocols are in Appendices[B\.3](https://arxiv.org/html/2609.36099#A2.SS3)and[B\.4](https://arxiv.org/html/2609.36099#A2.SS4)\. Unless noted, values are means and standard deviations over five evaluations; marked CIFAR\-10 score\-based values follow\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\)\.
Table 1:Main FM/SiD results on MNIST and CIFAR\-10\. FID uses full\-data references for Pretrain/Pure distillation and retained\-data references for Retrain/IDU; FGR is reported per forgotten class\. Values are mean±\\pmstandard deviation over five runs, except†\\daggervalues from\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\)\.MNISTCIFAR\-10ModeFID↓\\downarrowFGR \(%\)↓\\downarrowClass 3 / Class 7FID↓\\downarrowFGR \(%\)↓\\downarrowClass 1 / Class 9FMPretrain0\.88±0\.010\.88\\pm 0\.0110\.40±0\.1210\.40\\pm 0\.1210\.10±0\.2610\.10\\pm 0\.263\.66±0\.033\.66\\pm 0\.0312\.50±0\.0812\.50\\pm 0\.0811\.18±0\.1211\.18\\pm 0\.12Pure distillation3\.23±0\.033\.23\\pm 0\.0310\.50±0\.1510\.50\\pm 0\.1510\.53±0\.0910\.53\\pm 0\.094\.35±0\.054\.35\\pm 0\.057\.58±0\.097\.58\\pm 0\.098\.75±0\.218\.75\\pm 0\.21Forgotten classes\{3,7\}\\\{3,7\\\}\{1,9\}\\\{1,9\\\}Retrain teacher0\.340±0\.0010\.340\\pm 0\.001—3\.93±0\.023\.93\\pm 0\.02—Retrain distillation5\.72±0\.125\.72\\pm 0\.12—5\.12±0\.055\.12\\pm 0\.05—IDU\(ρMNIST=0\.4\)\(\\rho\_\{\\mathrm\{MNIST\}\}=0\.4\)\(ρCIFAR\-10=0\.6\)\(\\rho\_\{\\mathrm\{CIFAR\\text\{\-\}10\}\}=0\.6\)3\.57±0\.033\.57\\pm 0\.030\.16±0\.010\.16\\pm 0\.010\.16±0\.010\.16\\pm 0\.015\.81±0\.055\.81\\pm 0\.050\.56±0\.050\.56\\pm 0\.050\.39±0\.030\.39\\pm 0\.03SiDPretrain1\.35±0\.021\.35\\pm 0\.0210\.99±0\.1410\.99\\pm 0\.149\.65±0\.229\.65\\pm 0\.221\.97†1\.97^\{\\dagger\}11\.13±0\.1211\.13\\pm 0\.1210\.04±0\.1510\.04\\pm 0\.15Pure distillation1\.12±0\.011\.12\\pm 0\.0110\.51±0\.1210\.51\\pm 0\.1210\.07±0\.2410\.07\\pm 0\.241\.92±0\.02†1\.92\\pm 0\.02\\,^\{\\dagger\}10\.11±0\.1410\.11\\pm 0\.1410\.66±0\.1410\.66\\pm 0\.14Forgotten classes\{3,7\}\\\{3,7\\\}\{1,9\}\\\{1,9\\\}Retrain teacher0\.85±0\.020\.85\\pm 0\.02—2\.26±0\.022\.26\\pm 0\.02—Retrain distillation1\.20±0\.021\.20\\pm 0\.02—2\.82±0\.022\.82\\pm 0\.02—IDU\(ρ=0\.2\)\(\\rho=0\.2\)1\.65±0\.021\.65\\pm 0\.020\.63±0\.030\.63\\pm 0\.030\.32±0\.020\.32\\pm 0\.023\.15±0\.033\.15\\pm 0\.031\.11±0\.061\.11\\pm 0\.061\.25±0\.061\.25\\pm 0\.06
#### Main results\.
Table[1](https://arxiv.org/html/2609.36099#S4.T1)shows that IDU suppresses both forgotten modes under all four dataset–backbone combinations\. In the FM experiments, every forgotten\-class FGR is at most0\.56%0\.56\\%; in the SiD experiments, it is at most1\.25%1\.25\\%\. IDU also retains generation quality close to the corresponding retrain\-distillation oracle in three settings and improves on that reference for FM on MNIST, indicating no substantial Retain FID degradation relative to target\-matched retraining\. The FM/MNIST oracle comparison warrants caution owing to sensitivity of the retained\-only retraining baseline; see Appendix[B\.9](https://arxiv.org/html/2609.36099#A2.SS9)\. Visual results with Teacher–IDU sample grids appear in Appendix[C](https://arxiv.org/html/2609.36099#A3)\.Fine\-tuningexperiments on the purely distilled generators achieve similar metrics; see Appendix[B\.7](https://arxiv.org/html/2609.36099#A2.SS7)\.
#### Robustness across forget classes\.
We also evaluate single\-class forgetting for all MNIST classes with FM and all CIFAR\-10 classes with FM and SiD \(Table[2](https://arxiv.org/html/2609.36099#S4.T2)\)\. Across these ten tasks per setting, mean Retain FID/FGR is4\.54/0\.16%4\.54/0\.16\\%for FM–MNIST,7\.73/0\.84%7\.73/0\.84\\%for FM–CIFAR\-10, and3\.61/0\.75%3\.61/0\.75\\%for SiD–CIFAR\-10\. The per\-class results show that low FGR is not confined to the pairs used in the main experiments\.
Table 2:Single\-class robustness of IDU with FM on MNIST and CIFAR\-10 and with SiD on CIFAR\-10\. The final row averages the per\-class means across the ten forget\-class experiments\.FM MNISTFM CIFAR\-10SiD CIFAR\-10ForgottenClassRetain FID↓\\downarrowFGR \(%\)↓\\downarrowRetain FID↓\\downarrowFGR \(%\)↓\\downarrowRetain FID↓\\downarrowFGR \(%\)↓\\downarrow05\.05±0\.035\.05\\pm 0\.030\.10±0\.010\.10\\pm 0\.016\.52±0\.066\.52\\pm 0\.060\.76±0\.030\.76\\pm 0\.033\.54±0\.023\.54\\pm 0\.020\.58±0\.030\.58\\pm 0\.0313\.46±0\.043\.46\\pm 0\.040\.03±0\.010\.03\\pm 0\.016\.45±0\.096\.45\\pm 0\.091\.06±0\.071\.06\\pm 0\.072\.96±0\.012\.96\\pm 0\.010\.61±0\.030\.61\\pm 0\.0323\.84±0\.023\.84\\pm 0\.020\.12±0\.010\.12\\pm 0\.017\.83±0\.057\.83\\pm 0\.050\.73±0\.040\.73\\pm 0\.044\.07±0\.054\.07\\pm 0\.050\.87±0\.040\.87\\pm 0\.0435\.23±0\.035\.23\\pm 0\.030\.10±0\.010\.10\\pm 0\.017\.82±0\.087\.82\\pm 0\.081\.51±0\.041\.51\\pm 0\.044\.10±0\.024\.10\\pm 0\.021\.11±0\.021\.11\\pm 0\.0244\.71±0\.034\.71\\pm 0\.030\.13±0\.020\.13\\pm 0\.028\.53±0\.088\.53\\pm 0\.080\.82±0\.040\.82\\pm 0\.044\.13±0\.034\.13\\pm 0\.030\.78±0\.050\.78\\pm 0\.0554\.37±0\.064\.37\\pm 0\.060\.17±0\.020\.17\\pm 0\.029\.70±0\.059\.70\\pm 0\.051\.18±0\.061\.18\\pm 0\.064\.21±0\.034\.21\\pm 0\.030\.88±0\.040\.88\\pm 0\.0464\.39±0\.044\.39\\pm 0\.040\.07±0\.010\.07\\pm 0\.018\.45±0\.098\.45\\pm 0\.090\.23±0\.020\.23\\pm 0\.024\.16±0\.054\.16\\pm 0\.050\.53±0\.030\.53\\pm 0\.0375\.17±0\.045\.17\\pm 0\.040\.15±0\.010\.15\\pm 0\.015\.82±0\.065\.82\\pm 0\.061\.23±0\.051\.23\\pm 0\.052\.87±0\.022\.87\\pm 0\.020\.90±0\.050\.90\\pm 0\.0584\.22±0\.044\.22\\pm 0\.040\.24±0\.030\.24\\pm 0\.038\.88±0\.138\.88\\pm 0\.130\.39±0\.020\.39\\pm 0\.023\.22±0\.023\.22\\pm 0\.020\.54±0\.050\.54\\pm 0\.0594\.93±0\.054\.93\\pm 0\.050\.45±0\.030\.45\\pm 0\.037\.26±0\.057\.26\\pm 0\.050\.47±0\.030\.47\\pm 0\.032\.86±0\.022\.86\\pm 0\.020\.72±0\.010\.72\\pm 0\.01Mean4\.544\.540\.160\.167\.737\.730\.840\.843\.613\.610\.750\.75
#### Effect of the forgetting weight\.
Table[3](https://arxiv.org/html/2609.36099#S4.T3)varies the forget\-mixture weightρ\\rhoon CIFAR\-10 for both FM and SiD\. For FM,ρ=0\.9\\rho=0\.9diverges immediately\. Among the stable runs, lowρ\\rhopreserves FID but leaves higher FGR, whereasρ=0\.8\\rho=0\.8improves forgetting at a substantial FID cost\. We therefore useρ=0\.6\\rho=0\.6as the best empirical balance for FM on CIFAR\-10\. The FM coefficient for MNIST,ρ=0\.4\\rho=0\.4, is selected by the same quality–forgetting criterion\. For SiD, increasingρ\\rhofrom0\.20\.2to0\.80\.8progressively worsens Retain FID, whereas the two class\-wise FGR values vary non\-monotonically\. Atρ=0\.9\\rho=0\.9, SiD does not diverge, but its FGRs \(12\.05%12\.05\\%and9\.55%9\.55\\%\) approach those of the full\-data teacher, while Retain FID rises to10\.3610\.36\. Thus, excessiveρ\\rhocan degrade fidelity without forgetting\.
Table 3:Effect of the forget\-mixture weightρ\\rhoon CIFAR\-10 for FM\- and SiD\-based IDU when jointly forgetting automobile \(class 1\) and truck \(class 9\)\. The selected configuration for each setup is bold; dashes denote settings not evaluated with SiD\. FM training atρ=0\.9\\rho=0\.9diverges immediately\.Flow MatchingSiDρ\\rhoFID↓\\downarrowFGR 1 \(%\)↓\\downarrowFGR 9 \(%\)↓\\downarrowρ\\rhoFID↓\\downarrowFGR 1 \(%\)↓\\downarrowFGR 9 \(%\)↓\\downarrow0\.050\.056\.72±0\.076\.72\\pm 0\.073\.04±0\.083\.04\\pm 0\.085\.02±0\.085\.02\\pm 0\.08—0\.10\.16\.21±0\.056\.21\\pm 0\.053\.75±0\.093\.75\\pm 0\.093\.80±0\.083\.80\\pm 0\.08—0\.20\.25\.64±0\.045\.64\\pm 0\.043\.35±0\.033\.35\\pm 0\.033\.04±0\.103\.04\\pm 0\.100\.20\.23\.15±0\.03\\mathbf\{3\.15\\pm 0\.03\}1\.11±0\.06\\mathbf\{1\.11\\pm 0\.06\}1\.25±0\.06\\mathbf\{1\.25\\pm 0\.06\}0\.40\.45\.87±0\.075\.87\\pm 0\.070\.62±0\.060\.62\\pm 0\.060\.48±0\.020\.48\\pm 0\.020\.40\.43\.53±0\.043\.53\\pm 0\.040\.74±0\.050\.74\\pm 0\.050\.60±0\.010\.60\\pm 0\.010\.60\.65\.81±0\.05\\mathbf\{5\.81\\pm 0\.05\}0\.56±0\.05\\mathbf\{0\.56\\pm 0\.05\}0\.39±0\.03\\mathbf\{0\.39\\pm 0\.03\}0\.60\.64\.68±0\.064\.68\\pm 0\.061\.16±0\.051\.16\\pm 0\.050\.76±0\.020\.76\\pm 0\.020\.80\.88\.59±0\.118\.59\\pm 0\.110\.38±0\.030\.38\\pm 0\.030\.32±0\.030\.32\\pm 0\.030\.80\.86\.29±0\.076\.29\\pm 0\.070\.81±0\.050\.81\\pm 0\.050\.58±0\.020\.58\\pm 0\.020\.90\.9Diverged0\.90\.910\.36±0\.1010\.36\\pm 0\.1012\.05±0\.1812\.05\\pm 0\.189\.55±0\.129\.55\\pm 0\.12
## 5Discussion and comparison
#### How does our method work?
Our IDU leverages the teacher as guidance to steer generations away from unwanted data\. Specifically, the IDU’s generator loss is a tractable form of the squared norm of the difference between the teacherf∗f^\{\*\}and the fake modelfmixf^\{\\mathrm\{mix\}\}trained on the mixed forget and generated datap0mix:=ρ⋅p0F\+\(1−ρ\)⋅p0θp\_\{0\}^\{\\mathrm\{mix\}\}:=\\rho\\cdot p^\{F\}\_\{0\}\+\(1\-\\rho\)\\cdot p^\{\\theta\}\_\{0\}; see Appendix[A\.2](https://arxiv.org/html/2609.36099#A1.SS2):
ℒIDU\(fmix,p0θ\)=𝔼t,xt∼ptmix\[‖ft∗\(xt\)−ftmix\(xt\)‖2\],fmix:=argminℒUMf\(f,p0mix\)\.\\mathcal\{L\}\_\{\\text\{IDU\}\}\(f^\{\\mathrm\{mix\}\},p\_\{0\}^\{\\theta\}\)=\\mathbb\{E\}\_\{t,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\[\\left\\\|f\_\{t\}^\{\*\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\)\\right\\\|^\{2\}\],\\quad f^\{\\mathrm\{mix\}\}:=\\argmin\{\{\}\_\{f\}\}\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\)\.Since the supplied forget samples already cover the forget component ofp0mixp\_\{0\}^\{\\mathrm\{mix\}\}, further forget samples generated byGθG\_\{\\theta\}make the fake model depart from the teacher\. Minimizing the generator loss counteracts this excess while also discouraging low\-quality retained generations\. Similar ideas are proposed for distillation acceleration in\([Kornilov et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib9)\)\. There, the authors utilize real data to push the generator toward the teacher faster, whereas we use forget data to steer the model away\.
#### Training and evaluation requirements\.
IDU training uses the frozen teacher and forget samples only; retained images and external classifiers enter neither its losses nor its gradient updates\. On public MNIST and CIFAR\-10, retained\-data FID and classifier\-based FGR provide direct, reproducible measurements\. We use these metrics to selectρ\\rho; without the retained data,ρ\\rhocan be chosen from a wide range of values, starting from the forget\-set proportion, and still avoid significant degradation\. However, the ablations show that larger values do not necessarily improve forgetting and may degrade fidelity\. If direct evaluation is critical, retained data can be selected from the teacher samples via classification or manual selection\.
#### Data unlearning methods\.
Our IDU is suitable for both the unlearning of the multi\-step teacher matching models \(with additional distillation\) and the unlearning of the pretrained generators\.
We begin by comparing theunlearning of matching models\. ContinualFlow\([Simone et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib49)\)uses an energy function to guide a teacher flow model away from forget data, but it is limited to flow matching and relies on an image classifier to calculate the energy\. Other methods, such as NegGrad, SA\([Heng and Soh, 2023](https://arxiv.org/html/2609.36099#bib.bib36)\), SalUn\([Fan et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib37)\), SISS\([Alberti et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib43)\), Retrack\([Shi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib47)\), MGSM\([Jiang et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib46)\), EraseDiff\([Wu et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib40)\), and VDU\([Panda et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib5)\), unlearn teacher models using a combination of forget and remaining losses \([3](https://arxiv.org/html/2609.36099#S2.E3)\)\. Although our IDU also optimizes a similar combination of losses within the fake model, our working principle is different\. The forget and remaining losses in other methods are trained adversarially: the former degrades quality on the forget dataset, while the latter preserves it on the remaining one\. This is why such methods often resort to constrained optimization or more stable losses\. In our IDU, by contrast, the forget and generated data distributions are modeled jointly as a mixture\. As a result, we avoid multi\-objective optimization and train our losses coherently\. Moreover, distilled generators provide efficient one\-step inference rather than the multi\-step sampling for the teacher\.
For pretrained generators, UOT\-Unlearn\([Choi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib45)\)employs a cost function from unbalanced optimal transport to guide the unlearning\. This guidance requires an additional feature extractor and manual cost\-function tuning instead of a teacher model\. Our IDU uses more complex teacher guidance but yields better forgetting and retaining metrics\. As the UOT\-Unlearn code is unavailable, we compare against their best unlearned CTM model with the SiD architecture and similar initial FID of 1\.73\. For CIFAR\-10 unlearning classes 1, 6, and 8, their Retain FIDs \(↓\\downarrow\) are \(9\.90, 5\.11, 5\.88\) versus our \(2\.96, 4\.16, 3\.22\), and their FGRs \(%,↓\\downarrow\) are around \(2, 0\.9, 1\.5\) versus our \(0\.61, 0\.53, 0\.54\)\. Many of the above matching model unlearning methods can be generalized to one\-step consistency models\. These models enforce self\-consistency along the generation trajectory by mapping any point directly to the start\. Thus, they can optimize different losses on the forget and remaining data\. For other one\-step model types that do not take data samples as input, such as GANs, OT, and distillation, this generalization does not work\. Moreover, such methods lag significantly behind in terms of retaining and forgetting \(see Table 2 in\([Choi et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib45)\)\)\. In contrast, our IDU does not require self\-consistency or a particular generator model type\.
#### Class unlearning methods\.
The most relevant SFD approach\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\)also performs unlearning during distillation but operates in the class unlearning setup\. This method focuses on conditional models, where the data to erase is prompted through the input labels rather than through given samples\. In contrast, our IDU can erase any part of training data, offering more flexible forgetting opportunities\. The working mechanisms also differ dramatically\. SFD modifies the generator loss \([10](https://arxiv.org/html/2609.36099#A1.E10)\), writing safe teacher information into the generator’s unwanted classes\. We modify the fake model loss to make it remember the mixture of generated and forget data and then compare it with the teacher’s correct one, penalizing the generator for reproducing unwanted samples\. In the class\-defined experiments reported in Tables[1](https://arxiv.org/html/2609.36099#S4.T1),[2](https://arxiv.org/html/2609.36099#S4.T2), and[3](https://arxiv.org/html/2609.36099#S4.T3), the generator is never prompted with a class label: class annotations are used only to construct the forget subset and to evaluate FGR\. Nevertheless, under the same conditions, our IDU achieves comparable results\. In the CIFAR\-10 class\-0 forgetting experiment, SFD attains a Retain FID \(↓\\downarrow\) of around 3\.1 and an FGR \(%,↓\\downarrow\) of 0\.36 \(Figure 5\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\)\), versus 3\.54 and 0\.58 for our method \(Table[2](https://arxiv.org/html/2609.36099#S4.T2)\)\.
#### Optimization stability across backbones\.
In extended FM runs on MNIST and CIFAR\-10, IDU often suppresses the target classes early, but their generation frequency can rise again after roughly 20k generator updates; the onset depends onρ\\rho\. FM therefore requires joint selection ofρ\\rhoand checkpoint\. We did not observe this reversal over the evaluated SiD training horizon\. This appears to be a limitation of the current FM instantiation, not of the IDU objective; larger, more stable FM architectures may allowρ\\rhoto approach the forget\-set proportion, as it does for SiD, whereρ=0\.2\\rho=0\.2matches two forgotten classes out of ten\.
### AI use statement
Generative AI tools were used solely to improve grammar, clarity, concision, and academic style during manuscript preparation\. They were not used to formulate the research problem, develop the method, design or execute experiments, generate or analyze results, or determine the scientific claims and conclusions\. All AI\-assisted edits were reviewed and verified by the authors, who take full responsibility for the final content of this work\.
## References
- Albertiet al\.\(2025\)S\. Alberti, K\. Hasanaliyev, M\. Shah, and S\. ErmonData unlearning in diffusion models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SuHScQv5gP),2503\.01034Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p2.1),[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Bourtouleet al\.\(2021\)L\. Bourtoule, V\. Chandrasekaran, C\. A\. Choquette\-Choo, H\. Jia, A\. Travers, B\. Zhang, D\. Lie, and N\. PapernotMachine unlearning\.In2021 IEEE symposium on security and privacy \(SP\),pp\. 141–159\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Buiet al\.\(2025\)A\. Bui, T\. Vu, L\. Vuong, T\. Le, P\. Montague, T\. Abraham, J\. Kim, and D\. PhungFantastic targets for concept erasure in diffusion models and where to find them\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 64032–64074\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/a10946e1f46e1ffc0daf37cb2abfdcad-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1)\.
- Carliniet al\.\(2023\)N\. Carlini, J\. Hayes, M\. Nasr, M\. Jagielski, V\. Sehwag, F\. Tramer, B\. Balle, D\. Ippolito, and E\. WallaceExtracting training data from diffusion models\.In32nd USENIX security symposium \(USENIX Security 23\),pp\. 5253–5270\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Chavhanet al\.\(2024\)R\. Chavhan, D\. Li, and T\. HospedalesConceptprune: concept editing in diffusion models via skilled neuron pruning\.arXiv preprint arXiv:2405\.19237\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1)\.
- Chenet al\.\(2025\)T\. Chen, S\. Zhang, and M\. ZhouScore forgetting distillation: a swift, data\-free method for machine unlearning in diffusion models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gjwhDHeAsz),2409\.11219Cited by:[§A\.3](https://arxiv.org/html/2609.36099#A1.SS3.p1.1),[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§1](https://arxiv.org/html/2609.36099#S1.p6.1.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p2.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px4.p1.1)\.
- Choiet al\.\(2026\)H\. Choi, J\. An, J\. Park, and J\. ChoiUnlearning for one\-step generative models via unbalanced optimal transport\.InICML 2026 Workshop on Foundations of Deep Generative Models \(FoGen\),External Links:2603\.16489,[Link](https://arxiv.org/abs/2603.16489)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§1](https://arxiv.org/html/2609.36099#S1.p4.1.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p4.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p3.1)\.
- Fanet al\.\(2024\)C\. Fan, J\. Liu, Y\. Zhang, E\. Wong, D\. Wei, and S\. LiuSalun: empowering machine unlearning via gradient\-based weight saliency in both image classification and generation\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 53643–53673\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Fanet al\.\(2026\)Z\. Fan, N\. Jiang, D\. Gao, S\. Zhou, and W\. WuEraseAnything\+\+: enabling concept erasure in rectified flow transformers leveraging multi\-object optimization\.External Links:2603\.00978,[Link](https://arxiv.org/abs/2603.00978)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Franset al\.\(2025\)K\. Frans, D\. Hafner, S\. Levine, and P\. AbbeelOne step diffusion via shortcut models\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p4.1)\.
- Gandikotaet al\.\(2023\)R\. Gandikota, J\. Materzynska, J\. Fiotto\-Kaufman, and D\. BauErasing concepts from diffusion models\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 2426–2436\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p5.1),[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Gandikotaet al\.\(2024\)R\. Gandikota, H\. Orgad, Y\. Belinkov, J\. Materzyńska, and D\. BauUnified concept editing in diffusion models\.In2024 IEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\),pp\. 5099–5108\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Gaoet al\.\(2025a\)D\. Gao, S\. Lu, W\. Zhou, J\. Chu, J\. Zhang, M\. Jia, B\. Zhang, Z\. Fan, and W\. ZhangEraseAnything: enabling concept erasure in rectified flow transformers\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 18470–18494\.External Links:[Link](https://proceedings.mlr.press/v267/gao25j.html)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Gaoet al\.\(2025b\)R\. Gao, E\. Hoogeboom, J\. Heek, V\. De Bortoli, K\. P\. Murphy, and T\. SalimansDiffusion models and gaussian flow matching: two sides of the same coin\.InThe Fourth Blogpost Track at ICLR 2025,Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Genget al\.\(2025\)Z\. Geng, M\. Deng, X\. Bai, J\. Z\. Kolter, and K\. HeMean flows for one\-step generative modeling\.InAdvances in Neural Information Processing Systems,External Links:2505\.13447,[Link](https://arxiv.org/abs/2505.13447)Cited by:[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p4.1)\.
- Golatkaret al\.\(2020\)A\. Golatkar, A\. Achille, and S\. SoattoEternal sunshine of the spotless net: selective forgetting in deep networks\.In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 9301–9309\.Cited by:[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p1.1)\.
- Goldman \(2020\)E\. GoldmanAn introduction to the california consumer privacy act \(ccpa\)\.Santa Clara Univ\. Legal Studies Research Paper\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Goodfellowet al\.\(2014\)I\. J\. Goodfellow, J\. Pouget\-Abadie, M\. Mirza, B\. Xu, D\. Warde\-Farley, S\. Ozair, A\. Courville, and Y\. BengioGenerative adversarial nets\.Advances in neural information processing systems27\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Gushchinet al\.\(2025\)N\. Gushchin, D\. Li, D\. Selikhanovych, E\. Burnaev, D\. Baranchuk, and A\. KorotinInverse bridge matching distillation\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 21471–21496\.External Links:[Link](https://proceedings.mlr.press/v267/gushchin25b.html)Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Heng and Soh \(2023\)A\. Heng and H\. SohSelective amnesia: a continual learning approach to forgetting in deep generative models\.Advances in Neural Information Processing Systems36,pp\. 17170–17194\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Hoet al\.\(2020\)J\. Ho, A\. Jain, and P\. AbbeelDenoising diffusion probabilistic models\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 6840–6851\.External Links:[Link](https://arxiv.org/abs/2006.11239)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36099#S2.SS1.p1.1)\.
- Holderriethet al\.\(2024\)P\. Holderrieth, M\. Havasi, J\. Yim, N\. Shaul, I\. Gat, T\. Jaakkola, B\. Karrer, R\. T\. Chen, and Y\. LipmanGenerator matching: generative modeling with arbitrary markov processes\.arXiv preprint arXiv:2410\.20587\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Jianget al\.\(2025\)W\. Jiang, H\. Wang, X\. Zhang, D\. Guo, Z\. Fan, Y\. Diao, and R\. HongModerating the generalization of score\-based generative model\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 360–369\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§1](https://arxiv.org/html/2609.36099#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Karraset al\.\(2024\)T\. Karras, M\. Aittala, J\. Lehtinen, J\. Hellsten, T\. Aila, and S\. LaineAnalyzing and improving the training dynamics of diffusion models\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 24174–24184\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1),[§4](https://arxiv.org/html/2609.36099#S4.SS0.SSS0.Px1.p1.1)\.
- Khalafiet al\.\(2026\)S\. Khalafi, A\. Ribeiro, and D\. DingUnlearning in diffusion models: a unified framework with KL divergence and likelihood constraints\.InInternational Conference on Machine Learning,External Links:2605\.30825,[Link](https://arxiv.org/abs/2605.30825)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1)\.
- Kimet al\.\(2024\)D\. Kim, C\. Lai, W\. Liao, N\. Murata, Y\. Takida, T\. Uesaka, Y\. He, Y\. Mitsufuji, and S\. ErmonConsistency trajectory models: learning probability flow ODE trajectory of diffusion\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ymjI8feDTD),2310\.02279Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p4.1)\.
- Kingma and Welling \(2013\)D\. P\. Kingma and M\. WellingAuto\-encoding variational bayes\.arXiv preprint arXiv:1312\.6114\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Kornilovet al\.\(2026\)N\. Kornilov, D\. Li, T\. Mavrin, A\. Leonov, N\. Gushchin, E\. Burnaev, I\. Koshelev, and A\. KorotinUniversal inverse distillation for matching models with real\-data supervision \(no gans\)\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 67906–67948\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.36099#S3.SS2.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px1.p1.2),[Lemma 1](https://arxiv.org/html/2609.36099#Thmlemma1.3)\.
- Lipmanet al\.\(2023\)Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PqvMRDCJT9t),2210\.02747Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36099#S2.SS1.p1.1)\.
- Liu and Zhang \(2025\)P\. Liu and C\. ZhangErased or dormant? rethinking concept erasure through reversibility\.arXiv preprint arXiv:2505\.16174\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1)\.
- Liuet al\.\(2023\)X\. Liu, C\. Gong, and Q\. LiuFlow straight and fast: learning to generate and transfer data with rectified flow\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=XVjTT1nw5z),2209\.03003Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Luet al\.\(2022\)C\. Lu, Y\. Zhou, F\. Bao, J\. Chen, C\. Li, and J\. ZhuDpm\-solver: a fast ode solver for diffusion probabilistic model sampling in around 10 steps\.Advances in neural information processing systems35,pp\. 5775–5787\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Luet al\.\(2025\)K\. Lu, N\. Kriplani, R\. Gandikota, M\. Pham, D\. Bau, C\. Hegde, and N\. CohenWhen are concepts erased from diffusion models?\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 14751–14775\.External Links:[Document](https://dx.doi.org/10.52202/085713-0497),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/15bbe6ddfc88d8e7f59c8f7d4e2541f5-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1)\.
- Luet al\.\(2024\)S\. Lu, Z\. Wang, L\. Li, Y\. Liu, and A\. W\. KongMace: mass concept erasure in diffusion models\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 6430–6440\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Nguyenet al\.\(2025\)T\. T\. Nguyen, T\. T\. Huynh, Z\. Ren, P\. L\. Nguyen, A\. W\. Liew, H\. Yin, and Q\. V\. H\. NguyenA survey of machine unlearning\.ACM Transactions on Intelligent Systems and Technology16\(5\),pp\. 1–46\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Pandaet al\.\(2024\)S\. Panda, M\. Varun, S\. Jain, S\. K\. Maharana, and A\. PrathoshVariational diffusion unlearning: a variational inference framework for unlearning in diffusion models\.InNeurips Safe Generative AI Workshop 2024,Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p3.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Parmaret al\.\(2022\)G\. Parmar, R\. Zhang, and J\. ZhuOn aliased resizing and surprising subtleties in GAN evaluation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 11410–11420\.Cited by:[§B\.4](https://arxiv.org/html/2609.36099#A2.SS4.p1.1)\.
- Schuhmannet al\.\(2022\)C\. Schuhmann, R\. Beaumont, R\. Vencu, C\. Gordon, R\. Wightman, M\. Cherti, T\. Coombes, A\. Katta, C\. Mullis, M\. Wortsman,et al\.Laion\-5b: an open large\-scale dataset for training next generation image\-text models\.Advances in neural information processing systems35,pp\. 25278–25294\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Sharmaet al\.\(2024\)A\. S\. Sharma, N\. Sarkar, V\. Chundawat, A\. A\. Mali, and M\. MandalUnlearning or concealment? a critical analysis and evaluation metrics for unlearning in diffusion models\.arXiv preprint arXiv:2409\.05668\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1)\.
- Shiet al\.\(2026\)Q\. Shi, C\. Jin, J\. Zhang, and Y\. GuReTrack: data unlearning in diffusion models through redirecting the denoising trajectory\.InProceedings of The 29th International Conference on Artificial Intelligence and Statistics,E\. Khan, Y\. Li, A\. Solin, and A\. Ramdas \(Eds\.\),Proceedings of Machine Learning Research, Vol\.300,pp\. 2818–2826\.External Links:[Link](https://proceedings.mlr.press/v300/shi26a.html)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Simoneet al\.\(2025\)L\. Simone, D\. Bacciu, and S\. MaContinualFlow: learning and unlearning with neural flow matching\.InICML 2025 Workshop on Machine Unlearning for Generative AI \(MUGen\),External Links:2506\.18747,[Link](https://arxiv.org/abs/2506.18747)Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§1](https://arxiv.org/html/2609.36099#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p3.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Sohl\-Dicksteinet al\.\(2015\)J\. Sohl\-Dickstein, E\. Weiss, N\. Maheswaranathan, and S\. GanguliDeep unsupervised learning using nonequilibrium thermodynamics\.InInternational conference on machine learning,pp\. 2256–2265\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Songet al\.\(2021\)Y\. Song, J\. Sohl\-Dickstein, D\. P\. Kingma, A\. Kumar, S\. Ermon, and B\. PooleScore\-based generative modeling through stochastic differential equations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=PxTIG12RRHS),2011\.13456Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.36099#S2.SS1.p1.1)\.
- Tang and Khanna \(2026\)H\. Tang and R\. KhannaSharpness\-aware machine unlearning\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 79941–79985\.Cited by:[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p1.1)\.
- Thudiet al\.\(2022\)A\. Thudi, G\. Deza, V\. Chandrasekaran, and N\. PapernotUnrolling sgd: understanding factors influencing machine unlearning\.In2022 IEEE 7th European symposium on security and privacy \(EuroS&P\),pp\. 303–319\.Cited by:[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p1.1)\.
- Tonget al\.\(2023\)A\. Tong, K\. Fatras, N\. Malkin, G\. Huguet, Y\. Zhang, J\. Rector\-Brooks, G\. Wolf, and Y\. BengioImproving and generalizing flow\-based generative models with minibatch optimal transport\.arXiv preprint arXiv:2302\.00482\.Cited by:[§4](https://arxiv.org/html/2609.36099#S4.SS0.SSS0.Px1.p1.1)\.
- Voigt and Von dem Bussche \(2017\)P\. Voigt and A\. Von dem BusscheThe eu general data protection regulation \(gdpr\): a practical guide\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1)\.
- Wuet al\.\(2025\)J\. Wu, T\. Le, M\. Hayat, and M\. HarandiErasing undesirable influence in diffusion models\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 28263–28273\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p3.1),[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.3](https://arxiv.org/html/2609.36099#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1),[§5](https://arxiv.org/html/2609.36099#S5.SS0.SSS0.Px3.p2.1)\.
- Yinet al\.\(2024a\)T\. Yin, M\. Gharbi, T\. Park, R\. Zhang, E\. Shechtman, F\. Durand, and B\. FreemanImproved distribution matching distillation for fast image synthesis\.Advances in neural information processing systems37,pp\. 47455–47487\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Yinet al\.\(2024b\)T\. Yin, M\. Gharbi, R\. Zhang, E\. Shechtman, F\. Durand, W\. T\. Freeman, and T\. ParkOne\-step diffusion with distribution matching distillation\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 6613–6623\.Cited by:[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1)\.
- Zhanget al\.\(2024\)G\. Zhang, K\. Wang, X\. Xu, Z\. Wang, and H\. ShiForget\-me\-not: learning to forget in text\-to\-image diffusion models\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops \(CVPRW\),pp\. 1755–1764\.Cited by:[§1](https://arxiv.org/html/2609.36099#S1.p6.1),[§2\.4](https://arxiv.org/html/2609.36099#S2.SS4.p1.1)\.
- Zhouet al\.\(2024\)M\. Zhou, H\. Zheng, Z\. Wang, M\. Yin, and H\. HuangScore identity distillation: exponentially fast distillation of pretrained diffusion models for one\-step generation\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 62307–62331\.External Links:[Link](https://proceedings.mlr.press/v235/zhou24x.html)Cited by:[§A\.3](https://arxiv.org/html/2609.36099#A1.SS3.p1.1),[§B\.5](https://arxiv.org/html/2609.36099#A2.SS5.SSS0.Px3.p1.1),[Table 4](https://arxiv.org/html/2609.36099#A2.T4),[§1](https://arxiv.org/html/2609.36099#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.36099#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.36099#S3.SS2.SSS0.Px1.p3.1),[§3\.2](https://arxiv.org/html/2609.36099#S3.SS2.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.36099#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.36099#S4.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.36099#S4.T1)\.
###### Contents
1. [1Introduction](https://arxiv.org/html/2609.36099#S1)1. [1\.1Contributions](https://arxiv.org/html/2609.36099#S1.SS1)
2. [2Related Work](https://arxiv.org/html/2609.36099#S2)1. [2\.1Diffusion, flow and matching models](https://arxiv.org/html/2609.36099#S2.SS1) 2. [2\.2Distillation and one\-step models](https://arxiv.org/html/2609.36099#S2.SS2) 3. [2\.3Data unlearning](https://arxiv.org/html/2609.36099#S2.SS3) 4. [2\.4Class unlearning](https://arxiv.org/html/2609.36099#S2.SS4)
3. [3Inverse Distillation Unlearning](https://arxiv.org/html/2609.36099#S3)1. [3\.1Method Description](https://arxiv.org/html/2609.36099#S3.SS1) 2. [3\.2Practical details](https://arxiv.org/html/2609.36099#S3.SS2)
4. [4Experiments](https://arxiv.org/html/2609.36099#S4)
5. [5Discussion and comparison](https://arxiv.org/html/2609.36099#S5)
6. [References](https://arxiv.org/html/2609.36099#bib)
7. [AProofs and related methods](https://arxiv.org/html/2609.36099#A1)1. [A\.1Proof of Theorem](https://arxiv.org/html/2609.36099#A1.SS1) 2. [A\.2The minimized distance](https://arxiv.org/html/2609.36099#A1.SS2) 3. [A\.3SFD’s details](https://arxiv.org/html/2609.36099#A1.SS3)
8. [BExperimental details](https://arxiv.org/html/2609.36099#A2)1. [B\.1Architectures and teacher checkpoints](https://arxiv.org/html/2609.36099#A2.SS1) 2. [B\.2Teacher sampling](https://arxiv.org/html/2609.36099#A2.SS2) 3. [B\.3Evaluation classifiers](https://arxiv.org/html/2609.36099#A2.SS3) 4. [B\.4FID protocol and evaluation data](https://arxiv.org/html/2609.36099#A2.SS4) 5. [B\.5Optimization hyperparameters](https://arxiv.org/html/2609.36099#A2.SS5) 6. [B\.6Initialization](https://arxiv.org/html/2609.36099#A2.SS6) 7. [B\.7Fine\-tuning experiments](https://arxiv.org/html/2609.36099#A2.SS7) 8. [B\.8Code and checkpoint release](https://arxiv.org/html/2609.36099#A2.SS8) 9. [B\.9MNIST FM retraining and sensitivity toρ\\rho](https://arxiv.org/html/2609.36099#A2.SS9)
9. [CVisual Results](https://arxiv.org/html/2609.36099#A3)
## Appendix AProofs and related methods
### A\.1Proof of Theorem[1](https://arxiv.org/html/2609.36099#Thmtheorem1)
First, we need the formal convergence guarantees for the inverse distillation scheme \([5](https://arxiv.org/html/2609.36099#S3.E5)\)\.
###### Lemma 1\(Inverse distillation scheme’s optimum\([Kornilov et al\., 2026](https://arxiv.org/html/2609.36099#bib.bib9)\)\)
The inverse scheme \([5](https://arxiv.org/html/2609.36099#S3.E5)\) with the teacherf∗=argminfℒUM\(f,p0∗\)f^\{\*\}=\\argmin\_\{f\}\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f,p^\{\*\}\_\{0\}\)always attains its optimum00when and only when the teacher data is retrieved, i\.e\.,p0=p0∗p\_\{0\}=p\_\{0\}^\{\*\}\.
Proof\.According to Lemma[1](https://arxiv.org/html/2609.36099#Thmlemma1), after we optimize the inverse distillation scheme \([5](https://arxiv.org/html/2609.36099#S3.E5)\) over distributionp0p\_\{0\}, we get the teacher datap0∗p\_\{0\}^\{\*\}distilled into this optimized distribution, i\.e\.,p0=p0∗p\_\{0\}=p\_\{0\}^\{\*\}\. Since we parametrize the optimized distribution as a mixture of the generated and forget datap0=ρp0F\+\(1−ρ\)p0θp\_\{0\}=\\rho\\,p^\{F\}\_\{0\}\+\(1\-\\rho\)\\,p^\{\\theta\}\_\{0\}withρ=π\\rho=\\pi, then at the optimal parametersθopt\\theta\_\{\\text\{opt\}\}, we get:
p0=ρp0F\+\(1−ρ\)p0θopt=πp0F\+\(1−π\)p0θopt=p0∗\.\\displaystyle p\_\{0\}=\\rho\\,p^\{F\}\_\{0\}\+\(1\-\\rho\)\\,p^\{\\theta\_\{\\text\{opt\}\}\}\_\{0\}=\\pi\\,p^\{F\}\_\{0\}\+\(1\-\\pi\)\\,p^\{\\theta\_\{\\text\{opt\}\}\}\_\{0\}=p\_\{0\}^\{\*\}\.\(8\)Finally, considering the structure of the teacher data \([4](https://arxiv.org/html/2609.36099#S3.E4)\), we conclude that:
πp0F\+\(1−π\)p0θopt=\([8](https://arxiv.org/html/2609.36099#A1.E8)\)p0∗=\([4](https://arxiv.org/html/2609.36099#S3.E4)\)πp0F\+\(1−π\)p0R⇒p0θopt=p0R\.\\displaystyle\\pi\\,p^\{F\}\_\{0\}\+\(1\-\\pi\)\\,p^\{\\theta\_\{\\text\{opt\}\}\}\_\{0\}\\overset\{\\eqref\{eq: split for opt params\}\}\{=\}p\_\{0\}^\{\*\}\\overset\{\\eqref\{eq: data split\}\}\{=\}\\pi\\,p^\{F\}\_\{0\}\+\(1\-\\pi\)\\,p^\{R\}\_\{0\}\\quad\\Rightarrow\\quad p^\{\\theta\_\{\\text\{opt\}\}\}\_\{0\}=p^\{R\}\_\{0\}\.□\\square
### A\.2The minimized distance
First, we use the formula of the IDU loss \([6](https://arxiv.org/html/2609.36099#S3.E6)\) with the optimal fake modelfmix:=argminℒUMf\(f,p0mix\)f^\{\\mathrm\{mix\}\}:=\\argmin\{\{\}\_\{f\}\}\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\)on the current mixed datap0mix:=ρ⋅p0F\+\(1−ρ\)⋅p0θp\_\{0\}^\{\\mathrm\{mix\}\}:=\\rho\\cdot p^\{F\}\_\{0\}\+\(1\-\\rho\)\\cdot p^\{\\theta\}\_\{0\}:
ℒIDU\(fmix,p0θ\)=ℒUM\(f∗,p0mix\)−ℒUM\(fmix,p0mix\)\.\\mathcal\{L\}\_\{\\text\{IDU\}\}\(f^\{\\mathrm\{mix\}\},p\_\{0\}^\{\\theta\}\)=\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f^\{\*\},p^\{\\mathrm\{mix\}\}\_\{0\}\)\-\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f^\{\\mathrm\{mix\}\},p^\{\\mathrm\{mix\}\}\_\{0\}\)\.
Following the structure of UM loss \([1](https://arxiv.org/html/2609.36099#S2.E1)\) on the mixed data, we get:
ℒUM\(f,p0mix\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\)=\\displaystyle=𝔼t,x0∼p0mix,xt∼ptmix\(⋅∣x0\)\[‖ft\(xt\)−ftmix\(xt∣x0\)‖2\]\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\},\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{0\}\)\}\\left\[\\left\\\|f\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\\right\\\|^\{2\}\\right\]=\\displaystyle=𝔼t,xt∼ptmix,x0∼p0mix\(⋅∣xt\)\[‖ft\(xt\)−ftmix\(xt∣x0\)‖2\]\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\},\\,x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{t\}\)\}\\left\[\\left\\\|f\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\\right\\\|^\{2\}\\right\]=\\displaystyle=𝔼t,xt∼ptmix\[‖ft\(xt\)−𝔼x0∼p0mix\(⋅∣xt\)\[ftmix\(xt∣x0\)\]‖2\]\+C\(p0mix\),\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|f\_\{t\}\(x\_\{t\}\)\-\\mathbb\{E\}\_\{x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{t\}\)\}\[f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\]\\right\\\|^\{2\}\\right\]\+C\(p^\{\\mathrm\{mix\}\}\_\{0\}\),C\(p0mix\)\\displaystyle C\(p^\{\\mathrm\{mix\}\}\_\{0\}\)=\\displaystyle=𝔼t,xt∼ptmix\[𝔼x0∼p0mix\(⋅∣xt\)\[‖ftmix\(xt∣x0\)‖2\]\]\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\mathbb\{E\}\_\{x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{t\}\)\}\\left\[\\left\\\|f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\\right\\\|^\{2\}\\right\]\\right\]−\\displaystyle\-𝔼t,xt∼ptmix\[‖𝔼x0∼p0mix\(⋅∣xt\)\[ftmix\(xt∣x0\)\]‖2\],\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|\\mathbb\{E\}\_\{x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{t\}\)\}\\left\[f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\\right\]\\right\\\|^\{2\}\\right\],whereC\(p0mix\)C\(p^\{\\mathrm\{mix\}\}\_\{0\}\)is the bias–variance decomposition term independent offf\. For fixedt,xtt,x\_\{t\}, the UM loss is minimized by the conditional mean:
ftmix\(xt\)\\displaystyle f\_\{t\}^\{\{\\mathrm\{mix\}\}\}\(x\_\{t\}\)=\\displaystyle=𝔼x0∼p0mix\(⋅∣xt\)\[ftmix\(xt∣x0\)\]=argminℒUMf\(f,p0mix\),\\displaystyle\\mathbb\{E\}\_\{x\_\{0\}\\sim p\_\{0\}^\{\\mathrm\{mix\}\}\(\\cdot\\mid x\_\{t\}\)\}\[f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\\mid x\_\{0\}\)\]=\\argmin\{\{\}\_\{f\}\}\\mathcal\{L\}\_\{\\text\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\),ℒUM\(f,p0mix\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f,p\_\{0\}^\{\\mathrm\{mix\}\}\)=\\displaystyle=𝔼t,xt∼ptmix\[‖ft\(xt\)−ftmix\(xt\)‖2\]\+C\(p0mix\)\.\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|f\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\)\\right\\\|^\{2\}\\right\]\+C\(p^\{\\mathrm\{mix\}\}\_\{0\}\)\.
Thus, for the IDU loss with the optimal fake modelfmixf^\{\\mathrm\{mix\}\}, we have:
ℒIDU\(fmix,p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{IDU\}\}\(f^\{\\mathrm\{mix\}\},p\_\{0\}^\{\\theta\}\)=\\displaystyle=ℒUM\(f∗,p0mix\)−ℒUM\(fmix,p0mix\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f^\{\*\},p\_\{0\}^\{\\mathrm\{mix\}\}\)\-\\mathcal\{L\}\_\{\\mathrm\{UM\}\}\(f^\{\\mathrm\{mix\}\},p\_\{0\}^\{\\mathrm\{mix\}\}\)=\\displaystyle=𝔼t,xt∼ptmix\[‖ft∗\(xt\)−ftmix\(xt\)‖2\]\+C\(p0mix\)\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|f^\{\*\}\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\)\\right\\\|^\{2\}\\right\]\+C\(p^\{\\mathrm\{mix\}\}\_\{0\}\)−\\displaystyle\-𝔼t,xt∼ptmix\[‖ftmix\(xt\)−ftmix\(xt\)‖2\]−C\(p0mix\)\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|f^\{\\mathrm\{mix\}\}\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\)\\right\\\|^\{2\}\\right\]\-C\(p^\{\\mathrm\{mix\}\}\_\{0\}\)=\\displaystyle=𝔼t,xt∼ptmix\[‖ft∗\(xt\)−ftmix\(xt\)‖2\]\.\\displaystyle\\mathbb\{E\}\_\{t,\\,x\_\{t\}\\sim p\_\{t\}^\{\\mathrm\{mix\}\}\}\\left\[\\left\\\|f^\{\*\}\_\{t\}\(x\_\{t\}\)\-f\_\{t\}^\{\\mathrm\{mix\}\}\(x\_\{t\}\)\\right\\\|^\{2\}\\right\]\.
### A\.3SFD’s details
SFD\([Chen et al\., 2025](https://arxiv.org/html/2609.36099#bib.bib44)\)distills a conditional teacher model into a one\-step generator in parallel with class forgetting\. It splits the generator loss from the inverse distillation scheme \([2](https://arxiv.org/html/2609.36099#S2.E2)\) into the forget and remaining losses as in \([3](https://arxiv.org/html/2609.36099#S2.E3)\): in the forget loss, it aligns the conditional scores of the generator’s forget classescFc\_\{F\}with the teacher’s scores for the safe classescSc\_\{S\}; in the remaining loss, it leaves other classescRc\_\{R\}unswapped\. The method also modifies these losses for better convergence, following the SiD framework\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\):
ℒSFD\-gen\-remain\(p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{SFD\-gen\-remain\}\}\(p^\{\\theta\}\_\{0\}\)=\\displaystyle=𝔼t,xθ0∼pθ0\(⋅\|cR\),xθt∼pθt\(⋅\|xθ0,cR\)\[−2⋅αSiD∥ft∗\(xtθ\|cR\)−ft\(xtθ\|cR\)∥2\\displaystyle\\mathbb\{E\}\_\{t,x^\{\\theta\}\_\{0\}\\sim p^\{\\theta\}\_\{0\}\(\\cdot\|c\_\{R\}\),x^\{\\theta\}\_\{t\}\\sim p^\{\\theta\}\_\{t\}\(\\cdot\|x^\{\\theta\}\_\{0\},c\_\{R\}\)\}\[\-2\\cdot\\alpha\_\{\\text\{SiD\}\}\\\|f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{R\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{R\}\)\\\|^\{2\}\(9\)\+\\displaystyle\+2⟨ft∗\(xtθ\|cR\)−ft\(xtθ\|cR\),ft∗\(xtθ\|cR\)−ftθ\(xtθ\|x0θ,cR\)⟩\],\\displaystyle 2\\langle f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{R\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{R\}\),f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{R\}\)\-f\_\{t\}^\{\\theta\}\(x^\{\\theta\}\_\{t\}\|x^\{\\theta\}\_\{0\},c\_\{R\}\)\\rangle\],ℒSFD\-gen\-forget\(p0θ\)\\displaystyle\\mathcal\{L\}\_\{\\text\{SFD\-gen\-forget\}\}\(p^\{\\theta\}\_\{0\}\)=\\displaystyle=𝔼t,xθ0∼pθ0\(⋅\|cF\),xθt∼pθt\(⋅\|xθ0,cF\)\[−2⋅αSiD∥ft∗\(xtθ\|cS\)−ft\(xtθ\|cF\)∥2\\displaystyle\\mathbb\{E\}\_\{t,x^\{\\theta\}\_\{0\}\\sim p^\{\\theta\}\_\{0\}\(\\cdot\|c\_\{F\}\),x^\{\\theta\}\_\{t\}\\sim p^\{\\theta\}\_\{t\}\(\\cdot\|x^\{\\theta\}\_\{0\},c\_\{F\}\)\}\[\-2\\cdot\\alpha\_\{\\text\{SiD\}\}\\\|f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{S\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{F\}\)\\\|^\{2\}\(10\)\+\\displaystyle\+2⟨ft∗\(xtθ\|cS\)−ft\(xtθ\|cF\),ft∗\(xtθ\|cS\)−ftθ\(xtθ\|x0θ,cF\)⟩\],\\displaystyle 2\\langle f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{S\}\)\-f\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{F\}\),f^\{\*\}\_\{t\}\(x^\{\\theta\}\_\{t\}\|c\_\{S\}\)\-f\_\{t\}^\{\\theta\}\(x^\{\\theta\}\_\{t\}\|x^\{\\theta\}\_\{0\},c\_\{F\}\)\\rangle\],whereαSiD\\alpha\_\{\\text\{SiD\}\}is an arbitrary parameter, usually taken from the range\[0\.5,1\.2\]\[0\.5,1\.2\]\. The loss for the fake model remains the same for all classes\.
## Appendix BExperimental details
### B\.1Architectures and teacher checkpoints
For CIFAR\-10, the FM setup uses the TorchCFM U\-Net architecture from the public 400k\-step Independent Conditional Flow Matching \(I\-CFM\) checkpoint111[https://github\.com/atong01/conditional\-flow\-matching/tree/main/examples/images/cifar10](https://github.com/atong01/conditional-flow-matching/tree/main/examples/images/cifar10), whereas the SiD setup uses the DDPM\+\+ \(SongUNet\) architecture of the public unconditional EDM\-VP teacher adopted by the official SiD implementation222[https://github\.com/mingyuanzhou/SiD](https://github.com/mingyuanzhou/SiD)\. Both CIFAR\-10 teachers operate on32×3232\\times 32RGB images\. The public checkpoints arecfm\_cifar10\_weights\_step\_400000\.ptfor FM andedm\-cifar10\-32x32\-uncond\-vp\.pklfor SiD\.
For MNIST, we preserve each architecture family but adapt it to28×2828\\times 28grayscale inputs\. In the FM U\-Net, we change the input and output channels from 3 to 1, reduce the resolution hierarchy from\[1,2,2,2\]to\[1,2,2\], move self\-attention from resolution 16 to 14, reduce the number of channels per attention head from 64 to 32\. In the SiD DDPM\+\+ model, we change only the image resolution from 32 to 28, the image channels from 3 to 1, and the attention resolution from 16 to 14\.
### B\.2Teacher sampling
For the CIFAR\-10 FM teacher, we use the adaptive Dormand–Prince \(Dopri5\) ODE solver with relative and absolute tolerances of10−510^\{\-5\}, following the original TorchCFM setup\. For the MNIST FM teacher, we use a fixed\-step Euler solver with 100 integration steps\. For both the MNIST and CIFAR\-10 SiD teachers, we use the deterministic 18\-step EDM sampler with the second\-order correction prescribed by the original EDM and SiD implementations\. Every distilled generator produces a sample in one step\.
### B\.3Evaluation classifiers
To compute FGR, we classify 50,000 generated images and report the percentage assigned to each forgotten class\. For MNIST, we use the public LeNet\-5 checkpoint333[https://github\.com/hrfang/LeNet5\-code\-examples](https://github.com/hrfang/LeNet5-code-examples); its repository reports99\.13%99\.13\\%validation accuracy and98\.94%98\.94\\%test accuracy\. For CIFAR\-10, we use the public ResNet\-56 checkpoint444[https://github\.com/chenyaofo/pytorch\-cifar\-models](https://github.com/chenyaofo/pytorch-cifar-models), for which the repository reports94\.37%94\.37\\%top\-1 and99\.83%99\.83\\%top\-5 accuracy\. These classifiers are external measurement instruments: they are not part of IDU, are never queried by the training loop, and contribute no loss, gradient, feature representation, or conditioning signal\.
### B\.4FID protocol and evaluation data
All reported FID values useclean\-fidv0\.1\.35\([Parmar et al\., 2022](https://arxiv.org/html/2609.36099#bib.bib52)\)inlegacy\_tensorflowmode and 50,000 generated images\. Images are quantized before extracting 2048\-dimensional TensorFlow\-compatible Inception features; grayscale MNIST samples are replicated across three channels\. We use all available real training images from the distribution targeted by each row: the full training set for the full\-data teacher and pure distillation, and the corresponding training subset with the forgotten labels removed for IDU and retained\-only retraining\. Thus, the paired\-class references contain 60,000 full or 47,604 retained MNIST images and 50,000 full or 40,000 retained CIFAR\-10 images\. Single\-class experiments analogously use the complete training subset excluding the selected class\. Test images are not used as FID references\.
Under the same legacy feature pipeline, the only numerical convention that differs from the original SiD evaluator is covariance normalization: CleanFID uses the sample covariance denominatorN−1N\-1, whereas the original SiD/StyleGAN implementation uses the population denominatorNN\. At 50,000 samples this changes the covariance scale only by the factorN/\(N−1\)≈1\.00002N/\(N\-1\)\\approx 1\.00002\. Minor implementation\-level differences remain in covariance symmetrization and numerical stabilization\. Teacher FID uses the multi\-step samplers described above; every distilled or IDU generator is evaluated with one\-step inference\.
### B\.5Optimization hyperparameters
#### FM on MNIST\.
Full\-data and retained\-only teachers use learning rate10−410^\{\-4\}, total batch size 128, and 100k and 50k optimizer steps, respectively\. The retained\-only teacher required fewer iterations because it empirically converged faster, likely due to the smaller amount of training data\. Pure distillation, retained\-only distillation, and IDU use learning rate10−410^\{\-4\}, total batch size 256, and training horizons of approximately 30,000 generator iterations\. Teacher pretraining uses Adam withβ=\(0\.9,0\.999\)\\beta=\(0\.9,0\.999\); distillation and IDU use Adam withβ=\(0,0\.999\)\\beta=\(0,0\.999\)for both trainable networks\. All MNIST FM optimizers use cosine learning\-rate annealing over their respective horizons\.
#### FM on CIFAR\-10\.
The 400k\-step teacher recipe uses learning rate2×10−42\\times 10^\{\-4\}, global batch size 128, 5,000 warm\-up steps, and 400,000 optimizer updates\. Pure and retained\-only distillation use learning rate3×10−53\\times 10^\{\-5\}, total batch size 256, and approximately 50,000 generator iterations; IDU uses the same learning rate and horizon with configured per\-process batch size 256\. These runs use 500 warm\-up steps\. Across FM distillation and IDU runs, we use gradient clipping at 1, generator EMA 0\.999, andαSiD=0\.5\\alpha\_\{\\mathrm\{SiD\}\}=0\.5; the CIFAR\-10 teacher itself uses EMA 0\.9999\.
#### SiD experiments\.
MNIST teacher training uses learning rate10−410^\{\-4\}, global batch size 512, and a 64\-million\-image \(64 Mimg\) horizon\. The CIFAR\-10 teacher recipe uses learning rate10−310^\{\-3\}, global batch size 512, and a 200 Mimg training horizon; the full\-data result uses the public EDM\-VP checkpoint, while retained\-only teachers follow this recipe\. On both datasets, IDU, pure, and retained\-only distillation use learning rates10−510^\{\-5\}for the generator and fake score network, global batch size 512, and a 100 Mimg training horizon\. We retainαSiD=1\.2\\alpha\_\{\\mathrm\{SiD\}\}=1\.2,tmax=800t\_\{\\max\}=800, and initial noise standard deviation 2\.5 from the original SiD recipe\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\)\.
These values are maximum optimization horizons rather than a claim that later checkpoints are always preferable\. We select the reported checkpoint by the lowest observed FID; for FM, selection additionally precedes the late forgetting reversal discussed in Section[5](https://arxiv.org/html/2609.36099#S5)\.
### B\.6Initialization
Our IDU can be used for both distillation from scratch and fine\-tuning of an already pretrained generator\. The only difference between the setups, besides the hyperparameter values, is the initialization: for fine\-tuning, we initializeGθG\_\{\\theta\}from the pretrained generator; otherwise, we initialize it from the one\-step teacher inference scheme\.
### B\.7Fine\-tuning experiments
We evaluate fine\-tuning on CIFAR\-10 by initializing from the corresponding pure\-distillation checkpoint and continuing IDU optimization with learning rates5×10−65\\times 10^\{\-6\}for FM and10−610^\{\-6\}for SiD, retaining the forgetting weights from the main experiments:ρ=0\.6\\rho=0\.6for FM andρ=0\.2\\rho=0\.2for SiD\. For forgotten classes 1 and 9, FM fine\-tuning yields Retain FID5\.17±0\.085\.17\\pm 0\.08and FGRs0\.89±0\.02%0\.89\\pm 0\.02\\%and0\.42±0\.03%0\.42\\pm 0\.03\\%, respectively\. SiD fine\-tuning yields Retain FID3\.12±0\.063\.12\\pm 0\.06and FGRs2\.69±0\.06%2\.69\\pm 0\.06\\%and2\.60±0\.04%2\.60\\pm 0\.04\\%; see Table[4](https://arxiv.org/html/2609.36099#A2.T4)for comparison\. Both fine\-tuned models reduce forgotten\-class generation relative to pure distillation, but the trade\-off differs by backbone: compared with IDU trained from scratch, FM improves Retain FID while slightly increasing FGR, whereas SiD maintains comparable Retain FID but higher FGR\.
Table 4:Fine\-tune FM/SiD results for the purely distilled generators on CIFAR\-10\. FID uses full\-data references for Pretrain/Pure distillation and retained\-data references for IDU and fine\-tuning; FGR is reported per forgotten class\. Values are mean±\\pmstandard deviation over five runs, except†\\daggervalues from\([Zhou et al\., 2024](https://arxiv.org/html/2609.36099#bib.bib14)\)\.FMSiDModeFID↓\\downarrowFGR \(%\)↓\\downarrowClass 1 / Class 9FID↓\\downarrowFGR \(%\)↓\\downarrowClass 1 / Class 9Pretrain3\.66±0\.033\.66\\pm 0\.0312\.50±0\.0812\.50\\pm 0\.0811\.18±0\.1211\.18\\pm 0\.121\.97†1\.97^\{\\dagger\}11\.13±0\.1211\.13\\pm 0\.1210\.04±0\.1510\.04\\pm 0\.15Pure distillation4\.35±0\.054\.35\\pm 0\.057\.58±0\.097\.58\\pm 0\.098\.75±0\.218\.75\\pm 0\.211\.92±0\.02†1\.92\\pm 0\.02\\,^\{\\dagger\}10\.11±0\.1410\.11\\pm 0\.1410\.66±0\.1410\.66\\pm 0\.14Forgotten classes\{1,9\}\\\{1,9\\\}\{1,9\}\\\{1,9\\\}IDU from scratch\(ρCIFAR\-10=0\.6\)\(\\rho\_\{\\mathrm\{CIFAR\\text\{\-\}10\}\}=0\.6\)\(ρSiD=0\.2\)\(\\rho\_\{\\mathrm\{SiD\}\}=0\.2\)5\.81±0\.055\.81\\pm 0\.050\.56±0\.050\.56\\pm 0\.050\.39±0\.030\.39\\pm 0\.033\.15±0\.033\.15\\pm 0\.031\.11±0\.061\.11\\pm 0\.061\.25±0\.061\.25\\pm 0\.06IDU Fine\-tuning\(ρCIFAR\-10=0\.6\)\(\\rho\_\{\\mathrm\{CIFAR\\text\{\-\}10\}\}=0\.6\)\(ρSiD=0\.2\)\(\\rho\_\{\\mathrm\{SiD\}\}=0\.2\)5\.17±0\.085\.17\\pm 0\.080\.89±0\.020\.89\\pm 0\.020\.42±0\.030\.42\\pm 0\.033\.12±0\.063\.12\\pm 0\.062\.69±0\.062\.69\\pm 0\.062\.60±0\.042\.60\\pm 0\.04
### B\.8Code and checkpoint release
Upon publication, we will release the complete source code, exact configurations, evaluation scripts, and checkpoints used to produce the main reported results\.
### B\.9MNIST FM retraining and sensitivity toρ\\rho
The retained\-only FM baseline on MNIST exhibits a distinct sensitivity\. The exceptionally low Retain FID of the retrained teacher \(0\.340\.34\) may reflect strong overfitting to the smaller retained subset rather than uniformly better generalization\. In its distilled counterpart, we consistently observed poor generation of digit 2, which raises Retain FID to5\.725\.72\. This failure occurred despite using the same architecture and optimization settings as the stable full\-data pretraining and distillation runs, suggesting sensitivity of the retraining baseline to the altered data distribution rather than an intentional hyperparameter disadvantage\.
Table[5](https://arxiv.org/html/2609.36099#A2.T5)reports the FM/MNIST results for differentρ\\rho\. The main setting,ρ=0\.4\\rho=0\.4, gives the best observed Retain FID–FGR trade\-off\. Increasingρ\\rhoto0\.60\.6or above drives the target\-class FGR values to zero, but also suppresses additional digits—most consistently 2 and 5, with 8 or 9 affected in some runs—and raises Retain FID above 10\. Conversely, decreasingρ\\rhobelow0\.40\.4progressively weakens forgetting of digits 3 and 7 and does not improve Retain FID over the main setting\. This sensitivity appears specific to the low\-dimensional MNIST setting and the FM architecture and training recipe used here, and motivates evaluation with larger and more stable backbones\. In contrast, the more recent SiD distillation backbone remained stable atρ=0\.2\\rho=0\.2, which matches the nominal fraction of two forgotten classes among ten, without the same collateral class suppression\.
Table 5:Effect of the forget\-mixture weightρ\\rhoon FM\-based IDU for MNIST when jointly forgetting digits 3 and 7\. The main configuration is bold\. The final column lists qualitatively suppressed digits in addition to the target digits 3 and 7\.ρ\\rhoRetain FID↓\\downarrowFGR 3 \(%\)↓\\downarrowFGR 7 \(%\)↓\\downarrowAdditional suppressed digits0\.050\.054\.29±0\.024\.29\\pm 0\.027\.07±0\.067\.07\\pm 0\.068\.18±0\.068\.18\\pm 0\.06—0\.10\.13\.96±0\.043\.96\\pm 0\.043\.87±0\.093\.87\\pm 0\.094\.62±0\.104\.62\\pm 0\.10—0\.20\.24\.81±0\.084\.81\\pm 0\.082\.22±0\.072\.22\\pm 0\.071\.63±0\.071\.63\\pm 0\.07—0\.4\\mathbf\{0\.4\}3\.57±0\.03\\mathbf\{3\.57\\pm 0\.03\}0\.16±0\.01\\mathbf\{0\.16\\pm 0\.01\}0\.16±0\.01\\mathbf\{0\.16\\pm 0\.01\}—0\.60\.611\.20±0\.1011\.20\\pm 0\.1000002, 5, 80\.80\.810\.58±0\.0710\.58\\pm 0\.0700002, 50\.90\.910\.35±0\.0410\.35\\pm 0\.0400002, 5, 9
## Appendix CVisual Results
We qualitatively compare samples from each multi\-step teacher with samples from the corresponding one\-step IDU generator\. For MNIST, IDU is trained to forget digits 3 and 7; for CIFAR\-10, it is trained to forget automobile \(class 1\) and truck \(class 9\)\. Across both FM and SiD, the teacher grids contain the target classes, whereas the IDU grids visibly suppress them while preserving samples from the retained classes\. These finite grids are intended as qualitative illustrations; the corresponding 50,000\-sample FGR measurements are reported in Table[1](https://arxiv.org/html/2609.36099#S4.T1)\.

Teacher

IDU
Figure 2:FM on MNIST: teacher and IDU samples when forgetting digits 3 and 7\.
Teacher

IDU
Figure 3:SiD on MNIST: teacher and IDU samples when forgetting digits 3 and 7\.
Teacher

IDU
Figure 4:FM on CIFAR\-10: teacher and IDU samples when forgetting automobile \(class 1\) and truck \(class 9\)\.
Teacher

IDU
Figure 5:SiD on CIFAR\-10: teacher and IDU samples when forgetting automobile \(class 1\) and truck \(class 9\)\.Similar Articles
On-Policy Distillation (5 minute read)
This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.
Dataset Distillation by Influence Matching
This paper introduces Influence Matching (Inf-Match), a dataset distillation method that aligns the final training outcome by learning a compact synthetic set whose effect on converged parameters matches that of the full dataset. It achieves state-of-the-art accuracy on classification benchmarks and outperforms strong baselines on vision-language distillation tasks.
Rethinking Reverse KL as Adaptive Entropy Distillation
This paper proposes Adaptive Entropy Distillation (AED), a method that dynamically calibrates token-level imitation strength in knowledge distillation using teacher entropy, achieving superior performance on instruction-following and mathematical reasoning benchmarks.
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
DiffusionOPD proposes a multi-task training paradigm for diffusion models that uses online policy distillation to efficiently combine task-specific teachers into a unified student, achieving state-of-the-art results on all evaluated benchmarks.
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.