Tag
This paper investigates exact deletion from language-model memory, showing that subtractive methods work when record influence is addressable, while replay/rebuild is needed when influence is woven into recurrent state. Experiments on Gemma and Kimi hybrid models demonstrate trade-offs in utility and exactness.
Introduces LENS, a contextualization-based evaluation protocol for testing narrative unlearning in large language models, evaluating suppression across direct, attributed, contrastive, and abstract levels.
This paper introduces Stochastic Meta-Unlearning (SMU), a bilevel framework that uses VLM-level feedback to learn an unlearning-ready initialization for the language backbone, achieving better forget-retain trade-offs in multimodal unlearning.
OriginBlame is a record- and token-level data provenance system that propagates author identity through AI training data pipelines, enabling precise forget sets for machine unlearning. It eliminates over-deletion from dataset-level systems and improves unlearning effectiveness.
Proposes GSUO, a guidance-signal-aware optimization framework for machine unlearning that uses fine-grained signals to guide the forgetting process, avoiding over-unlearning and under-unlearning, and outperforms 14 baselines.
This paper frames machine unlearning in LLMs as an asymmetric generalization problem, introduces the SUITE evaluation protocol and training corpus to address under- and over-forgetting, and presents JensUn++, an algorithm achieving the best forget-retain utility trade-off across three LLMs.
A comprehensive survey of methods, datasets, and benchmarks for multimodal unlearning across vision, language, video, and audio, providing a taxonomy and highlighting open problems.
Proposes a practical auditor that uses membership inference attacks to compute data-dependent lower bounds on the unlearning parameter, finding a sharp separation between certified algorithms (e.g., model clipping, rewind-to-delete) that achieve tight bounds and empirical methods (e.g., Hessian-based unlearning, gradient ascent) that exhibit large bounds, indicating poor unlearning.
CBD introduces an API-only black-box unlearning framework for LLMs that uses two auxiliary models to create controlled behavioral divergence between retained and target data, achieving a better unlearning-utility trade-off compared to existing methods.
This position paper argues that the term 'machine unlearning' is overused in LLM research, advocating for stricter terminology tied to dataset-defined deletion and retraining-equivalence guarantees.
This paper argues that standard output-level evaluations of machine unlearning overestimate success, showing that methods can appear successful at the output layer while retaining structured representation-level discrepancies relative to retrained models. The authors propose retraining-consistent representation forgetting as a stronger evaluative lens.
Proposes BindingSubspace (BSU), a representation-level framework that isolates and attenuates intent-conditioned directions in end-to-end spoken language understanding models to prevent capability persistence, where suppressing an intent still allows slot generation under forced prefixes. The method reduces forced-prefix recoverability while preserving retained performance on SLU benchmarks.
This paper proposes PreUnlearn, a framework for auditing collateral knowledge damage in LLM unlearning before execution, using data-centric analysis to predict downstream damage across semantic layers.
Proposes SAGE, a post-hoc method to sanitize the final unlearning vector in LLMs, improving the retain-forget trade-off without rerunning the unlearning pipeline.
RepSelect introduces a method for robust LLM unlearning that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50× better robustness against relearning attacks compared to existing baselines across multiple model families.
This paper introduces SPACE, the first source-free unlearning framework for multimodal large language models (MLLMs), which uses text-guided proxy anchor selection and dual-constraint semantic isolation to erase target concepts without requiring access to original training data, achieving performance comparable to data-dependent methods.
The paper proposes TRACE, a method for machine unlearning in Mixture-of-Experts language models that calibrates retain regularization by reweighting token-level retain losses to address forget-retain routing mismatch. Experiments show improved forget-utility trade-off across multiple MoE LLMs.
This paper formalizes exact unlearning in reinforcement learning, proposing a ρ-TV-stable RL algorithm for tabular MDPs that efficiently removes a user's data influence at a fraction of retraining cost, achieving near-minimax-optimal regret bounds. The work is accepted at ICML and establishes both upper and lower bounds for ρ-TV-stable RL algorithms.
Introduces MASC (Margin Self-Correction), an efficient unlearning method for LLMs that uses an online stopping rule to achieve competitive forget–retain trade-offs at reduced computational cost, validated on TOFU and MUSE benchmarks.
AMNESIA is the first large-scale open-source benchmark for medical unlearning, comprising 70,560 QA pairs from 8,820 patient notes across 11 diseases, designed to evaluate forgetting of both factual and reasoning knowledge in LLMs.