machine-unlearning

Tag

Cards List
#machine-unlearning

Subtract or Replay? Exact Deletion from Language-Model Memory

arXiv cs.LG ↗ · 2026-07-31 Cached

This paper investigates exact deletion from language-model memory, showing that subtractive methods work when record influence is addressable, while replay/rebuild is needed when influence is woven into recurrent state. Experiments on Gemma and Kimi hybrid models demonstrate trade-offs in utility and exactness.

0 favorites 0 likes
#machine-unlearning

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

arXiv cs.CL ↗ · 2026-07-28 Cached

Introduces LENS, a contextualization-based evaluation protocol for testing narrative unlearning in large language models, evaluating suppression across direct, attributed, contrastive, and abstract levels.

0 favorites 0 likes
#machine-unlearning

Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper introduces Stochastic Meta-Unlearning (SMU), a bilevel framework that uses VLM-level feedback to learn an unlearning-ready initialization for the language backbone, achieving better forget-retain trade-offs in multimodal unlearning.

0 favorites 0 likes
#machine-unlearning

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

arXiv cs.AI ↗ · 2026-07-16 Cached

OriginBlame is a record- and token-level data provenance system that propagates author identity through AI training data pipelines, enabling precise forget sets for machine unlearning. It eliminates over-deletion from dataset-level systems and improves unlearning effectiveness.

0 favorites 0 likes
#machine-unlearning

Signal-Guided Optimization for Machine Unlearning

arXiv cs.LG ↗ · 2026-07-15 Cached

Proposes GSUO, a guidance-signal-aware optimization framework for machine unlearning that uses fine-grained signals to guide the forgetting process, avoiding over-unlearning and under-unlearning, and outperforms 14 baselines.

0 favorites 0 likes
#machine-unlearning

Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem

arXiv cs.LG ↗ · 2026-07-13 Cached

This paper frames machine unlearning in LLMs as an asymmetric generalization problem, introduces the SUITE evaluation protocol and training corpus to address under- and over-forgetting, and presents JensUn++, an algorithm achieving the best forget-retain utility trade-off across three LLMs.

0 favorites 0 likes
#machine-unlearning

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

arXiv cs.LG ↗ · 2026-07-10 Cached

A comprehensive survey of methods, datasets, and benchmarks for multimodal unlearning across vision, language, video, and audio, providing a taxonomy and highlighting open problems.

0 favorites 0 likes
#machine-unlearning

Auditing of Unlearning Algorithms

arXiv cs.LG ↗ · 2026-07-08 Cached

Proposes a practical auditor that uses membership inference attacks to compute data-dependent lower bounds on the unlearning parameter, finding a sharp separation between certified algorithms (e.g., model clipping, rewind-to-delete) that achieve tight bounds and empirical methods (e.g., Hessian-based unlearning, gradient ascent) that exhibit large bounds, indicating poor unlearning.

0 favorites 0 likes
#machine-unlearning

CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence

arXiv cs.LG ↗ · 2026-06-29 Cached

CBD introduces an API-only black-box unlearning framework for LLMs that uses two auxiliary models to create controlled behavioral divergence between retained and target data, achieving a better unlearning-utility trade-off compared to existing methods.

0 favorites 0 likes
#machine-unlearning

Position: The Term "Machine Unlearning" Is Overused in LLMs

arXiv cs.CL ↗ · 2026-06-29 Cached

This position paper argues that the term 'machine unlearning' is overused in LLM research, advocating for stricter terminology tied to dataset-defined deletion and retraining-equivalence guarantees.

0 favorites 0 likes
#machine-unlearning

Erased, but Not Gone: Output Forgetting Is Not True Forgetting

arXiv cs.LG ↗ · 2026-06-25 Cached

This paper argues that standard output-level evaluations of machine unlearning overestimate success, showing that methods can appear successful at the output layer while retaining structured representation-level discrepancies relative to retrained models. The authors propose retraining-consistent representation forgetting as a stronger evaluative lens.

0 favorites 0 likes
#machine-unlearning

Selective Capability Unlearning in End-to-End Spoken Language Understanding

arXiv cs.CL ↗ · 2026-06-24 Cached

Proposes BindingSubspace (BSU), a representation-level framework that isolates and attenuates intent-conditioned directions in end-to-end spoken language understanding models to prevent capability persistence, where suppressing an intent still allows slot generation under forced prefixes. The method reduces forced-prefix recoverability while preserving retained performance on SLU benchmarks.

0 favorites 0 likes
#machine-unlearning

PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper proposes PreUnlearn, a framework for auditing collateral knowledge damage in LLM unlearning before execution, using data-centric analysis to predict downstream damage across semantic layers.

0 favorites 0 likes
#machine-unlearning

SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector

arXiv cs.LG ↗ · 2026-06-18 Cached

Proposes SAGE, a post-hoc method to sanitize the final unlearning vector in LLMs, improving the retain-forget trade-off without rerunning the unlearning pipeline.

0 favorites 0 likes
#machine-unlearning

RepSelect: Robust LLM Unlearning via Representation Selectivity

arXiv cs.CL ↗ · 2026-06-17 Cached

RepSelect introduces a method for robust LLM unlearning that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50× better robustness against relearning attacks compared to existing baselines across multiple model families.

0 favorites 0 likes
#machine-unlearning

SPACE: Source-free Proxy Anchor Concept Erasure for MLLMs

arXiv cs.LG ↗ · 2026-06-10 Cached

This paper introduces SPACE, the first source-free unlearning framework for multimodal large language models (MLLMs), which uses text-guided proxy anchor selection and dual-constraint semantic isolation to erase target concepts without requiring access to original training data, achieving performance comparable to data-dependent methods.

0 favorites 0 likes
#machine-unlearning

Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models

arXiv cs.CL ↗ · 2026-06-10 Cached

The paper proposes TRACE, a method for machine unlearning in Mixture-of-Experts language models that calibrates retain regularization by reweighting token-level retain losses to address forget-retain routing mismatch. Experiments show improved forget-utility trade-off across multiple MoE LLMs.

0 favorites 0 likes
#machine-unlearning

Exact Unlearning in Reinforcement Learning

arXiv cs.LG ↗ · 2026-06-04 Cached

This paper formalizes exact unlearning in reinforcement learning, proposing a ρ-TV-stable RL algorithm for tabular MDPs that efficiently removes a user's data influence at a fraction of retraining cost, achieving near-minimax-optimal regret bounds. The work is accepted at ICML and establishes both upper and lower bounds for ρ-TV-stable RL algorithms.

0 favorites 0 likes
#machine-unlearning

Fast Unlearning at Scale via Margin Self-Correction

arXiv cs.LG ↗ · 2026-06-03 Cached

Introduces MASC (Margin Self-Correction), an efficient unlearning method for LLMs that uses an online stopping rule to achieve competitive forget–retain trade-offs at reduced computational cost, validated on TOFU and MUSE benchmarks.

0 favorites 0 likes
#machine-unlearning

AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis

arXiv cs.LG ↗ · 2026-06-01 Cached

AMNESIA is the first large-scale open-source benchmark for medical unlearning, comprising 70,560 QA pairs from 8,820 patient notes across 11 diseases, designed to evaluate forgetting of both factual and reasoning knowledge in LLMs.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback