Tag
Cascade is a hierarchical framework for LLM unlearning that minimizes recoverability through multi-level controls, improving upon existing methods by reducing residual knowledge in intermediate representations while maintaining utility.
The paper introduces forget-set misalignment in LLM unlearning and proposes a data-blind framework called CONFS to address it, achieving a competitive forgetting-utility balance.
This paper identifies tool-mediated recovery as a failure mode in LLM unlearning and proposes Agentic Tool Unlearning (ATU) to reduce both parametric recall and tool-based recovery while preserving normal tool use.
The paper introduces ConceptGuard, a benchmark for evaluating context-sensitive unlearning in large language models using dual-use concepts, revealing that current unlearning techniques perform poorly under this practical evaluation framework.
The paper proposes AdaPop, an adaptive popularity-based method for LLM unlearning that adjusts gradient pressure based on fact frequency to improve forgetting effectiveness and reduce leakage under queries.
Proposes J-Access, an inference-time audit using the Jacobian lens to measure residual knowledge accessibility in unlearned LLMs, finding that accessibility predicts recovery speed but that directly minimizing it fails to promote genuine deletion.
This survey examines LLM unlearning methods for cyber defense, introducing a three-level framework to distinguish behavioral suppression, representation-level attenuation, and true forgetting, and analyzing gradient-based, influence-based, and localized editing approaches.
CBD introduces an API-only black-box unlearning framework for LLMs that uses two auxiliary models to create controlled behavioral divergence between retained and target data, achieving a better unlearning-utility trade-off compared to existing methods.
RepSelect introduces a method for robust LLM unlearning that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50× better robustness against relearning attacks compared to existing baselines across multiple model families.
The paper proposes the Unlearning Depth Score (UDS), a metric that uses activation patching to quantify how thoroughly target knowledge is erased from LLMs, achieving state-of-the-art faithfulness and robustness across multiple unlearning methods.
This paper introduces Minor Component Unlearning (MCU), a novel approach to LLM unlearning that targets minor components in representations to resist relearning attacks. It addresses the vulnerability of existing methods by focusing on robust directions within the model's spectral structure.