llm-unlearning

Tag

Cards List
#llm-unlearning

Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning

arXiv cs.CL ↗ · 2026-09-16 Cached

Cascade is a hierarchical framework for LLM unlearning that minimizes recoverability through multi-level controls, improving upon existing methods by reducing residual knowledge in intermediate representations while maintaining utility.

0 favorites 0 likes
#llm-unlearning

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

arXiv cs.LG ↗ · 2026-09-02 Cached

The paper introduces forget-set misalignment in LLM unlearning and proposes a data-blind framework called CONFS to address it, achieving a competitive forgetting-utility balance.

0 favorites 0 likes
#llm-unlearning

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

arXiv cs.CL ↗ · 2026-08-25 Cached

This paper identifies tool-mediated recovery as a failure mode in LLM unlearning and proposes Agentic Tool Unlearning (ATU) to reduce both parametric recall and tool-based recovery while preserving normal tool use.

0 favorites 0 likes
#llm-unlearning

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

arXiv cs.CL ↗ · 2026-08-21 Cached

The paper introduces ConceptGuard, a benchmark for evaluating context-sensitive unlearning in large language models using dual-use concepts, revealing that current unlearning techniques perform poorly under this practical evaluation framework.

0 favorites 0 likes
#llm-unlearning

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

arXiv cs.CL ↗ · 2026-08-17 Cached

The paper proposes AdaPop, an adaptive popularity-based method for LLM unlearning that adjusts gradient pressure based on fact frequency to improve forgetting effectiveness and reduce leakage under queries.

0 favorites 0 likes
#llm-unlearning

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

arXiv cs.CL ↗ · 2026-08-13 Cached

Proposes J-Access, an inference-time audit using the Jacobian lens to measure residual knowledge accessibility in unlearned LLMs, finding that accessibility predicts recovery speed but that directly minimizing it fails to promote genuine deletion.

0 favorites 0 likes
#llm-unlearning

LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats

arXiv cs.LG ↗ · 2026-07-21 Cached

This survey examines LLM unlearning methods for cyber defense, introducing a three-level framework to distinguish behavioral suppression, representation-level attenuation, and true forgetting, and analyzing gradient-based, influence-based, and localized editing approaches.

0 favorites 0 likes
#llm-unlearning

CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence

arXiv cs.LG ↗ · 2026-06-29 Cached

CBD introduces an API-only black-box unlearning framework for LLMs that uses two auxiliary models to create controlled behavioral divergence between retained and target data, achieving a better unlearning-utility trade-off compared to existing methods.

0 favorites 0 likes
#llm-unlearning

RepSelect: Robust LLM Unlearning via Representation Selectivity

arXiv cs.CL ↗ · 2026-06-17 Cached

RepSelect introduces a method for robust LLM unlearning that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50× better robustness against relearning attacks compared to existing baselines across multiple model families.

0 favorites 0 likes
#llm-unlearning

Measuring the Depth of LLM Unlearning via Activation Patching

arXiv cs.CL ↗ · 2026-05-26 Cached

The paper proposes the Unlearning Depth Score (UDS), a metric that uses activation patching to quantify how thoroughly target knowledge is erased from LLMs, achieving state-of-the-art faithfulness and robustness across multiple unlearning methods.

0 favorites 0 likes
#llm-unlearning

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

arXiv cs.CL ↗ · 2026-05-13 Cached

This paper introduces Minor Component Unlearning (MCU), a novel approach to LLM unlearning that targets minor components in representations to resist relearning attacks. It addresses the vulnerability of existing methods by focusing on robust directions within the model's spectral structure.

0 favorites 0 likes
← Back to home

Submit Feedback