knowledge-removal

Tag

Cards List
#knowledge-removal

LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats

arXiv cs.LG · 2026-07-21 Cached

This survey examines LLM unlearning methods for cyber defense, introducing a three-level framework to distinguish behavioral suppression, representation-level attenuation, and true forgetting, and analyzing gradient-based, influence-based, and localized editing approaches.

0 favorites 0 likes
#knowledge-removal

Anthropic published research on GRAM: a technique to surgically remove dangerous knowledge from AI models at the weight level

Reddit r/artificial · 2026-07-09

Anthropic published research on GRAM, a technique for surgically removing dangerous knowledge from AI models at the weight level, advancing AI safety.

0 favorites 0 likes
#knowledge-removal

Model Unlearning Objectives Vary for Distinct Language Functions

arXiv cs.CL · 2026-05-27 Cached

The paper argues that unlearning in LLMs should be goal-dependent, proposing a cosine-based meta-learned variant of RMU for dangerous knowledge and a multi-layer objective with probe directions for toxicity, achieving strong results across four 7-8B models.

0 favorites 0 likes
← Back to home

Submit Feedback