Tag
This survey examines LLM unlearning methods for cyber defense, introducing a three-level framework to distinguish behavioral suppression, representation-level attenuation, and true forgetting, and analyzing gradient-based, influence-based, and localized editing approaches.
Anthropic published research on GRAM, a technique for surgically removing dangerous knowledge from AI models at the weight level, advancing AI safety.
The paper argues that unlearning in LLMs should be goal-dependent, proposing a cosine-based meta-learned variant of RMU for dangerous knowledge and a multi-layer objective with probe directions for toxicity, achieving strong results across four 7-8B models.