Anthropic published research on GRAM: a technique to surgically remove dangerous knowledge from AI models at the weight level
Summary
Anthropic published research on GRAM, a technique for surgically removing dangerous knowledge from AI models at the weight level, advancing AI safety.
Similar Articles
@AnthropicAI: We’re pleased to have collaborated with AE Studio on this research. Read more here: https://anthropic.com/research/off-…
Anthropic and AE Studio propose GRAM (Gradient-Routed Auxiliary Modules), a method to surgically confine dual-use knowledge in AI models to removable modules, enabling selective access without retraining multiple models. Preliminary results show promise for more flexible safety controls.
Jul 8, 2026AlignmentAn off switch for dual use knowledge in AI models
Anthropic and AE Studio introduce GRAM, a method that confines dual-use knowledge in AI models to removable modules, enabling surgical control over dangerous capabilities without retraining the entire model. Preliminary results suggest potential for safer deployment of frontier models.
Modular Pretraining Enables Access Control
This paper introduces GRAM (gradient-routed auxiliary modules), a modular pretraining method that enables access control by selectively adding and ablating modules to limit dual-use capabilities in AI models, showing cost reductions compared to data filtering.
Anthropic guardrails does it again
Anthropic's guardrails have reportedly been tested again, highlighting ongoing developments in AI safety.
AI guardrails stripped from Meta and Google models in minutes
Researchers rapidly removed safety protections from widely deployed AI models, eliciting dangerous outputs and raising concerns about robustness and release practices.