Anthropic published research on GRAM: a technique to surgically remove dangerous knowledge from AI models at the weight level

Reddit r/artificial Papers

Summary

Anthropic published research on GRAM, a technique for surgically removing dangerous knowledge from AI models at the weight level, advancing AI safety.

No content available
Original Article

Similar Articles

Jul 8, 2026AlignmentAn off switch for dual use knowledge in AI models

Anthropic Research

Anthropic and AE Studio introduce GRAM, a method that confines dual-use knowledge in AI models to removable modules, enabling surgical control over dangerous capabilities without retraining the entire model. Preliminary results suggest potential for safer deployment of frontier models.

Modular Pretraining Enables Access Control

arXiv cs.LG

This paper introduces GRAM (gradient-routed auxiliary modules), a modular pretraining method that enables access control by selectively adding and ablating modules to limit dual-use capabilities in AI models, showing cost reductions compared to data filtering.