Tag
This paper introduces GRAM (gradient-routed auxiliary modules), a modular pretraining method that enables access control by selectively adding and ablating modules to limit dual-use capabilities in AI models, showing cost reductions compared to data filtering.
Anthropic and AE Studio propose GRAM (Gradient-Routed Auxiliary Modules), a method to surgically confine dual-use knowledge in AI models to removable modules, enabling selective access without retraining multiple models. Preliminary results show promise for more flexible safety controls.
Anthropic and AE Studio introduce GRAM, a method that confines dual-use knowledge in AI models to removable modules, enabling surgical control over dangerous capabilities without retraining the entire model. Preliminary results suggest potential for safer deployment of frontier models.