标签
Anthropic and AE Studio propose GRAM (Gradient-Routed Auxiliary Modules), a method to surgically confine dual-use knowledge in AI models to removable modules, enabling selective access without retraining multiple models. Preliminary results show promise for more flexible safety controls.
Anthropic和AE Studio推出了GRAM方法,该方法将AI模型中的双重用途知识限制在可移除模块中,从而无需重新训练整个模型即可对危险能力进行精准控制。初步结果表明,这有望实现前沿模型的更安全部署。