knowledge-control

Tag

Cards List
#knowledge-control

@AnthropicAI: We’re pleased to have collaborated with AE Studio on this research. Read more here: https://anthropic.com/research/off-…

X AI KOLs · 2026-07-08 Cached

Anthropic and AE Studio propose GRAM (Gradient-Routed Auxiliary Modules), a method to surgically confine dual-use knowledge in AI models to removable modules, enabling selective access without retraining multiple models. Preliminary results show promise for more flexible safety controls.

0 favorites 0 likes
#knowledge-control

Jul 8, 2026AlignmentAn off switch for dual use knowledge in AI models

Anthropic Research · 2026-07-09 Cached

Anthropic and AE Studio introduce GRAM, a method that confines dual-use knowledge in AI models to removable modules, enabling surgical control over dangerous capabilities without retraining the entire model. Preliminary results suggest potential for safer deployment of frontier models.

0 favorites 0 likes
← Back to home

Submit Feedback