modular-pretraining

Tag

Cards List
#modular-pretraining

Modular Pretraining Enables Access Control

arXiv cs.LG · 2026-07-10 Cached

This paper introduces GRAM (gradient-routed auxiliary modules), a modular pretraining method that enables access control by selectively adding and ablating modules to limit dual-use capabilities in AI models, showing cost reductions compared to data filtering.

0 favorites 0 likes
#modular-pretraining

@AnthropicAI: We’re pleased to have collaborated with AE Studio on this research. Read more here: https://anthropic.com/research/off-…

X AI KOLs · 2026-07-08 Cached

Anthropic and AE Studio propose GRAM (Gradient-Routed Auxiliary Modules), a method to surgically confine dual-use knowledge in AI models to removable modules, enabling selective access without retraining multiple models. Preliminary results show promise for more flexible safety controls.

0 favorites 0 likes
#modular-pretraining

Jul 8, 2026AlignmentAn off switch for dual use knowledge in AI models

Anthropic Research · 2026-07-09 Cached

Anthropic and AE Studio introduce GRAM, a method that confines dual-use knowledge in AI models to removable modules, enabling surgical control over dangerous capabilities without retraining the entire model. Preliminary results suggest potential for safer deployment of frontier models.

0 favorites 0 likes
← Back to home

Submit Feedback