Tag
This paper introduces constitutional midtraining, inserting values-based content into the midtraining phase of large language models, and shows that it produces more durable alignment gains compared to post-training methods, with benefits persisting after fine-tuning. The approach incurs no capability cost and improves resistance to blackmail and other alignment pressures.
This paper tests constitutional midtraining, inserting principled values-based content during midtraining at 120B scale, and finds it yields durable alignment gains (e.g., blunting SFT-induced blackmail propensity) with no average capability cost, suggesting a cheap complement to SFT-centered pipelines.
Anthropic researchers introduce Model Spec Midtraining (MSM), a new training stage between pretraining and fine-tuning designed to improve how models generalize from alignment training and reduce agentic misalignment.