Constitutional Midtraining: Content Presence Drives Alignment Gains
Summary
This paper tests constitutional midtraining, inserting principled values-based content during midtraining at 120B scale, and finds it yields durable alignment gains (e.g., blunting SFT-induced blackmail propensity) with no average capability cost, suggesting a cheap complement to SFT-centered pipelines.
View Cached Full Text
Cached at: 08/03/26, 09:35 PM
Paper page - Constitutional Midtraining: Content Presence Drives Alignment Gains
Source: https://huggingface.co/papers/2607.26654
Abstract
Post-trainingalignmentisoftenshallow,erodingunderfine-tuning.Whethermidtraininginterventions,cleanlyisolatedfrompost-training,canproducedurablealignmentremainsuntested.Wetestthisviaconstitutionalmidtraining:insertingprincipled,values-basedcontentintomidtrainingagainstareplay-onlycontrolat120Bscale.Our394M-tokenconstitutionalcorpus,builtfromAnthropic’sConstitution,usesa2x2factorialdesign(curriculumorderingxdeliberativereasoning)toproducefourconstitutionallymidtrainedconditionsplusacontrol,evaluatedonself-generatedandestablishedbenchmarksincludingalignmentunderpressure,valueconflictresolution,blackmail,andemergentmisalignmentacrossthreestages:post-midtraining,post-SFT,andpost-benignfine-tuning.Constitutionallymidtrainedmodelsoutperformthecontrolonalignmentgeneralizationanddurability,notablyonblackmail:SFTinstillsablackmailpropensityinallmodels,butconstitutionalmidtrainingbluntsit,withtheadvantagesurvivingbenignfine-tuning(-17.5pp).Thisdurabilitydoesnotextendtosettingsrequiringactiveresistancetoin-contextpressureorconflict,wheretheadvantageattenuatesafterSFT.Thepresenceofconstitutionalcontentatmidtrainingalsomattersmorethanitsstructure,andconstitutionalmidtrainingincursnocost,onaverage,onthecapabilitieswetest(MMLU,ARC-Easy,piqa,GSM8K)atanystage.Amodestamountofconstitutionalcontentatmidtrainingcouldthereforeyieldbroad,persistentalignmentgains,offeringacheap,complementaryadditiontoSFT-centeredpipelines.Code,data,andmodelsareavailable.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2607\.26654
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
Datasets citing this paper1
#### cho-ai/constitutional-mt-data Viewer• Updated3 days ago • 2.35M • 192 • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.26654 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Constitutional Midtraining: Content Presence Drives Alignment Gains
This paper introduces constitutional midtraining, inserting values-based content into the midtraining phase of large language models, and shows that it produces more durable alignment gains compared to post-training methods, with benefits persisting after fine-tuning. The approach incurs no capability cost and improves resistance to blackmail and other alignment pressures.
Contextual Value Alignment via Multilayer Combinatorial Fusion
This paper proposes MCF-CVA, a multilayer combinatorial fusion framework for contextual value alignment of LLMs, which uses multiple moral agents and an expansion-reduction process to better capture ethical pluralism and outperform single-agent baselines.
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
This paper investigates emergent and subliminal misalignment in LLMs through a data-centric lens, showing that harmful fine-tuning effects depend on structural properties of the data, task difficulty, pretraining composition, and training channels, with experiments comparing off-policy and on-policy distillation.
CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness
Proposes CASE, a framework combining training-time causal alignment and inference-time structural enforcement to improve faithfulness of chain-of-thought reasoning in large language models, achieving a 37% average improvement in CoT faithfulness across benchmarks.
Bypassing LLM Guardrails: How Plain Text Shifts Latent Trajectories Without Jailbreaks
The article presents a research finding that saturating an LLM's context window with benign narrative text can dominate the attention mechanism and shift latent trajectories, potentially bypassing alignment guardrails without traditional jailbreaks. It argues that current alignment methods are a superficial fix for a fundamentally fluid architecture.