Constitutional Midtraining: Content Presence Drives Alignment Gains

Hugging Face Daily Papers Papers

Summary

This paper tests constitutional midtraining, inserting principled values-based content during midtraining at 120B scale, and finds it yields durable alignment gains (e.g., blunting SFT-induced blackmail propensity) with no average capability cost, suggesting a cheap complement to SFT-centered pipelines.

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensity in all models, but constitutional midtraining blunts it, with the advantage surviving benign fine-tuning (-17.5pp). This durability does not extend to settings requiring active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also matters more than its structure, and constitutional midtraining incurs no cost, on average, on the capabilities we test (MMLU, ARC-Easy, piqa, GSM8K) at any stage. A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
Original Article
View Cached Full Text

Cached at: 08/03/26, 09:35 PM

Paper page - Constitutional Midtraining: Content Presence Drives Alignment Gains

Source: https://huggingface.co/papers/2607.26654

Abstract

Post-trainingalignmentisoftenshallow,erodingunderfine-tuning.Whethermidtraininginterventions,cleanlyisolatedfrompost-training,canproducedurablealignmentremainsuntested.Wetestthisviaconstitutionalmidtraining:insertingprincipled,values-basedcontentintomidtrainingagainstareplay-onlycontrolat120Bscale.Our394M-tokenconstitutionalcorpus,builtfromAnthropic’sConstitution,usesa2x2factorialdesign(curriculumorderingxdeliberativereasoning)toproducefourconstitutionallymidtrainedconditionsplusacontrol,evaluatedonself-generatedandestablishedbenchmarksincludingalignmentunderpressure,valueconflictresolution,blackmail,andemergentmisalignmentacrossthreestages:post-midtraining,post-SFT,andpost-benignfine-tuning.Constitutionallymidtrainedmodelsoutperformthecontrolonalignmentgeneralizationanddurability,notablyonblackmail:SFTinstillsablackmailpropensityinallmodels,butconstitutionalmidtrainingbluntsit,withtheadvantagesurvivingbenignfine-tuning(-17.5pp).Thisdurabilitydoesnotextendtosettingsrequiringactiveresistancetoin-contextpressureorconflict,wheretheadvantageattenuatesafterSFT.Thepresenceofconstitutionalcontentatmidtrainingalsomattersmorethanitsstructure,andconstitutionalmidtrainingincursnocost,onaverage,onthecapabilitieswetest(MMLU,ARC-Easy,piqa,GSM8K)atanystage.Amodestamountofconstitutionalcontentatmidtrainingcouldthereforeyieldbroad,persistentalignmentgains,offeringacheap,complementaryadditiontoSFT-centeredpipelines.Code,data,andmodelsareavailable.

View arXiv pageView PDFProject pageGitHub2Add to collection

Get this paper in your agent:

hf papers read 2607\.26654

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### peterstran/nemotron-super-120b-cc-mt-graft-currdr-on-baseline-sft 121B• Updatedabout 2 hours ago

Datasets citing this paper1

#### cho-ai/constitutional-mt-data Viewer• Updated3 days ago • 2.35M • 192 • 2

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.26654 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv cs.CL

This paper introduces constitutional midtraining, inserting values-based content into the midtraining phase of large language models, and shows that it produces more durable alignment gains compared to post-training methods, with benefits persisting after fine-tuning. The approach incurs no capability cost and improves resistance to blackmail and other alignment pressures.

Contextual Value Alignment via Multilayer Combinatorial Fusion

arXiv cs.AI

This paper proposes MCF-CVA, a multilayer combinatorial fusion framework for contextual value alignment of LLMs, which uses multiple moral agents and an expansion-reduction process to better capture ethical pluralism and outperform single-agent baselines.

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

arXiv cs.LG

This paper investigates emergent and subliminal misalignment in LLMs through a data-centric lens, showing that harmful fine-tuning effects depend on structural properties of the data, task difficulty, pretraining composition, and training channels, with experiments comparing off-policy and on-policy distillation.

Bypassing LLM Guardrails: How Plain Text Shifts Latent Trajectories Without Jailbreaks

Reddit r/AI_Agents

The article presents a research finding that saturating an LLM's context window with benign narrative text can dominate the attention mechanism and shift latent trajectories, potentially bypassing alignment guardrails without traditional jailbreaks. It argues that current alignment methods are a superficial fix for a fundamentally fluid architecture.