Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Summary
The Imprint Reader is a model trained to describe frozen weight updates in language models, enabling behavioral intervention to improve safety and reasoning. It uses SMaRT for training and MetaEdit for intervention, demonstrating feasibility in natural language readout.
View Cached Full Text
Cached at: 09/29/26, 08:10 AM
Paper page - Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Source: https://huggingface.co/papers/2609.35261
Abstract
AslanguagemodelstakeagrowingroleinAIdevelopment,anaturalaspirationisforthemtoreflectontheirownlearningprocess,ashumansdo,andusethatreflectiontoimprovethemselves.Atthesametime,thesemodelshaveanadvantagethathumanlearnerslack,sincetrainingleavesparameter-leveltracesthatcan,inprinciple,beinspecteddirectly.However,currentmodelscannotdecodethesetracesintoanexplicitaccountofwhattheyhavelearned.Tothisend,weintroducetheImprintReader,amodeltrainedwithSemanticMount-and-ReadTuning(SaRT)todescribefrozenweightupdates.SMaRTmountseachupdateontotheReaderandusesananchor-freemeta-querytoelicitanatural-languagedescription,whileno-changeandrandom-perturbationcontrolsdiscourageunsupportedclaims.Onheld-outupdates,thejointReaderreachesjudge-basedPass@100of2%forknowledgeand16%forbehavior.Theseresultsdemonstratethefeasibilityofnatural-languagereadoutwhilepointingtoreliabilityacrossupdatesasthenextstep.Beyondfree-formgeneration,theReaderprovidesadifferentiableproxyforthegapbetweenaspecifiedtargetbehaviorandacandidateweightupdate.Itscoordinate-alignedgradientssupportinterventionthroughMetaEdit.Ata0.5%pruningrate,Reader-guidedselectionraisesmeasuredharmful-promptrefusalfrom57.9%to64.1%underasafety-maintenancetarget.Usingbehaviordescriptionswithouttarget-tasktrainingdata,MetaEditincreasesthefrequencyofbacktrackingandsub-goalexpressionsinmathematicalreasoningtracesandraisesBFCLOverallfrom41.69%to44.60%.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.35261
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35261 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35261 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35261 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@LiorOnAI: New model from Thinking Machines: - Full weights available - Native text, image, and audio reasoning - 975B total param…
Thinking Machines releases Inkling, an MoE model with 975B total / 41B active parameters, supporting native text, image, and audio reasoning, up to 1M-token context, and full weights availability.
PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs
Introduces PromptPrint, a systematic study showing that users' habitual vocabulary and syntax in LLM prompts form a learnable behavioral biometric, with lexical features outperforming semantic encoders and revealing a uniqueness–consistency paradox.
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
This paper introduces influence-guided response rewriting to intervene on influential training examples, showing that rewriting responses creates stronger and more persistent behavioral shifts in language models than conventional reweighting.
@Marktechpost: Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With a Trained Effort Dial. No RoP…
Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal MoE model with 41B active parameters and a controllable thinking effort feature. It uses an encoder-free approach for multimodality and achieves state-of-the-art on several benchmarks.
@levie: Fantastic to see more open weights innovation happening right now, especially coming from a US Lab. The future of AI is…
Thinking Machines released Inkling, a multimodal AI model with open weights, capable of reasoning across text, image, and audio modalities. The model is available for fine-tuning on Tinker and via the Inkling Playground.