Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

arXiv cs.LG Papers

Summary

This paper introduces causal retention in interactive agents, examining how frozen learned states can answer intervention queries independently of training objectives, supported by theoretical analysis and experiments on finite causal systems and AI models like Qwen.

arXiv:2609.30650v1 Announce Type: new Abstract: Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:41 AM

# Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation
Source: [https://arxiv.org/html/2609.30650](https://arxiv.org/html/2609.30650)
Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, and Cheng Zeng

Shengjun Zhang sj\.zhang@hubu\.edu\.cnAffiliation:School of Artificial Intelligence, Hubei UniversityAffiliation:Key Laboratory of Intelligent Sensing System and Security \(Hubei University\), Ministry of EducationDong Xie xiedong04@baidu\.comAffiliation:Baidu Inc\.Yunlong Dong yunlongdong@outlook\.comAffiliation:Independent Researcher\*Correspondence: xiedong04@baidu\.com; wangxiang\.whu@whu\.edu\.cnXiang Wang wangxiang\.whu@whu\.edu\.cnAffiliation:School of Artificial Intelligence, Hubei UniversityAffiliation:Key Laboratory of Intelligent Sensing System and Security \(Hubei University\), Ministry of EducationCheng Zeng zc@hubu\.edu\.cnAffiliation:School of Artificial Intelligence, Hubei UniversityAffiliation:Key Laboratory of Intelligent Sensing System and Security \(Hubei University\), Ministry of Education

###### Abstract

Task performance need not determine which intervention mechanism an agent retains\. We study causal retention: whether a frozen learned state answers a mechanism\-probe map fixed independently of training, including action, context, direct target, value, and delay\. For finite structural causal model classes, the optimal probe error is a Bayes decision risk\. It vanishes exactly when every learning\-interface fiber lies within one probe\-answer fiber; any state obtained by post\-processing that interface inherits the same lower bound\. A posterior\-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error\-free target update\. Causal Core implements these conditions through evidence\-gated writing, readout filtering, temporal credit, hidden\-context setup, and local diagnostic updates\. Experiments cover finite causal systems, continuous simulators, an official TD\-MPC2 world model, and Qwen2\.5\-7B\-Instruct\. A frozen Qwen last\-layer probe reaches 0\.958 balanced accuracy on source mechanisms but 0\.583 on changed delays; the gated mechanism state reaches 1\.000 and accepts only 0\.056 of synchronized\-readout candidates\. In TD\-MPC2, five target states per actuator recover effect\-sign accuracy from 0\.057 to 0\.948 without degrading stable responses\. Causal retention is therefore distinct from task sufficiency and source\-domain decodability\.

††heading:Preprint 2026 1–September 2026††shortheadings:Causal Retention in Interactive Agents / Zhang, Liu, Xie, Dong, Wang, and Zeng††firstpage:1###### keywords

causal retention, interactive agents, mechanism memory, causal adaptation, structural causal models

## 1Introduction and Related Work

Training objectives identify equivalence classes of environments, not complete intervention mechanisms\. Two interactive systems may induce the same reward process, preferred action, or one\-step prediction target while disagreeing about the direct target of an action, its active context, or its delay\. A state that is sufficient for the training task can therefore be insufficient for an intervention query that the objective never had to answer\.

The ability of a frozen learned state to answer such queries is called*causal retention*\. The evaluator fixes the probe semantics independently of the learned state, freezes that state after source learning, and then asks which action changes which target, under which context, to what value, and after what delay\. Under a local mechanism shift, the same state must identify the changed entry without overwriting invariant ones\. This is a property of what learning retained, not only of the behavior it produced\.

Several causal\-learning traditions address neighboring but different estimands\. Structural causal models define interventions and counterfactuals\([Pearl, 2009](https://arxiv.org/html/2609.30650#bib.bib17);[Peters et al\., 2017](https://arxiv.org/html/2609.30650#bib.bib19)\)\. Treatment\-effect methods estimate average, conditional, or individual effects on a designated outcome\([Rubin, 1974](https://arxiv.org/html/2609.30650#bib.bib20);[Imbens and Rubin, 2015](https://arxiv.org/html/2609.30650#bib.bib13);[Chernozhukov et al\., 2018](https://arxiv.org/html/2609.30650#bib.bib4);[Wager and Athey, 2018](https://arxiv.org/html/2609.30650#bib.bib30);[Künzel et al\., 2019](https://arxiv.org/html/2609.30650#bib.bib14);[Nie and Wager, 2021](https://arxiv.org/html/2609.30650#bib.bib16);[Shalit et al\., 2017](https://arxiv.org/html/2609.30650#bib.bib23)\)\. Causal discovery instead estimates a graph or an interventional equivalence class\([Spirtes et al\., 2000](https://arxiv.org/html/2609.30650#bib.bib26);[Chickering, 2002](https://arxiv.org/html/2609.30650#bib.bib5);[Hauser and Bühlmann, 2012](https://arxiv.org/html/2609.30650#bib.bib11);[Shimizu et al\., 2006](https://arxiv.org/html/2609.30650#bib.bib24);[Zheng et al\., 2018](https://arxiv.org/html/2609.30650#bib.bib34)\)\. A correct outcome effect or graph can remain silent about the action\-indexed value, delay, and context that must be edited after transfer\.

Invariant prediction seeks predictors stable across environments\([Peters et al\., 2016](https://arxiv.org/html/2609.30650#bib.bib18);[Arjovsky et al\., 2019](https://arxiv.org/html/2609.30650#bib.bib1)\)\. Causal representation learning instead studies recovery of latent causal variables or graphs from indirect observations\([Locatello et al\., 2019](https://arxiv.org/html/2609.30650#bib.bib15);[Schölkopf et al\., 2021](https://arxiv.org/html/2609.30650#bib.bib22);[Gamella et al\., 2025](https://arxiv.org/html/2609.30650#bib.bib7);[Varici et al\., 2025](https://arxiv.org/html/2609.30650#bib.bib29)\)\. Intervention extrapolation uses an identifiable representation to predict the effect of unseen actions on an outcome\([Saengkyongam et al\., 2024](https://arxiv.org/html/2609.30650#bib.bib21)\)\. Causal retention fixes the downstream answer map and asks whether a completed learning process preserved it\. It neither requires recovery of the latent SCM nor reduces the answer to one designated outcome\.

Reinforcement learning, world models, and language agents optimize return, transition prediction, or action quality\([Sutton and Barto, 2018](https://arxiv.org/html/2609.30650#bib.bib27);[Ha and Schmidhuber, 2018](https://arxiv.org/html/2609.30650#bib.bib8);[Hafner et al\., 2025](https://arxiv.org/html/2609.30650#bib.bib9);[Hansen et al\., 2024](https://arxiv.org/html/2609.30650#bib.bib10);[Yao et al\., 2023](https://arxiv.org/html/2609.30650#bib.bib32)\)\. Causal reinforcement learning and data\-fusion methods improve decisions under interventions\([Bareinboim and Pearl, 2016](https://arxiv.org/html/2609.30650#bib.bib2);[Zhang et al\., 2021](https://arxiv.org/html/2609.30650#bib.bib33)\), while ranking\-style dynamics losses such as Hybrid2supervise which intervention should be preferred\([Zou et al\., 2024](https://arxiv.org/html/2609.30650#bib.bib35)\)\. These objectives can be useful and exactly optimized without requiring the frozen state to expose every coordinate of the mechanism needed by a later probe\. The distinction studied here is therefore between an objective’s behavioral sufficiency and a learned state’s causal sufficiency\.

Concurrent work on Explicit Symbolic Behavioral Models uses executable “mechanism memory,” adaptive questions, and active world\-model branches to train editable symbolic policies\([Shindo et al\., 2026](https://arxiv.org/html/2609.30650#bib.bib25)\)\. There, mechanism memory predicts symbolic events and rewards and is optimized jointly with policy\-specific questions\. Here, the probe map and its direct\-target semantics are fixed independently of the learner, scoring occurs after the state is frozen, and the main object is the smallest probe error permitted by a given learning interface\. The two formulations therefore use similar language for different statistical questions\.

One realization of causal retention is a frozen, queryable mechanism memory with entries of the form

θ=\(a,c,y,v,δ\),\\theta=\(a,c,y,v,\\delta\),whereaais an action,cca context predicate,yythe direct target,vvthe effect value, andδ\\deltathe delay\. After learning, the memory is frozen and queried on held\-out mechanism probes\. Equal reward laws, rankings, graphs, or action choices may still induce different probe answers\.

For finite structural causal model \(SCM\) classes, the optimal population probe error is an explicit Bayes risk\. It is zero exactly when the interface quotient refines the probe quotient, is monotone under refinement, and lower\-bounds every decoder of a post\-processed state\. The same construction with a metric loss covers continuous intervention responses\. A posterior\-coverage identity characterizes restricted target\-domain retesting, and an exact edit decomposition characterizes selective adaptation\. Standard concentration and testing inequalities are used only for finite\-sample rates\.

Causal Core is the corresponding evidence\-gated memory layer, with target admission, readout filtering, temporal credit, hidden\-context setup, diagnostic localization, and entrywise updating\. Experiments compare policy, ranking, conditional\-discovery, tuple, latent\-world\-model, neural\-probe, transfer, language\-model proposal, readout\-unsafe, and ungated baselines\. Continuous responses are scored over held\-out state, action\-direction, and horizon distributions\. An official pretrained TD\-MPC2 checkpoint supplies a public implicit\-world\-model comparison\. Frozen Qwen hidden states are evaluated with linear and nonlinear decoders under externally fixed probe semantics\.

Figure 1:Interface collapse and causal retention\. A task interface can merge probe\-distinct SCMs\. A refined interface retains the answer map and supports selective adaptation\.
## 2Problem Formulation

Let an interactive structural causal model \(SCM\) be

Xt\+1=F⁡\(Xt,Ht,At,εt\+1\),Ot=ψ⁡\(Xt,Ht,ηt\),X\_\{t\+1\}=F\(X\_\{t\},H\_\{t\},A\_\{t\},\\varepsilon\_\{t\+1\}\),\\qquad O\_\{t\}=\\psi\(X\_\{t\},H\_\{t\},\\eta\_\{t\}\),with physical stateXt∈𝒳X\_\{t\}\\in\\mathcal\{X\}, optional latent contextHt∈ℋH\_\{t\}\\in\\mathcal\{H\}, actionAt∈𝒜A\_\{t\}\\in\\mathcal\{A\}, and observationOtO\_\{t\}\. WriteΞt=\(Xt,Ht\)∈𝒳×ℋ\\Xi\_\{t\}=\(X\_\{t\},H\_\{t\}\)\\in\\mathcal\{X\}\\\!\\times\\\!\\mathcal\{H\}for the full evaluator state; the learner need not observeHtH\_\{t\}\. A protocolΠ\\Pimaps histories to actions and induces a trajectory lawPMΠP\_\{M\}^\{\\Pi\}for each SCMMM\.

For an action intervention at timett, writePMd​o​\(At=a\),ΠP\_\{M\}^\{do\(A\_\{t\}=a\),\\Pi\}for the law obtained by replacing the protocol action withaaatttand leaving the remaining rollout policy fixed\. A context predicate is a mapc:𝒳×ℋ→\{0,1\}c:\\mathcal\{X\}\\times\\mathcal\{H\}\\to\\\{0,1\\\}\. From the equations ofMM, the evaluator fixes a structural index

σM​\(ξ,a\)∈\(𝒞×𝒴×\{0,…,D\}\)∪\{⊥\}\.\\sigma\_\{M\}\(\\xi,a\)\\in\\bigl\(\\mathcal\{C\}\\times\\mathcal\{Y\}\\times\\\{0,\\ldots,D\\\}\\bigr\)\\cup\\\{\\bot\\\}\.When non\-null,σM​\(ξ,a\)=\(c,y,δ\)\\sigma\_\{M\}\(\\xi,a\)=\(c,y,\\delta\)records the active context class, the variable directly acted on by the mechanism, and its first structural response delay\. A deterministic descendant or sensor readout is not a direct target\. In finite worldsσM\\sigma\_\{M\}is read from the structural assignment; in continuous simulators it is defined by the paired\-pulse functional in Appendix[B\.4](https://arxiv.org/html/2609.30650#A2.SS4)\.

Forz=\(ξ,a\)z=\(\\xi,a\)andr=\(c,y,v,δ\)r=\(c,y,v,\\delta\), define

pM\(r∣z\):=𝟏\{σM\(ξ,a\)=\(c,y,δ\)\}ℙMd​o​\(At=a\),Π\{Xt\+δy=v∣Ξt=ξ\}\.p\_\{M\}\(r\\mid z\):=\\mathbf\{1\}\\\{\\sigma\_\{M\}\(\\xi,a\)=\(c,y,\\delta\)\\\}\\mathbb\{P\}\_\{M\}^\{do\(A\_\{t\}=a\),\\Pi\}\\\{X\_\{t\+\\delta\}^\{y\}=v\\mid\\Xi\_\{t\}=\\xi\\\}\.Letℛ\\mathcal\{R\}be the finite set of candidate quadruples\(c,y,v,δ\)\(c,y,v,\\delta\)\. IfσM\(ξ,a\)=⊥\\sigma\_\{M\}\(\\xi,a\)=\\bot, setGM​\(z\)=∅G\_\{M\}\(z\)=\\emptyset; otherwise define

GM​\(z\)=arg​maxr∈ℛ⁡pM​\(r∣z\),G\_\{M\}\(z\)=\\operatorname\*\{arg\\,max\}\_\{r\\in\\mathcal\{R\}\}p\_\{M\}\(r\\mid z\),restricted to a probe class on which the maximizer is unique\. In deterministic SCMs a tupleθ=\(a,c,y,v,δ\)\\theta=\(a,c,y,v,\\delta\)is correct exactly whenpM​\(\(c,y,v,δ\)∣z\)=1p\_\{M\}\(\(c,y,v,\\delta\)\\mid z\)=1\. In stochastic SCMs uniqueness is enforced by a marginκ\>0\\kappa\>0:

pM​\(r⋆∣z\)−maxr≠r⋆⁡pM​\(r∣z\)≥κ\.p\_\{M\}\(r^\{\\star\}\\mid z\)\-\\max\_\{r\\neq r^\{\\star\}\}p\_\{M\}\(r\\mid z\)\\geq\\kappa\.
### 2\.1Mechanism Space

Let

Θ=𝒜×𝒞×𝒴×𝒱×\{0,…,D\}\.\\Theta=\\mathcal\{A\}\\times\\mathcal\{C\}\\times\\mathcal\{Y\}\\times\\mathcal\{V\}\\times\\\{0,\\ldots,D\\\}\.A mechanism memory is a finite setμ⊂Θ\\mu\\subset\\Theta\. For a queryz=\(ξ,a\)z=\(\\xi,a\), define

Tμ\(ξ,a\)=\{\(c,y,v,δ\):\(a,c,y,v,δ\)∈μ,c\(ξ\)=1\}\.T\_\{\\mu\}\(\\xi,a\)=\\\{\(c,y,v,\\delta\):\(a,c,y,v,\\delta\)\\in\\mu,\\ c\(\\xi\)=1\\\}\.Two answers are equivalent, writtenTμ​\(ξ,a\)≡Tμ′​\(ξ,a\)T\_\{\\mu\}\(\\xi,a\)\\equiv T\_\{\\mu^\{\\prime\}\}\(\\xi,a\), when they agree on the active context, direct target, value, and delay classes fixed by the evaluator\. Exact tuple evaluation uses ordinary set equality\.

The constructive and empirical results use a representable mechanism schema: for every evaluated SCMMM, there is a finite ground\-truth memoryμM⋆⊂Θ\\mu\_\{M\}^\{\\star\}\\subset\\Thetasuch thatTμM⋆​\(z\)=GM​\(z\)T\_\{\\mu\_\{M\}^\{\\star\}\}\(z\)=G\_\{M\}\(z\)forQQ\-almost every probezz\. The lower bounds are stated directly forGMG\_\{M\}and do not require such a representation\.

### 2\.2Mechanism Risk

For probe distributionQQand ground\-truth memoryμ⋆=μM⋆\\mu^\{\\star\}=\\mu\_\{M\}^\{\\star\}, the exact mechanism risk is

RQ\(μ^,μ⋆\)=ℙ\(ξ,a\)∼Q\[Tμ^\(ξ,a\)≢Tμ⋆\(ξ,a\)\]\.R\_\{Q\}\(\\widehat\{\\mu\},\\mu^\{\\star\}\)=\\mathbb\{P\}\_\{\(\\xi,a\)\\sim Q\}\\\!\\left\[T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)\\not\\equiv T\_\{\\mu^\{\\star\}\}\(\\xi,a\)\\right\]\.Graph error, reward regret, top\-action accuracy, and prediction error are not substitutes forRQR\_\{Q\}: each marginalizes at least one coordinate of\(a,c,y,v,δ\)\(a,c,y,v,\\delta\)\.

Coordinate risks localize the source of an error\. Let𝒦=\{ctx,tar,val,del\}\\mathcal\{K\}=\\\{\\mathrm\{ctx\},\\mathrm\{tar\},\\mathrm\{val\},\\mathrm\{del\}\\\}and letπk​Tμ​\(ξ,a\)\\pi\_\{k\}T\_\{\\mu\}\(\\xi,a\)be the projection of the answer set onto coordinatekk\. Define

RQk\(μ^,μ⋆\)=ℙQ\[πkTμ^\(ξ,a\)≠πkTμ⋆\(ξ,a\)\]\.R\_\{Q\}^\{k\}\(\\widehat\{\\mu\},\\mu^\{\\star\}\)=\\mathbb\{P\}\_\{Q\}\\\!\\left\[\\pi\_\{k\}T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)\\neq\\pi\_\{k\}T\_\{\\mu^\{\\star\}\}\(\\xi,a\)\\right\]\.When equivalence means equality of all four projections,

maxk∈𝒦⁡RQk≤RQ≤∑k∈𝒦RQk\.\\max\_\{k\\in\\mathcal\{K\}\}R\_\{Q\}^\{k\}\\leq R\_\{Q\}\\leq\\sum\_\{k\\in\\mathcal\{K\}\}R\_\{Q\}^\{k\}\.Thus the exact risk is the main target, while coordinate risks diagnose which part of the mechanism failed\. Two additional operational risks are

Rfp\\displaystyle R\_\{\\mathrm\{fp\}\}=ℙQ​\[Tμ^​\(ξ,a\)≠∅,Tμ⋆​\(ξ,a\)=∅\],\\displaystyle=\\mathbb\{P\}\_\{Q\}\[T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)\\neq\\emptyset,\\ T\_\{\\mu^\{\\star\}\}\(\\xi,a\)=\\emptyset\],Rfn\\displaystyle R\_\{\\mathrm\{fn\}\}=ℙQ​\[Tμ^​\(ξ,a\)=∅,Tμ⋆​\(ξ,a\)≠∅\]\.\\displaystyle=\\mathbb\{P\}\_\{Q\}\[T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)=\\emptyset,\\ T\_\{\\mu^\{\\star\}\}\(\\xi,a\)\\neq\\emptyset\]\.They distinguish spurious mechanisms from missed mechanisms\.

The fixed\-memory protocol is essential\. Exploration may use rewards, observations, auxiliary descriptions, or rankings; evaluation fixesμ^\\widehat\{\\mu\}and queries only the answer mapTμ^T\_\{\\widehat\{\\mu\}\}\. Thus a behaviorally competent learner receives no credit unless its stored mechanism answers the probe: writeT^i=Tμ^​\(ξi,ai\)\\widehat\{T\}\_\{i\}=T\_\{\\widehat\{\\mu\}\}\(\\xi\_\{i\},a\_\{i\}\)andTi⋆=Tμ⋆​\(ξi,ai\)T\_\{i\}^\{\\star\}=T\_\{\\mu^\{\\star\}\}\(\\xi\_\{i\},a\_\{i\}\)\. Then

μ^↦\{Tμ^​\(ξi,ai\)\}i=1m,\\widehat\{\\mu\}\\mapsto\\\{T\_\{\\widehat\{\\mu\}\}\(\\xi\_\{i\},a\_\{i\}\)\\\}\_\{i=1\}^\{m\},score=1−1m∑i=1mℓi,ℓi=𝟏\{T^i≢Ti⋆\}\.\\mathrm\{score\}=1\-\\frac\{1\}\{m\}\\sum\_\{i=1\}^\{m\}\\ell\_\{i\},\\quad\\ell\_\{i\}=\\mathbf\{1\}\\\{\\widehat\{T\}\_\{i\}\\not\\equiv T\_\{i\}^\{\\star\}\\\}\.

### 2\.3Selective Adaptation

Letμs=\{θjs\}j=1n\\mu^\{s\}=\\\{\\theta\_\{j\}^\{s\}\\\}\_\{j=1\}^\{n\}be a source memory\. A schema mapϕ=\(ϕA,ϕC,ϕY\)\\phi=\(\\phi\_\{A\},\\phi\_\{C\},\\phi\_\{Y\}\)induces mapped entries

ϕ⁡\(θjs\)=\(ϕA​\(ajs\),ϕC​\(cjs\),ϕY​\(yjs\),vjs,δjs\)\.\\phi\(\\theta\_\{j\}^\{s\}\)=\(\\phi\_\{A\}\(a\_\{j\}^\{s\}\),\\phi\_\{C\}\(c\_\{j\}^\{s\}\),\\phi\_\{Y\}\(y\_\{j\}^\{s\}\),v\_\{j\}^\{s\},\\delta\_\{j\}^\{s\}\)\.The target memory isμt=\{θjt\}j=1n\\mu^\{t\}=\\\{\\theta\_\{j\}^\{t\}\\\}\_\{j=1\}^\{n\}\. Define the stable and shifted index sets

S=\{j:θjt=ϕ⁡\(θjs\)\},B=\[n\]∖S\.S=\\\{j:\\theta\_\{j\}^\{t\}=\\phi\(\\theta\_\{j\}^\{s\}\)\\\},\\qquad B=\[n\]\\setminus S\.For an adapted memoryμ^t\\widehat\{\\mu\}^\{t\},

Ladapt\\displaystyle L\_\{\\mathrm\{adapt\}\}=1n∑j=1n𝟏\{θ^jt≠θjt\},\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\\neq\\theta\_\{j\}^\{t\}\\\},Lstab\\displaystyle L\_\{\\mathrm\{stab\}\}=1\|S\|∑j∈S𝟏\{θ^jt≠θjt\},\\displaystyle=\\frac\{1\}\{\|S\|\}\\sum\_\{j\\in S\}\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\\neq\\theta\_\{j\}^\{t\}\\\},andLshiftL\_\{\\mathrm\{shift\}\}is defined analogously overBB; an empty\-set loss is defined as zero\. Selective adaptation requiresLstab=Lshift=0L\_\{\\mathrm\{stab\}\}=L\_\{\\mathrm\{shift\}\}=0under a small target budget\. For exploration and sampling randomness conditional on a fixed source\-target pair, define

ℛadapt=𝔼\[Ladapt∣μs,μt\],ℛstab=𝔼\[Lstab∣μs,μt\],\\mathcal\{R\}\_\{\\mathrm\{adapt\}\}=\\mathbb\{E\}\[L\_\{\\mathrm\{adapt\}\}\\mid\\mu^\{s\},\\mu^\{t\}\],\\qquad\\mathcal\{R\}\_\{\\mathrm\{stab\}\}=\\mathbb\{E\}\[L\_\{\\mathrm\{stab\}\}\\mid\\mu^\{s\},\\mu^\{t\}\],ℛshift=𝔼\[Lshift∣μs,μt\]\.\\mathcal\{R\}\_\{\\mathrm\{shift\}\}=\\mathbb\{E\}\[L\_\{\\mathrm\{shift\}\}\\mid\\mu^\{s\},\\mu^\{t\}\]\.The expectation decomposes as

ℛadapt=\|S\|n​ℛstab\+\|B\|n​ℛshift\.\\mathcal\{R\}\_\{\\mathrm\{adapt\}\}=\\frac\{\|S\|\}\{n\}\\mathcal\{R\}\_\{\\mathrm\{stab\}\}\+\\frac\{\|B\|\}\{n\}\\mathcal\{R\}\_\{\\mathrm\{shift\}\}\.The operational failure probability is

ℛop=ℙ⁡\(Lstab\>0∨Lshift\>0\),\\mathcal\{R\}\_\{\\mathrm\{op\}\}=\\mathbb\{P\}\(L\_\{\\mathrm\{stab\}\}\>0\\ \\vee\\ L\_\{\\mathrm\{shift\}\}\>0\),because a single missed shifted mechanism can be invisible in average loss when\|B\|≪n\|B\|\\ll n\.

Two limiting cases illustrate the requirement\. Pure transfer keepsμ^t=ϕ⁡\(μs\)\\widehat\{\\mu\}^\{t\}=\\phi\(\\mu^\{s\}\)and therefore fails onBB\. Broad re\-learning may fixBBbut can corruptSS\. The desired operation is an entrywise map

θ^jt=\{ϕ⁡\(θjs\),j∈S,θjt,j∈B,\\widehat\{\\theta\}\_\{j\}^\{t\}=\\begin\{cases\}\\phi\(\\theta\_\{j\}^\{s\}\),&j\\in S,\\\\ \\theta\_\{j\}^\{t\},&j\\in B,\\end\{cases\}with evidence deciding membership inBB\.

### 2\.4Standing Conditions

Each lower bound uses only the assumptions in its statement\. The constructive guarantees for Causal Core use the following grouped conditions\.

1. \(C1\)The source and target memories containnnaligned keys\. The evaluator’s probe distributionQQis fixed independently of target\-domain samples\. The schema mapϕ\\phiis supplied or estimated from a disjoint alignment sample; its error is represented by the eventFϕF\_\{\\phi\}rather than assumed away\. Target\-coordinate values and visible context predicates in the memory schema are measurable from the learner’s observation history\. Latent predicates are never supplied to the learner and are handled only through assigned setup and control regimes\.
2. \(C2\)A gate is estimated from independent, bounded outcomes collected under paired action and reference interventions at the same pre\-intervention context\. Equivalently, an observational implementation must satisfy consistency, conditional exchangeability, and positivity and use the corresponding propensity correction\. An adaptive randomized policy may use the logged inverse\-propensity contrast in Lemma[11](https://arxiv.org/html/2609.30650#Thmtheorem11)\. The finite agents execute interventions and are scored descriptively; the continuous and simulator protocols use explicit paired pulses\.
3. \(C3\)Every finite decision used by a gate is separated from its threshold\. Writeγg\>0\\gamma\_\{g\}\>0for target, context, readout, and temporal margins,Δ=p−q\>0\\Delta=p\-q\>0for diagnostic separation, andηj\>0\\eta\_\{j\}\>0for the excess loss of a wrong local\-update candidate\. The correct local candidate belongs to the finite class𝒩j\\mathcal\{N\}\_\{j\}\.
4. \(C4\)Metric\-valued probe answers lie in a compact space and use a bounded, continuous distortion\. The probe distribution is fixed before fitting\. A discrete implementation may additionally fix a measurable partition; its paired contrast routine fails with probability at mostβM\\beta\_\{M\}\.

Conditions[C2](https://arxiv.org/html/2609.30650#S2.I1.i2)–[C3](https://arxiv.org/html/2609.30650#S2.I1.i3)are sufficient, not necessary\. When they fail, the quotient and decision lower bounds still apply, but the finite\-sample upper bounds need not\.

### 2\.5Population Interfaces

An interfaceℐ\\mathcal\{I\}induces a population descriptorτℐ,Π​\(M\)=PMℐ,Π\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\)=P\_\{M\}^\{\\mathcal\{I\},\\Pi\}: the data law generated by the protocol\. Examples include reward trajectories, observations, rankings, prediction errors, diagnostic signals, or mechanism probes\. For probe class𝒵\\mathcal\{Z\}with answer mapGM​\(z\)G\_\{M\}\(z\), write

M0∼ℐ,ΠM1\\displaystyle M\_\{0\}\\sim\_\{\\mathcal\{I\},\\Pi\}M\_\{1\}⟺PM0ℐ,Π=PM1ℐ,Π,\\displaystyle\\Longleftrightarrow P\_\{M\_\{0\}\}^\{\\mathcal\{I\},\\Pi\}=P\_\{M\_\{1\}\}^\{\\mathcal\{I\},\\Pi\},M0∼𝒵M1\\displaystyle M\_\{0\}\\sim\_\{\\mathcal\{Z\}\}M\_\{1\}⟺GM0\(z\)=GM1\(z\)∀z∈𝒵\.\\displaystyle\\Longleftrightarrow G\_\{M\_\{0\}\}\(z\)=G\_\{M\_\{1\}\}\(z\)\\quad\\forall z\\in\\mathcal\{Z\}\.The relevant condition is whether∼ℐ,Π\\sim\_\{\\mathcal\{I\},\\Pi\}refines∼𝒵\\sim\_\{\\mathcal\{Z\}\}for mechanism probes\. If not, the interface is too coarse\. For the finite analysis, fix full\-support distributionsν\\nuonℳ\\mathcal\{M\}andQQon𝒵\\mathcal\{Z\}\. Let

W∼ν,Z∼Q,𝖳ℐ=τℐ,Π​\(W\),U=GW​\(Z\),W\\sim\\nu,\\qquad Z\\sim Q,\\qquad\\mathsf\{T\}\_\{\\mathcal\{I\}\}=\\tau\_\{\\mathcal\{I\},\\Pi\}\(W\),\\qquad U=G\_\{W\}\(Z\),whereZZis independent ofWW\. The*causal\-retention risk*of the interface is the probe Bayes risk

𝔅ν,Q​\(ℐ,Π\)=𝔼⁡\[1−maxu∈𝒰⁡ℙ⁡\(U=u∣𝖳ℐ,Z\)\]\.\\mathfrak\{B\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)=\\mathbb\{E\}\\\!\\left\[1\-\\max\_\{u\\in\\mathcal\{U\}\}\\mathbb\{P\}\(U=u\\mid\\mathsf\{T\}\_\{\\mathcal\{I\}\},Z\)\\right\]\.\(1\)It will equal the smallest error of any decoder that sees the complete population interface\. ForU=\(U1,…,Ud\)U=\(U\_\{1\},\\ldots,U\_\{d\}\), define𝔅j\\mathfrak\{B\}\_\{j\}by replacingUUwithUjU\_\{j\}in Eq\. \([1](https://arxiv.org/html/2609.30650#S2.E1)\)\. These risks depend on the interface and the fixed probe map rather than on a chosen decoder\. The population minimax interface risk is

ℜ\(ℐ,𝒵\)=infG^supM∈ℳ𝟏\{G^\(τℐ,Π\(M\)\)≠GM\}\.\\mathfrak\{R\}\(\\mathcal\{I\},\\mathcal\{Z\}\)=\\inf\_\{\\widehat\{G\}\}\\sup\_\{M\\in\\mathcal\{M\}\}\\mathbf\{1\}\\\{\\widehat\{G\}\(\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\)\)\\neq G\_\{M\}\\\}\.Theorem[2](https://arxiv.org/html/2609.30650#Thmtheorem2)gives the exact operational condition:ℜ⁡\(ℐ,𝒵\)=0\\mathfrak\{R\}\(\\mathcal\{I\},\\mathcal\{Z\}\)=0only when the interface separates all probe\-distinct SCMs\.

For a compact metric answer space\(𝒰,ρ\)\(\\mathcal\{U\},\\rho\), define the expected and worst\-case metric risks

𝔅ν,Qρ​\(ℐ,Π\)\\displaystyle\\mathfrak\{B\}^\{\\rho\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)=infg𝔼​ρ​\(U,g⁡\(𝖳ℐ,Z\)\),\\displaystyle=\\inf\_\{g\}\\mathbb\{E\}\\rho\\bigl\(U,g\(\\mathsf\{T\}\_\{\\mathcal\{I\}\},Z\)\\bigr\),\(2\)ℜρ​\(ℐ,𝒵\)\\displaystyle\\mathfrak\{R\}^\{\\rho\}\(\\mathcal\{I\},\\mathcal\{Z\}\)=infgsupM∈ℳ,z∈𝒵ρ⁡\(GM​\(z\),g⁡\(τℐ,Π​\(M\),z\)\)\.\\displaystyle=\\inf\_\{g\}\\sup\_\{M\\in\\mathcal\{M\},\\,z\\in\\mathcal\{Z\}\}\\rho\\bigl\(G\_\{M\}\(z\),g\(\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\),z\)\\bigr\)\.For Hamming loss,𝔅ρ=𝔅\\mathfrak\{B\}^\{\\rho\}=\\mathfrak\{B\}\. The metric version therefore changes the loss geometry, not the semantics of retention\.

### 2\.6Exact Causal Retention

A possibly random learned stateSShas exact causal retention for𝒵\\mathcal\{Z\}onℳ\\mathcal\{M\}if there exists a single decoderddsuch that

ℙM\{d\(S,z\)=GM\(z\)for everyz∈𝒵\}=1∀M∈ℳ\.\\mathbb\{P\}\_\{M\}\\\!\\left\\\{d\(S,z\)=G\_\{M\}\(z\)\\text\{ for every \}z\\in\\mathcal\{Z\}\\right\\\}=1\\qquad\\forall M\\in\\mathcal\{M\}\.The probe class and its answer semantics are fixed independently of the learned state\. This makes causal retention independent of the internal implementation: a memory table, recurrent state, or external module qualifies only to the extent that it supports the same decoder after learning has stopped\.

For ranking interfaces, the observed object is a distribution over ordered pairs

𝒟rank=\{\(i,j\):UM​\(i\)\>UM​\(j\)\},\\mathcal\{D\}\_\{\\mathrm\{rank\}\}=\\\{\(i,j\):U\_\{M\}\(i\)\>U\_\{M\}\(j\)\\\},whereUM​\(i\)U\_\{M\}\(i\)is the intervention utility\. For reward interfaces it is a trajectory marginal of∑trt\\sum\_\{t\}r\_\{t\}\. For graph interfaces it is an edge set\. For diagnostic adaptation it is instead an indexed surprise channel\{Xj,r\}j,r\\\{X\_\{j,r\}\\\}\_\{j,r\}tied to transferred memory entries\. If two SCMs have the same interface law but different probe answers, no estimator using that interface can guarantee zero mechanism risk\.

## 3Causal Core

Causal Core maintains mechanism memory independently of the policy optimizer\. Its state consists of a candidate setCtC\_\{t\}, readout setℛt\\mathcal\{R\}\_\{t\}, temporal bufferBtB\_\{t\}, diagnostic surprise vectorStS\_\{t\}, and memoryμt\\mu\_\{t\}\. It writes\(a,c,y,v,δ\)\(a,c,y,v,\\delta\)only when intervention evidence supports every coordinate\. For a candidateθ=\(a,c,y,v,δ\)\\theta=\(a,c,y,v,\\delta\), define its memory keyκ⁡\(θ\)=\(a,c,y,δ\)\\kappa\(\\theta\)=\(a,c,y,\\delta\)\. Leta0a^\{0\}be the reference intervention foraa\. For a learner\-visible context predicatecc, let

𝒟t1​\(a,c,δ\)\\displaystyle\\mathcal\{D\}\_\{t\}^\{1\}\(a,c,\\delta\)=\{s:s\+δ≤t,As=a,c\(Ξs\)=1\},\\displaystyle=\\\{s:s\+\\delta\\leq t,\\ A\_\{s\}=a,\\ c\(\\Xi\_\{s\}\)=1\\\},𝒟t0​\(a,c,δ\)\\displaystyle\\mathcal\{D\}\_\{t\}^\{0\}\(a,c,\\delta\)=\{s:s\+δ≤t,As=a0,c\(Ξs\)=1\}\.\\displaystyle=\\\{s:s\+\\delta\\leq t,\\ A\_\{s\}=a^\{0\},\\ c\(\\Xi\_\{s\}\)=1\\\}\.Writentu=\|𝒟tu\|n\_\{t\}^\{u\}=\|\\mathcal\{D\}\_\{t\}^\{u\}\|andYsθ=𝟏\{Xs\+δy=v\}Y\_\{s\}^\{\\theta\}=\\mathbf\{1\}\\\{X\_\{s\+\\delta\}^\{y\}=v\\\}\. A candidate does not mature untilnt0,nt1\>0n\_\{t\}^\{0\},n\_\{t\}^\{1\}\>0\. Its paired empirical frequencies are

p^tu​\(θ\)=1ntu​∑s∈𝒟tu​\(a,c,δ\)Ysθ,u∈\{0,1\}\.\\widehat\{p\}\_\{t\}^\{u\}\(\\theta\)=\\frac\{1\}\{n\_\{t\}^\{u\}\}\\sum\_\{s\\in\\mathcal\{D\}\_\{t\}^\{u\}\(a,c,\\delta\)\}Y\_\{s\}^\{\\theta\},\\qquad u\\in\\\{0,1\\\}\.\(3\)Under Condition[C2](https://arxiv.org/html/2609.30650#S2.I1.i2), their difference estimates the interventional contrast betweend​o​\(a\)do\(a\)andd​o​\(a0\)do\(a^\{0\}\)rather than an unadjusted action association\. The target gate is

Δ^t​\(θ\)\\displaystyle\\widehat\{\\Delta\}\_\{t\}\(\\theta\)=p^t1​\(θ\)−p^t0​\(θ\),\\displaystyle=\\widehat\{p\}\_\{t\}^\{1\}\(\\theta\)\-\\widehat\{p\}\_\{t\}^\{0\}\(\\theta\),\(4\)Γttar​\(θ\)\\displaystyle\\Gamma\_\{t\}^\{\\mathrm\{tar\}\}\(\\theta\)=𝟏​\{nt0​nt1\>0,Δ^t​\(θ\)≥λt\}\.\\displaystyle=\\mathbf\{1\}\\\{n\_\{t\}^\{0\}n\_\{t\}^\{1\}\>0,\\ \\widehat\{\\Delta\}\_\{t\}\(\\theta\)\\geq\\lambda\_\{t\}\\\}\.The full write gate is

Gt​\(θ\)\\displaystyle G\_\{t\}\(\\theta\)=∏g∈\{tar,time,readout,ctx\}Γtg​\(θ\),\\displaystyle=\\prod\_\{g\\in\\\{\\mathrm\{tar\},\\mathrm\{time\},\\mathrm\{readout\},\\mathrm\{ctx\}\\\}\}\\Gamma\_\{t\}^\{g\}\(\\theta\),\(5\)Ctadm\\displaystyle C\_\{t\}^\{\\mathrm\{adm\}\}=\{θ∈Ct:Gt​\(θ\)=1\}\.\\displaystyle=\\\{\\theta\\in C\_\{t\}:G\_\{t\}\(\\theta\)=1\\\}\.The memory transition is the key\-preserving operator

𝖶⁡\(μ,θ\)=\{θ′∈μ:κ⁡\(θ′\)≠κ⁡\(θ\)\}∪\{θ\}\.\\mathsf\{W\}\(\\mu,\\theta\)=\\\{\\theta^\{\\prime\}\\in\\mu:\\kappa\(\\theta^\{\\prime\}\)\\neq\\kappa\(\\theta\)\\\}\\cup\\\{\\theta\\\}\.\(6\)If several candidates mature at the same time, they are applied in decreasingΔ^t\\widehat\{\\Delta\}\_\{t\}with a fixed lexicographic tie\-breaker\. Thus the memory contains at most one value for each action\-context\-target\-delay key\. Admission is determined by the four gates; planner design does not enter the definition of mechanism memory\. For a latent context, the sets in Eq\. \([3](https://arxiv.org/html/2609.30650#S3.E3)\) are instead indexed by the learner\-known assignment to the setup or control regime; membership never usesc⁡\(Ξs\)c\(\\Xi\_\{s\}\)itself\. The resulting context gate is defined in Section[3\.4](https://arxiv.org/html/2609.30650#S3.SS4)\.

### 3\.1Readout Filtering

IfYYandRRare synchronized, the two SCMs

a→Y,R:=Yanda→R,Y:=Ra\\to Y,\\ R:=Y\\qquad\\text\{and\}\\qquad a\\to R,\\ Y:=Rcan produce the same observations\. Causal Core therefore excludes readout candidates from direct\-target writing unless an intervention separates them\. Formally, for a candidate targetuu, define the direct\-target set

𝒴tdir=\{u∈𝒴:u∉ℛt​or​Sept​\(u\)=1\},\\mathcal\{Y\}\_\{t\}^\{\\mathrm\{dir\}\}=\\\{u\\in\\mathcal\{Y\}:u\\notin\\mathcal\{R\}\_\{t\}\\ \\text\{or\}\\ \\mathrm\{Sep\}\_\{t\}\(u\)=1\\\},Sept​\(u\)=𝟏​\{∃b:P⁡\(Xt\+1u∣d​o​\(b\)\)≠P⁡\(Xt\+1r⁡\(u\)∣d​o​\(b\)\)\},\\mathrm\{Sep\}\_\{t\}\(u\)=\\mathbf\{1\}\\\{\\exists b:P\(X\_\{t\+1\}^\{u\}\\mid do\(b\)\)\\neq P\(X\_\{t\+1\}^\{r\(u\)\}\\mid do\(b\)\)\\\},wherer⁡\(u\)r\(u\)denotes the synchronized readout paired withuuwhen known\. For formula readouts, the finite implementation uses, for each candidate coordinateuu, a classℱu\\mathcal\{F\}\_\{u\}of formulas whose inputs excludeuu\. It computes

A^t\(f,u\)=1mt∑r=1mt𝟏\{f\(Xr\)=Xru\},u∈ℛt⟺maxf∈ℱuA^t\(f,u\)≥cro\.\\widehat\{A\}\_\{t\}\(f,u\)=\\frac\{1\}\{m\_\{t\}\}\\sum\_\{r=1\}^\{m\_\{t\}\}\\mathbf\{1\}\\\{f\(X\_\{r\}\)=X\_\{r\}^\{u\}\\\},\\qquad u\\in\\mathcal\{R\}\_\{t\}\\Longleftrightarrow\\max\_\{f\\in\\mathcal\{F\}\_\{u\}\}\\widehat\{A\}\_\{t\}\(f,u\)\\geq c\_\{\\mathrm\{ro\}\}\.Proposition[12](https://arxiv.org/html/2609.30650#Thmtheorem12)controls both directions of this decision under a finite\-class margin\. A variable classified as a readout can re\-enter the direct\-target set only through a separating intervention\. The readout gate is

Γtreadout\(a,c,y,v,δ\)=𝟏\{y∈𝒴tdir\}\.\\Gamma\_\{t\}^\{\\mathrm\{readout\}\}\(a,c,y,v,\\delta\)=\\mathbf\{1\}\\\{y\\in\\mathcal\{Y\}\_\{t\}^\{\\mathrm\{dir\}\}\\\}\.

### 3\.2Temporal Credit

For delayed effects, Causal Core tests stability of the lag:

δ^​\(a,c,y,v\)∈arg⁡max0≤δ≤D​Δ^t​\(a,c,y,v,δ\)\.\\widehat\{\\delta\}\(a,c,y,v\)\\in\\arg\\max\_\{0\\leq\\delta\\leq D\}\\widehat\{\\Delta\}\_\{t\}\(a,c,y,v,\\delta\)\.An entry is written only if the selected lag beats competing recent actions by a margin\. The analyzed margin form is

Δ^t​\(a,c,y,v,δ\)−max\(a′,v′,δ′\)≠\(a,v,δ\)⁡Δ^t​\(a′,c,y,v′,δ′\)≥γtime,\\widehat\{\\Delta\}\_\{t\}\(a,c,y,v,\\delta\)\-\\max\_\{\(a^\{\\prime\},v^\{\\prime\},\\delta^\{\\prime\}\)\\neq\(a,v,\\delta\)\}\\widehat\{\\Delta\}\_\{t\}\(a^\{\\prime\},c,y,v^\{\\prime\},\\delta^\{\\prime\}\)\\geq\\gamma\_\{\\mathrm\{time\}\},where every contrast uses only matured paired trials from Eq\. \([3](https://arxiv.org/html/2609.30650#S3.E3)\)\. This prevents a waiting action or a post\-effect observation from being written as the cause\.

### 3\.3Diagnostic Adaptation

After transfer, mechanismjjproduces binary surprise samplesXj,r∈\{0,1\}X\_\{j,r\}\\in\\\{0,1\\\}\. Causal Core computes

X¯j​\(m\)=1m​∑r=1mXj,r\\overline\{X\}\_\{j\}\(m\)=\\frac\{1\}\{m\}\\sum\_\{r=1\}^\{m\}X\_\{j,r\}and retests entries above a threshold\. The edit is local:θ^jt\\widehat\{\\theta\}\_\{j\}^\{t\}may change only for selectedjj\. Forτ=\(p\+q\)/2\\tau=\(p\+q\)/2,

B^m=\{j:X¯j​\(m\)≥τ\}\.\\widehat\{B\}\_\{m\}=\\\{j:\\overline\{X\}\_\{j\}\(m\)\\geq\\tau\\\}\.\(7\)θ^jt=\{ϕ⁡\(θjs\),j∉B^m,θ~j,j∈B^m\.\\widehat\{\\theta\}\_\{j\}^\{t\}=\\begin\{cases\}\\phi\(\\theta\_\{j\}^\{s\}\),&j\\notin\\widehat\{B\}\_\{m\},\\\\ \\widetilde\{\\theta\}\_\{j\},&j\\in\\widehat\{B\}\_\{m\}\.\\end\{cases\}\(8\)For a selected entry, the identifying retest gives an independent local sampleDjre=\{\(Zj,r,Uj,r\)\}r=1rjD\_\{j\}^\{\\mathrm\{re\}\}=\\\{\(Z\_\{j,r\},U\_\{j,r\}\)\\\}\_\{r=1\}^\{r\_\{j\}\}, whereUj,rU\_\{j,r\}is the observed response statistic specified by the intervention protocol, and a finite candidate class𝒩j\\mathcal\{N\}\_\{j\}\. Letgθg\_\{\\theta\}be the probe\-answer map predicted by candidateθ\\theta\. Define

L^j​\(θ\)\\displaystyle\\widehat\{L\}\_\{j\}\(\\theta\)=1rj∑r=1rj𝟏\{gθ\(Zj,r\)≠Uj,r\},\\displaystyle=\\frac\{1\}\{r\_\{j\}\}\\sum\_\{r=1\}^\{r\_\{j\}\}\\mathbf\{1\}\\\{g\_\{\\theta\}\(Z\_\{j,r\}\)\\neq U\_\{j,r\}\\\},\(9\)θ~j\\displaystyle\\widetilde\{\\theta\}\_\{j\}∈arg⁡minθ∈𝒩j​L^j​\(θ\)\.\\displaystyle\\in\\arg\\min\_\{\\theta\\in\\mathcal\{N\}\_\{j\}\}\\widehat\{L\}\_\{j\}\(\\theta\)\.The corresponding local edit is

𝖴j​\(μ,θ~j\)=\{θℓ∈μ:ℓ≠j\}∪\{θ~j\}\.\\mathsf\{U\}\_\{j\}\(\\mu,\\widetilde\{\\theta\}\_\{j\}\)=\\\{\\theta\_\{\\ell\}\\in\\mu:\\ell\\neq j\\\}\\cup\\\{\\widetilde\{\\theta\}\_\{j\}\\\}\.\(10\)All entries outsideB^m\\widehat\{B\}\_\{m\}are held fixed\. Under Eq\. \([10](https://arxiv.org/html/2609.30650#S3.E10)\), stable corruption can occur only through the set\-selection eventB^m≠B\\widehat\{B\}\_\{m\}\\neq B\.

### 3\.4Hidden\-Context Setup

A hidden gate is not written from passive frequency alone\. The module actively searches setup actions that increase eligibility, then tests a visible\-separator class𝒢\\mathcal\{G\}\. A hidden label survives only when a controlled setup regime and its matched control have separated success rates and no visible separator explains the pattern\. For candidateii, letEiE\_\{i\}be latent eligibility andZiZ\_\{i\}success\. The learner observes the assigned regime, notEiE\_\{i\}\. It seeks a setup policyπiset\\pi\_\{i\}^\{\\mathrm\{set\}\}and a matched controlπi0\\pi\_\{i\}^\{0\}such that

ℙπiset​\(Ei=1\)≥1−ϵset,ℙπi0​\(Ei=1\)≤ϵ0,\\mathbb\{P\}\_\{\\pi\_\{i\}^\{\\mathrm\{set\}\}\}\(E\_\{i\}=1\)\\geq 1\-\\epsilon\_\{\\mathrm\{set\}\},\\qquad\\mathbb\{P\}\_\{\\pi\_\{i\}^\{0\}\}\(E\_\{i\}=1\)\\leq\\epsilon\_\{0\},𝔼πiset​Zi−𝔼πi0​Zi≥2​γ\.\\mathbb\{E\}\_\{\\pi\_\{i\}^\{\\mathrm\{set\}\}\}Z\_\{i\}\-\\mathbb\{E\}\_\{\\pi\_\{i\}^\{0\}\}Z\_\{i\}\\geq 2\\gamma\.Withmiset,mi0\>0m\_\{i\}^\{\\mathrm\{set\}\},m\_\{i\}^\{0\}\>0outcomes assigned to the two regimes, the context gate uses

d^i=1miset​∑r∈𝒟isetZi,r−1mi0​∑r∈𝒟i0Zi,r,Γictx=𝟏​\{d^i≥γ,g^i=∅\},\\widehat\{d\}\_\{i\}=\\frac\{1\}\{m\_\{i\}^\{\\mathrm\{set\}\}\}\\sum\_\{r\\in\\mathcal\{D\}\_\{i\}^\{\\mathrm\{set\}\}\}Z\_\{i,r\}\-\\frac\{1\}\{m\_\{i\}^\{0\}\}\\sum\_\{r\\in\\mathcal\{D\}\_\{i\}^\{0\}\}Z\_\{i,r\},\\qquad\\Gamma\_\{i\}^\{\\mathrm\{ctx\}\}=\\mathbf\{1\}\\\{\\widehat\{d\}\_\{i\}\\geq\\gamma,\\ \\widehat\{g\}\_\{i\}=\\varnothing\\\},whereg^i\\widehat\{g\}\_\{i\}is a visible separator found in𝒢\\mathcal\{G\}\. The label is retained only when the contrast passes and no visible predicate explains it\. Lemma[10](https://arxiv.org/html/2609.30650#Thmtheorem10)controls the contrast error, while Proposition[13](https://arxiv.org/html/2609.30650#Thmtheorem13)converts repeated setup reachability into action cost when setup completion is observable\. If passive access givesℙ⁡\(Ei=1\)=qi\\mathbb\{P\}\(E\_\{i\}=1\)=q\_\{i\}, then even an oracle that marks eligible trials needsm/qim/q\_\{i\}attempts on average to collectmmeligible outcomes\.

Algorithm 1Causal Core memory update and adaptation1:Initialize

μ\\mu, readout set

ℛ\\mathcal\{R\}, buffer

BB, scores

dj=0d\_\{j\}=0\.

2:foreach interaction step

ttdo

3:Observe

oto\_\{t\}, execute

ata\_\{t\}, and observe the next response\.

4:Update

ℛ\\mathcal\{R\}using readout formulas and separating interventions\.

5:Add candidate effects to

BBwith action, context, target, value, and delay fields\.

6:foreach buffered candidate

\(a,c,y,v,δ\)\(a,c,y,v,\\delta\)whose delay has matureddo

7:if

\(a,c,y,v,δ\)∈Ctadm\(a,c,y,v,\\delta\)\\in C\_\{t\}^\{\\mathrm\{adm\}\}from Eq\. \([5](https://arxiv.org/html/2609.30650#S3.E5)\)then

8:Update

μ\\muby Eq\. \([6](https://arxiv.org/html/2609.30650#S3.E6)\)\.

9:endif

10:endfor

11:For transferred entries, accumulate diagnostic surprise

dj←dj\+Xj,td\_\{j\}\\leftarrow d\_\{j\}\+X\_\{j,t\}\.

12:endfor

13:Select shifted candidates by Eq\. \([7](https://arxiv.org/html/2609.30650#S3.E7)\) or top\-

ssdiagnostic scores\.

14:foreach selected mechanism index

jjdo

15:Retest entry

jjand compute

θ~j\\widetilde\{\\theta\}\_\{j\}by Eq\. \([9](https://arxiv.org/html/2609.30650#S3.E9)\)\.

16:Update

μ\\muby Eq\. \([10](https://arxiv.org/html/2609.30650#S3.E10)\)\.

17:endfor

18:returnfixed memory

μ\\mu\.

## 4Theoretical Results

The results connect interface fibers, frozen\-state probe error, and the number of target\-domain retests\. Theorem[2](https://arxiv.org/html/2609.30650#Thmtheorem2)gives an exact factorization criterion and the optimal probe risk\. Corollary[3](https://arxiv.org/html/2609.30650#Thmtheorem3)gives the finite\-sample two\-model bound\. Theorem[4](https://arxiv.org/html/2609.30650#Thmtheorem4)gives the exact posterior coverage of budgeted retesting, while Propositions[5](https://arxiv.org/html/2609.30650#Thmtheorem5)and[6](https://arxiv.org/html/2609.30650#Thmtheorem6)give constructive rates under the stated margins\. Theorem[7](https://arxiv.org/html/2609.30650#Thmtheorem7)identifies the unique support of an exact target update and composes the finite\-sample events\. Theorem[8](https://arxiv.org/html/2609.30650#Thmtheorem8)gives the metric\-valued extension\. The core results follow by conditional decision minimization, fiber factorization, and an exact edit identity\. Standard testing inequalities enter only the finite\-sample corollaries\. Auxiliary separations and consequences appear in Appendix[A](https://arxiv.org/html/2609.30650#A1)\. LetM0∼ℐ,ΠM1M\_\{0\}\\sim\_\{\\mathcal\{I\},\\Pi\}M\_\{1\}denote equality of the interface law under protocolΠ\\Pi, and letM0∼𝒵M1M\_\{0\}\\sim\_\{\\mathcal\{Z\}\}M\_\{1\}denote equality of all fixed probe answers on𝒵\\mathcal\{Z\}\. All logarithms are natural\.

###### Theorem 2\.

Letℳ\\mathcal\{M\}and𝒵\\mathcal\{Z\}be finite, letν\\nuandQQhave full support, and use the variables in Eq\. \([1](https://arxiv.org/html/2609.30650#S2.E1)\)\. Then the following statements hold\.

1. \(i\)The probe Bayes risk has the exact representation 𝔅ν,Q\(ℐ,Π\)=infgℙ\{g\(𝖳ℐ,Z\)≠U\}\.\\mathfrak\{B\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)=\\inf\_\{g\}\\mathbb\{P\}\\\{g\(\\mathsf\{T\}\_\{\\mathcal\{I\}\},Z\)\\neq U\\\}\.The infimum is attained by a deterministic decoder\.
2. \(ii\)𝔅ν,Q​\(ℐ,Π\)=0\\mathfrak\{B\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)=0if and only if M0∼ℐ,ΠM1⟹M0∼𝒵M1∀M0,M1∈ℳ\.M\_\{0\}\\sim\_\{\\mathcal\{I\},\\Pi\}M\_\{1\}\\Longrightarrow M\_\{0\}\\sim\_\{\\mathcal\{Z\}\}M\_\{1\}\\qquad\\forall M\_\{0\},M\_\{1\}\\in\\mathcal\{M\}\.Hence a zero\-risk population decoder exists exactly when the interface quotient refines the probe quotient\.
3. \(iii\)If two interface\-equivalent models disagree on a set of probes ofQQ\-massρ\\rho, every population decoder has worst\-case mechanism risk at leastρ/2\\rho/2\.
4. \(iv\)LetDND\_\{N\}be a training transcript whose conditional law givenWWdepends onWWonly through𝖳ℐ\\mathsf\{T\}\_\{\\mathcal\{I\}\}, and letSSbe any randomized function ofDND\_\{N\}\. Every frozen\-state decoder satisfies ℙ\{d\(S,Z\)≠U\}≥𝔅ν,Q\(ℐ,Π\)\.\\mathbb\{P\}\\\{d\(S,Z\)\\neq U\\\}\\geq\\mathfrak\{B\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)\.
5. \(v\)If𝖳1=h⁡\(𝖳2\)\\mathsf\{T\}\_\{1\}=h\(\\mathsf\{T\}\_\{2\}\), then𝔅⁡\(𝖳2\)≤𝔅⁡\(𝖳1\)\\mathfrak\{B\}\(\\mathsf\{T\}\_\{2\}\)\\leq\\mathfrak\{B\}\(\\mathsf\{T\}\_\{1\}\)\. For a vector answer, maxj∈\[d\]⁡𝔅j≤𝔅ν,Q​\(ℐ,Π\)≤∑j=1d𝔅j\.\\max\_\{j\\in\[d\]\}\\mathfrak\{B\}\_\{j\}\\leq\\mathfrak\{B\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)\\leq\\sum\_\{j=1\}^\{d\}\\mathfrak\{B\}\_\{j\}\.\(11\)

###### Proof\.

Write𝖳=𝖳ℐ\\mathsf\{T\}=\\mathsf\{T\}\_\{\\mathcal\{I\}\}\. Conditional on\(𝖳,Z\)=\(t,z\)\(\\mathsf\{T\},Z\)=\(t,z\), a decoder returninguuhas error1−ℙ⁡\(U=u∣t,z\)1\-\\mathbb\{P\}\(U=u\\mid t,z\)\. Pointwise minimization overuu, followed by expectation over\(𝖳,Z\)\(\\mathsf\{T\},Z\), proves \(i\)\.

The risk in \(i\) is zero exactly when every conditional distributionP⁡\(U∣𝖳=t,Z=z\)P\(U\\mid\\mathsf\{T\}=t,Z=z\)is a point mass\. Full support ofν\\nuandQQturns this into a pointwise statement\. Thus, wheneverτℐ,Π​\(M0\)=τℐ,Π​\(M1\)\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\_\{0\}\)=\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\_\{1\}\),

GM0​\(z\)=g⁡\(τℐ,Π​\(M0\),z\)=g⁡\(τℐ,Π​\(M1\),z\)=GM1​\(z\)G\_\{M\_\{0\}\}\(z\)=g\(\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\_\{0\}\),z\)=g\(\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\_\{1\}\),z\)=G\_\{M\_\{1\}\}\(z\)for everyz∈𝒵z\\in\\mathcal\{Z\}\. This proves the forward implication in \(ii\)\. Conversely, suppose the quotient implication holds\. Every interface classC=\{M:τℐ,Π​\(M\)=t\}C=\\\{M:\\tau\_\{\\mathcal\{I\},\\Pi\}\(M\)=t\\\}has a unique probe mapgCg\_\{C\}, so defineg​\(t,z\)=gC​\(z\)g\(t,z\)=g\_\{C\}\(z\)\. ThenU=g⁡\(𝖳,Z\)U=g\(\\mathsf\{T\},Z\)and the Bayes risk is zero\.

For \(iii\), letA=\{z:GM0​\(z\)≠GM1​\(z\)\}A=\\\{z:G\_\{M\_\{0\}\}\(z\)\\neq G\_\{M\_\{1\}\}\(z\)\\\}withQ⁡\(A\)=ρQ\(A\)=\\rho\. Under the common interface descriptor, any randomized decoder has the same output law in the two models\. Pointwise onAA,

𝟏\{G^\(z\)≠GM0\(z\)\}\+𝟏\{G^\(z\)≠GM1\(z\)\}≥1\.\\mathbf\{1\}\\\{\\widehat\{G\}\(z\)\\neq G\_\{M\_\{0\}\}\(z\)\\\}\+\\mathbf\{1\}\\\{\\widehat\{G\}\(z\)\\neq G\_\{M\_\{1\}\}\(z\)\\\}\\geq 1\.Expectation over decoder randomness andz∼Qz\\sim Qgives

RQ​\(G^,M0\)\+RQ​\(G^,M1\)≥ρ,R\_\{Q\}\(\\widehat\{G\},M\_\{0\}\)\+R\_\{Q\}\(\\widehat\{G\},M\_\{1\}\)\\geq\\rho,and hencemaxi⁡RQ​\(G^,Mi\)≥ρ/2\\max\_\{i\}R\_\{Q\}\(\\widehat\{G\},M\_\{i\}\)\\geq\\rho/2\.

For \(iv\), the transcript condition makesSSa randomized garbling of𝖳\\mathsf\{T\}\. Composingd⁡\(S,Z\)d\(S,Z\)with the garbling kernel produces a randomized decoder from\(𝖳,Z\)\(\\mathsf\{T\},Z\)\. Randomization cannot improve the pointwise minimum in \(i\), proving the bound\.

For \(v\), every decoder based on𝖳1=h⁡\(𝖳2\)\\mathsf\{T\}\_\{1\}=h\(\\mathsf\{T\}\_\{2\}\)is also a decoder based on𝖳2\\mathsf\{T\}\_\{2\}, so the latter infimum is no larger\. Any joint decoder induces a decoder for coordinatejj; therefore its vector error is at least𝔅j\\mathfrak\{B\}\_\{j\}, proving the lower bound in Eq\. \([11](https://arxiv.org/html/2609.30650#S4.E11)\)\. Conversely, concatenate the coordinatewise Bayes decoders\. Its vector error is contained in the union of their coordinate errors, and the union bound proves the upper inequality\. ∎

For a uniform two\-model fiber and a point\-mass probe on which the answers differ, the probe Bayes risk is1/21/2\. A probe\-revealing interface has risk zero\.

With finite samples, total variation controls the residual indistinguishability of two probe\-distinct models\.

###### Corollary 3\.

LetP0NP\_\{0\}^\{N\}andP1NP\_\{1\}^\{N\}be theNN\-sample transcript laws of two SCMs whose probe answers differ on a set ofQQ\-massρ\\rho\. Then every estimator obeys

maxi∈\{0,1\}⁡RQ​\(G^,Mi\)≥ρ2​\(1−∥P0N−P1N∥TV\)\.\\max\_\{i\\in\\\{0,1\\\}\}R\_\{Q\}\(\\widehat\{G\},M\_\{i\}\)\\geq\\frac\{\\rho\}\{2\}\\left\(1\-\\lVert P\_\{0\}^\{N\}\-P\_\{1\}^\{N\}\\rVert\_\{\\mathrm\{TV\}\}\\right\)\.If the samples are conditionally independent with one\-sample lawsP0,P1P\_\{0\},P\_\{1\}andd=min\{KL\(P0∥P1\),KL\(P1∥P0\)\}d=\\min\\\{\\mathrm\{KL\}\(P\_\{0\}\\\|P\_\{1\}\),\\mathrm\{KL\}\(P\_\{1\}\\\|P\_\{0\}\)\\\}, then

maxi⁡RQ​\(G^,Mi\)≥ρ2​\[1−N​d2\]\+\.\\max\_\{i\}R\_\{Q\}\(\\widehat\{G\},M\_\{i\}\)\\geq\\frac\{\\rho\}\{2\}\\left\[1\-\\sqrt\{\\frac\{Nd\}\{2\}\}\\right\]\_\{\+\}\.

###### Proof\.

Fix a probezzon which the answersg0​\(z\)g\_\{0\}\(z\)andg1​\(z\)g\_\{1\}\(z\)differ\. LetAi\(z\)=\{G^\(z\)=gi\(z\)\}A\_\{i\}\(z\)=\\\{\\widehat\{G\}\(z\)=g\_\{i\}\(z\)\\\}\. The eventsA0​\(z\)A\_\{0\}\(z\)andA1​\(z\)A\_\{1\}\(z\)are disjoint, hence

P0N​\(A0​\(z\)c\)\+P1N​\(A1​\(z\)c\)\\displaystyle P\_\{0\}^\{N\}\(A\_\{0\}\(z\)^\{c\}\)\+P\_\{1\}^\{N\}\(A\_\{1\}\(z\)^\{c\}\)≥P0N​\(A1​\(z\)\)\+P1N​\(A1​\(z\)c\)\\displaystyle\\geq P\_\{0\}^\{N\}\(A\_\{1\}\(z\)\)\+P\_\{1\}^\{N\}\(A\_\{1\}\(z\)^\{c\}\)=1−\{P1N​\(A1​\(z\)\)−P0N​\(A1​\(z\)\)\}\\displaystyle=1\-\\\{P\_\{1\}^\{N\}\(A\_\{1\}\(z\)\)\-P\_\{0\}^\{N\}\(A\_\{1\}\(z\)\)\\\}≥1−∥P0N−P1N∥TV\.\\displaystyle\\geq 1\-\\lVert P\_\{0\}^\{N\}\-P\_\{1\}^\{N\}\\rVert\_\{\\mathrm\{TV\}\}\.Integrating over the disagreement set and dividing the sum of the two risks by two proves the first inequality\. For product laws,KL\(P0N∥P1N\)=NKL\(P0∥P1\)\\mathrm\{KL\}\(P\_\{0\}^\{N\}\\\|P\_\{1\}^\{N\}\)=N\\mathrm\{KL\}\(P\_\{0\}\\\|P\_\{1\}\)\. Pinsker’s inequality\([Tsybakov, 2009](https://arxiv.org/html/2609.30650#bib.bib28)\), applied in the better of the two directions, gives∥P0N−P1N∥TV≤N​d/2\\lVert P\_\{0\}^\{N\}\-P\_\{1\}^\{N\}\\rVert\_\{\\mathrm\{TV\}\}\\leq\\sqrt\{Nd/2\}; truncation at zero gives the second display\. ∎

The proof is decision\-theoretic: it does not pass through entropy, mutual information, or a decoder\-class restriction\.

Risk decompositions, standard\-target lower bounds, and explicit two\-world constructions for ranking, readout, and hidden\-context interfaces appear in Appendix[A](https://arxiv.org/html/2609.30650#A1)\.

###### Theorem 4\.

Letℬs=\{b⊆\[n\]:\|b\|=s\}\\mathcal\{B\}\_\{s\}=\\\{b\\subseteq\[n\]:\|b\|=s\\\}, letB∈ℬsB\\in\\mathcal\{B\}\_\{s\}, and letZZbe any pre\-retest diagnostic signal\. Fors≤k<ns\\leq k<n, define

Γk\(z\)=maxL⊆\[n\]:\|L\|≤k∑b∈ℬs:b⊆Lℙ\(B=b∣Z=z\)\.\\Gamma\_\{k\}\(z\)=\\max\_\{L\\subseteq\[n\]:\\,\|L\|\\leq k\}\\sum\_\{b\\in\\mathcal\{B\}\_\{s\}:\\,b\\subseteq L\}\\mathbb\{P\}\(B=b\\mid Z=z\)\.\(12\)Then the largest probability that aZZ\-measurable retest list contains every shifted entry is

supℒ:\|ℒ⁡\(Z\)\|≤kℙ\{B⊆ℒ\(Z\)\}=𝔼Γk\(Z\)\.\\sup\_\{\\mathcal\{L\}:\\,\|\\mathcal\{L\}\(Z\)\|\\leq k\}\\mathbb\{P\}\\\{B\\subseteq\\mathcal\{L\}\(Z\)\\\}=\\mathbb\{E\}\\Gamma\_\{k\}\(Z\)\.\(13\)Internal randomization cannot increase this value\. Zero omission is possible if and only if, for almost everyzz,

\|⋃\{b:ℙ⁡\(B=b∣Z=z\)\>0\}\|≤k\.\\left\|\\bigcup\\\{b:\\mathbb\{P\}\(B=b\\mid Z=z\)\>0\\\}\\right\|\\leq k\.\(14\)IfBBis uniform onℬs\\mathcal\{B\}\_\{s\}and independent ofZZ, the optimal coverage is exactly\(ks\)/\(ns\)\\binom\{k\}\{s\}/\\binom\{n\}\{s\}\. Finally, ifZ1=h⁡\(Z2\)Z\_\{1\}=h\(Z\_\{2\}\), then𝔼​Γk​\(Z1\)≤𝔼​Γk​\(Z2\)\\mathbb\{E\}\\Gamma\_\{k\}\(Z\_\{1\}\)\\leq\\mathbb\{E\}\\Gamma\_\{k\}\(Z\_\{2\}\)\.

###### Proof\.

For a deterministic listL⁡\(z\)L\(z\), conditional coverage is

ℙ⁡\{B⊆L⁡\(z\)∣Z=z\}=∑b⊆L⁡\(z\)ℙ⁡\(B=b∣Z=z\)\.\\mathbb\{P\}\\\{B\\subseteq L\(z\)\\mid Z=z\\\}=\\sum\_\{b\\subseteq L\(z\)\}\\mathbb\{P\}\(B=b\\mid Z=z\)\.Maximizing this finite sum separately for everyzzgives Eq\. \([12](https://arxiv.org/html/2609.30650#S4.E12)\); averaging gives Eq\. \([13](https://arxiv.org/html/2609.30650#S4.E13)\)\. A randomized rule is a mixture of deterministic lists, so its conditional coverage is a convex combination of values no larger thanΓk​\(z\)\\Gamma\_\{k\}\(z\)\.

The equalityΓk​\(z\)=1\\Gamma\_\{k\}\(z\)=1holds precisely when one list of size at mostkkcontains everybbwith positive posterior mass\. Such a list exists precisely when the union of those posterior\-support sets has size at mostkk, proving Eq\. \([14](https://arxiv.org/html/2609.30650#S4.E14)\)\. Under the uniform independent law, a list of sizeℓ\\ellcontains exactly\(ℓs\)\\binom\{\\ell\}\{s\}of the\(ns\)\\binom\{n\}\{s\}possible shifted sets\. This is maximized atℓ=k\\ell=k\. Finally, every list rule based onZ1=h⁡\(Z2\)Z\_\{1\}=h\(Z\_\{2\}\)is also available to a rule based onZ2Z\_\{2\}; taking suprema in Eq\. \([13](https://arxiv.org/html/2609.30650#S4.E13)\) proves monotonicity\. ∎

###### Proposition 5\.

For each transferred entryj∈\[n\]j\\in\[n\], suppose the diagnostic samplesXj,1,…,Xj,m∈\[0,1\]X\_\{j,1\},\\ldots,X\_\{j,m\}\\in\[0,1\]are independent\. Suppose also that𝔼​Xj,r≥p\\mathbb\{E\}X\_\{j,r\}\\geq pfor shifted entriesj∈Bj\\in Band𝔼​Xj,r≤q\\mathbb\{E\}X\_\{j,r\}\\leq qfor stable entriesj∉Bj\\notin B, withΔ=p−q\>0\\Delta=p\-q\>0\. Letτ=\(p\+q\)/2\\tau=\(p\+q\)/2and

X¯j=1m​∑r=1mXj,r,B^=\{j:X¯j≥τ\}\.\\overline\{X\}\_\{j\}=\\frac\{1\}\{m\}\\sum\_\{r=1\}^\{m\}X\_\{j,r\},\\qquad\\widehat\{B\}=\\\{j:\\overline\{X\}\_\{j\}\\geq\\tau\\\}\.Ifm≥2​Δ−2​log⁡\(n/δ\)m\\geq 2\\Delta^\{\-2\}\\log\(n/\\delta\), thenℙ⁡\(B^=B\)≥1−δ\\mathbb\{P\}\(\\widehat\{B\}=B\)\\geq 1\-\\delta\. Moreover, in the symmetric one\-shift Bernoulli family withp=1/2\+Δ/2p=1/2\+\\Delta/2,q=1/2−Δ/2q=1/2\-\\Delta/2, and0<Δ≤1/20<\\Delta\\leq 1/2, every estimator with localization error at mostε\\varepsilonrequires

m≥\(1−ε\)​log⁡n−log⁡28​Δ2\.m\\geq\\frac\{\(1\-\\varepsilon\)\\log n\-\\log 2\}\{8\\Delta^\{2\}\}\.Thus theΔ−2​log⁡n\\Delta^\{\-2\}\\log nscaling is necessary and sufficient up to constants\.

###### Proof\.

Forj∈Bj\\in B,

ℙ\{X¯j<τ\}≤exp\{−2m\(p−τ\)2\}=exp\{−mΔ2/2\}\.\\displaystyle\\mathbb\{P\}\\\{\\overline\{X\}\_\{j\}<\\tau\\\}\\leq\\exp\\\{\-2m\(p\-\\tau\)^\{2\}\\\}=\\exp\\\{\-m\\Delta^\{2\}/2\\\}\.Forj∉Bj\\notin B,

ℙ\{X¯j≥τ\}≤exp\{−2m\(τ−q\)2\}=exp\{−mΔ2/2\}\.\\displaystyle\\mathbb\{P\}\\\{\\overline\{X\}\_\{j\}\\geq\\tau\\\}\\leq\\exp\\\{\-2m\(\\tau\-q\)^\{2\}\\\}=\\exp\\\{\-m\\Delta^\{2\}/2\\\}\.These two applications of Hoeffding’s inequality\([Hoeffding, 1963](https://arxiv.org/html/2609.30650#bib.bib12)\), followed by a union bound over thennentries, giveℙ\(B^≠B\)≤nexp\{−mΔ2/2\}≤δ\\mathbb\{P\}\(\\widehat\{B\}\\neq B\)\\leq n\\exp\\\{\-m\\Delta^\{2\}/2\\\}\\leq\\delta\. This is the selection guarantee used by Eq\. \([7](https://arxiv.org/html/2609.30650#S3.E7)\); the top\-ssvariant is proved in Appendix[A\.6](https://arxiv.org/html/2609.30650#A1.SS6)\. The matching lower\-bound calculation is given in Appendix[A\.5](https://arxiv.org/html/2609.30650#A1.SS5)\. ∎

Theorem[4](https://arxiv.org/html/2609.30650#Thmtheorem4)characterizes every retest list of sizekkwithout reducing the diagnostic signal to a scalar information budget\. In contrast, Proposition[5](https://arxiv.org/html/2609.30650#Thmtheorem5)recovers the shifted set withm=O⁡\(Δ−2​log⁡\(n/δ\)\)m=O\(\\Delta^\{\-2\}\\log\(n/\\delta\)\)diagnostic samples per entry\.

###### Proposition 6\.

Fix a selected entryjjand a finite candidate class𝒩j\\mathcal\{N\}\_\{j\}\. Letθj⋆\\theta\_\{j\}^\{\\star\}be the target\-domain mechanism and define the population retest loss

Lj\(θ\)=𝔼1\{gθ\(Zj\)≠Uj\},θ∈𝒩j\.L\_\{j\}\(\\theta\)=\\mathbb\{E\}\\,\\mathbf\{1\}\\\{g\_\{\\theta\}\(Z\_\{j\}\)\\neq U\_\{j\}\\\},\\qquad\\theta\\in\\mathcal\{N\}\_\{j\}\.Supposeθj⋆∈𝒩j\\theta\_\{j\}^\{\\star\}\\in\\mathcal\{N\}\_\{j\}andLj​\(θ\)−Lj​\(θj⋆\)≥ηj\>0L\_\{j\}\(\\theta\)\-L\_\{j\}\(\\theta\_\{j\}^\{\\star\}\)\\geq\\eta\_\{j\}\>0for everyθ≠θj⋆\\theta\\neq\\theta\_\{j\}^\{\\star\}, and suppose therjr\_\{j\}retest pairs are independent and identically distributed\. If the retest sample size is

rj≥2ηj2​log⁡\|𝒩j\|δj,r\_\{j\}\\geq\\frac\{2\}\{\\eta\_\{j\}^\{2\}\}\\log\\frac\{\|\\mathcal\{N\}\_\{j\}\|\}\{\\delta\_\{j\}\},then the empirical minimizer in Eq\. \([9](https://arxiv.org/html/2609.30650#S3.E9)\) satisfiesℙ⁡\(θ~j=θj⋆\)≥1−δj\\mathbb\{P\}\(\\widetilde\{\\theta\}\_\{j\}=\\theta\_\{j\}^\{\\star\}\)\\geq 1\-\\delta\_\{j\}\.

###### Proof\.

For a wrong candidateθ\\theta, set

Dθ=𝟏\{gθ\(Zj\)≠Uj\}−𝟏\{gθj⋆\(Zj\)≠Uj\}\.D\_\{\\theta\}=\\mathbf\{1\}\\\{g\_\{\\theta\}\(Z\_\{j\}\)\\neq U\_\{j\}\\\}\-\\mathbf\{1\}\\\{g\_\{\\theta\_\{j\}^\{\\star\}\}\(Z\_\{j\}\)\\neq U\_\{j\}\\\}\.ThenDθ∈\[−1,1\]D\_\{\\theta\}\\in\[\-1,1\]and𝔼​Dθ=Lj​\(θ\)−Lj​\(θj⋆\)≥ηj\\mathbb\{E\}D\_\{\\theta\}=L\_\{j\}\(\\theta\)\-L\_\{j\}\(\\theta\_\{j\}^\{\\star\}\)\\geq\\eta\_\{j\}\. IfL^j​\(θ\)≤L^j​\(θj⋆\)\\widehat\{L\}\_\{j\}\(\\theta\)\\leq\\widehat\{L\}\_\{j\}\(\\theta\_\{j\}^\{\\star\}\), then the empirical meanD¯θ\\overline\{D\}\_\{\\theta\}is nonpositive\. ForAθ=\{L^j\(θ\)≤L^j\(θj⋆\)\}A\_\{\\theta\}=\\\{\\widehat\{L\}\_\{j\}\(\\theta\)\\leq\\widehat\{L\}\_\{j\}\(\\theta\_\{j\}^\{\\star\}\)\\\}, Hoeffding’s inequality\([Hoeffding, 1963](https://arxiv.org/html/2609.30650#bib.bib12)\)gives

ℙ\(Aθ\)≤ℙ\{D¯θ−𝔼Dθ≤−ηj\}≤e−rjηj2/2\.\\mathbb\{P\}\(A\_\{\\theta\}\)\\leq\\mathbb\{P\}\\\{\\overline\{D\}\_\{\\theta\}\-\\mathbb\{E\}D\_\{\\theta\}\\leq\-\\eta\_\{j\}\\\}\\leq e^\{\-r\_\{j\}\\eta\_\{j\}^\{2\}/2\}\.A union bound over𝒩j∖\{θj⋆\}\\mathcal\{N\}\_\{j\}\\setminus\\\{\\theta\_\{j\}^\{\\star\}\\\}yieldsℙ\(θ~j≠θj⋆\)≤\|𝒩j\|e−rjηj2/2≤δj\\mathbb\{P\}\(\\widetilde\{\\theta\}\_\{j\}\\neq\\theta\_\{j\}^\{\\star\}\)\\leq\|\\mathcal\{N\}\_\{j\}\|e^\{\-r\_\{j\}\\eta\_\{j\}^\{2\}/2\}\\leq\\delta\_\{j\}\. ∎

###### Theorem 7\.

Fix an aligned source\-target pair satisfying Condition[C1](https://arxiv.org/html/2609.30650#S2.I1.i1)\. Writemj=ϕ⁡\(θjs\)m\_\{j\}=\\phi\(\\theta\_\{j\}^\{s\}\)andB=\{j:mj≠θjt\}B=\\\{j:m\_\{j\}\\neq\\theta\_\{j\}^\{t\}\\\}, withS=\[n\]∖BS=\[n\]\\setminus B\. For any selected setD⊆\[n\]D\\subseteq\[n\]and local estimates\{θ~j:j∈D\}\\\{\\widetilde\{\\theta\}\_\{j\}:j\\in D\\\}, let

θ^jt​\(D\)=\{mj,j∉D,θ~j,j∈D\.\\widehat\{\\theta\}\_\{j\}^\{t\}\(D\)=\\begin\{cases\}m\_\{j\},&j\\notin D,\\\\ \\widetilde\{\\theta\}\_\{j\},&j\\in D\.\\end\{cases\}Then

n​Ladapt\\displaystyle nL\_\{\\mathrm\{adapt\}\}=\|B∖D\|\+∑j∈D𝟏\{θ~j≠θjt\},\\displaystyle=\|B\\setminus D\|\+\\sum\_\{j\\in D\}\\mathbf\{1\}\\\{\\widetilde\{\\theta\}\_\{j\}\\neq\\theta\_\{j\}^\{t\}\\\},\(15\)\|S\|​Lstab\\displaystyle\|S\|L\_\{\\mathrm\{stab\}\}=∑j∈D∩S𝟏\{θ~j≠mj\},\\displaystyle=\\sum\_\{j\\in D\\cap S\}\\mathbf\{1\}\\\{\\widetilde\{\\theta\}\_\{j\}\\neq m\_\{j\}\\\},\(16\)\|B\|​Lshift\\displaystyle\|B\|L\_\{\\mathrm\{shift\}\}=\|B∖D\|\+∑j∈D∩B𝟏\{θ~j≠θjt\}\.\\displaystyle=\|B\\setminus D\|\+\\sum\_\{j\\in D\\cap B\}\\mathbf\{1\}\\\{\\widetilde\{\\theta\}\_\{j\}\\neq\\theta\_\{j\}^\{t\}\\\}\.\(17\)Consequently the adapted memory is exact if and only ifB⊆DB\\subseteq Dandθ~j=θjt\\widetilde\{\\theta\}\_\{j\}=\\theta\_\{j\}^\{t\}for everyj∈Dj\\in D\. Moreover, for any exact target memoryθ¯=θt\\bar\{\\theta\}=\\theta^\{t\},

\{j:θ¯j≠mj\}=B;\\\{j:\\bar\{\\theta\}\_\{j\}\\neq m\_\{j\}\\\}=B;\(18\)thusBBis the unique realized edit support of an exact update relative to the mapped source memory\.

LetFsrcF\_\{\\mathrm\{src\}\}be the event that the frozen source memory has a missing entry or a wrong target, context, value, or delay, including a target error caused by readout contamination\. LetFϕF\_\{\\phi\}be schema\-map failure,Fdiag=\{B^≠B\}F\_\{\\mathrm\{diag\}\}=\\\{\\widehat\{B\}\\neq B\\\}, andFupdF\_\{\\mathrm\{upd\}\}be failure of any local retest onBB\. If their probabilities are bounded byδsrc,δϕ,δdiag\\delta\_\{\\mathrm\{src\}\},\\delta\_\{\\phi\},\\delta\_\{\\mathrm\{diag\}\}, andδupd\\delta\_\{\\mathrm\{upd\}\}, then the update in Eq\. \([8](https://arxiv.org/html/2609.30650#S3.E8)\) satisfies

ℙ⁡\(Ladapt\>0\)≤δsrc\+δϕ\+δdiag\+δupd\.\\mathbb\{P\}\(L\_\{\\mathrm\{adapt\}\}\>0\)\\leq\\delta\_\{\\mathrm\{src\}\}\+\\delta\_\{\\phi\}\+\\delta\_\{\\mathrm\{diag\}\}\+\\delta\_\{\\mathrm\{upd\}\}\.On the complement of these events,Lstab=Lshift=Ladapt=0L\_\{\\mathrm\{stab\}\}=L\_\{\\mathrm\{shift\}\}=L\_\{\\mathrm\{adapt\}\}=0\. For every realization,

𝔼​Ladapt=\|S\|n​𝔼​Lstab\+\|B\|n​𝔼​Lshift\.\\mathbb\{E\}L\_\{\\mathrm\{adapt\}\}=\\frac\{\|S\|\}\{n\}\\mathbb\{E\}L\_\{\\mathrm\{stab\}\}\+\\frac\{\|B\|\}\{n\}\\mathbb\{E\}L\_\{\\mathrm\{shift\}\}\.

###### Proof\.

Forj∉Dj\\notin D, the update retainsmjm\_\{j\}, and therefore

𝟏\{θ^jt\(D\)≠θjt\}=𝟏\{j∈B\}\.\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\(D\)\\neq\\theta\_\{j\}^\{t\}\\\}=\\mathbf\{1\}\\\{j\\in B\\\}\.Forj∈Dj\\in D, the same error indicator is𝟏\{θ~j≠θjt\}\\mathbf\{1\}\\\{\\widetilde\{\\theta\}\_\{j\}\\neq\\theta\_\{j\}^\{t\}\\\}\. Summing the first identity overj∉Dj\\notin Dand the second overj∈Dj\\in Dproves Eq\. \([15](https://arxiv.org/html/2609.30650#S4.E15)\)\. OnSS,mj=θjtm\_\{j\}=\\theta\_\{j\}^\{t\}; onBB,mj≠θjtm\_\{j\}\\neq\\theta\_\{j\}^\{t\}\. Splitting the two sums accordingly proves Eqs\. \([16](https://arxiv.org/html/2609.30650#S4.E16)\) and \([17](https://arxiv.org/html/2609.30650#S4.E17)\)\.

Every term on the right of Eq\. \([15](https://arxiv.org/html/2609.30650#S4.E15)\) is nonnegative\. It vanishes precisely whenB∖D=∅B\\setminus D=\\varnothingand every selected estimate equals its target value, giving the stated necessary and sufficient condition\. Ifθ¯=θt\\bar\{\\theta\}=\\theta^\{t\}, thenθ¯j≠mj\\bar\{\\theta\}\_\{j\}\\neq m\_\{j\}holds precisely whenθjt≠mj\\theta\_\{j\}^\{t\}\\neq m\_\{j\}, which is the definition ofj∈Bj\\in B\. This proves Eq\. \([18](https://arxiv.org/html/2609.30650#S4.E18)\)\.

Now letE=\(Fsrc∪Fϕ∪Fdiag∪Fupd\)cE=\(F\_\{\\mathrm\{src\}\}\\cup F\_\{\\phi\}\\cup F\_\{\\mathrm\{diag\}\}\\cup F\_\{\\mathrm\{upd\}\}\)^\{c\}\. OnEE, the mapped source entries are correct,D=B^=BD=\\widehat\{B\}=B, andθ~j=θjt\\widetilde\{\\theta\}\_\{j\}=\\theta\_\{j\}^\{t\}for every selectedjj\. The exact condition just proved givesLadapt=0L\_\{\\mathrm\{adapt\}\}=0\. Therefore

\{Ladapt\>0\}⊆Fsrc∪Fϕ∪Fdiag∪Fupd,\\\{L\_\{\\mathrm\{adapt\}\}\>0\\\}\\subseteq F\_\{\\mathrm\{src\}\}\\cup F\_\{\\phi\}\\cup F\_\{\\mathrm\{diag\}\}\\cup F\_\{\\mathrm\{upd\}\},and the probability bound follows by the union bound\. The decomposition of the mean loss follows from partitioning the exact coordinate errors:

Ladapt\\displaystyle L\_\{\\mathrm\{adapt\}\}=1n∑j∈S𝟏\{θ^jt≠θjt\}\+1n∑j∈B𝟏\{θ^jt≠θjt\}\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{j\\in S\}\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\\neq\\theta\_\{j\}^\{t\}\\\}\+\\frac\{1\}\{n\}\\sum\_\{j\\in B\}\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\\neq\\theta\_\{j\}^\{t\}\\\}=\|S\|n​Lstab\+\|B\|n​Lshift\.\\displaystyle=\\frac\{\|S\|\}\{n\}L\_\{\\mathrm\{stab\}\}\+\\frac\{\|B\|\}\{n\}L\_\{\\mathrm\{shift\}\}\.Taking expectation gives the displayed identity\. Appendix[A\.7](https://arxiv.org/html/2609.30650#A1.SS7)gives the coordinate\-level loss expansion and decomposesFsrcF\_\{\\mathrm\{src\}\}into its target, context, readout, and temporal gates\. ∎

Substitution of the target\-contrast, readout, hidden\-setup, temporal, diagnostic, schema, and local\-update bounds into Theorem[7](https://arxiv.org/html/2609.30650#Thmtheorem7)is given in Appendix[A\.7](https://arxiv.org/html/2609.30650#A1.SS7)\. The corresponding top\-ssdiagnostic result appears in Appendix[A\.6](https://arxiv.org/html/2609.30650#A1.SS6); its selected empirical means recover the shifted set under the sameΔ−2​log⁡\(n/δ\)\\Delta^\{\-2\}\\log\(n/\\delta\)scaling\.

###### Theorem 8\.

Letℳ\\mathcal\{M\}and𝒵\\mathcal\{Z\}be finite, letν\\nuandQQhave full support, and let\(𝒰,ρ\)\(\\mathcal\{U\},\\rho\)be a compact metric space with continuousρ\\rho\. Use Eq\. \([2](https://arxiv.org/html/2609.30650#S2.E2)\) and writeV=\(𝖳ℐ,Z\)V=\(\\mathsf\{T\}\_\{\\mathcal\{I\}\},Z\)\. Then:

1. \(i\)The expected metric risk has the pointwise representation 𝔅ν,Qρ​\(ℐ,Π\)=𝔼⁡\[mina∈𝒰⁡𝔼⁡\{ρ⁡\(U,a\)∣V\}\],\\mathfrak\{B\}^\{\\rho\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)=\\mathbb\{E\}\\\!\\left\[\\min\_\{a\\in\\mathcal\{U\}\}\\mathbb\{E\}\\\{\\rho\(U,a\)\\mid V\\\}\\right\],and a deterministic decoder attains the minimum\.
2. \(ii\)𝔅ν,Qρ=0\\mathfrak\{B\}^\{\\rho\}\_\{\\nu,Q\}=0andℜρ=0\\mathfrak\{R\}^\{\\rho\}=0if and only if every interface fiber has one probe answer map; equivalently,M0∼ℐ,ΠM1M\_\{0\}\\sim\_\{\\mathcal\{I\},\\Pi\}M\_\{1\}impliesM0∼𝒵M1M\_\{0\}\\sim\_\{\\mathcal\{Z\}\}M\_\{1\}\.
3. \(iii\)IfDND\_\{N\}andSSsatisfy the transcript condition in Theorem[2](https://arxiv.org/html/2609.30650#Thmtheorem2), every frozen\-state decoder obeys 𝔼​ρ​\{U,d⁡\(S,Z\)\}≥𝔅ν,Qρ​\(ℐ,Π\)\.\\mathbb\{E\}\\rho\\\{U,d\(S,Z\)\\\}\\geq\\mathfrak\{B\}^\{\\rho\}\_\{\\nu,Q\}\(\\mathcal\{I\},\\Pi\)\.If𝖳1=h⁡\(𝖳2\)\\mathsf\{T\}\_\{1\}=h\(\\mathsf\{T\}\_\{2\}\), then𝔅ρ​\(𝖳2\)≤𝔅ρ​\(𝖳1\)\\mathfrak\{B\}^\{\\rho\}\(\\mathsf\{T\}\_\{2\}\)\\leq\\mathfrak\{B\}^\{\\rho\}\(\\mathsf\{T\}\_\{1\}\)andℜρ​\(𝖳2\)≤ℜρ​\(𝖳1\)\\mathfrak\{R\}^\{\\rho\}\(\\mathsf\{T\}\_\{2\}\)\\leq\\mathfrak\{R\}^\{\\rho\}\(\\mathsf\{T\}\_\{1\}\)\.
4. \(iv\)If two interface\-equivalent modelsM0,M1M\_\{0\},M\_\{1\}have answer mapsG0,G1G\_\{0\},G\_\{1\}, every possibly randomized decoder of their common interface satisfies maxi∈\{0,1\}⁡𝔼Z∼Q​ρ​\{G^​\(Z\),Gi​\(Z\)\}≥12​𝔼Z∼Q​ρ​\{G0​\(Z\),G1​\(Z\)\}\.\\max\_\{i\\in\\\{0,1\\\}\}\\mathbb\{E\}\_\{Z\\sim Q\}\\rho\\\{\\widehat\{G\}\(Z\),G\_\{i\}\(Z\)\\\}\\geq\\frac\{1\}\{2\}\\mathbb\{E\}\_\{Z\\sim Q\}\\rho\\\{G\_\{0\}\(Z\),G\_\{1\}\(Z\)\\\}\.In particular, separation by at least2​ε2\\varepsilonon a set ofQQ\-massqqforces worst\-case metric risk at leastq​εq\\varepsilon\.

###### Proof\.

Condition onV=vV=v\. A deterministic answeraahas conditional loss𝔼​\{ρ⁡\(U,a\)∣V=v\}\\mathbb\{E\}\\\{\\rho\(U,a\)\\mid V=v\\\}\. Compactness and continuity give a minimizer; pointwise minimization and averaging prove \(i\)\. Randomization only averages these conditional losses and cannot improve their minimum\.

The risk in \(i\) is zero exactly whenUUis constant conditional on every\(𝖳ℐ,Z\)\(\\mathsf\{T\}\_\{\\mathcal\{I\}\},Z\)value with positive probability\. Full support and the metric property turn this into equality of the complete probe maps inside each interface fiber\. The same condition is plainly necessary and sufficient for zero worst\-case risk, proving \(ii\)\.

For \(iii\), composing a decoder ofSSwith the random kernel from𝖳ℐ\\mathsf\{T\}\_\{\\mathcal\{I\}\}toSSgives a randomized decoder of the population interface\. Part \(i\) lower\-bounds its loss\. If𝖳1=h⁡\(𝖳2\)\\mathsf\{T\}\_\{1\}=h\(\\mathsf\{T\}\_\{2\}\), every decoder based on𝖳1\\mathsf\{T\}\_\{1\}is available from𝖳2\\mathsf\{T\}\_\{2\}; taking the two infima proves both monotonicity statements\.

For \(iv\), interface equivalence makes the output law ofG^​\(z\)\\widehat\{G\}\(z\)the same underM0M\_\{0\}andM1M\_\{1\}\. The triangle inequality gives, pointwise inzzand in decoder randomness,

ρ⁡\{G0​\(z\),G1​\(z\)\}≤ρ⁡\{G0​\(z\),G^​\(z\)\}\+ρ⁡\{G^​\(z\),G1​\(z\)\}\.\\rho\\\{G\_\{0\}\(z\),G\_\{1\}\(z\)\\\}\\leq\\rho\\\{G\_\{0\}\(z\),\\widehat\{G\}\(z\)\\\}\+\\rho\\\{\\widehat\{G\}\(z\),G\_\{1\}\(z\)\\\}\.Expectation overZZand decoder randomness shows that the sum of the two risks is at least𝔼Q​ρ​\{G0​\(Z\),G1​\(Z\)\}\\mathbb\{E\}\_\{Q\}\\rho\\\{G\_\{0\}\(Z\),G\_\{1\}\(Z\)\\\}\. Their maximum is at least half this quantity, which proves Eq\. \([\(iv\)](https://arxiv.org/html/2609.30650#S4.Ex66)\)\. The final statement follows by integrating the assumed separation over its probe set\. ∎

## 5Experiments

The evaluation covers finite SCM families, readout and hidden\-context variants, few\-shot renamed adaptation, metric\-valued control responses, language\-model hidden states, and language\-model proposal policies\. Every learned state is frozen before the mechanism probes are scored\. Exact protocols, seeds, and auxiliary tables are in Appendices[B\.1](https://arxiv.org/html/2609.30650#A2.SS1)–[B\.5](https://arxiv.org/html/2609.30650#A2.SS5)\.

Proposition[9](https://arxiv.org/html/2609.30650#Thmtheorem9)supplies exact indistinguishable model pairs\. The empirical runs measure the corresponding fixed\-probe failures and the finite\-sample behavior of the constructive rules\. Renamed transfer uses shifted\-set localization and entrywise retesting\. Continuous responses are integrated over held\-out state, action\-direction, and horizon distributions; language\-model results use a separately fixed probe map\.

Held\-out state\-action probes query target, value, delay, context, and stable preservation\. Reward and top\-action accuracy are not terms in the score\. For probe set𝒫\\mathcal\{P\},

MechAcc=1\|𝒫\|∑\(ξ,a\)∈𝒫𝟏\{Tμ^\(ξ,a\)≡Tμ⋆\(ξ,a\)\}\.\\mathrm\{MechAcc\}=\\frac\{1\}\{\|\\mathcal\{P\}\|\}\\sum\_\{\(\\xi,a\)\\in\\mathcal\{P\}\}\\mathbf\{1\}\\\{T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)\\equiv T\_\{\\mu^\{\\star\}\}\(\\xi,a\)\\\}\.\(19\)ThusR^mech=1−MechAcc\\widehat\{R\}\_\{\\mathrm\{mech\}\}=1\-\\mathrm\{MechAcc\}estimatesRQR\_\{Q\}\. The coordinate estimates are

R^k=1\|𝒫\|∑\(ξ,a\)∈𝒫𝟏\{πkTμ^\(ξ,a\)≠πkTμ⋆\(ξ,a\)\},\\widehat\{R\}\_\{k\}=\\frac\{1\}\{\|\\mathcal\{P\}\|\}\\sum\_\{\(\\xi,a\)\\in\\mathcal\{P\}\}\\mathbf\{1\}\\\{\\pi\_\{k\}T\_\{\\widehat\{\\mu\}\}\(\\xi,a\)\\neq\\pi\_\{k\}T\_\{\\mu^\{\\star\}\}\(\\xi,a\)\\\},\(20\)with false\-write risks obtained by replacing the indicator by the corresponding empty\-set event\. Few\-shot renamed adaptation reports

AdaptScore=𝟏​\{Lstab=0,Lshift=0\},\\mathrm\{AdaptScore\}=\\mathbf\{1\}\\\{L\_\{\\mathrm\{stab\}\}=0,\\ L\_\{\\mathrm\{shift\}\}=0\\\},\(21\)so a method must both preserve stable entries and correct shifted ones\. The coordinate risks are empirical counterparts of𝔅j\\mathfrak\{B\}\_\{j\}; they identify which answer coordinate remains unresolved in the frozen state\.

The finite discovery and language\-model tables additionally report the direct target projection

E\(μ\)=\{\(a,y\):∃\(c,v,δ\),\(a,c,y,v,δ\)∈μ\},E\(\\mu\)=\\\{\(a,y\):\\exists\(c,v,\\delta\),\(a,c,y,v,\\delta\)\\in\\mu\\\},\(22\)using set precision, recall, and F1 againstE⁡\(μ⋆\)E\(\\mu^\{\\star\}\)\. This projected F1 is accompanied by context recall and false\-write counts; it is not presented as the exact tuple accuracy in Eq\. \([19](https://arxiv.org/html/2609.30650#S5.E19)\)\. The reported false\-edge count is\|E⁡\(μ^\)∖E⁡\(μ⋆\)\|\|E\(\\widehat\{\\mu\}\)\\setminus E\(\\mu^\{\\star\}\)\|; readout false edges are the subset whose target coordinate is a designated deterministic proxy\.

Figure 2:Fixed\-probe target and adaptation outcomes\.### 5\.1Protocols and Baselines

The suite contains finite worlds, readout and hidden\-context variants, continuous paired interventions, and language\-model proposal policies\. Baselines match the available interface: passive and reward learners, ranking loss, conditional discovery, latent world models, tuple and neural probes, the official TD\-MPC2 world model, ungated proposals, and transfer ablations\. All are evaluated after freezing, using Eqs\. \([19](https://arxiv.org/html/2609.30650#S5.E19)\)–\([22](https://arxiv.org/html/2609.30650#S5.E22)\) or the metric loss in Eq\. \([2](https://arxiv.org/html/2609.30650#S2.E2)\)\.

### 5\.2Coarse Interfaces

The complex noisy\-hidden family realizes positive probe Bayes risk under Theorem[2](https://arxiv.org/html/2609.30650#Thmtheorem2)\. Passive correlation gives target\-edge F10\.0000\.000, while a reward\-optimized learner has target\-edge F10\.3260\.326\. A ranking\-loss predictor receives supervised intervention\-order labels; nevertheless, in the noisy\-hidden ranking task it has target\-edge F10\.0000\.000even though Causal Core reaches0\.7640\.764\. Stronger trajectory\-matched baselines control for output format\. On the expanded matched run, a thresholded conditional\-discovery baseline has target\-edge F10\.640\.64, and its readout\-oracle variant reaches0\.760\.76but still writes hidden labels with false\-positive count1\.501\.50\. The readout\-oracle tuple writer has target\-edge F10\.790\.79but6\.196\.19hidden\-label false positives\. A latent world model reaches next\-bit accuracy0\.9050\.905yet has target\-edge F10\.240\.24; with a readout oracle it reaches0\.320\.32and writes39\.8839\.88hidden\-label false positives\. Causal Core has target\-edge F10\.790\.79with zero readout and hidden\-label false positives\. High task or predictive performance therefore coexists with poor fixed\-probe recovery\.

Under the uniform prior, each exact two\-world separation in Proposition[9](https://arxiv.org/html/2609.30650#Thmtheorem9)has probe Bayes risk1/21/2and is decoder\-independent\. The neural rows support a narrower empirical statement: high transition accuracy does not make the mechanism answer recoverable by the specified frozen\-state decoders, including readout\-oracle variants\. The hidden\-state experiment below adds linear and nonlinear decoders of a public language model rather than inferring omission from its generated answer alone\.

### 5\.3Write Validity and Local Adaptation

Indiscriminate writing raises recall at the cost of validity\. The ungated Qwen explorer writes92\.392\.3false entries and37\.037\.0readout false positives; the context\-search and control\-planner gates reduce these counts to6\.0/0\.06\.0/0\.0and2\.0/0\.02\.0/0\.0, respectively\. In renamed and semantic\-readout adaptation, non\-diagnostic and readout\-unsafe baselines have AdaptScore00, while Causal Core has AdaptScore1\.001\.00\. The evaluation uses the diagnostic selection and local update in Eqs\. \([7](https://arxiv.org/html/2609.30650#S3.E7)\)–\([10](https://arxiv.org/html/2609.30650#S3.E10)\), whose finite\-sample guarantees are Propositions[5](https://arxiv.org/html/2609.30650#Thmtheorem5)and[6](https://arxiv.org/html/2609.30650#Thmtheorem6)\.

### 5\.4Hidden Contexts and Continuous Responses

Hidden gates test the setup part of the interface\. Active setup improves hidden recall from0\.1250\.125after 80 interactions to0\.6250\.625after 160 interactions at zero hidden false positives in the four\-family horizon sweep\. Rare eligibility events leave some hidden contexts outside the observed quotient\. In a continuous metric SCM, global linear regression reaches F10\.6670\.667because it writes a deterministic readout and extends a context\-specific effect to the wrong cell; metric Causal Core has F11\.0001\.000with no readout false positives\.

We next evaluate the official5\.25\.2M\-parameter TD\-MPC2 checkpoint\([Hansen et al\., 2024](https://arxiv.org/html/2609.30650#bib.bib10)\)on cheetah\-run, walker\-walk, and reacher\-easy\. For saved statexx, unit action directionuu, and horizonh∈\{1,2,4\}h\\in\\\{1,2,4\\\}, the evaluator records the continuous response

Δh​\(x,u\)=Oh​\(x,\+ϵ​u\)−Oh​\(x,−ϵ​u\)2​ϵ,ϵ=0\.25\.\\Delta\_\{h\}\(x,u\)=\\frac\{O\_\{h\}\(x,\+\\epsilon u\)\-O\_\{h\}\(x,\-\\epsilon u\)\}\{2\\epsilon\},\\qquad\\epsilon=0\.25\.\(23\)Linear and two\-layer nonlinear decoders receive the frozen source latent state, latent intervention difference,uu, andhh\. Thus a failure is not tied to a linear readout\. The nonlinear decoder has source NRMSE0\.6290\.629and effect\-sign accuracy0\.9190\.919over nine task\-seed runs\.

At transfer, one actuator changes sign and its index varies across runs\. The frozen model’s shifted\-actuator sign accuracy falls to0\.0570\.057\. A global probe\-head update using the target pairs raises it to0\.6940\.694but increases stable\-response NRMSE from0\.6730\.673to1\.4601\.460\. The selective update compares the two actuator hypotheses on five target states per actuator, localizes the shift in all nine runs, and reaches sign accuracy0\.9480\.948\. Its random\-direction NRMSE is0\.6300\.630, equal to the oracle actuator map within reported precision, while its stable\-response NRMSE remains0\.6730\.673\. Appendix[B\.4](https://arxiv.org/html/2609.30650#A2.SS4)gives per\-task results, decoder training, and diagnostic\-margin sensitivity\.

Figure 3:Distributional TD\-MPC2 responses under one actuator remapping\.
### 5\.5Language\-Model States and Proposal Policies

A frozen\-state experiment tests whether mechanism answers are retained inside Qwen2\.5\-7B\-Instruct\([Yang et al\., 2024](https://arxiv.org/html/2609.30650#bib.bib31)\), rather than judging generated explanations alone\. Logistic probes are fit at five layers on 12 procedural families and then frozen for six disjoint families\. A two\-layer decoder and an RBF decoder test nonlinear accessibility at the final layer\. Action and state names are independently replaced by neutral tokens in every family and condition\.

The final\-layer linear decoder reaches balanced accuracy0\.9580\.958on source mechanisms and0\.8470\.847after renaming, but0\.5830\.583on the changed\-delay subset\. The nonlinear RBF decoder obtains0\.9440\.944,0\.8750\.875, and0\.2500\.250, respectively\. Thus the source result cannot be dismissed as complete absence of mechanism information, while that information does not support the held\-out changed delay\. A paired\-evidence rule recovers every changed delay but accepts0\.9440\.944of synchronized readouts\. Applying the admissible\-write gate preserves changed\-delay accuracy1\.0001\.000and lowers readout FPR to0\.0560\.056\. Table[5](https://arxiv.org/html/2609.30650#A2.T5)gives all controls\.

Figure 4:Qwen hidden\-state retention on held\-out mechanism families\.In the proposal\-policy experiment, the same Qwen model chooses legal interventions and the evidence gate processes the resulting transitions\. Causal prompting alone has target\-edge F10\.1630\.163and32\.032\.0readout false positives\. Gating raises target\-edge F1 to0\.6780\.678–0\.7270\.727and eliminates readout false positives\. Table[6](https://arxiv.org/html/2609.30650#A2.T6)gives the full breakdown\.

Figure[2](https://arxiv.org/html/2609.30650#S5.F2)localizes four distinct errors in the frozen answer map: ranking loses target identity, ungated writing loses validity, source dynamics loses the shifted sign, and prompting loses readout control\. These errors persist under relabeling, hidden gates, or readout corruption even when the corresponding task objective remains useful\.

### 5\.6Reproducibility

Code for generating the SCM families, running the baselines, reproducing the MuJoCo and language\-model experiments, and rebuilding every figure is available at[https://github\.com/ShengjunZhang/mechanism\-memory](https://github.com/ShengjunZhang/mechanism-memory)\. The source\-data tables and raw result files are included with the submission\. Appendix[B](https://arxiv.org/html/2609.30650#A2)gives the random seeds, budgets, model variants, probe definitions, and implementation details used in the reported results\.

## 6Conclusion

Causal retention asks whether a frozen learned state preserves an interventional answer map beyond what task success requires\. Probe Bayes risk, posterior retest coverage, and the exact edit decomposition connect interface identification to selective adaptation without assuming a particular decoder\. Finite SCMs, continuous control, a pretrained world model, and frozen Qwen states expose target, context, value, and delay errors despite useful task or source\-decoding performance\. The continuous results concern a fixed distribution of state, intervention, and horizon queries rather than recovery of an unrestricted latent SCM\.

## Appendix AAuxiliary Theoretical Results

The finite constructions below establish the interface separations and the extensions used by the selective\-adaptation analysis\. All logarithms are natural\.

### A\.1Coarse Behavioral and Predictive Interfaces

###### Proposition 9\.

There are finite deterministic structural causal model classes in which reward, optimal\-policy, intervention\-ranking, graph, synchronized\-readout, and passive hidden\-context interfaces are strictly coarser than the exact mechanism\-probe interface\. For each interface, a point\-mass probe distribution has minimax mechanism risk at least1/21/2\.

###### Proof\.

Each construction consists of two models with the same coarse interface and different answers to one mechanism probe\. The lower bound then follows from Theorem[2](https://arxiv.org/html/2609.30650#Thmtheorem2)withρ=1\\rho=1\.

For reward and optimal\-policy interfaces, letU,Y∈\{0,1\}U,Y\\in\\\{0,1\\\}and letaabe the only nontrivial action\. InM0M\_\{0\},aadirectly setsY←1Y\\leftarrow 1and leavesUUunchanged\. InM1M\_\{1\},aadirectly setsU←1U\\leftarrow 1and the structural equation isY←UY\\leftarrow U\. InitializeU=Y=0U=Y=0and let reward beYY\. Every policy induces the same reward sequence in the two models, and the optimal value and optimal\-policy set agree\. The direct target ofaaisYYinM0M\_\{0\}andUUinM1M\_\{1\}\.

For intervention ranking, add a null actionbbwith utility zero\. In both models above,aayields utility one andbbyields utility zero\. Hence every pairwise ranking label is identical, while the direct target of the preferred action still differs\.

For a graph interface, use the same graphA→YA\\to Yin both models and letd​o​\(A=1\)do\(A=1\)setY←1Y\\leftarrow 1inM0M\_\{0\}andY←0Y\\leftarrow 0inM1M\_\{1\}from a common baselineY=1/2Y=1/2\. The graph is identical, whereas the value coordinate of the mechanism answer differs\. A delay or context coordinate can be varied in the same way without changing the graph\.

For a synchronized\-readout interface, letRRbe an observed proxy\. InM0M\_\{0\},aadirectly setsY←1Y\\leftarrow 1andR←YR\\leftarrow Y; inM1M\_\{1\},aadirectly setsR←1R\\leftarrow 1andY←RY\\leftarrow R\. Both produce the observed pair\(Y,R\)=\(1,1\)\(Y,R\)=\(1,1\)afteraa, but the direct target isYYinM0M\_\{0\}andRRinM1M\_\{1\}\.

For passive hidden\-context observation, fixp∈\(0,1\)p\\in\(0,1\)\. InM0M\_\{0\}, actionaasetsY←1Y\\leftarrow 1with probabilityppindependently on each trial\. InM1M\_\{1\}, a latentH∼Bernoulli⁡\(p\)H\\sim\\mathrm\{Bernoulli\}\(p\)is drawn before each trial andaadeterministically setsY←1Y\\leftarrow 1if and only ifH=1H=1\. The passive law of\(a,Y\)\(a,Y\)is Bernoulli with parameterppin both models, while onlyM1M\_\{1\}has a hidden eligibility context\. These pairs establish all claimed strict coarsenings\. ∎

The proposition does not assert that reward, graph estimation, or ranking is uninformative\. It states that these population objects need not determine all coordinates of the mechanism answer\. Additional intervention contrasts can refine the interface and remove a particular ambiguity\.

### A\.2Finite\-Sample Contrast Gates

The constructive results use paired intervention contrasts rather than raw action correlations\. The following calculation supplies the common finite\-sample term for the target, context, and temporal gates\.

###### Lemma 10\.

Foru∈\{0,1\}u\\in\\\{0,1\\\}, letY1u,…,YnuuY\_\{1\}^\{u\},\\ldots,Y\_\{n\_\{u\}\}^\{u\}be independent random variables in\[0,1\]\[0,1\]with meanspup\_\{u\}\. Define

Δ^=1n1​∑r=1n1Yr1−1n0​∑r=1n0Yr0,Δ=p1−p0\.\\widehat\{\\Delta\}=\\frac\{1\}\{n\_\{1\}\}\\sum\_\{r=1\}^\{n\_\{1\}\}Y\_\{r\}^\{1\}\-\\frac\{1\}\{n\_\{0\}\}\\sum\_\{r=1\}^\{n\_\{0\}\}Y\_\{r\}^\{0\},\\qquad\\Delta=p\_\{1\}\-p\_\{0\}\.Then, for everyx\>0x\>0,

ℙ\{\|Δ^−Δ\|≥x\}≤2exp\{−2​x2n1−1\+n0−1\}\.\\mathbb\{P\}\\\{\|\\widehat\{\\Delta\}\-\\Delta\|\\geq x\\\}\\leq 2\\exp\\\!\\left\\\{\-\\frac\{2x^\{2\}\}\{n\_\{1\}^\{\-1\}\+n\_\{0\}^\{\-1\}\}\\right\\\}\.If a gate accepts whenΔ^≥λ\\widehat\{\\Delta\}\\geq\\lambdaand every population contrast is at leastλ\+γ\\lambda\+\\gammaor at mostλ−γ\\lambda\-\\gamma, the probability of any error amongKKcandidates is at most

2​K​exp⁡\{−2​γ2n1−1\+n0−1\}\.2K\\exp\\\!\\left\\\{\-\\frac\{2\\gamma^\{2\}\}\{n\_\{1\}^\{\-1\}\+n\_\{0\}^\{\-1\}\}\\right\\\}\.For a balanced designn0=n1=mn\_\{0\}=n\_\{1\}=m, this becomes2​K​e−m​γ22Ke^\{\-m\\gamma^\{2\}\}\.

###### Proof\.

Center all observations and write

Δ^−Δ=∑r=1n1Yr1−p1n1−∑r=1n0Yr0−p0n0\.\\widehat\{\\Delta\}\-\\Delta=\\sum\_\{r=1\}^\{n\_\{1\}\}\\frac\{Y\_\{r\}^\{1\}\-p\_\{1\}\}\{n\_\{1\}\}\-\\sum\_\{r=1\}^\{n\_\{0\}\}\\frac\{Y\_\{r\}^\{0\}\-p\_\{0\}\}\{n\_\{0\}\}\.Each summand in the first sum has an interval of possible values of length1/n11/n\_\{1\}, and each summand in the second has interval length1/n01/n\_\{0\}\. Hence the sum of squared interval lengths is

n1​\(1n1\)2\+n0​\(1n0\)2=1n1\+1n0\.n\_\{1\}\\left\(\\frac\{1\}\{n\_\{1\}\}\\right\)^\{2\}\+n\_\{0\}\\left\(\\frac\{1\}\{n\_\{0\}\}\\right\)^\{2\}=\\frac\{1\}\{n\_\{1\}\}\+\\frac\{1\}\{n\_\{0\}\}\.Hoeffding’s inequality\([Hoeffding, 1963](https://arxiv.org/html/2609.30650#bib.bib12)\)gives the stated exponential bound for each tail; adding the two tails proves the first display\. A population contrast on either side of the threshold can be misclassified only if\|Δ^−Δ\|≥γ\|\\widehat\{\\Delta\}\-\\Delta\|\\geq\\gamma\. Applying the first bound and then a union bound overKKcandidates proves the remaining claims\. ∎

The independent\-pair calculation extends to adaptive interaction when action probabilities are logged\.

###### Lemma 11\.

Let\(ℱt\)t=0N\(\\mathcal\{F\}\_\{t\}\)\_\{t=0\}^\{N\}be the interaction filtration\. At roundtt, letYt​\(1\),Yt​\(0\)∈\[0,1\]Y\_\{t\}\(1\),Y\_\{t\}\(0\)\\in\[0,1\]be potential responses and suppose

At⟂⟂\(Yt​\(1\),Yt​\(0\)\)\|ℱt−1,et=ℙ⁡\(At=1∣ℱt−1\)∈\[ϵ,1−ϵ\]A\_\{t\}\\perp\\\!\\\!\\\!\\perp\(Y\_\{t\}\(1\),Y\_\{t\}\(0\)\)\\mid\\mathcal\{F\}\_\{t\-1\},\\qquad e\_\{t\}=\\mathbb\{P\}\(A\_\{t\}=1\\mid\\mathcal\{F\}\_\{t\-1\}\)\\in\[\\epsilon,1\-\\epsilon\]almost surely for some0<ϵ≤1/20<\\epsilon\\leq 1/2\. ObserveYt=At​Yt​\(1\)\+\(1−At\)​Yt​\(0\)Y\_\{t\}=A\_\{t\}Y\_\{t\}\(1\)\+\(1\-A\_\{t\}\)Y\_\{t\}\(0\)and define

Δ^N=1N​∑t=1N\{At​Ytet−\(1−At\)​Yt1−et\},\\widehat\{\\Delta\}\_\{N\}=\\frac\{1\}\{N\}\\sum\_\{t=1\}^\{N\}\\left\\\{\\frac\{A\_\{t\}Y\_\{t\}\}\{e\_\{t\}\}\-\\frac\{\(1\-A\_\{t\}\)Y\_\{t\}\}\{1\-e\_\{t\}\}\\right\\\},Δ¯N=1N​∑t=1N𝔼⁡\[Yt​\(1\)−Yt​\(0\)∣ℱt−1\]\.\\overline\{\\Delta\}\_\{N\}=\\frac\{1\}\{N\}\\sum\_\{t=1\}^\{N\}\\mathbb\{E\}\[Y\_\{t\}\(1\)\-Y\_\{t\}\(0\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\.Then, for everyx\>0x\>0,

ℙ\{\|Δ^N−Δ¯N\|≥x\}≤2exp\{−Nϵ2x2/2\}\.\\mathbb\{P\}\\\{\|\\widehat\{\\Delta\}\_\{N\}\-\\overline\{\\Delta\}\_\{N\}\|\\geq x\\\}\\leq 2\\exp\\\{\-N\\epsilon^\{2\}x^\{2\}/2\\\}\.Consequently,KKadaptive gates whose average contrasts are separated from their thresholds byγ\\gammahave joint error probability at most2Kexp\{−Nϵ2γ2/2\}2K\\exp\\\{\-N\\epsilon^\{2\}\\gamma^\{2\}/2\\\}\.

###### Proof\.

Set

Vt=At​Ytet−\(1−At\)​Yt1−et,dt=𝔼⁡\[Yt​\(1\)−Yt​\(0\)∣ℱt−1\]\.V\_\{t\}=\\frac\{A\_\{t\}Y\_\{t\}\}\{e\_\{t\}\}\-\\frac\{\(1\-A\_\{t\}\)Y\_\{t\}\}\{1\-e\_\{t\}\},\\qquad d\_\{t\}=\\mathbb\{E\}\[Y\_\{t\}\(1\)\-Y\_\{t\}\(0\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\.Consistency and sequential randomization imply

𝔼⁡\[Vt∣ℱt−1\]\\displaystyle\\mathbb\{E\}\[V\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\]=𝔼⁡\[At​Yt​\(1\)∣ℱt−1\]et−𝔼⁡\[\(1−At\)​Yt​\(0\)∣ℱt−1\]1−et\\displaystyle=\\frac\{\\mathbb\{E\}\[A\_\{t\}Y\_\{t\}\(1\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\}\{e\_\{t\}\}\-\\frac\{\\mathbb\{E\}\[\(1\-A\_\{t\}\)Y\_\{t\}\(0\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\}\{1\-e\_\{t\}\}=𝔼⁡\[Yt​\(1\)∣ℱt−1\]−𝔼⁡\[Yt​\(0\)∣ℱt−1\]=dt\.\\displaystyle=\\mathbb\{E\}\[Y\_\{t\}\(1\)\\mid\\mathcal\{F\}\_\{t\-1\}\]\-\\mathbb\{E\}\[Y\_\{t\}\(0\)\\mid\\mathcal\{F\}\_\{t\-1\}\]=d\_\{t\}\.ThusDt=Vt−dtD\_\{t\}=V\_\{t\}\-d\_\{t\}is a martingale difference\. Conditional onℱt−1\\mathcal\{F\}\_\{t\-1\},VtV\_\{t\}lies between−1/\(1−et\)\-1/\(1\-e\_\{t\}\)and1/et1/e\_\{t\}; subtractingdtd\_\{t\}does not change the interval width, and

1et\+11−et≤2ϵ\.\\frac\{1\}\{e\_\{t\}\}\+\\frac\{1\}\{1\-e\_\{t\}\}\\leq\\frac\{2\}\{\\epsilon\}\.Hoeffding’s conditional lemma therefore gives

𝔼⁡\[eλ​Dt∣ℱt−1\]≤exp⁡\{λ2/\(2​ϵ2\)\}\.\\mathbb\{E\}\[e^\{\\lambda D\_\{t\}\}\\mid\\mathcal\{F\}\_\{t\-1\}\]\\leq\\exp\\\{\\lambda^\{2\}/\(2\\epsilon^\{2\}\)\\\}\.Iterating conditional expectation yields

𝔼​exp⁡\{λ​∑t=1NDt\}≤exp⁡\{N​λ2/\(2​ϵ2\)\}\.\\mathbb\{E\}\\exp\\\!\\left\\\{\\lambda\\sum\_\{t=1\}^\{N\}D\_\{t\}\\right\\\}\\leq\\exp\\\{N\\lambda^\{2\}/\(2\\epsilon^\{2\}\)\\\}\.Forλ\>0\\lambda\>0, Markov’s inequality gives

ℙ\{∑t=1NDt≥Nx\}≤exp\{−λNx\+Nλ2/\(2ϵ2\)\}\.\\mathbb\{P\}\\\!\\left\\\{\\sum\_\{t=1\}^\{N\}D\_\{t\}\\geq Nx\\right\\\}\\leq\\exp\\\{\-\\lambda Nx\+N\\lambda^\{2\}/\(2\\epsilon^\{2\}\)\\\}\.The minimizing value isλ=ϵ2​x\\lambda=\\epsilon^\{2\}x, which givesexp\{−Nϵ2x2/2\}\\exp\\\{\-N\\epsilon^\{2\}x^\{2\}/2\\\}\. Applying the same argument to−Dt\-D\_\{t\}and adding the tails proves the first claim\. The gate bound follows by settingx=γx=\\gammaand taking a union bound overKKcandidates\. ∎

###### Proposition 12\.

Letℱ\\mathcal\{F\}be a finite class of candidate readout formulas and letA\(f\)=ℙ\{f\(X\)=R\}A\(f\)=\\mathbb\{P\}\\\{f\(X\)=R\\\}\. Frommmindependent observations defineA^m\(f\)=m−1∑r=1m𝟏\{f\(Xr\)=Rr\}\\widehat\{A\}\_\{m\}\(f\)=m^\{\-1\}\\sum\_\{r=1\}^\{m\}\\mathbf\{1\}\\\{f\(X\_\{r\}\)=R\_\{r\}\\\}\. The rule declaresRRa formula readout whenmaxf∈ℱ⁡A^m​\(f\)≥c\\max\_\{f\\in\\mathcal\{F\}\}\\widehat\{A\}\_\{m\}\(f\)\\geq c\. If somef⋆f^\{\\star\}satisfiesA⁡\(f⋆\)≥c\+γA\(f^\{\\star\}\)\\geq c\+\\gamma, its false\-negative probability is at moste−2​m​γ2e^\{\-2m\\gamma^\{2\}\}\. If everyffsatisfiesA⁡\(f\)≤c−γA\(f\)\\leq c\-\\gamma, the false\-positive probability is at most\|ℱ\|​e−2​m​γ2\|\\mathcal\{F\}\|e^\{\-2m\\gamma^\{2\}\}\.

###### Proof\.

In the first case, a false negative impliesA^m​\(f⋆\)<c\\widehat\{A\}\_\{m\}\(f^\{\\star\}\)<c\. Therefore

ℙ​\{false negative\}\\displaystyle\\mathbb\{P\}\\\{\\text\{false negative\}\\\}≤ℙ\{A^m\(f⋆\)−A\(f⋆\)≤−γ\}\\displaystyle\\leq\\mathbb\{P\}\\\{\\widehat\{A\}\_\{m\}\(f^\{\\star\}\)\-A\(f^\{\\star\}\)\\leq\-\\gamma\\\}≤e−2​m​γ2\.\\displaystyle\\leq e^\{\-2m\\gamma^\{2\}\}\.In the second case, a false positive implies that at least onef∈ℱf\\in\\mathcal\{F\}satisfiesA^m​\(f\)−A⁡\(f\)≥γ\\widehat\{A\}\_\{m\}\(f\)\-A\(f\)\\geq\\gamma\. Hoeffding’s inequality gives probability at moste−2​m​γ2e^\{\-2m\\gamma^\{2\}\}for each fixedff; a union bound overℱ\\mathcal\{F\}proves the result\. ∎

### A\.3Hidden\-Context Acquisition

The no\-signal bound is exact for a uniformly hidden label\. IfB∼Unif⁡\(\{0,1\}m\)B\\sim\\operatorname\{Unif\}\(\\\{0,1\\\}^\{m\}\)andI⁡\(B,S\)=0I\(B;S\)=0, thenBBremains uniform conditional onSS\. Consequently every deterministic or randomized decoderB^\\widehat\{B\}based onSSsatisfies

𝔼dH\(B^,B\)=m/2,ℙ\{B^=B\}=2−m\.\\mathbb\{E\}d\_\{\\mathrm\{H\}\}\(\\widehat\{B\},B\)=m/2,\\qquad\\mathbb\{P\}\\\{\\widehat\{B\}=B\\\}=2^\{\-m\}\.Active setup changes the rate at which informative eligible outcomes arrive\.

###### Proposition 13\.

Fix a hidden\-gated candidate\. Suppose a setup policy reaches an eligible state within at mosthhactions with probability at leastβ\\beta, independently across reset attempts\. To collectmmeligible process outcomes with probability at least1−δ1\-\\delta, it is sufficient to run

N≥8β​\(m\+log⁡1δ\)N\\geq\\frac\{8\}\{\\beta\}\\left\(m\+\\log\\frac\{1\}\{\\delta\}\\right\)setup attempts, for total action cost at mostN⁡\(h\+1\)N\(h\+1\)\. Passive access is the same calculation with the ambient eligibility probabilityqqin place ofβ\\beta; its expected number of attempts ism/qm/q\.

###### Proof\.

LetKNK\_\{N\}be the number of successful setup attempts\. It stochastically dominates aBinomial⁡\(N,β\)\\mathrm\{Binomial\}\(N,\\beta\)variable with meanν=N​β\\nu=N\\beta\. The stated condition implies

ν≥8​\(m\+log⁡1δ\),m≤ν/2\.\\nu\\geq 8\\left\(m\+\\log\\frac\{1\}\{\\delta\}\\right\),\\qquad m\\leq\\nu/2\.The multiplicative Chernoff lower\-tail inequality\([Boucheron et al\., 2013](https://arxiv.org/html/2609.30650#bib.bib3)\)therefore gives

ℙ⁡\(KN<m\)\\displaystyle\\mathbb\{P\}\(K\_\{N\}<m\)≤ℙ\(KN≤ν/2\)≤e−ν/8\\displaystyle\\leq\\mathbb\{P\}\(K\_\{N\}\\leq\\nu/2\)\\leq e^\{\-\\nu/8\}≤e−m−log⁡\(1/δ\)=e−m​δ≤δ\.\\displaystyle\\leq e^\{\-m\-\\log\(1/\\delta\)\}=e^\{\-m\}\\delta\\leq\\delta\.Each successful attempt contributes one process outcome and uses at mosth\+1h\+1actions\. Under passive access, eligible arrivals are Bernoulli with parameterqq, so the waiting time formmarrivals is negative binomial with expectationm/qm/q; the same Chernoff calculation gives the high\-probability1/q1/qfactor\. ∎

### A\.4Randomized Lists and Adaptive Stopping

LetRRbe internal randomness independent of\(B,Z\)\(B,Z\)\. Conditional onZ=zZ=z, the coverage ofℒ⁡\(z,R\)\\mathcal\{L\}\(z,R\)is a convex combination of the coverage values of deterministic lists and is therefore at mostΓk​\(z\)\\Gamma\_\{k\}\(z\)\. An adaptive procedure that stops after retesting at mostkkdistinct entries induces the list of all entries it examined\. If its final memory is exact, this list must contain every shifted entry unless an untested entry was supplied by an external oracle\. Hence its success probability is at most𝔼​Γk​\(Z\)\\mathbb\{E\}\\Gamma\_\{k\}\(Z\)as well\. Adaptive ordering can reduce the number of tests on favorable paths, but cannot increase the coverage permitted by a hard cap ofkkdistinct entries\.

### A\.5Diagnostic Localization Lower Bound

Consider the symmetric one\-shift family in Proposition[5](https://arxiv.org/html/2609.30650#Thmtheorem5)\. LetJJbe uniform on\[n\]\[n\]\. Conditional onJ=jJ=j, themmsamples at coordinatejjareBernoulli⁡\(1/2\+Δ/2\)\\mathrm\{Bernoulli\}\(1/2\+\\Delta/2\)and those at every other coordinate areBernoulli⁡\(1/2−Δ/2\)\\mathrm\{Bernoulli\}\(1/2\-\\Delta/2\)\. Denote the joint law byPjP\_\{j\}\.

Foru=\(1\+Δ\)/2u=\(1\+\\Delta\)/2andv=\(1−Δ\)/2v=\(1\-\\Delta\)/2,

kl⁡\(u,v\)=Δ​log⁡1\+Δ1−Δ\.\\operatorname\{kl\}\(u,v\)=\\Delta\\log\\frac\{1\+\\Delta\}\{1\-\\Delta\}\.When0<Δ≤1/20<\\Delta\\leq 1/2,

log⁡1\+Δ1−Δ=log⁡\(1\+Δ\)−log⁡\(1−Δ\)≤Δ\+Δ1−Δ≤3​Δ,\\log\\frac\{1\+\\Delta\}\{1\-\\Delta\}=\\log\(1\+\\Delta\)\-\\log\(1\-\\Delta\)\\leq\\Delta\+\\frac\{\\Delta\}\{1\-\\Delta\}\\leq 3\\Delta,so both directed Bernoulli divergences are at most4​Δ24\\Delta^\{2\}\. Forj≠kj\\neq k,PjP\_\{j\}andPkP\_\{k\}differ only at coordinatesjjandkk; product additivity therefore gives

KL\(Pj∥Pk\)≤8mΔ2\.\\mathrm\{KL\}\(P\_\{j\}\\\|P\_\{k\}\)\\leq 8m\\Delta^\{2\}\.IfP∘=n−1​∑k=1nPkP\_\{\\circ\}=n^\{\-1\}\\sum\_\{k=1\}^\{n\}P\_\{k\}, convexity in the second argument yields

I⁡\(J,X\)\\displaystyle I\(J;X\)=1n∑j=1nKL\(Pj∥P∘\)\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{j=1\}^\{n\}\\mathrm\{KL\}\(P\_\{j\}\\\|P\_\{\\circ\}\)≤1n2∑j,k=1nKL\(Pj∥Pk\)≤8mΔ2\.\\displaystyle\\leq\\frac\{1\}\{n^\{2\}\}\\sum\_\{j,k=1\}^\{n\}\\mathrm\{KL\}\(P\_\{j\}\\\|P\_\{k\}\)\\leq 8m\\Delta^\{2\}\.Fano’s inequality\([Cover and Thomas, 2006](https://arxiv.org/html/2609.30650#bib.bib6)\)now implies

ℙ⁡\(J^≠J\)≥1−8​m​Δ2\+log⁡2log⁡n\.\\mathbb\{P\}\(\\widehat\{J\}\\neq J\)\\geq 1\-\\frac\{8m\\Delta^\{2\}\+\\log 2\}\{\\log n\}\.Requiring this error to be at mostε\\varepsilonand rearranging proves

m≥\(1−ε\)​log⁡n−log⁡28​Δ2\.m\\geq\\frac\{\(1\-\\varepsilon\)\\log n\-\\log 2\}\{8\\Delta^\{2\}\}\.

### A\.6Top\-ssDiagnostic Selection

Suppose\|B\|=s\|B\|=s, shifted entries have means at leastpp, stable entries have means at mostq<pq<p, andΔ=p−q\\Delta=p\-q\. LetB^s\\widehat\{B\}\_\{s\}contain thesslargest empirical means\. If

minj∈B⁡X¯j\>p\+q2andmaxj∉B⁡X¯j<p\+q2,\\min\_\{j\\in B\}\\overline\{X\}\_\{j\}\>\\frac\{p\+q\}\{2\}\\quad\\text\{and\}\\quad\\max\_\{j\\notin B\}\\overline\{X\}\_\{j\}<\\frac\{p\+q\}\{2\},thenB^s=B\\widehat\{B\}\_\{s\}=B\. Hoeffding’s inequality and a union bound give

ℙ\{B^s≠B\}≤nexp\(−mΔ2/2\)\.\\mathbb\{P\}\\\{\\widehat\{B\}\_\{s\}\\neq B\\\}\\leq n\\exp\(\-m\\Delta^\{2\}/2\)\.Consequentlym≥2​Δ−2​log⁡\(n/δ\)m\\geq 2\\Delta^\{\-2\}\\log\(n/\\delta\)is sufficient for exact top\-ssrecovery with probability at least1−δ1\-\\delta\.

### A\.7Coordinate Loss and End\-to\-End Sample Composition

Letej,ke\_\{j,k\}indicate an error in coordinatek∈\{ctx,tar,val,del\}k\\in\\\{\\mathrm\{ctx\},\\mathrm\{tar\},\\mathrm\{val\},\\mathrm\{del\}\\\}of entryjj\. Exact entry error satisfies

maxkej,k≤𝟏\{θ^jt≠θjt\}≤∑kej,k\.\\max\_\{k\}e\_\{j,k\}\\leq\\mathbf\{1\}\\\{\\widehat\{\\theta\}\_\{j\}^\{t\}\\neq\\theta\_\{j\}^\{t\}\\\}\\leq\\sum\_\{k\}e\_\{j,k\}\.Summing first overSSandBBand then dividing bynnyields

Ladapt=\|S\|n​Lstab\+\|B\|n​Lshift,L\_\{\\mathrm\{adapt\}\}=\\frac\{\|S\|\}\{n\}L\_\{\\mathrm\{stab\}\}\+\\frac\{\|B\|\}\{n\}L\_\{\\mathrm\{shift\}\},and bounds each term by its coordinate losses\. LetFtar,Fctx,Fro,FtimeF\_\{\\mathrm\{tar\}\},F\_\{\\mathrm\{ctx\}\},F\_\{\\mathrm\{ro\}\},F\_\{\\mathrm\{time\}\}, andFhidF\_\{\\mathrm\{hid\}\}denote, respectively, a target\-contrast error, a visible context error, a readout\-classification error, a delay error, and a hidden setup or hidden\-label error in the source memory\. Then

Fsrc⊆Ftar∪Fctx∪Fro∪Ftime∪Fhid\.F\_\{\\mathrm\{src\}\}\\subseteq F\_\{\\mathrm\{tar\}\}\\cup F\_\{\\mathrm\{ctx\}\}\\cup F\_\{\\mathrm\{ro\}\}\\cup F\_\{\\mathrm\{time\}\}\\cup F\_\{\\mathrm\{hid\}\}\.Diagnostic selection controls whether the appropriate target entry is tested; local retesting controls its replacement coordinates\. This is the component factorization used by Theorem[7](https://arxiv.org/html/2609.30650#Thmtheorem7)\.

For balanced paired designs, Lemma[10](https://arxiv.org/html/2609.30650#Thmtheorem10)gives

ℙ⁡\(Ftar\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{tar\}\}\)≤2​Ktar​e−mtar​γtar2,\\displaystyle\\leq 2K\_\{\\mathrm\{tar\}\}e^\{\-m\_\{\\mathrm\{tar\}\}\\gamma\_\{\\mathrm\{tar\}\}^\{2\}\},ℙ⁡\(Fctx\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{ctx\}\}\)≤2​Kctx​e−mctx​γctx2,\\displaystyle\\leq 2K\_\{\\mathrm\{ctx\}\}e^\{\-m\_\{\\mathrm\{ctx\}\}\\gamma\_\{\\mathrm\{ctx\}\}^\{2\}\},ℙ⁡\(Ftime\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{time\}\}\)≤2​Ktime​e−mtime​γtime2\.\\displaystyle\\leq 2K\_\{\\mathrm\{time\}\}e^\{\-m\_\{\\mathrm\{time\}\}\\gamma\_\{\\mathrm\{time\}\}^\{2\}\}\.Proposition[12](https://arxiv.org/html/2609.30650#Thmtheorem12)and the hidden\-label threshold test give the following bounds, whereFmaxF\_\{\\max\}is the largest readout formula\-class size:

ℙ⁡\(Fro\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{ro\}\}\)≤Kro​\(Fmax\+1\)​e−2​mro​γro2,\\displaystyle\\leq K\_\{\\mathrm\{ro\}\}\(F\_\{\\max\}\+1\)e^\{\-2m\_\{\\mathrm\{ro\}\}\\gamma\_\{\\mathrm\{ro\}\}^\{2\}\},ℙ⁡\(Fhid\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{hid\}\}\)≤δset\+2​Khid​e−mhid​γhid2,\\displaystyle\\leq\\delta\_\{\\mathrm\{set\}\}\+2K\_\{\\mathrm\{hid\}\}e^\{\-m\_\{\\mathrm\{hid\}\}\\gamma\_\{\\mathrm\{hid\}\}^\{2\}\},where Proposition[13](https://arxiv.org/html/2609.30650#Thmtheorem13)supplies an attempted\-action budget forδset\\delta\_\{\\mathrm\{set\}\}\. Diagnostic localization and local retesting obey

ℙ⁡\(Fdiag\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{diag\}\}\)≤ne−mdiagΔ2/2,\\displaystyle\\leq ne^\{\-m\_\{\\mathrm\{diag\}\}\\Delta^\{2\}/2\},ℙ⁡\(Fupd\)\\displaystyle\\mathbb\{P\}\(F\_\{\\mathrm\{upd\}\}\)≤∑j∈B\|𝒩j\|e−rjηj2/2\.\\displaystyle\\leq\\sum\_\{j\\in B\}\|\\mathcal\{N\}\_\{j\}\|e^\{\-r\_\{j\}\\eta\_\{j\}^\{2\}/2\}\.For example, the sufficient allocations

mg\\displaystyle m\_\{g\}≥γg−2​log⁡2​Kgδg,\\displaystyle\\geq\\gamma\_\{g\}^\{\-2\}\\log\\frac\{2K\_\{g\}\}\{\\delta\_\{g\}\},g∈\{tar,ctx,time\},\\displaystyle g\\in\\\{\\mathrm\{tar\},\\mathrm\{ctx\},\\mathrm\{time\}\\\},mro\\displaystyle m\_\{\\mathrm\{ro\}\}≥12​γro2​log⁡Kro​\(Fmax\+1\)δro,\\displaystyle\\geq\\frac\{1\}\{2\\gamma\_\{\\mathrm\{ro\}\}^\{2\}\}\\log\\frac\{K\_\{\\mathrm\{ro\}\}\(F\_\{\\max\}\+1\)\}\{\\delta\_\{\\mathrm\{ro\}\}\},mhid\\displaystyle m\_\{\\mathrm\{hid\}\}≥1γhid2​log⁡2​Khidδlabel,\\displaystyle\\geq\\frac\{1\}\{\\gamma\_\{\\mathrm\{hid\}\}^\{2\}\}\\log\\frac\{2K\_\{\\mathrm\{hid\}\}\}\{\\delta\_\{\\mathrm\{label\}\}\},mdiag\\displaystyle m\_\{\\mathrm\{diag\}\}≥2Δ2​log⁡nδdiag,\\displaystyle\\geq\\frac\{2\}\{\\Delta^\{2\}\}\\log\\frac\{n\}\{\\delta\_\{\\mathrm\{diag\}\}\},rj\\displaystyle r\_\{j\}≥2ηj2​log⁡\|B\|​\|𝒩j\|δupd\\displaystyle\\geq\\frac\{2\}\{\\eta\_\{j\}^\{2\}\}\\log\\frac\{\|B\|\\,\|\\mathcal\{N\}\_\{j\}\|\}\{\\delta\_\{\\mathrm\{upd\}\}\}make the corresponding terms no larger than their assigned failure budgets\. Set

δsrc=δtar\+δctx\+δro\+δtime\+δhid,δhid=δset\+δlabel\.\\delta\_\{\\mathrm\{src\}\}=\\delta\_\{\\mathrm\{tar\}\}\+\\delta\_\{\\mathrm\{ctx\}\}\+\\delta\_\{\\mathrm\{ro\}\}\+\\delta\_\{\\mathrm\{time\}\}\+\\delta\_\{\\mathrm\{hid\}\},\\qquad\\delta\_\{\\mathrm\{hid\}\}=\\delta\_\{\\mathrm\{set\}\}\+\\delta\_\{\\mathrm\{label\}\}\.Ifδsrc\+δϕ\+δdiag\+δupd≤δ\\delta\_\{\\mathrm\{src\}\}\+\\delta\_\{\\phi\}\+\\delta\_\{\\mathrm\{diag\}\}\+\\delta\_\{\\mathrm\{upd\}\}\\leq\\delta, substitution into Theorem[7](https://arxiv.org/html/2609.30650#Thmtheorem7)yieldsℙ⁡\(Ladapt\>0\)≤δ\\mathbb\{P\}\(L\_\{\\mathrm\{adapt\}\}\>0\)\\leq\\delta\.

## Appendix BExperimental Protocols and Results

### B\.1Fixed\-Memory Evaluation

During interaction, each method may use its allowed internal state to choose actions\. Learning then stops\. The evaluator queries only the frozen answer mapTμ^T\_\{\\widehat\{\\mu\}\}on held\-out probes and compares it with the SCM answerGMG\_\{M\}\. Ground\-truth mechanisms are used to construct probes and scores after the run; they are never supplied to the writer during interaction\.

Exact mechanism accuracy uses tuple matches\. The finite discovery and language tables report target\-edge precision, recall, and F1 from Eq\. \([22](https://arxiv.org/html/2609.30650#S5.E22)\), while context, value, delay, and readout coordinates are probed separately\. The FP column counts\|E⁡\(μ^\)∖E⁡\(μ⋆\)\|\|E\(\\widehat\{\\mu\}\)\\setminus E\(\\mu^\{\\star\}\)\|\. Readout false positives are the subset whose direct\-target field is a designated sensor or deterministic proxy\. Hidden recall counts true hidden\-gated mechanisms whose context type is retained\. AdaptScore equals one only when all shifted entries are corrected and all stable entries are preserved\.

### B\.2Baselines

The reward learner stores effects associated with high\-return actions\. Passive correlation stores observed changes without action contrasts\. Conditional discovery thresholds conditional transition differences; its readout\-oracle variant removes known readouts before scoring\. Tuple lift converts discovered dependencies into the full output schema, and its oracle variant receives the same readout information\. The latent world model and neural transition probe fit source transitions and are decoded only after their states are frozen\. TD\-MPC2 uses the released MT30 checkpoint; linear and nonlinear response decoders are fitted to source intervention pairs and frozen before the actuator remapping\. Its global and selective updates receive identical target pairs\. The ranking predictor minimizes pairwise intervention\-ranking loss\. Ungated writers admit every proposed tuple; gated variants use the same proposals but apply Eqs\. \([5](https://arxiv.org/html/2609.30650#S3.E5)\) and \([6](https://arxiv.org/html/2609.30650#S3.E6)\)\. These controls separate objective and interface effects from output\-format effects\.

Table 1:Matched finite\-SCM baselines\.All rows in Table[1](https://arxiv.org/html/2609.30650#A2.T1)receive the same trajectories\. The latent world model attains next\-bit accuracy0\.9050\.905\(and0\.9070\.907with the readout oracle\), so its probe error is not explained by failed source prediction\. Treating each generated SCM family as the independent unit, the mean target\-edge F1 and family\-clustered 95%ttintervals are0\.7930\.793\[0\.749,0\.8370\.749,0\.837\] for Causal Core,0\.6400\.640\[0\.606,0\.6730\.606,0\.673\] for conditional discovery,0\.7890\.789\[0\.740,0\.8390\.740,0\.839\] for oracle tuple lift,0\.2390\.239\[0\.185,0\.2920\.185,0\.292\] for latent modeling, and0\.3960\.396\[0\.387,0\.4040\.387,0\.404\] for the transition probe\. The intervals are descriptive across the eight generated families; the exact family\-seed values are in the source\-data archive\.

Table 2:Ranking and selective\-adaptation results\.
### B\.3Continuous Metric Cells

The continuous SCM has two physical state coordinates, a two\-dimensional action, a deterministic pressure readout, and a context\-specific vent effect\. The evaluator fixes metric cells before training\. A pulse is correct only when its direct target, sign, and context cell match the reference answer\.

Table 3:Continuous metric\-cell results\.For noise levels0\.02,0\.05,0\.10,0\.150\.02,0\.05,0\.10,0\.15, metric Causal Core retains F11\.0001\.000\. The corresponding random\-correlation F1 scores are0\.402,0\.321,0\.265,0\.2520\.402,0\.321,0\.265,0\.252; global linear scores are0\.667,0\.667,0\.667,0\.6950\.667,0\.667,0\.667,0\.695\.

### B\.4Distributional World\-Model Probes

The public TD\-MPC2 MT30 checkpoint contains5,241,0665\{,\}241\{,\}066parameters and is evaluated without weight updates\. Its SHA\-256 digest, checkpoint metadata, and all per\-run outputs are included with the source data\. The tasks are cheetah\-run, walker\-walk, and reacher\-easy, with three state\-sampling seeds per task\. States are collected under the checkpoint policy with Gaussian action noise\. The evaluator draws unit action directions and horizonsh∈\{1,2,4\}h\\in\\\{1,2,4\\\}, then branches the MuJoCo physics state to compute Eq\. \([23](https://arxiv.org/html/2609.30650#S5.E23)\)\. Each source decoder receives 900 training queries; source and target evaluation use 240 held\-out queries each\.

The linear decoder is ridge regression\. The nonlinear decoder has two 256\-unit GELU layers, uses a disjoint 15% validation split, and stops after 24 epochs without improvement\. Both receive the encoded state, the difference of positive and negative latent rollouts, the action direction, and the horizon\. The normalized response error is

NRMSE=\(∑i‖Δ^i−Δi‖22∑i‖Δi‖22\)1/2\.\\operatorname\{NRMSE\}=\\left\(\\frac\{\\sum\_\{i\}\\\|\\widehat\{\\Delta\}\_\{i\}\-\\Delta\_\{i\}\\\|\_\{2\}^\{2\}\}\{\\sum\_\{i\}\\\|\\Delta\_\{i\}\\\|\_\{2\}^\{2\}\}\\right\)^\{1/2\}\.Effect\-sign accuracy evaluates the sign at the largest\-magnitude coordinate of the true response\.

For seedss, actuator\(s−1\)modda\(s\-1\)\\bmod d\_\{a\}changes sign\. Adaptation uses five states per actuator\. The selective update compares positive and negative actuator hypotheses and changes an entry only when the alternative halves its paired loss\. The global update fits an unrestricted residual probe head to the same target pairs\. Thresholds from0\.450\.45through0\.900\.90all localize the same single actuator in every run\.

Table 4:Distributional intervention responses\.On random target directions, NRMSE is1\.1381\.138for the frozen model,1\.5481\.548for the global update, and0\.6300\.630for the selective update\. On directions that exclude the shifted actuator, the corresponding stable\-response errors are0\.6730\.673,1\.4601\.460, and0\.6730\.673\. The selective update therefore matches the oracle actuator map at the reported precision without changing stable responses\. The fixed finite query support together with a bounded truncation of Euclidean loss instantiates Eq\. \([2](https://arxiv.org/html/2609.30650#S2.E2)\); untruncated NRMSE is reported to expose large errors\. The experiment does not infer an unrestricted latent SCM\.

### B\.5Language\-Model States and Proposal Policies

The hidden\-state experiment presents Qwen2\.5\-7B\-Instruct with a controlled intervention record followed by a binary direct\-target or delay probe\. Twelve procedural families supply 504 balanced training probes; six disjoint families supply 360 balanced test probes\. Each family and condition receives a fresh random permutation of neutral action and state tokens, so a decoder cannot identify a mechanism from a recurring name\. Every record contains two controlled repetitions per action\. Prompt length is 697 tokens on average and 812 at maximum, with no truncation\.

The language model is frozen\. At the answer position, hidden vectors are extracted from layers0,7,14,21,0,7,14,21,and2828\. A logistic decoder is fit at each layer, and a two\-layer multilayer perceptron and an RBF support\-vector decoder are fit to the last layer\. The controls are query\-only TF–IDF, the model’s calibrated Yes/No logit margin, and a paired\-evidence rule\. The gated mechanism state admits a candidate only when the controlled pair supports its lag and unrelated interventions do not support the same target\. All decoders and thresholds are frozen before evaluation on the six test families\. Balanced accuracy is reported for mixed\-label probes; readout FPR is the acceptance rate on synchronized\-readout candidates\. Standard errors in Figure[4](https://arxiv.org/html/2609.30650#S5.F4)use the six test families as clusters\.

Table 5:Frozen language\-model state probes\.The source score of the last\-layer linear decoder shows that substantial mechanism information is present in the frozen hidden state\. Its shifted\-delay score of0\.5830\.583, together with the RBF decoder’s score of0\.2500\.250, shows that source decodability does not yield a stable answer under a held\-out mechanism change\. Paired evidence recovers the changed delay but accepts94\.4%94\.4\\%of synchronized readouts\. The gate retains its perfect shifted\-delay score while reducing that false\-positive rate to5\.6%5\.6\\%\.

Qwen2\.5\-7B\-Instruct is used only to propose the next legal intervention from the current transcript\. Runs use deterministic decoding, 100 interaction steps, and three seeds\. The action is executed in the same complex noisy\-hidden family used by the non\-language comparison\. Explanations and reward predictions receive no probe credit\. Invalid outputs are mapped by the same deterministic action parser in every language condition\.

Table 6:Language\-model proposal policies\.The causal prompt does not remove readout contamination\. The gated variants use the same language model as a proposal source but apply the evidence gate after observing the transition\. Their advantage therefore comes from the memory update, not from treating generated explanations as causal labels\.

### B\.6Candidate Budgets and Implementation

Candidate schedules cap computation but do not enter the probe metric\. Small, default, and wide schedules yield the same target\-edge F1 at 80 and 160 steps \(0\.6980\.698and0\.7630\.763\), zero readout false positives, and hidden recall rising from0\.1250\.125to0\.6250\.625\. All thresholds, target budgets, metric cells, saved MuJoCo states, language prompts, and parsers are fixed before held\-out scoring; malformed language proposals count as failures\.

Finite\-SCM actions are interventions\. The writer estimates Eq\. \([4](https://arxiv.org/html/2609.30650#S3.E4)\) from action\-conditioned lift and repeated context visits\. Because these trajectories are adaptive, the iid bounds in Appendix[A\.2](https://arxiv.org/html/2609.30650#A1.SS2)are not reported as confidence intervals for the empirical tables\. Lemma[11](https://arxiv.org/html/2609.30650#Thmtheorem11)covers randomized adaptive logging; the continuous protocols use the paired design directly\. Candidate caps determine which contrasts are attempted, not how frozen probes are scored\. Numerical tables and raw JSON files are included as source data\.

## References

- Arjovsky et al\. \(2019\)M\. Arjovsky, L\. Bottou, I\. Gulrajani, and D\. Lopez\-Paz\.Invariant risk minimization\.*arXiv preprint arXiv:1907\.02893*, 2019\.
- Bareinboim and Pearl \(2016\)E\. Bareinboim and J\. Pearl\.Causal inference and the data\-fusion problem\.*Proceedings of the National Academy of Sciences*, 113\(27\):7345–7352, 2016\.
- Boucheron et al\. \(2013\)S\. Boucheron, G\. Lugosi, and P\. Massart\.*Concentration Inequalities: A Nonasymptotic Theory of Independence*\.Oxford University Press, 2013\.
- Chernozhukov et al\. \(2018\)V\. Chernozhukov, D\. Chetverikov, M\. Demirer, E\. Duflo, C\. Hansen, W\. Newey, and J\. Robins\.Double/debiased machine learning for treatment and structural parameters\.*The Econometrics Journal*, 21\(1\):C1–C68, 2018\.
- Chickering \(2002\)D\. M\. Chickering\.Optimal structure identification with greedy search\.*Journal of Machine Learning Research*, 3:507–554, 2002\.
- Cover and Thomas \(2006\)T\. M\. Cover and J\. A\. Thomas\.*Elements of Information Theory*\.Wiley\-Interscience, 2 edition, 2006\.
- Gamella et al\. \(2025\)J\. L\. Gamella, S\. Bing, and J\. Runge\.Sanity checking causal representation learning on a simple real\-world system\.In*Proceedings of the 42nd International Conference on Machine Learning*, 2025\.
- Ha and Schmidhuber \(2018\)D\. Ha and J\. Schmidhuber\.World models\.*arXiv preprint arXiv:1803\.10122*, 2018\.
- Hafner et al\. \(2025\)D\. Hafner, J\. Pasukonis, J\. Ba, and T\. Lillicrap\.Mastering diverse control tasks through world models\.*Nature*, 2025\.
- Hansen et al\. \(2024\)N\. Hansen, H\. Su, and X\. Wang\.TD\-MPC2: Scalable, robust world models for continuous control\.In*International Conference on Learning Representations*, 2024\.
- Hauser and Bühlmann \(2012\)A\. Hauser and P\. Bühlmann\.Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs\.*Journal of Machine Learning Research*, 13:2409–2464, 2012\.
- Hoeffding \(1963\)W\. Hoeffding\.Probability inequalities for sums of bounded random variables\.*Journal of the American Statistical Association*, 58\(301\):13–30, 1963\.
- Imbens and Rubin \(2015\)G\. W\. Imbens and D\. B\. Rubin\.*Causal Inference in Statistics, Social, and Biomedical Sciences*\.Cambridge University Press, 2015\.
- Künzel et al\. \(2019\)S\. R\. Künzel, J\. S\. Sekhon, P\. J\. Bickel, and B\. Yu\.Metalearners for estimating heterogeneous treatment effects using machine learning\.*Proceedings of the National Academy of Sciences*, 116\(10\):4156–4165, 2019\.
- Locatello et al\. \(2019\)F\. Locatello, S\. Bauer, M\. Lucic, S\. Gelly, B\. Schölkopf, and O\. Bachem\.Challenging common assumptions in the unsupervised learning of disentangled representations\.*Proceedings of Machine Learning Research*, 97:4114–4124, 2019\.
- Nie and Wager \(2021\)X\. Nie and S\. Wager\.Quasi\-oracle estimation of heterogeneous treatment effects\.*Biometrika*, 108\(2\):299–319, 2021\.
- Pearl \(2009\)J\. Pearl\.*Causality: Models, Reasoning, and Inference*\.Cambridge University Press, 2 edition, 2009\.
- Peters et al\. \(2016\)J\. Peters, P\. Bühlmann, and N\. Meinshausen\.Causal inference by using invariant prediction: Identification and confidence intervals\.*Journal of the Royal Statistical Society: Series B*, 78\(5\):947–1012, 2016\.
- Peters et al\. \(2017\)J\. Peters, D\. Janzing, and B\. Schölkopf\.*Elements of Causal Inference: Foundations and Learning Algorithms*\.MIT Press, 2017\.
- Rubin \(1974\)D\. B\. Rubin\.Estimating causal effects of treatments in randomized and nonrandomized studies\.*Journal of Educational Psychology*, 66\(5\):688–701, 1974\.
- Saengkyongam et al\. \(2024\)S\. Saengkyongam, E\. Rosenfeld, P\. Ravikumar, N\. Pfister, and J\. Peters\.Identifying representations for intervention extrapolation\.In*International Conference on Learning Representations*, 2024\.
- Schölkopf et al\. \(2021\)B\. Schölkopf, F\. Locatello, S\. Bauer, N\. R\. Ke, N\. Kalchbrenner, A\. Goyal, and Y\. Bengio\.Toward causal representation learning\.*Proceedings of the IEEE*, 109\(5\):612–634, 2021\.
- Shalit et al\. \(2017\)U\. Shalit, F\. D\. Johansson, and D\. Sontag\.Estimating individual treatment effect: Generalization bounds and algorithms\.*Proceedings of Machine Learning Research*, 70:3076–3085, 2017\.
- Shimizu et al\. \(2006\)S\. Shimizu, P\. O\. Hoyer, A\. Hyvärinen, and A\. Kerminen\.A linear non\-gaussian acyclic model for causal discovery\.*Journal of Machine Learning Research*, 7:2003–2030, 2006\.
- Shindo et al\. \(2026\)H\. Shindo, Y\. Deng, T\. Cao, Q\. Delfosse, C\. Tauchmann, J\. Blüml, G\. Sudhakaran, and K\. Kersting\.Learning explicit behavioral models with adaptive questions and world\-model probes\.*arXiv preprint arXiv:2606\.07127*, 2026\.
- Spirtes et al\. \(2000\)P\. Spirtes, C\. N\. Glymour, and R\. Scheines\.*Causation, Prediction, and Search*\.MIT Press, 2 edition, 2000\.
- Sutton and Barto \(2018\)R\. S\. Sutton and A\. G\. Barto\.*Reinforcement Learning: An Introduction*\.MIT Press, 2 edition, 2018\.
- Tsybakov \(2009\)A\. B\. Tsybakov\.*Introduction to Nonparametric Estimation*\.Springer, 2009\.
- Varici et al\. \(2025\)B\. Varici, E\. Acartürk, K\. Shanmugam, A\. Kumar, and A\. Tajer\.Score\-based causal representation learning: Linear and general transformations\.*Journal of Machine Learning Research*, 26\(112\):1–90, 2025\.
- Wager and Athey \(2018\)S\. Wager and S\. Athey\.Estimation and inference of heterogeneous treatment effects using random forests\.*Journal of the American Statistical Association*, 113\(523\):1228–1242, 2018\.
- Yang et al\. \(2024\)A\. Yang, B\. Yang, B\. Hui, B\. Zheng, B\. Yu, C\. Zhou, C\. Li, C\. Li, D\. Liu, F\. Huang, et al\.Qwen2 technical report\.*arXiv preprint arXiv:2407\.10671*, 2024\.
- Yao et al\. \(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\.ReAct: Synergizing reasoning and acting in language models\.*International Conference on Learning Representations*, 2023\.
- Zhang et al\. \(2021\)A\. Zhang, C\. Lyle, S\. Sodhani, A\. Filos, M\. Kwiatkowska, J\. Pineau, Y\. Gal, and D\. Precup\.A survey on causal reinforcement learning\.*arXiv preprint arXiv:2002\.05209*, 2021\.
- Zheng et al\. \(2018\)X\. Zheng, B\. Aragam, P\. K\. Ravikumar, and E\. P\. Xing\.DAGs with NO TEARS: Continuous optimization for structure learning\.*Advances in Neural Information Processing Systems*, 31, 2018\.
- Zou et al\. \(2024\)B\. J\. Zou, M\. E\. Levine, D\. P\. Zaharieva, R\. Johari, and E\. B\. Fox\.Hybrid2neural ODE causal modeling and an application to glycemic response\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, 2024\.

Similar Articles

Selective Memory Retention for Long-Horizon LLM Agents

arXiv cs.AI

This paper presents TraceRetain, a lightweight framework for bounded external memory in frozen LLM agents, demonstrating that selective retention differentiates from cache heuristics primarily when memory streams contain noise, offering task-success and efficiency benefits.

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Hugging Face Daily Papers

CausaLab is a scalable environment for evaluating LLM agents on interactive causal discovery, assessing both predictive accuracy and faithful recovery of underlying causal mechanisms. Experiments reveal a gap between prediction and mechanism recovery, highlighting limits in current LLM agents as experimental causal reasoners.