@Memoirs: On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models Yang Yu https://arxiv…
Summary
This paper analyzes the capability separation between world-model policy learning and imitated world-action models, demonstrating that world-model learning differs in its decision rule and information requirements under certain conditions.
View Cached Full Text
Cached at: 08/26/26, 01:27 PM
On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models
Yang Yu https://t.co/ilwLEoLrRw [𝚌𝚜.𝙻𝙶] https://t.co/ir1TdTcw7q
On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models
Source: https://arxiv.org/html/2608.22197
Abstract
World-action models predict a future world outcome and then infer an action associated with that outcome. This factorization differs from direct action prediction and may provide practical benefits through video pretraining, temporal supervision, and future-conditioned representations. It remains unclear, however, whether these benefits imply a stronger control capability when both systems are trained only from the same observational demonstrations.
We study this question at the level of controller classes and population learning targets. A population objective is an expectation under the true demonstration distribution rather than an empirical average over a finite dataset. The main analysis therefore abstracts away from finite-sample estimation error, optimization error, and model misspecification. We define the control-capability class of a policy class through the closed-loop trajectory distributions that its members can induce.
We compare a direct behavior-cloning policyπA\pi_{\mathrm{A}}, an imitated world-action policyπWA\pi_{\mathrm{WA}}, and a policyπWM\pi_{\mathrm{WM}}optimized using an action-conditioned world model. Every world-action controller can be flattened into a direct stochastic policy with the same closed-loop behavior. Under unrestricted stochastic-kernel classes, the two architectures consequently have the same external control-capability class. Moreover, under realizability, exact population optimization, common deployment information, and distribution-preserving world-action deployment,
πWA∗(a∣h)=πA∗(a∣h)=πμ(a∣h),\pi_{\mathrm{WA}}^{*}(a\mid h)=\pi_{\mathrm{A}}^{*}(a\mid h)=\pi_{\mu}(a\mid h),whereπμ(a∣h):=pμ(a∣h)\pi_{\mu}(a\mid h):=p_{\mu}(a\mid h)is the observational behavior-action conditional. Under an additional behavior-sufficiency condition, these policies induce the same closed-loop trajectory distribution as the demonstrator.
World-model policy learning differs in its decision rule and, when observational identification fails, in the information required to select a policy. It predicts the consequences of specified candidate actions,
P(Y∣H,do(A=a)),P(Y\mid H,\operatorname{do}(A=a)),and compares those consequences through a control objective. We establish an irreducible action-specific prediction gap for future predictors that do not condition on the current candidate action. We then characterize when an exact world-action joint determines a forward model and show why causal identification and action-clamping deployment remain separate requirements. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative action intervention permits zero regret.
The relevant distinction is therefore not whether a controller predicts the future, but whether it identifies and compares futures under specified alternative actions.
1Introduction
Future prediction has become an important component of embodied decision-making systems. Instead of mapping the current observation, instruction, and interaction history directly to an action, a model may first predict a future image, video, state, or latent representation and then infer an action associated with that future. Such world-action models can exploit large-scale video data, dense temporal supervision, and representations of geometry, contact, and motion[16,19].
This development raises a basic question: if a world-action model and a direct action model are trained only from the same observational demonstrations, does future prediction give the world-action model a stronger control capability, or does it provide a different factorization for learning the same behavior policy?
Several effects are easily conflated in empirical comparisons. A world-action system may use additional pretraining data, receive future-frame supervision, have a more suitable inductive bias, or be easier to optimize. It may therefore outperform a direct behavior-cloning baseline even if both methods have the same ideal action target. Conversely, a system that predicts the consequences ofspecified candidate actionsuses a decision interface that standard behavior-reproduction deployment does not provide. When those action effects are not observationally identified, it also requires additional information, causal assumptions, or interventions. Improved task performance alone does not distinguish among these possibilities.
A direct behavior-cloning policy predicts the next actionAAfrom the controller’s available historyHH:
We denote this policy byπA\pi_{\mathrm{A}}. Under a realizable population objective, it recovers the observational behavior-action conditionalpμ(A∣H)p_{\mu}(A\mid H). Here, a population objective is an expected loss under the true demonstration distribution rather than an empirical average over a finite dataset. Many vision-language-action systems are trained primarily through behavior-cloning objectives, but the term “vision-language-action” describes an interface or architecture and does not by itself determine the learning objective. The claims in this paper concern systems trained through observational behavior reproduction.
An imitated world-action model uses a different factorization:
H⟶Y⟶A,H\longrightarrow Y\longrightarrow A,whereYYis a future observation, state, video chunk, or latent outcome. We denote its externally executed policy byπWA\pi_{\mathrm{WA}}. When the model is trained to fit observational tuples(H,A,Y)(H,A,Y), its future component models a future under the behavior distribution, and its inverse component predicts an action associated with that future. The future variable can provide useful training supervision even though it is not yet observed when the action must be selected at deployment.
World-model policy learning uses future prediction differently. An action-conditioned world model represents
(H,A)⟶Y(H,A)\longrightarrow Yand is intended to answer:
What outcome would occur if the controller selected this candidate action?
Formally,do(A=a)\operatorname{do}(A=a)denotes an intervention that externally sets the current action toaa. A planner, value estimator, verifier, or policy optimizer can compare the predicted consequences of alternative actions and produce a policyπWM\pi_{\mathrm{WM}}. This is the decision mechanism underlying model-based reinforcement learning and model-predictive control[10].
The three paradigms can be summarized as
πA:H→A,direct behavior reproduction,πWA:H→Y→A,structured behavior reproduction,πWM:(H,A)→Y→utility,action comparison and optimization.\begin{array}[]{lll}\pi_{\mathrm{A}}:&H\rightarrow A,&\text{direct behavior reproduction},\\[2.5pt] \pi_{\mathrm{WA}}:&H\rightarrow Y\rightarrow A,&\text{structured behavior reproduction},\\[2.5pt] \pi_{\mathrm{WM}}:&(H,A)\rightarrow Y\rightarrow\text{utility},&\text{action comparison and optimization}.\end{array} All three ultimately execute stochastic mappings from histories to actions. Their distinction therefore cannot be established merely by comparing their final interfaces. A comparison between trained systems mixes at least four questions:
- 1.External expressivity:which closed-loop behaviors can the controller architecture represent?
- 2.Population target:which conditional distribution minimizes the expected training objective under the true demonstration distribution?
- 3.Information and identification:which action effects are determined by the information available during training?
- 4.Learning error:how do finite data, restricted model classes, optimization, and computation affect the learned controller?
This paper focuses on the first three questions. We compare policy and learner classes rather than the test errors of particular finite neural networks. In the main exact results, the model classes contain the relevant true conditionals and the population objectives are optimized exactly. The approximate result is a sensitivity statement rather than a finite-sample learning bound.
We define a control-capability class through the closed-loop trajectory distributions representable by a policy class. We separately define observational and interventional learner classes according to the information available to the learning rule. The distinction established below is always one of deployment and decision rule and, when observational identification fails, also one of available information. It is not a separation in unrestricted external policy expressivity.
Main results.
The analysis gives five main conclusions.
- 1.Architecture-level equivalence.Every distributional world-action controller induces a direct stochastic action kernel with exactly the same closed-loop trajectory distribution. Conversely, every direct stochastic policy has a degenerate world-action representation. Direct policies and world-action controllers therefore have the same external control-capability class when both range over unrestricted stochastic kernels.
- 2.Population equivalence under observational imitation.Under realizability, exact population optimization, common observational data, and distribution-preserving deployment, πWA∗(a∣h)=πA∗(a∣h)=πμ(a∣h)\pi_{\mathrm{WA}}^{*}(a\mid h)=\pi_{\mathrm{A}}^{*}(a\mid h)=\pi_{\mu}(a\mid h)forpμ(H)p_{\mu}(H)-almost every history, whereπμ(a∣h):=pμ(a∣h)\pi_{\mu}(a\mid h):=p_{\mu}(a\mid h). Under behavior sufficiency, this equality extends to complete closed-loop trajectory distributions. We also give an approximate version that bounds trajectory-level disagreement in terms of errors in future prediction, inverse-action prediction, and direct action prediction.
- 3.Action-specific prediction separation.A future predictor that does not condition on the current candidate action must average over candidate-action branches. We characterize its irreducible interventional prediction error by Iν,ρint(A;Y∣H),I_{\nu,\rho}^{\mathrm{int}}(A;Y\mid H),which is positive whenever alternative actions induce different outcome distributions on the evaluation support.
- 4.Separation of representation, identification, and deployment.An exact world-action joint determines the observational conditionalpμ(Y∣H,A)p_{\mu}(Y\mid H,A)wherever the behavior distribution has positive action support. Under consistency, conditional exchangeability, and positivity, this conditional equals P(Y∣H,do(A)).P(Y\mid H,\operatorname{do}(A)).The identified joint can then support action comparison if candidate actions are explicitly clamped. Standard future-then-inverse deployment, however, marginalizes the same joint and recovers the behavior policy.
- 5.Decision and information separation.Observational demonstrations do not identify unsupported action effects and need not identify action effects in the presence of hidden confounding. When an exact interventional model is available, model-based optimization can improve beyond a suboptimal demonstrator. Moreover, we construct a two-action environment family in which every observational population learner has positive worst-case regret, whereas one informative action intervention permits zero regret.
These results locate the difference among the three paradigms. Direct behavior cloning and standard world-action imitation differ in their internal factorization but not in their unrestricted external policy class or population action target. World-model policy learning differs in its deployment rule and, when action effects are not observationally identified, in the information required for policy selection. The practical benefits of world-action modeling through representation learning, temporal supervision, and finite-sample efficiency remain compatible with this population-level equivalence.
Organization.
Section2reviews the most closely related work.Section3defines the environment, population objectives, policy classes, learner classes, and assumptions.Section4establishes the architecture-level and population-level equivalence of direct behavior cloning and imitated world-action modeling.Section5formalizes the interventional identification and decision requirements of world-model policy learning.Section6discusses limitations and concludes. Proofs and additional technical remarks are collected in the appendix.
2Related Work
The terms behavior cloning, world-action modeling, world modeling, and model-based control are sometimes applied to systems with overlapping architectural components. The distinction relevant here is whether future prediction is used to reproduce observational behavior or to compare the consequences of specified actions.
2.1Behavior cloning and imitation learning
Behavior cloning estimates a demonstrator’s conditional action distribution from observational trajectories. A central sequential difficulty is that action errors can move the learned policy to histories poorly covered by the demonstrations[14]. This has motivated trajectory-level analyses of imitation error[17,18], imitation with misspecified simulators[5], and learning from imperfect demonstrations[6,1].
These works primarily study finite-sample error, optimization, and distribution shift. Our question is complementary: before those errors are introduced, does an observational future variable change the population action target or the externally realizable policy class?
2.2World-action models and world-model control
World-action architectures use future prediction as an intermediate representation for action inference[16,19]. This factorization may exploit video pretraining and temporal structure even when the final objective remains behavior reproduction.
World models are commonly used for planning, policy optimization, value expansion, and imagined experience[10]. Work on model–Bellman inconsistency and reward-consistent dynamics emphasizes that predictive accuracy and decision-relevant accuracy need not coincide[15,9]. Policy-conditioned and long-horizon models further study how predictive structure interacts with control[2,7].
In visual control, future models have been used for viewpoint-invariant prediction and world-model-based policy learning[12,20]. These systems illustrate that future prediction can enter a controller through different decision rules. Our analysis isolates the distinction between using a future as an observational latent variable and using action-conditioned futures for optimization.
2.3Causal identification and offline model-based control
An observational conditional
is not automatically equal to
P(Y∣H,do(A)).P(Y\mid H,\operatorname{do}(A)).The equality requires causal identification conditions such as consistency, conditional exchangeability, and positivity[13,4].
Invariant action-effect models, identifiable world-model factorizations, and counterfactual environment models seek to recover stable or causal action effects[21,3,8,22]. Model-predictive control from observational and interventional data likewise depends on the distinction between correlation and intervention[11].
Offline model-based reinforcement learning faces a related support problem: an optimized policy may select actions whose consequences are poorly represented in the dataset
3Framework and Scope
This section defines the sequential environment, demonstration distribution, population objectives, controller classes, learner classes, and assumptions used in the comparison.
The total-variation distance between probability measuresPPandQQis
TV(P,Q):=supB|P(B)−Q(B)|,\operatorname{TV}(P,Q):=\sup_{B}|P(B)-Q(B)|,where the supremum is over measurable events. The Kullback–Leibler divergence is denoted by
DKL(P∥Q):=∫log(dPdQ)dPD_{\mathrm{KL}}(P\|Q):=\int\log\left(\frac{dP}{dQ}\right)dPwhenPPis absolutely continuous with respect toQQ, and is+∞+\inftyotherwise.
We use density notation such asp(y∣h)p(y\mid h)for readability. For discrete variables, integrals should be read as sums. For continuous variables, the displayed expressions are interpreted using densities with respect to fixed reference measures or, more generally, regular conditional probabilities.
3.1Sequential environment and trajectories
Consider a finite-horizon partially observed controlled process
ℳ=(𝒮,𝒪,𝒜,P,Ω,ρ0,T),\mathcal{M}=(\mathcal{S},\mathcal{O},\mathcal{A},P,\Omega,\rho_{0},T),where𝒮\mathcal{S},𝒪\mathcal{O}, and𝒜\mathcal{A}are standard Borel spaces. The initial state, observations, and controlled transitions satisfy
S0∼ρ0,Ot∼Ω(⋅∣St),St+1∼P(⋅∣St,At).S_{0}\sim\rho_{0},\qquad O_{t}\sim\Omega(\cdot\mid S_{t}),\qquad S_{t+1}\sim P(\cdot\mid S_{t},A_{t}). At timett, the controller receives
Ht=(G,O0,A0,…,At−1,Ot)∈ℋt,H_{t}=(G,O_{0},A_{0},\ldots,A_{t-1},O_{t})\in\mathcal{H}_{t},(1)whereGGmay contain an instruction, goal, task identifier, or other deployment-time context. The context may be fixed or sampled from a task distribution included in the initial distribution. Time may be included inHtH_{t}, and we write
ℋ=⋃t=0T−1ℋt.\mathcal{H}=\bigcup_{t=0}^{T-1}\mathcal{H}_{t}. Let𝔎(𝒳∣𝒵)\mathfrak{K}(\mathcal{X}\mid\mathcal{Z})denote the set of stochastic kernels from𝒵\mathcal{Z}to𝒳\mathcal{X}. A policy is an element
π∈𝔎(𝒜∣ℋ).\pi\in\mathfrak{K}(\mathcal{A}\mid\mathcal{H}). A complete trajectory is
τ=(G,O0,A0,…,AT−1,OT)∈𝒯T.\tau=(G,O_{0},A_{0},\ldots,A_{T-1},O_{T})\in\mathcal{T}_{T}.The trajectory distribution induced by policyπ\piin environmentℳ\mathcal{M}is denoted by
Pℳπ∈𝒫(𝒯T),P_{\mathcal{M}}^{\pi}\in\mathcal{P}(\mathcal{T}_{T}),where𝒫(𝒯T)\mathcal{P}(\mathcal{T}_{T})is the set of probability measures on the trajectory space.
3.2Demonstrations, future labels, and population objectives
At a demonstrated decision point, let
- •HHbe the history available to the learned controller;
- •AAbe the next action or action chunk;
- •Y∈𝒴Y\in\mathcal{Y}be a future observation, state, latent, or video chunk.
The future label is available in a recorded trajectory but is not observed when the action must be selected at deployment. It may be defined by a measurable trajectory map
Yt=ψt(τ).Y_{t}=\psi_{t}(\tau).Examples includeOt+1O_{t+1}, a future image sequence, a latent future representation, or the nextKKobservations.
Letμ\mudenote the demonstration-generating mechanism. At each time step, it induces
pμ,t(h,a,y)=dμ,t(h)pμ,t(a,y∣h),p_{\mu,t}(h,a,y)=d_{\mu,t}(h)p_{\mu,t}(a,y\mid h),wheredμ,td_{\mu,t}is the distribution of histories visited by the demonstration process at timett.
When the training objective combines decision points from different time steps, define
pμ(h,a,y)=∑t=0T−1wtpμ,t(h,a,y),wt>0,∑t=0T−1wt=1.p_{\mu}(h,a,y)=\sum_{t=0}^{T-1}w_{t}p_{\mu,t}(h,a,y),\qquad w_{t}>0,\qquad\sum_{t=0}^{T-1}w_{t}=1.(2)Because time may be included inHH, this mixture preserves the relevant time-dependent conditionals.
Definition 3.1(Population risk and population optimum).
LetZZbe a training example with true distributionpμp_{\mu}, letq∈𝒬q\in\mathcal{Q}be a model, and letℓ(q,Z)\ell(q;Z)be its loss. The population risk is
ℒ(q):=𝔼Z∼pμ[ℓ(q,Z)].\mathcal{L}(q):=\mathbb{E}_{Z\sim p_{\mu}}[\ell(q;Z)].(3)A population optimum is any
q∗∈argminq∈𝒬ℒ(q).q^{*}\in\operatorname*{arg\,min}_{q\in\mathcal{Q}}\mathcal{L}(q).(4)
A population objective uses the underlying distribution rather than a finite sample. For a dataset𝒟n={Zi}i=1n\mathcal{D}_{n}=\{Z_{i}\}_{i=1}^{n}, the corresponding empirical risk is
ℒ^n(q)=1n∑i=1nℓ(q,Zi).\widehat{\mathcal{L}}_{n}(q)=\frac{1}{n}\sum_{i=1}^{n}\ell(q;Z_{i}). A conditional model is exact at the population optimum if it equals the true conditional distribution for almost every input under the relevant marginal ofpμp_{\mu}. Equality almost everywhere allows disagreement only on sets of probability zero.
Define the behavior-action conditional visible to the learner by
πμ(a∣h):=pμ(a∣h).\pi_{\mu}(a\mid h):=p_{\mu}(a\mid h).(5)If the demonstrator selects actions using onlyHH, thenπμ\pi_{\mu}is the actual demonstrator policy. If the demonstrator uses private information omitted fromHH, thenπμ\pi_{\mu}is only the observational action conditional from the learner’s perspective.
3.3Control-capability and learner classes
Let𝔐\mathfrak{M}be a family of environments sharing the same history and action interfaces.
Definition 3.2(Control equivalence).
Two policiesπ\piandπ′\pi^{\prime}are control-equivalent over(𝔐,T)(\mathfrak{M},T), written
π≡𝔐,Tπ′,\pi\equiv_{\mathfrak{M},T}\pi^{\prime},if
Pℳπ=Pℳπ′for everyℳ∈𝔐.P_{\mathcal{M}}^{\pi}=P_{\mathcal{M}}^{\pi^{\prime}}\qquad\text{for every }\mathcal{M}\in\mathfrak{M}.(6)
Let[π]𝔐,T[\pi]_{\mathfrak{M},T}denote the equivalence class containingπ\pi.
Definition 3.3(Control-capability class).
For a policy classΠ⊆𝔎(𝒜∣ℋ)\Pi\subseteq\mathfrak{K}(\mathcal{A}\mid\mathcal{H}), define
ℭT(Π,𝔐):={[π]𝔐,T:π∈Π}.\mathfrak{C}_{T}(\Pi;\mathfrak{M}):=\left\{[\pi]_{\mathfrak{M},T}:\pi\in\Pi\right\}.(7)
The control-capability class records the externally distinguishable closed-loop behaviors representable byΠ\Pi. It does not characterize sample efficiency, optimization difficulty, parameter count, or utility on a particular task.
Controller classes must also be distinguished from learner classes. A controller class describes which policies can be represented; a learner class describes the information available for selecting one of those policies.
Definition 3.4(Information-restricted population learners).
LetU:𝒯T→[0,1]U:\mathcal{T}_{T}\rightarrow[0,1]be a known decision objective. An observational population learner is a mapping
Lobs:(pμ(H,A,Y),U)⟼π∈ΠL_{\mathrm{obs}}:\bigl(p_{\mu}(H,A,Y),U\bigr)\longmapsto\pi\in\Pithat receives the true observational distribution and the utility but no outcomes generated under additional action interventions. The class of such learners is denoted by
𝔏obs(Π).\mathfrak{L}_{\mathrm{obs}}(\Pi). An interventional population learner additionally receives data𝒟int\mathcal{D}_{\mathrm{int}}containing outcomes under specified action interventions:
Lint:(pμ(H,A,Y),𝒟int,U)⟼π∈Π.L_{\mathrm{int}}:\bigl(p_{\mu}(H,A,Y),\mathcal{D}_{\mathrm{int}},U\bigr)\longmapsto\pi\in\Pi.The corresponding class is denoted by
𝔏int(Π).\mathfrak{L}_{\mathrm{int}}(\Pi).
The environment family, utility, and candidate policy class are common to both learner classes. Their difference is the observational or interventional information available for selecting a policy.
3.4Scope of the comparison
Assumption 3.5(Population-level comparison).
Unless otherwise stated, the main equivalence results use the following conditions:
- (i)Common deployment information.The compared policies receive the same historyHH.
- (ii)Common observational data.Direct behavior cloning and world-action training use trajectories from the same observational process. The world-action learner may use future labelsYYcontained in those trajectories but receives no additional action interventions.
- (iii)Realizability.The relevant policy and conditional-distribution classes contain their population targets.
- (iv)Exact population optimization.The population objectives are globally minimized.
- (v)Distribution-preserving deployment.A world-action controller samples from, or exactly marginalizes over, its learned future distribution. Nonlinear point decoding is excluded unless explicitly discussed.
These conditions remove finite-sample estimation error, approximation error, and optimization error from the primary comparison. Approximation parameters introduced later are used only to describe sensitivity of the exact result.
Assumption 3.6(Behavior sufficiency).
Wheneverπμ\pi_{\mu}is identified with the data-generating demonstrator and equality with the demonstrator’s trajectory distribution is claimed, the demonstration action mechanism is a stochastic kernel ofHHalone:
At∼πμ(⋅∣Ht).A_{t}\sim\pi_{\mu}(\cdot\mid H_{t}).(8)After conditioning onHtH_{t}, no additional demonstrator-only variable affects action selection.
Behavior sufficiency is not required for the probability identity
∫pμ(y∣h)pμ(a∣h,y)𝑑y=pμ(a∣h).\int p_{\mu}(y\mid h)p_{\mu}(a\mid h,y)\,dy=p_{\mu}(a\mid h).It is required whenpμ(a∣h)p_{\mu}(a\mid h)is interpreted as the actual policy that generated the full demonstration trajectory distribution.
3.5The three policy-learning paradigms
3.5.1Direct behavior cloning
Let
ΠA⊆𝔎(𝒜∣ℋ)\Pi_{\mathrm{A}}\subseteq\mathfrak{K}(\mathcal{A}\mid\mathcal{H})be a direct policy class. Its population behavior-cloning objective is
ℒA(πA)=𝔼(H,A)∼pμ[−logπA(A∣H)].\mathcal{L}_{\mathrm{A}}(\pi_{\mathrm{A}})=\mathbb{E}_{(H,A)\sim p_{\mu}}\left[-\log\pi_{\mathrm{A}}(A\mid H)\right].(9) Ifπμ∈ΠA\pi_{\mu}\in\Pi_{\mathrm{A}}and the objective is optimized exactly, then
πA∗(a∣h)=pμ(a∣h)=πμ(a∣h)\pi_{\mathrm{A}}^{*}(a\mid h)=p_{\mu}(a\mid h)=\pi_{\mu}(a\mid h)(10)forpμ(H)p_{\mu}(H)-almost every history.
3.5.2Imitated world-action policies
Let
ℱ⊆𝔎(𝒴∣ℋ)\mathcal{F}\subseteq\mathfrak{K}(\mathcal{Y}\mid\mathcal{H})be a class of future predictors, and let
ℐ⊆𝔎(𝒜∣ℋ×𝒴)\mathcal{I}\subseteq\mathfrak{K}(\mathcal{A}\mid\mathcal{H}\times\mathcal{Y})be a class of inverse-action predictors. The future predictor does not condition on the current candidate action, althoughHHmay contain past actions.
The corresponding world-action joint class is
𝒬WA={qWA(y,a∣h)=qF(y∣h)qI(a∣h,y):qF∈ℱ,qI∈ℐ}.\mathcal{Q}_{\mathrm{WA}}=\left\{q_{\mathrm{WA}}(y,a\mid h)=q_{F}(y\mid h)q_{I}(a\mid h,y):q_{F}\in\mathcal{F},\ q_{I}\in\mathcal{I}\right\}.(11) This factorization also covers architectures in which the future and action are generated by one network, provided that the induced joint distribution admits the displayed conditionals.
A natural observational objective is
ℒWA(qF,qI)=𝔼(H,A,Y)∼pμ[−logqF(Y∣H)−logqI(A∣H,Y)].\begin{split}\mathcal{L}_{\mathrm{WA}}(q_{F},q_{I})=\mathbb{E}_{(H,A,Y)\sim p_{\mu}}\left[-\log q_{F}(Y\mid H)-\log q_{I}(A\mid H,Y)\right].\end{split}(12) Under distribution-preserving deployment, the executed policy is
πWA(a∣h)=∫𝒴qF(y∣h)qI(a∣h,y)𝑑y.\pi_{\mathrm{WA}}(a\mid h)=\int_{\mathcal{Y}}q_{F}(y\mid h)q_{I}(a\mid h,y)\,dy.(13) The external policy class induced by(ℱ,ℐ)(\mathcal{F},\mathcal{I})is
ΠWA(ℱ,ℐ):={π:π(a∣h)=∫qF(y∣h)qI(a∣h,y)dy,qF∈ℱ,qI∈ℐ}.\begin{split}\Pi_{\mathrm{WA}}(\mathcal{F},\mathcal{I}):=\left\{\pi:\pi(a\mid h)=\int q_{F}(y\mid h)q_{I}(a\mid h,y)\,dy,\right.\\ \left.q_{F}\in\mathcal{F},\ q_{I}\in\mathcal{I}\right\}.\end{split}(14) The joint class𝒬WA\mathcal{Q}_{\mathrm{WA}}describes internal future-action representations, whereasΠWA(ℱ,ℐ)\Pi_{\mathrm{WA}}(\mathcal{F},\mathcal{I})describes externally observable control behavior.
We use
ΠWAall:=ΠWA(𝔎(𝒴∣ℋ),𝔎(𝒜∣ℋ×𝒴))\begin{split}\Pi_{\mathrm{WA}}^{\mathrm{all}}:=\Pi_{\mathrm{WA}}\left(\mathfrak{K}(\mathcal{Y}\mid\mathcal{H}),\mathfrak{K}(\mathcal{A}\mid\mathcal{H}\times\mathcal{Y})\right)\end{split}(15)for the induced policy class when both components range over all admissible stochastic kernels.
3.5.3World-model policy learning
For a specified candidate actionaa, define the interventional outcome kernel by
𝖳a(B∣h):=Pℳ(Y∈B∣H=h,do(A=a)),B∈ℬ(𝒴),\mathsf{T}_{a}(B\mid h):=P_{\mathcal{M}}\bigl(Y\in B\mid H=h,\operatorname{do}(A=a)\bigr),\qquad B\in\mathcal{B}(\mathcal{Y}),(16)whereℬ(𝒴)\mathcal{B}(\mathcal{Y})denotes the measurable subsets of𝒴\mathcal{Y}. When the kernel appears inside an integral, we write
𝖳a(dy∣h).\mathsf{T}_{a}(dy\mid h).Thus, for any bounded measurable functionff,
𝔼[f(Y)∣H=h,do(A=a)]=∫𝒴f(y)𝖳a(dy∣h).\mathbb{E}\left[f(Y)\mid H=h,\operatorname{do}(A=a)\right]=\int_{\mathcal{Y}}f(y)\mathsf{T}_{a}(dy\mid h). Let
𝖳^∈𝒲Y⊆𝔎(𝒴∣ℋ×𝒜)\widehat{\mathsf{T}}\in\mathcal{W}_{Y}\subseteq\mathfrak{K}(\mathcal{Y}\mid\mathcal{H}\times\mathcal{A})be a learned action-conditioned outcome model, where
𝖳^(B∣h,a)\widehat{\mathsf{T}}(B\mid h,a)approximates𝖳a(B∣h)\mathsf{T}_{a}(B\mid h). For a one-step decision problem,YYmay be any future variable sufficient to evaluate a bounded utilityr(h,a,Y)r(h,a,Y). The model-based action value is then
Q^(h,a):=∫𝒴r(h,a,y)𝖳^(𝑑y∣h,a).\widehat{Q}(h,a):=\int_{\mathcal{Y}}r(h,a,y)\widehat{\mathsf{T}}(dy\mid h,a).(17) For finite-horizon planning, let
𝖪ℳ(C∣h,a):=Pℳ(Ht+1∈C∣Ht=h,do(At=a)),C∈ℬ(ℋ),\mathsf{K}_{\mathcal{M}}(C\mid h,a):=P_{\mathcal{M}}\bigl(H_{t+1}\in C\mid H_{t}=h,\operatorname{do}(A_{t}=a)\bigr),\qquad C\in\mathcal{B}(\mathcal{H}),(18)denote the true controlled next-history kernel. Under an integral, we write
𝖪ℳ(dh′∣h,a).\mathsf{K}_{\mathcal{M}}(dh^{\prime}\mid h,a). Let
𝖪^∈𝒲H⊆𝔎(ℋ∣ℋ×𝒜)\widehat{\mathsf{K}}\in\mathcal{W}_{H}\subseteq\mathfrak{K}(\mathcal{H}\mid\mathcal{H}\times\mathcal{A})be a learned next-history model. Letη0\eta_{0}be the initial-history distribution. A policyπ\piand model𝖪^\widehat{\mathsf{K}}induce a rollout according to
H0∼η0,At∼π(⋅∣Ht),Ht+1∼𝖪^(⋅∣Ht,At).H_{0}\sim\eta_{0},\qquad A_{t}\sim\pi(\cdot\mid H_{t}),\qquad H_{t+1}\sim\widehat{\mathsf{K}}(\cdot\mid H_{t},A_{t}).Denote the resulting trajectory distribution byP𝖪^πP_{\widehat{\mathsf{K}}}^{\pi}. For a bounded trajectory utilityU:𝒯T→[0,1]U:\mathcal{T}_{T}\rightarrow[0,1], define
J𝖪^,UT(π):=𝔼τ∼P𝖪^π[U(τ)].J_{\widehat{\mathsf{K}},U}^{T}(\pi):=\mathbb{E}_{\tau\sim P_{\widehat{\mathsf{K}}}^{\pi}}[U(\tau)].(19) Given a candidate policy class𝚷\boldsymbol{\Pi}, a world-model policy satisfies
πWM∈argmaxπ∈𝚷J𝖪^,UT(π).\pi_{\mathrm{WM}}\in\operatorname*{arg\,max}_{\pi\in\boldsymbol{\Pi}}J_{\widehat{\mathsf{K}},U}^{T}(\pi).(20)Whenever anargmax\operatorname*{arg\,max}is displayed, we assume that the maximum is attained.
Table 1:The three policy-learning paradigms.An observational predictorp(Y∣H,A)p(Y\mid H,A)is not automatically an interventional world model. Interpreting it asP(Y∣H,do(A))P(Y\mid H,\operatorname{do}(A))requires sufficient action support and a valid causal identification argument.
4Observational Equivalence of Direct and World-Action Policies
We first ask whether the internal future variable enlarges the externally observable controller class. We then compare the policies selected by the direct and world-action population objectives.
4.1A decision-complete control metric
Equality under one benchmark reward is too weak to establish general control equivalence. We instead compare complete trajectory distributions.
Definition 4.1(Control ability).
For a measurable utilityU:𝒯T→[0,1]U:\mathcal{T}_{T}\rightarrow[0,1], define
Abilityℳ,UT(π):=Jℳ,UT(π):=𝔼τ∼Pℳπ[U(τ)].\operatorname{Ability}_{\mathcal{M},U}^{T}(\pi):=J_{\mathcal{M},U}^{T}(\pi):=\mathbb{E}_{\tau\sim P_{\mathcal{M}}^{\pi}}[U(\tau)].(21)
Definition 4.2(Decision-complete control distance).
For two policies in a fixed environmentℳ\mathcal{M}, define
dctrlℳ,T(π,π′):=supU:𝒯T→[0,1]|Jℳ,UT(π)−Jℳ,UT(π′)|.\begin{split}d_{\mathrm{ctrl}}^{\mathcal{M},T}(\pi,\pi^{\prime}):=\sup_{U:\mathcal{T}_{T}\rightarrow[0,1]}\left|J_{\mathcal{M},U}^{T}(\pi)-J_{\mathcal{M},U}^{T}(\pi^{\prime})\right|.\end{split}(22)
Proposition 4.3(Trajectory characterization).
dctrlℳ,T(π,π′)=TV(Pℳπ,Pℳπ′).d_{\mathrm{ctrl}}^{\mathcal{M},T}(\pi,\pi^{\prime})=\operatorname{TV}\left(P_{\mathcal{M}}^{\pi},P_{\mathcal{M}}^{\pi^{\prime}}\right).(23)
The proof is given inSectionA.1. Zero control distance means that no bounded return, success indicator, safety statistic, verifier, or other trajectory-level evaluation can distinguish the policies.
4.2Architecture-level flattening
Theorem 4.4(World-action flattening).
For every world-action controller(qF,qI)(q_{F},q_{I}), define
π¯(a∣h)=∫𝒴qF(y∣h)qI(a∣h,y)𝑑y.\bar{\pi}(a\mid h)=\int_{\mathcal{Y}}q_{F}(y\mid h)q_{I}(a\mid h,y)\,dy.(24)Thenπ¯∈𝔎(𝒜∣ℋ)\bar{\pi}\in\mathfrak{K}(\mathcal{A}\mid\mathcal{H}). LetPℳ(qF,qI)P_{\mathcal{M}}^{(q_{F},q_{I})}denote the trajectory distribution obtained by deploying the two-stage world-action controller. In every environment with the same history and action interfaces,
Pℳ(qF,qI)=Pℳπ¯.P_{\mathcal{M}}^{(q_{F},q_{I})}=P_{\mathcal{M}}^{\bar{\pi}}.(25) Consequently,
ΠWA(ℱ,ℐ)⊆𝔎(𝒜∣ℋ).\Pi_{\mathrm{WA}}(\mathcal{F},\mathcal{I})\subseteq\mathfrak{K}(\mathcal{A}\mid\mathcal{H}).(26) Conversely, if the world-action class permits a point-mass future and arbitrary inverse kernels at that future, then every direct stochastic policy has a degenerate world-action representation. Hence
ℭT(ΠWAall,𝔐)=ℭT(𝔎(𝒜∣ℋ),𝔐).\begin{split}\mathfrak{C}_{T}\bigl(\Pi_{\mathrm{WA}}^{\mathrm{all}};\mathfrak{M}\bigr)=\mathfrak{C}_{T}\bigl(\mathfrak{K}(\mathcal{A}\mid\mathcal{H});\mathfrak{M}\bigr).\end{split}(27)
The proof is inSectionA.2. The theorem shows that the internal future variable does not enlarge the unrestricted external policy class. It may still provide a more efficient representation within restricted parametric classes.
4.3Population equivalence under observational training
Architecture-level flattening does not determine which policy is selected by training. We next compare the population optima of direct behavior cloning and world-action maximum likelihood on the same observational distribution.
Assumption 4.5(Ideal observational objectives).
In addition to3.5, assume that
πμ∈ΠA,\pi_{\mu}\in\Pi_{\mathrm{A}},and that the world-action classes contain the true observational conditionals:
pμ(Y∣H)∈ℱ,pμ(A∣H,Y)∈ℐ.p_{\mu}(Y\mid H)\in\mathcal{F},\qquad p_{\mu}(A\mid H,Y)\in\mathcal{I}.
Theorem 4.6(Population action-marginal equivalence).
Under4.5, every exact population optimum satisfies
πWA∗(a∣h)=πA∗(a∣h)=πμ(a∣h)\pi_{\mathrm{WA}}^{*}(a\mid h)=\pi_{\mathrm{A}}^{*}(a\mid h)=\pi_{\mu}(a\mid h)(28)forpμ(H)p_{\mu}(H)-almost every history.
If3.6also holds, then the three action kernels agree atdμ,td_{\mu,t}-almost every history for each timett, and
PℳπWA∗=PℳπA∗=Pℳπμ.P_{\mathcal{M}}^{\pi_{\mathrm{WA}}^{*}}=P_{\mathcal{M}}^{\pi_{\mathrm{A}}^{*}}=P_{\mathcal{M}}^{\pi_{\mu}}.(29)Consequently,
dctrlℳ,T(πWA∗,πA∗)=0.d_{\mathrm{ctrl}}^{\mathcal{M},T}\bigl(\pi_{\mathrm{WA}}^{*},\pi_{\mathrm{A}}^{*}\bigr)=0.(30)
The proof is inSectionA.3. The central identity is
πWA∗(a∣h)=∫pμ(y∣h)pμ(a∣h,y)𝑑y=pμ(a∣h).\begin{split}\pi_{\mathrm{WA}}^{*}(a\mid h)&=\int p_{\mu}(y\mid h)p_{\mu}(a\mid h,y)\,dy\\ &=p_{\mu}(a\mid h).\end{split}Thus, sampling a future under the behavior distribution and then sampling its posterior behavior action recovers the behavior-action conditional.
The theorem concerns distribution-preserving deployment. A system that selects a single future by MAP decoding, reranks futures with a verifier, or otherwise modifies the learned joint may induce a different action marginal.
4.4Approximate equivalence
The exact theorem describes the population limit. The following continuity bound shows how discrepancies in the learned components translate into closed-loop disagreement. It does not specify how those discrepancies depend on sample size.
Suppose that, uniformly over the relevant deployment histories,
TV(qF(⋅∣h),pμ(⋅∣h))≤ϵF,\operatorname{TV}\bigl(q_{F}(\cdot\mid h),p_{\mu}(\cdot\mid h)\bigr)\leq\epsilon_{F},(31)𝔼Y∼pμ(⋅∣h)[TV(qI(⋅∣h,Y),pμ(⋅∣h,Y))]≤ϵI,\mathbb{E}_{Y\sim p_{\mu}(\cdot\mid h)}\left[\operatorname{TV}\bigl(q_{I}(\cdot\mid h,Y),p_{\mu}(\cdot\mid h,Y)\bigr)\right]\leq\epsilon_{I},(32)and
TV(πA(⋅∣h),πμ(⋅∣h))≤ϵA.\operatorname{TV}\bigl(\pi_{\mathrm{A}}(\cdot\mid h),\pi_{\mu}(\cdot\mid h)\bigr)\leq\epsilon_{A}.(33)
Proposition 4.8(Approximate control equivalence).
Let
ϵ=min{1,ϵF+ϵI+ϵA}.\epsilon=\min\{1,\epsilon_{F}+\epsilon_{I}+\epsilon_{A}\}.Then
TV(πWA(⋅∣h),πA(⋅∣h))≤ϵ.\operatorname{TV}\bigl(\pi_{\mathrm{WA}}(\cdot\mid h),\pi_{\mathrm{A}}(\cdot\mid h)\bigr)\leq\epsilon.(34)If the same bound holds at every history that may be reached by either policy, then
dctrlℳ,T(πWA,πA)≤1−(1−ϵ)T≤Tϵ.\begin{split}d_{\mathrm{ctrl}}^{\mathcal{M},T}\bigl(\pi_{\mathrm{WA}},\pi_{\mathrm{A}}\bigr)&\leq 1-(1-\epsilon)^{T}\\ &\leq T\epsilon.\end{split}(35)
The proof is inSectionA.4.
4.5Practical advantages of the world-action factorization
Equal external capability classes and equal population action targets remain compatible with substantial practical differences. At training time, the futureYYmay contain information about the demonstrated action. For discrete variables under conditional log loss, the Bayes risks of direct and future-conditioned action prediction are
Hμ(A∣H)andHμ(A∣H,Y),H_{\mu}(A\mid H)\qquad\text{and}\qquad H_{\mu}(A\mid H,Y),whereHμH_{\mu}denotes conditional entropy underpμp_{\mu}. Their difference is
Hμ(A∣H)−Hμ(A∣H,Y)=Iμ(A;Y∣H)≥0.H_{\mu}(A\mid H)-H_{\mu}(A\mid H,Y)=I_{\mu}(A;Y\mid H)\geq 0.(36) For Euclidean actions under squared loss,
𝔼[∥A−𝔼[A∣H]∥22]−𝔼[∥A−𝔼[A∣H,Y]∥22]=𝔼[∥𝔼[A∣H,Y]−𝔼[A∣H]∥22]≥0.\begin{split}&\mathbb{E}\left[\|A-\mathbb{E}[A\mid H]\|_{2}^{2}\right]-\mathbb{E}\left[\|A-\mathbb{E}[A\mid H,Y]\|_{2}^{2}\right]\\ &=\mathbb{E}\left[\|\mathbb{E}[A\mid H,Y]-\mathbb{E}[A\mid H]\|_{2}^{2}\right]\geq 0.\end{split}(37) Action prediction may therefore be easier when the realized future is available as a training-time conditioning variable. At deployment, however, that future has not yet occurred and must itself be predicted. The factorization trades lower future-conditioned action uncertainty against future-prediction error.
Future prediction may also improve video pretraining, temporal representations, multimodal behavior modeling, parameter sharing, and cross-task transfer. These are learning and representation advantages; they do not show that standard future-then-inverse deployment performs interventional action comparison.
5Interventional Identification and Decision Separation
The previous section rules out an unrestricted external-policy-class separation. The remaining distinction concerns whether a learner can identify and use action-specific outcome information.
For the explicit Bayes-ratio arguments in this section, we assume that𝒜\mathcal{A}and𝒴\mathcal{Y}are finite. Accordingly, we write
𝖳a(y∣h):=𝖳a({y}∣h)\mathsf{T}_{a}(y\mid h):=\mathsf{T}_{a}(\{y\}\mid h)for the probability mass assigned toyy. Integrals with respect to𝖳a(dy∣h)\mathsf{T}_{a}(dy\mid h)therefore become sums overy∈𝒴y\in\mathcal{Y}.
5.1Inverse dynamics as a behavior posterior
Assume temporarily that the observational conditional is causally identified on the evaluated support:
pμ(y∣h,a)=𝖳a(y∣h).p_{\mu}(y\mid h,a)=\mathsf{T}_{a}(y\mid h).(38) The future predictor then learns the behavior mixture
𝖳μ(y∣h)=∑a∈𝒜πμ(a∣h)𝖳a(y∣h).\mathsf{T}_{\mu}(y\mid h)=\sum_{a\in\mathcal{A}}\pi_{\mu}(a\mid h)\mathsf{T}_{a}(y\mid h).(39) For everyyywith𝖳μ(y∣h)>0\mathsf{T}_{\mu}(y\mid h)>0, the exact inverse conditional is
pμ(a∣h,y)=πμ(a∣h)𝖳a(y∣h)∑b∈𝒜πμ(b∣h)𝖳b(y∣h).p_{\mu}(a\mid h,y)=\frac{\pi_{\mu}(a\mid h)\mathsf{T}_{a}(y\mid h)}{\sum_{b\in\mathcal{A}}\pi_{\mu}(b\mid h)\mathsf{T}_{b}(y\mid h)}.(40) The inverse conditional depends on both the environmental effect𝖳a(y∣h)\mathsf{T}_{a}(y\mid h)and the behavior priorπμ(a∣h)\pi_{\mu}(a\mid h). Summing this posterior against the behavior-mixture future recovers the behavior policy:
∑y∈𝒴𝖳μ(y∣h)pμ(a∣h,y)=πμ(a∣h).\sum_{y\in\mathcal{Y}}\mathsf{T}_{\mu}(y\mid h)p_{\mu}(a\mid h,y)=\pi_{\mu}(a\mid h).(41) Inverse dynamics asks which behavior action is associated with a realized transition. An interventional forward model asks what outcome would result from a specified action. The two conditionals support different deployment rules.
5.2An action-specific prediction metric
The likelihood of a future under the behavior distribution does not measure whether a model can answer action-specific “what if” questions. We instead evaluate histories and candidate actions independently.
Letν\nube an evaluation distribution over histories, and letρ(⋅∣h)\rho(\cdot\mid h)be a distribution over candidate actions. Define
ℰint(q):=𝔼H∼νA∼ρ(⋅∣H)[DKL(𝖳A(⋅∣H)∥q(⋅∣H,A))].\begin{split}\mathcal{E}_{\mathrm{int}}(q):=\mathbb{E}_{\begin{subarray}{c}H\sim\nu\\ A\sim\rho(\cdot\mid H)\end{subarray}}\left[D_{\mathrm{KL}}\left(\mathsf{T}_{A}(\cdot\mid H)\,\|\,q(\cdot\mid H,A)\right)\right].\end{split}(42) For a future predictorg(y∣h)g(y\mid h)that does not condition on the current candidate action, define
qg(y∣h,a):=g(y∣h).q_{g}(y\mid h,a):=g(y\mid h).Also define
𝖳ρ(y∣h)=∑a∈𝒜ρ(a∣h)𝖳a(y∣h).\mathsf{T}_{\rho}(y\mid h)=\sum_{a\in\mathcal{A}}\rho(a\mid h)\mathsf{T}_{a}(y\mid h).(43) For a fixed historyhh, let
Iρint(A;Y∣H=h):=∑a∈𝒜ρ(a∣h)DKL(𝖳a(⋅∣h)∥𝖳ρ(⋅∣h)).\begin{split}I_{\rho}^{\mathrm{int}}(A;Y\mid H=h):=\sum_{a\in\mathcal{A}}\rho(a\mid h)D_{\mathrm{KL}}\left(\mathsf{T}_{a}(\cdot\mid h)\,\|\,\mathsf{T}_{\rho}(\cdot\mid h)\right).\end{split}(44)Averaging overH∼νH\sim\nugives
Iν,ρint(A;Y∣H):=𝔼H∼ν[Iρint(A;Y∣H)].\begin{split}I_{\nu,\rho}^{\mathrm{int}}(A;Y\mid H)&:=\mathbb{E}_{H\sim\nu}\left[I_{\rho}^{\mathrm{int}}(A;Y\mid H)\right].\end{split}(45)
Theorem 5.1(Irreducible action-unconditioned prediction gap).
For every predictorg(y∣h)g(y\mid h)that does not condition on the current candidate action,
ℰint(qg)=Iν,ρint(A;Y∣H)+𝔼H∼ν[DKL(𝖳ρ(⋅∣H)∥g(⋅∣H))].\begin{split}\mathcal{E}_{\mathrm{int}}(q_{g})=I_{\nu,\rho}^{\mathrm{int}}(A;Y\mid H)+\mathbb{E}_{H\sim\nu}\left[D_{\mathrm{KL}}\left(\mathsf{T}_{\rho}(\cdot\mid H)\,\|\,g(\cdot\mid H)\right)\right].\end{split}(46)Consequently,
infgℰint(qg)=Iν,ρint(A;Y∣H),\inf_{g}\mathcal{E}_{\mathrm{int}}(q_{g})=I_{\nu,\rho}^{\mathrm{int}}(A;Y\mid H),(47)whereas an exact action-conditioned modelq(y∣h,a)=𝖳a(y∣h)q(y\mid h,a)=\mathsf{T}_{a}(y\mid h)has zero risk.
The proof is inSectionA.5. The lower bound is positive whenever candidate actions produce different outcome distributions on a set of histories with positiveν\nu-probability. A future under the behavior distribution therefore cannot substitute for an action-specific prediction.
5.3Converting a world-action joint into a forward model
The preceding theorem concerns the future predictor in isolation. The complete world-action joint may contain more information. An exact joint determines the observational conditionalpμ(Y∣H,A)p_{\mu}(Y\mid H,A)wherever the behavior distribution assigns positive probability to the action. Whether that conditional has a causal interpretation, and whether the controller uses it for action comparison, are separate questions.
Proposition 5.2(Extracting an observational forward conditional).
Suppose that the learned world-action joint is exact at the population optimum:
qF∗(y∣h)qI∗(a∣h,y)=pμ(y,a∣h)q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)=p_{\mu}(y,a\mid h)(48)forpμ(H)p_{\mu}(H)-almost everyhh. Define
π¯q(a∣h):=∑y∈𝒴qF∗(y∣h)qI∗(a∣h,y).\bar{\pi}_{q}(a\mid h):=\sum_{y\in\mathcal{Y}}q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y).(49)Then
π¯q(a∣h)=pμ(a∣h).\bar{\pi}_{q}(a\mid h)=p_{\mu}(a\mid h).(50) For every pair(h,a)(h,a)withπ¯q(a∣h)>0\bar{\pi}_{q}(a\mid h)>0, define
𝖳qWA(y∣h,a):=qF∗(y∣h)qI∗(a∣h,y)π¯q(a∣h).\mathsf{T}^{\mathrm{WA}}_{q}(y\mid h,a):=\frac{q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)}{\bar{\pi}_{q}(a\mid h)}.(51)Then
𝖳qWA(y∣h,a)=pμ(y∣h,a).\mathsf{T}^{\mathrm{WA}}_{q}(y\mid h,a)=p_{\mu}(y\mid h,a).(52)
Proof.
SummingEquation48overyygives
π¯q(a∣h)=∑y∈𝒴qF∗(y∣h)qI∗(a∣h,y)=∑y∈𝒴pμ(y,a∣h)=pμ(a∣h).\begin{split}\bar{\pi}_{q}(a\mid h)&=\sum_{y\in\mathcal{Y}}q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)\\ &=\sum_{y\in\mathcal{Y}}p_{\mu}(y,a\mid h)\\ &=p_{\mu}(a\mid h).\end{split}Forpμ(a∣h)>0p_{\mu}(a\mid h)>0,
𝖳qWA(y∣h,a)=qF∗(y∣h)qI∗(a∣h,y)π¯q(a∣h)=pμ(y,a∣h)pμ(a∣h)=pμ(y∣h,a).\begin{split}\mathsf{T}^{\mathrm{WA}}_{q}(y\mid h,a)&=\frac{q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)}{\bar{\pi}_{q}(a\mid h)}\\ &=\frac{p_{\mu}(y,a\mid h)}{p_{\mu}(a\mid h)}\\ &=p_{\mu}(y\mid h,a).\end{split}∎
The proposition is probabilistic rather than causal. To identifypμ(Y∣H,A)p_{\mu}(Y\mid H,A)with an interventional distribution, letY(a)Y(a)denote the potential outcome under actionaaand suppose that the following conditions hold:
- 1.Consistency:ifA=aA=a, thenY=Y(a)Y=Y(a);
- 2.Conditional exchangeability: Y(a)⟂⟂A|H;Y(a)\perp\!\!\!\perp A\mid H;(53)
- 3.Positivity: pμ(a∣h)>0p_{\mu}(a\mid h)>0(54)for every evaluated history-action pair.
Under these conditions,
𝖳qWA(y∣h,a)=pμ(y∣h,a)=Pℳ(Y=y∣H=h,do(A=a))=𝖳a(y∣h).\begin{split}\mathsf{T}^{\mathrm{WA}}_{q}(y\mid h,a)&=p_{\mu}(y\mid h,a)\\ &=P_{\mathcal{M}}\bigl(Y=y\mid H=h,\operatorname{do}(A=a)\bigr)\\ &=\mathsf{T}_{a}(y\mid h).\end{split}(55) Equivalently,
𝖳a(y∣h)=qF∗(y∣h)qI∗(a∣h,y)∑y′∈𝒴qF∗(y′∣h)qI∗(a∣h,y′).\mathsf{T}_{a}(y\mid h)=\frac{q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)}{\sum_{y^{\prime}\in\mathcal{Y}}q_{F}^{*}(y^{\prime}\mid h)q_{I}^{*}(a\mid h,y^{\prime})}.(56)At the population optimum,
𝖳a(y∣h)=pμ(y∣h)pμ(a∣h,y)pμ(a∣h).\mathsf{T}_{a}(y\mid h)=\frac{p_{\mu}(y\mid h)p_{\mu}(a\mid h,y)}{p_{\mu}(a\mid h)}.(57)The causal derivation is given inSectionA.6.
Three requirements are involved: the world-action factors must recover the observational joint, the action must have positive support so that the observational conditional is defined, and that conditional must be causally identified. None of these requirements determines how the model is deployed.
Indeed, standard world-action deployment still gives
πWA∗(a∣h)=∑y∈𝒴qF∗(y∣h)qI∗(a∣h,y)=pμ(a∣h).\begin{split}\pi_{\mathrm{WA}}^{*}(a\mid h)&=\sum_{y\in\mathcal{Y}}q_{F}^{*}(y\mid h)q_{I}^{*}(a\mid h,y)\\ &=p_{\mu}(a\mid h).\end{split}(58) A planning rule uses the same joint differently. For a one-step utilityr(h,a,y)∈[0,1]r(h,a,y)\in[0,1], define
QWA(h,a):=𝔼Y∼𝖳WAq(⋅∣h,a)[r(h,a,Y)]Q_{\mathrm{WA}}(h,a):=\mathbb{E}_{Y\sim\mathsf{T}^{\mathrm{WA}}_{q}(\cdot\mid h,a)}\left[r(h,a,Y)\right](59)and select
πplan(h)∈argmaxaQWA(h,a).\pi_{\mathrm{plan}}(h)\in\operatorname*{arg\,max}_{a}Q_{\mathrm{WA}}(h,a).(60)
Corollary 5.3(One-step planning with an identified world-action joint).
Suppose thatProposition5.2holds and that consistency, conditional exchangeability, and positivity identify
𝖳qWA(⋅∣h,a)=𝖳a(⋅∣h)\mathsf{T}^{\mathrm{WA}}_{q}(\cdot\mid h,a)=\mathsf{T}_{a}(\cdot\mid h)for every action considered by the planner. Then
πplan(h)∈argmaxa𝔼Y∼𝖳a(⋅∣h)[r(h,a,Y)].\pi_{\mathrm{plan}}(h)\in\operatorname*{arg\,max}_{a}\mathbb{E}_{Y\sim\mathsf{T}_{a}(\cdot\mid h)}[r(h,a,Y)].(61)
Proof.
Under the assumptions,
QWA(h,a)=𝔼Y∼𝖳a(⋅∣h)[r(h,a,Y)]Q_{\mathrm{WA}}(h,a)=\mathbb{E}_{Y\sim\mathsf{T}_{a}(\cdot\mid h)}[r(h,a,Y)]for every candidate action. MaximizingQWA(h,a)Q_{\mathrm{WA}}(h,a)therefore maximizes the true interventional expected utility. ∎
The same exact joint can thus support either behavior reproduction or action comparison. The distinction lies in causal identification and deployment, not merely in the represented joint distribution.
5.4Observational non-identification
The recovery result depends on support and causal assumptions. Observational fitting alone does not guarantee either condition.
Theorem 5.4(Two sources of observational non-identification).
Interventional action effects are not identified bypμ(H,A,Y)p_{\mu}(H,A,Y)in general.
- (i)Without positivity, two causal environments can have the same observational distribution while differing under an unsupported action intervention.
- (ii)Even when every observed action has positive probability, two causal environments can have the same observational distribution while differing interventionally ifHHomits a variable that affects both action selection and outcomes.
The proof is inSectionA.7. The first construction uses an action that is never selected by the behavior policy. The second uses a demonstrator-only variable omitted fromHH. Thus, action coverage and causal sufficiency address different identification failures.
5.5Strict separation of observational and interventional learners
The strongest separation is not between the external policy classes. It is between learners that receive different information about action effects.
For an environmentℳ\mathcal{M}and bounded utilityUU, define regret relative to a candidate class𝚷\boldsymbol{\Pi}by
Regℳ,U(π):=supπ′∈𝚷Jℳ,UT(π′)−Jℳ,UT(π).\operatorname{Reg}_{\mathcal{M},U}(\pi):=\sup_{\pi^{\prime}\in\boldsymbol{\Pi}}J_{\mathcal{M},U}^{T}(\pi^{\prime})-J_{\mathcal{M},U}^{T}(\pi).(62) Under the assumptions ofTheorem4.6,
πWA∗=πA∗=πμ.\pi_{\mathrm{WA}}^{*}=\pi_{\mathrm{A}}^{*}=\pi_{\mu}.
Corollary 5.5(Suboptimal-demonstrator separation).
Suppose that behavior sufficiency holds, thatπμ∈𝚷\pi_{\mu}\in\boldsymbol{\Pi}, and that
Δ:=supπ∈𝚷Jℳ,UT(π)−Jℳ,UT(πμ)>0.\Delta:=\sup_{\pi\in\boldsymbol{\Pi}}J_{\mathcal{M},U}^{T}(\pi)-J_{\mathcal{M},U}^{T}(\pi_{\mu})>0.(63)Assume that an optimal policy in𝚷\boldsymbol{\Pi}exists, that the learned next-history model equals the true interventional kernel on every history-action pair reachable by policies in𝚷\boldsymbol{\Pi}, and that model-based optimization is exact. Then
Jℳ,UT(πWM∗)=Jℳ,UT(πWA∗)+Δ=Jℳ,UT(πA∗)+Δ.\begin{split}J_{\mathcal{M},U}^{T}(\pi_{\mathrm{WM}}^{*})&=J_{\mathcal{M},U}^{T}(\pi_{\mathrm{WA}}^{*})+\Delta\\ &=J_{\mathcal{M},U}^{T}(\pi_{\mathrm{A}}^{*})+\Delta.\end{split}(64)
Proof.
Exactness of the interventional model on the relevant history-action pairs implies
J𝖪^,UT(π)=Jℳ,UT(π)for everyπ∈𝚷.J_{\widehat{\mathsf{K}},U}^{T}(\pi)=J_{\mathcal{M},U}^{T}(\pi)\qquad\text{for every }\pi\in\boldsymbol{\Pi}.Exact optimization therefore gives the best true value in𝚷\boldsymbol{\Pi}. ByTheorem4.6and behavior sufficiency,
Jℳ,UT(πWA∗)=Jℳ,UT(πA∗)=Jℳ,UT(πμ).J_{\mathcal{M},U}^{T}(\pi_{\mathrm{WA}}^{*})=J_{\mathcal{M},U}^{T}(\pi_{\mathrm{A}}^{*})=J_{\mathcal{M},U}^{T}(\pi_{\mu}).Substituting the definition ofΔ\Deltaproves the result. ∎
The corollary assumes that the relevant action effects are already known or identified. Interventions can also create a strict information advantage.
Theorem 5.6(Strict value of interventional information).
There exists a one-step, two-action environment family with a common known utility and candidate class
𝚷=𝔎({0,1}∣{h})\boldsymbol{\Pi}=\mathfrak{K}(\{0,1\}\mid\{h\})such that:
- (i)the observational distributionpμ(H,A,Y)p_{\mu}(H,A,Y)is identical across the environments;
- (ii)every observational population learnerLobs∈𝔏obs(𝚷)L_{\mathrm{obs}}\in\mathfrak{L}_{\mathrm{obs}}(\boldsymbol{\Pi})has worst-case regret at least1/41/4;
- (iii)the exact population optimaπA∗\pi_{\mathrm{A}}^{*}andπWA∗\pi_{\mathrm{WA}}^{*}have worst-case regret1/21/2;
- (iv)one informative action intervention identifies the environment and permits an interventional learner, representable as an exact world-model learner, to achieve zero regret.
The proof is inSectionA.8. The theorem isolates the source of the separation: the observational learners output ordinary stochastic policies, but the observational distribution does not determine which policy is optimal across the environment family.
6Conclusion
A direct behavior-cloning policy and an imitation-trained world-action policy use different internal factorizations, but every distributional world-action controller induces an ordinary stochastic action kernel. Under unrestricted kernel classes, the two therefore have the same external control-capability class. Under realizability, exact population optimization, matched deployment information, and distribution-preserving deployment, they also have the same population action target:
πWA∗=πA∗=πμ.\pi_{\mathrm{WA}}^{*}=\pi_{\mathrm{A}}^{*}=\pi_{\mu}.With behavior sufficiency, this equality extends to their complete closed-loop trajectory distributions.
These results do not deny the practical value of world-action modeling. Future prediction may improve temporal representations, video pretraining, multimodal behavior modeling, parameter sharing, and finite-sample performance. Gains of this kind concern how the observational target is represented and learned; they do not by themselves establish interventional reasoning.
The complete world-action joint can contain more information than the future predictor alone. With positive action support, an exact joint determinespμ(Y∣H,A)p_{\mu}(Y\mid H,A). Under additional causal assumptions, that conditional may equal
P(Y∣H,do(A=a)).P(Y\mid H,\operatorname{do}(A=a)).Standard future-then-inverse deployment nevertheless marginalizes the joint and reproduces the behavior policy. Using the same representation for world-model control requires a different operation: candidate actions must be specified, their consequences evaluated, and their utilities compared.
Observational demonstrations do not identify these action effects in general. Unsupported actions and hidden confounding can make different causal environments observationally indistinguishable. Interventions, exploratory coverage, or valid causal assumptions are required whenever the desired decision depends on effects not identified by the demonstration distribution.
This distinction also suggests how empirical claims should be evaluated. Imitation quality should be separated from action-effect prediction. Evidence for interventional control should test predictions from the same state under multiple specified actions and determine whether explicit action comparison improves value beyond future-then-inverse deployment. A practical system may share a common representation while using separate components for
qF(Y∣H),qI(A∣H,Y),qdo(Y∣H,A).q_{F}(Y\mid H),\qquad q_{I}(A\mid H,Y),\qquad q_{\mathrm{do}}(Y\mid H,A).The first two support structured imitation; the third supports action-specific evaluation.
The essential distinction is therefore between predicting a future associated with observed behavior and predicting the consequences of specified actions for policy optimization:
future prediction for behavior decoding≠interventional action-conditioned prediction for control.\text{future prediction for behavior decoding}\;\neq\;\text{interventional action-conditioned prediction for control}. While we expect this paper to shed some light on the understanding of world-action models and world models, it comes with several obvious limitations:
Class-level rather than finite-network equivalence.The flattening theorem concerns unrestricted stochastic kernels, or restricted classes closed under the required marginalization. Under parameter, memory, latency, or compute constraints, a world-action factorization may represent a useful policy more efficiently than a direct architecture.
Population rather than statistical analysis.The main comparison assumes realizability and exact population optimization. It does not provide sample-complexity bounds or optimization guarantees for large neural models. The approximate result only describes sensitivity to specified distributional errors.
Matched deployment information.The equivalence results require the compared policies to receive the same history. If one architecture receives a longer context, privileged state, future frame, external memory, or additional instruction at deployment, the policy classes are not directly comparable.
Behavior sufficiency.Equality with the demonstrator’s trajectory distribution requires the recorded history to contain the information used for action selection. When a demonstrator uses omitted private information, direct and world-action training still recover the same observational action marginal, but deploying that marginal need not reproduce the original demonstrator–environment coupling.
Distribution-preserving deployment.MAP future selection, deterministic decoding, temperature changes, best-of-NNsampling, and verifier-based future selection may alter the action marginal. Such controllers remain flattenable into direct policies, but they need not equal the behavior-cloning population optimum.
Conditional identification.A perfect observational joint identifies interventional dynamics only under valid support and causal assumptions. Greater model capacity cannot substitute for missing interventions or unrecorded confounders.
Decision relevance of the future variable.The predicted variableYYmust preserve information relevant to the utility and subsequent decisions. Accurate prediction of visually salient but decision-irrelevant features does not guarantee useful control.
Acknowledgments
The author used a large language model as a writing assistant during the preparation of this manuscript. The LLM assisted with language editing, symbol consistency checking, and formatting.
Appendix AProofs
A.1Proof ofProposition4.3
For probability measuresPPandQQ,
TV(P,Q)=sup0≤f≤1|𝔼P[f]−𝔼Q[f]|.\operatorname{TV}(P,Q)=\sup_{0\leq f\leq 1}\left|\mathbb{E}_{P}[f]-\mathbb{E}_{Q}[f]\right|.Applying this variational characterization toPℳπP_{\mathcal{M}}^{\pi}andPℳπ′P_{\mathcal{M}}^{\pi^{\prime}}, withf=Uf=U, gives
TV(Pℳπ,Pℳπ′)=supU:𝒯T→[0,1]|𝔼τ∼Pℳπ[U(τ)]−𝔼τ∼Pℳπ′[U(τ)]|=dctrlℳ,T(π,π′).\begin{split}\operatorname{TV}\left(P_{\mathcal{M}}^{\pi},P_{\mathcal{M}}^{\pi^{\prime}}\right)&=\sup_{U:\mathcal{T}_{T}\rightarrow[0,1]}\left|\mathbb{E}_{\tau\sim P_{\mathcal{M}}^{\pi}}[U(\tau)]-\mathbb{E}_{\tau\sim P_{\mathcal{M}}^{\pi^{\prime}}}[U(\tau)]\right|\\ &=d_{\mathrm{ctrl}}^{\mathcal{M},T}(\pi,\pi^{\prime}).\end{split}
A.2Proof ofTheorem4.4
Define
π¯(a∣h)=∫qF(y∣h)qI(a∣h,y)𝑑y.\bar{\pi}(a\mid h)=\int q_{F}(y\mid h)q_{I}(a\mid h,y)\,dy.By construction,
π¯(⋅∣h)=πWA(⋅∣h)\bar{\pi}(\cdot\mid h)=\pi_{\mathrm{WA}}(\cdot\mid h)at every history.
Starting from the same initial-history distribution, the two controllers induce the same action distribution conditional on every common history. Because they also share the same environment transition and observation kernels, they induce the same next-history distribution. Induction over time gives
Pℳ(qF,qI)=Pℳπ¯.P_{\mathcal{M}}^{(q_{F},q_{I})}=P_{\mathcal{M}}^{\bar{\pi}}. Conversely, letπ∈𝔎(𝒜∣ℋ)\pi\in\mathfrak{K}(\mathcal{A}\mid\mathcal{H})and choosey0∈𝒴y_{0}\in\mathcal{Y}. Letδy0\delta_{y_{0}}be the Dirac probability measure concentrated aty0y_{0}, and define
qF(⋅∣h)=δy0,qI(a∣h,y0)=π(a∣h).q_{F}(\cdot\mid h)=\delta_{y_{0}},\qquad q_{I}(a\mid h,y_{0})=\pi(a\mid h).Then
πWA(a∣h)=∫δy0(dy)qI(a∣h,y)=π(a∣h).\begin{split}\pi_{\mathrm{WA}}(a\mid h)&=\int\delta_{y_{0}}(dy)q_{I}(a\mid h,y)\\ &=\pi(a\mid h).\end{split}Thus, every direct stochastic policy has a degenerate world-action representation.
A.3Proof ofTheorem4.6
LetHμH_{\mu}denote conditional entropy under the demonstration distribution. The direct behavior-cloning objective decomposes as
ℒA(πA)=Hμ(A∣H)+𝔼H[DKL(pμ(A∣H)∥πA(A∣H))].\begin{split}\mathcal{L}_{\mathrm{A}}(\pi_{\mathrm{A}})&=H_{\mu}(A\mid H)\\ &\quad+\mathbb{E}_{H}\left[D_{\mathrm{KL}}\left(p_{\mu}(A\mid H)\,\|\,\pi_{\mathrm{A}}(A\mid H)\right)\right].\end{split}Under realizability, the KL term can be minimized to zero, so
πA∗(a∣h)=pμ(a∣h)=πμ(a∣h)\pi_{\mathrm{A}}^{*}(a\mid h)=p_{\mu}(a\mid h)=\pi_{\mu}(a\mid h)forpμ(H)p_{\mu}(H)-almost everyhh.
Similarly,
ℒWA(qF,qI)=Hμ(Y∣H)+Hμ(A∣H,Y)+𝔼H[DKL(pμ(Y∣H)∥qF(Y∣H))]+𝔼H,Y[DKL(pμ(A∣H,Y)∥qI(A∣H,Y))].\begin{split}\mathcal{L}_{\mathrm{WA}}(q_{F},q_{I})&=H_{\mu}(Y\mid H)+H_{\mu}(A\mid H,Y)\\ &\quad+\mathbb{E}_{H}\left[D_{\mathrm{KL}}\left(p_{\mu}(Y\mid H)\,\|\,q_{F}(Y\mid H)\right)\right]\\ &\quad+\mathbb{E}_{H,Y}\left[D_{\mathrm{KL}}\left(p_{\mu}(A\mid H,Y)\,\|\,q_{I}(A\mid H,Y)\right)\right].\end{split}At an exact population optimum,
qF∗(y∣h)=pμ(y∣h),qI∗(a∣h,y)=pμ(a∣h,y)q_{F}^{*}(y\mid h)=p_{\mu}(y\mid h),\qquad q_{I}^{*}(a\mid h,y)=p_{\mu}(a\mid h,y)almost everywhere. Therefore,
πWA∗(a∣h)=∫pμ(y∣h)pμ(a∣h,y)𝑑y=∫pμ(a,y∣h)𝑑y=pμ(a∣h)=πμ(a∣h).\begin{split}\pi_{\mathrm{WA}}^{*}(a\mid h)&=\int p_{\mu}(y\mid h)p_{\mu}(a\mid h,y)\,dy\\ &=\int p_{\mu}(a,y\mid h)\,dy\\ &=p_{\mu}(a\mid h)\\ &=\pi_{\mu}(a\mid h).\end{split} Under behavior sufficiency,dμ,td_{\mu,t}is generated by deployingπμ\pi_{\mu}inℳ\mathcal{M}. Atdμ,td_{\mu,t}-almost every history, the three policies have the same action kernel. Induction over time yields
PℳπWA∗=PℳπA∗=Pℳπμ.P_{\mathcal{M}}^{\pi_{\mathrm{WA}}^{*}}=P_{\mathcal{M}}^{\pi_{\mathrm{A}}^{*}}=P_{\mathcal{M}}^{\pi_{\mu}}.
A.4Proof ofProposition4.8
Fix a historyhhand define
Qh(dy,da)=qF(dy∣h)qI(da∣h,y),Q_{h}(dy,da)=q_{F}(dy\mid h)q_{I}(da\mid h,y),Ph(dy,da)=pμ(dy∣h)pμ(da∣h,y),P_{h}(dy,da)=p_{\mu}(dy\mid h)p_{\mu}(da\mid h,y),and
Rh(dy,da)=pμ(dy∣h)qI(da∣h,y).R_{h}(dy,da)=p_{\mu}(dy\mid h)q_{I}(da\mid h,y). Marginalization contracts total variation, so
TV(πWA(⋅∣h),πμ(⋅∣h))≤TV(Qh,Ph).\operatorname{TV}\left(\pi_{\mathrm{WA}}(\cdot\mid h),\pi_{\mu}(\cdot\mid h)\right)\leq\operatorname{TV}(Q_{h},P_{h}).By the triangle inequality,
TV(Qh,Ph)≤TV(Qh,Rh)+TV(Rh,Ph).\operatorname{TV}(Q_{h},P_{h})\leq\operatorname{TV}(Q_{h},R_{h})+\operatorname{TV}(R_{h},P_{h}).The two terms are bounded byϵF\epsilon_{F}andϵI\epsilon_{I}, respectively. Hence
TV(πWA(⋅∣h),πμ(⋅∣h))≤ϵF+ϵI.\operatorname{TV}\left(\pi_{\mathrm{WA}}(\cdot\mid h),\pi_{\mu}(\cdot\mid h)\right)\leq\epsilon_{F}+\epsilon_{I}.Applying the triangle inequality throughπμ\pi_{\mu}gives
TV(πWA(⋅∣h),πA(⋅∣h))≤ϵ.\operatorname{TV}\left(\pi_{\mathrm{WA}}(\cdot\mid h),\pi_{\mathrm{A}}(\cdot\mid h)\right)\leq\epsilon. For the sequential bound, use maximal coupling at each common history. As long as the trajectory prefixes agree, the actions can be coupled to agree with probability at least1−ϵ1-\epsilon. Conditional on equal actions, the same environment kernel can be used to couple the next observations. Therefore,
TV(PℳπWA,PℳπA)≤1−(1−ϵ)T≤Tϵ.\operatorname{TV}\left(P_{\mathcal{M}}^{\pi_{\mathrm{WA}}},P_{\mathcal{M}}^{\pi_{\mathrm{A}}}\right)\leq 1-(1-\epsilon)^{T}\leq T\epsilon.The conclusion follows fromProposition4.3.
A.5Proof ofTheorem5.1
For a fixed historyhh,
∑a∈𝒜ρ(a∣h)DKL(𝖳a(⋅∣h)∥g(⋅∣h))=∑a∈𝒜ρ(a∣h)∑y∈𝒴𝖳a(y∣h)log𝖳a(y∣h)g(y∣h).\begin{split}&\sum_{a\in\mathcal{A}}\rho(a\mid h)D_{\mathrm{KL}}\left(\mathsf{T}_{a}(\cdot\mid h)\,\|\,g(\cdot\mid h)\right)\\ &=\sum_{a\in\mathcal{A}}\rho(a\mid h)\sum_{y\in\mathcal{Y}}\mathsf{T}_{a}(y\mid h)\log\frac{\mathsf{T}_{a}(y\mid h)}{g(y\mid h)}.\end{split}Insert𝖳ρ\mathsf{T}_{\rho}:
log𝖳ag=log𝖳a𝖳ρ+log𝖳ρg.\log\frac{\mathsf{T}_{a}}{g}=\log\frac{\mathsf{T}_{a}}{\mathsf{T}_{\rho}}+\log\frac{\mathsf{T}_{\rho}}{g}.The first term gives
Iρint(A;Y∣H=h).I_{\rho}^{\mathrm{int}}(A;Y\mid H=h).Using
∑a∈𝒜ρ(a∣h)𝖳a(y∣h)=𝖳ρ(y∣h),\sum_{a\in\mathcal{A}}\rho(a\mid h)\mathsf{T}_{a}(y\mid h)=\mathsf{T}_{\rho}(y\mid h),the second term gives
DKL(𝖳ρ(⋅∣h)∥g(⋅∣h)).D_{\mathrm{KL}}\left(\mathsf{T}_{\rho}(\cdot\mid h)\,\|\,g(\cdot\mid h)\right).Averaging overH∼νH\sim\nuprovesEquation46. The second KL term is minimized to zero by
g(⋅∣h)=𝖳ρ(⋅∣h),g(\cdot\mid h)=\mathsf{T}_{\rho}(\cdot\mid h),which provesEquation47.
A.6Derivation ofEquation57
Conditional exchangeability gives
Y(a)⟂⟂A|H.Y(a)\perp\!\!\!\perp A\mid H.Therefore,
P(Y(a)=y∣H=h)=P(Y(a)=y∣H=h,A=a).P(Y(a)=y\mid H=h)=P(Y(a)=y\mid H=h,A=a).By consistency,
P(Y(a)=y∣H=h,A=a)=P(Y=y∣H=h,A=a).P(Y(a)=y\mid H=h,A=a)=P(Y=y\mid H=h,A=a).Thus,
𝖳a(y∣h)=pμ(y∣h,a).\mathsf{T}_{a}(y\mid h)=p_{\mu}(y\mid h,a).Under positivity,
pμ(y∣h,a)=pμ(y,a∣h)pμ(a∣h)=pμ(y∣h)pμ(a∣h,y)πμ(a∣h).\begin{split}p_{\mu}(y\mid h,a)&=\frac{p_{\mu}(y,a\mid h)}{p_{\mu}(a\mid h)}\\ &=\frac{p_{\mu}(y\mid h)p_{\mu}(a\mid h,y)}{\pi_{\mu}(a\mid h)}.\end{split}
A.7Proof ofTheorem5.4
We give separate constructions for support failure and hidden confounding.
Support failure.
LetHHbe constant and let𝒜={0,1}\mathcal{A}=\{0,1\}. The behavior policy always selectsA=0A=0. Define two environments by
Y(0)Y(1)ℳ100ℳ201.\begin{array}[]{c|cc}&Y(0)&Y(1)\\ \hline\cr\mathcal{M}_{1}&0&0\\ \mathcal{M}_{2}&0&1.\end{array}Both environments produce
P(A=0,Y=0)=1P(A=0,Y=0)=1under the behavior policy, so their observational distributions are identical. Underdo(A=1)\operatorname{do}(A=1),
Pℳ1(Y=1∣do(A=1))=0,P_{\mathcal{M}_{1}}(Y=1\mid\operatorname{do}(A=1))=0,whereas
Pℳ2(Y=1∣do(A=1))=1.P_{\mathcal{M}_{2}}(Y=1\mid\operatorname{do}(A=1))=1.
Hidden confounding under positive action support.
LetHHbe constant and let
Z∼Bernoulli(1/2)Z\sim\operatorname{Bernoulli}(1/2)be observed by the demonstrator but omitted fromHH. Let the demonstrator choose
Both actions have positive observational probability.
In environmentℳ1\mathcal{M}_{1}, define
In environmentℳ2\mathcal{M}_{2}, define
Under the demonstration mechanism, both environments produce
P(A=0,Y=0)=12,P(A=1,Y=1)=12.P(A=0,Y=0)=\frac{1}{2},\qquad P(A=1,Y=1)=\frac{1}{2}.Their observational distributions are therefore identical. Underdo(A=1)\operatorname{do}(A=1),
Pℳ1(Y=1∣do(A=1))=1,P_{\mathcal{M}_{1}}(Y=1\mid\operatorname{do}(A=1))=1,whereas
Pℳ2(Y=1∣do(A=1))=P(Z=1)=12.P_{\mathcal{M}_{2}}(Y=1\mid\operatorname{do}(A=1))=P(Z=1)=\frac{1}{2}.Positive observational action support is therefore insufficient whenHHomits an action–outcome confounder.
A.8Proof ofTheorem5.6
Consider a one-step problem with a single history and action space{0,1}\{0,1\}. The behavior policy always selects
Define two environments:
Y(0)Y(1)ℳ+1/21ℳ−1/20.\begin{array}[]{c|cc}&Y(0)&Y(1)\\ \hline\cr\mathcal{M}_{+}&1/2&1\\ \mathcal{M}_{-}&1/2&0.\end{array}Let the common known utility be
The observational demonstration is identical in the two environments:
A=0,Y=12.A=0,\qquad Y=\frac{1}{2}.Every observational learner therefore receives the same observational distribution and utility in both environments and must produce the same stochastic action distribution.
Letppbe the probability of selectingA=1A=1. Inℳ+\mathcal{M}_{+}, the regret is
Reg+(p)=12(1−p).\operatorname{Reg}_{+}(p)=\frac{1}{2}(1-p).Inℳ−\mathcal{M}_{-}, the regret is
Reg−(p)=p2.\operatorname{Reg}_{-}(p)=\frac{p}{2}.Therefore,
max{Reg+(p),Reg−(p)}=12max{1−p,p}≥14.\begin{split}\max\{\operatorname{Reg}_{+}(p),\operatorname{Reg}_{-}(p)\}&=\frac{1}{2}\max\{1-p,p\}\\ &\geq\frac{1}{4}.\end{split} The exact behavior-cloning and world-action population optima selectA=0A=0with probability one, sop=0p=0. Their worst-case regret is1/21/2.
Finally, one intervention selectingA=1A=1yields
Y=1inℳ+,Y=1\quad\text{in }\mathcal{M}_{+},and
Y=0inℳ−.Y=0\quad\text{in }\mathcal{M}_{-}.The intervention identifies the environment. SelectingA=1A=1inℳ+\mathcal{M}_{+}andA=0A=0inℳ−\mathcal{M}_{-}then achieves zero regret.
Appendix BAdditional Technical Remarks
B.1Point decoding
Suppose that, for a fixed history,
P(Y=0,A=0)=0.6,P(Y=1,A=1)=0.4.P(Y=0,A=0)=0.6,\qquad P(Y=1,A=1)=0.4.A distributional world-action model reproduces
P(A=0)=0.6,P(A=1)=0.4.P(A=0)=0.6,\qquad P(A=1)=0.4.A MAP-future decoder instead selectsY=0Y=0and always executesA=0A=0. The resulting controller can still be flattened into a direct deterministic policy, but it is not equal to the distributional behavior-cloning optimum.
For continuous actions under squared loss,
𝔼Y|H=h[𝔼[A∣H=h,Y]]=𝔼[A∣H=h]\mathbb{E}_{Y\mid H=h}\left[\mathbb{E}[A\mid H=h,Y]\right]=\mathbb{E}[A\mid H=h]by the tower property. Sampling one future and applying its conditional mean preserves equality in expectation but not necessarily equality of the full action distribution.
B.2Privileged demonstrator information
The hidden-confounding construction inSectionA.7also shows why matchingpμ(A∣H)p_{\mu}(A\mid H)need not reproduce the original demonstration trajectory distribution when behavior sufficiency fails.
In that construction,
πμ(A=1∣H)=12,\pi_{\mu}(A=1\mid H)=\frac{1}{2},but the demonstrated action satisfiesA=ZA=Z. Deploying an independent Bernoulli action with the same marginal does not reproduce the dependence betweenAAand the omitted variableZZ. IncludingZZin the history,
H′=(H,Z),H^{\prime}=(H,Z),restores the information used by the demonstrator.
B.3Action chunks and language conditioning
If
is an action chunk and
is the corresponding future chunk, the same marginalization identity holds:
∫pμ(yt+1:t+K∣ht)pμ(at:t+K−1∣ht,yt+1:t+K)dy=pμ(at:t+K−1∣ht).\begin{split}&\int p_{\mu}(y_{t+1:t+K}\mid h_{t})p_{\mu}(a_{t:t+K-1}\mid h_{t},y_{t+1:t+K})\,dy\\ &=p_{\mu}(a_{t:t+K-1}\mid h_{t}).\end{split} Goals and language instructions are handled by includingGGin the history. The results then apply conditionally onGG, provided that the compared policies receive the same instruction and context at deployment.
B.4Restricted parametric classes
LetΠAparam\Pi_{\mathrm{A}}^{\mathrm{param}}be a restricted direct policy class and let
ΠWAparam=ΠWA(ℱ,ℐ).\Pi_{\mathrm{WA}}^{\mathrm{param}}=\Pi_{\mathrm{WA}}(\mathcal{F},\mathcal{I}).The flattening theorem guarantees
ΠWAparam⊆𝔎(𝒜∣ℋ),\Pi_{\mathrm{WA}}^{\mathrm{param}}\subseteq\mathfrak{K}(\mathcal{A}\mid\mathcal{H}),but not necessarily
ΠWAparam⊆ΠAparam.\Pi_{\mathrm{WA}}^{\mathrm{param}}\subseteq\Pi_{\mathrm{A}}^{\mathrm{param}}. A restricted direct class reproduces the world-action class only if it is closed under the required marginalization:
qF∈ℱ,qI∈ℐ⟹[h↦∫qF(y∣h)qI(⋅∣h,y)dy]∈ΠAparam.q_{F}\in\mathcal{F},\ q_{I}\in\mathcal{I}\quad\Longrightarrow\quad\left[h\mapsto\int q_{F}(y\mid h)q_{I}(\cdot\mid h,y)\,dy\right]\in\Pi_{\mathrm{A}}^{\mathrm{param}}.Finite architectures may therefore exhibit a genuine representational or computational advantage from the world-action factorization even though no such advantage exists relative to unrestricted stochastic policies.
B.5Observational recovery and model exploitation
Even when
pμ(Y∣H,A)=P(Y∣H,do(A))p_{\mu}(Y\mid H,A)=P(Y\mid H,\operatorname{do}(A))on the behavior support, an optimized policy may select actions or reach histories outside that support. Its performance then depends on extrapolation of the learned model. Causal identification on the observed support and accuracy under the optimized policy are different requirements.
This distinction motivates support constraints, pessimistic or robust planning, uncertainty-aware action selection, and online model correction.
References
- [1]X. Cao, F. Luo, J. Ye, T. Xu, Z. Zhang, and Y. Yu(2024)Limited preference aided imitation learning from imperfect demonstrations.InProceedings of the International Conference on Machine Learning,Cited by:§2.1.
- [2]R. Chen, X. Chen, Y. Sun, S. Xiao, M. Li, and Y. Yu(2024)Policy-conditioned environment models are more generalizable.InProceedings of the International Conference on Machine Learning,Cited by:§2.2.
- [3]X. Chen, Y. Yu, Z. Zhu, Z. Yu, Z. Chen, C. Wang, Y. Wu, H. Wu, R. Qin, R. Ding, and F. Huang(2023)Adversarial counterfactual environment model learning.InAdvances in Neural Information Processing Systems,Vol.36.Cited by:§2.3.
- [4]P. de Haan, D. Jayaraman, and S. Levine(2019)Causal confusion in imitation learning.InAdvances in Neural Information Processing Systems,Vol.32.Cited by:§2.3.
- [5]S. Jiang, J. Pang, and Y. Yu(2020)Offline imitation learning with a misspecified simulator.InAdvances in Neural Information Processing Systems,Vol.33.Cited by:§2.1.
- [6]Z. Li, T. Xu, Z. Qin, Y. Yu, and Z. Luo(2023)Imitation learning from imperfection: theoretical justifications and algorithms.InAdvances in Neural Information Processing Systems,Vol.36.Cited by:§2.1.
- [7]H. Lin, Y. Xu, Y. Sun, Z. Zhang, Y. Li, C. Jia, J. Ye, J. Zhang, and Y. Yu(2025)Any-step dynamics model improves future predictions for online and offline reinforcement learning.InProceedings of the International Conference on Learning Representations,Cited by:§2.2.
- [8]Y. Liu, B. Huang, Z. Zhu, H. Tian, M. Gong, Y. Yu, and K. Zhang(2023)Learning world models with identifiable factorization.InAdvances in Neural Information Processing Systems,Vol.36.Cited by:§2.3.
- [9]F. Luo, T. Xu, X. Cao, and Y. Yu(2024)Reward-consistent dynamics models are strongly generalizable for offline reinforcement learning.InProceedings of the International Conference on Learning Representations,Cited by:§2.2.
- [10]F. Luo, T. Xu, H. Lai, X. Chen, W. Zhang, and Y. Yu(2024)A survey on model-based reinforcement learning.Science China Information Sciences67(2),pp. 121101.Cited by:§1,§2.2.
- [11]M. Mou, Y. Guo, F. Luo, Y. Yu, and J. Zhang(2024)Model predictive complex system control from observational and interventional data.Chaos34,pp. 093125.Cited by:§2.3.
- [12]J. Pang, N. Tang, K. Li, Y. Tang, X. Cai, Z. Zhang, G. Niu, M. Sugiyama, and Y. Yu(2025)Learning view-invariant world models for visual robotic manipulation.InProceedings of the International Conference on Learning Representations,Cited by:§2.2.
- [13]J. Pearl(2009)Causality: models, reasoning, and inference.2 edition,Cambridge University Press.Cited by:§2.3.
- [14]S. Ross, G. Gordon, and D. Bagnell(2011)A reduction of imitation learning and structured prediction to no-regret online learning.InProceedings of the International Conference on Artificial Intelligence and Statistics,pp. 627–635.Cited by:§2.1.
- [15]Y. Sun, J. Zhang, C. Jia, H. Lin, J. Ye, and Y. Yu(2023)Model-bellman inconsistency for model-based offline reinforcement learning.InProceedings of the International Conference on Machine Learning,Cited by:§2.2.
- [16]Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang(2025)Predictive inverse dynamics models are scalable learners for robotic manipulation.InInternational Conference on Learning Representations,Cited by:§1,§2.2.
- [17]T. Xu, Z. Li, and Y. Yu(2020)Error bounds of imitating policies and environments.InAdvances in Neural Information Processing Systems,Vol.33.Cited by:§2.1.
- [18]T. Xu, Z. Li, and Y. Yu(2022)Error bounds of imitating policies and environments for reinforcement learning.IEEE Transactions on Pattern Analysis and Machine Intelligence44(10),pp. 6968–6980.Cited by:§2.1.
- [19]S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian,et al.(2026)World action models are zero-shot policies.External Links:2602.15922Cited by:§1,§2.2.
- [20]Z. Zhang, H. Ren, Y. Sun, Y. Sheng, H. Wang, Z. Wu, H. Lin, P. Bacon, and Y. Yu(2026)Towards practical world model-based reinforcement learning for vision-language-action models.InProceedings of the International Conference on Machine Learning,Cited by:§2.2.
- [21]Z. Zhu, S. Jiang, Y. Liu, Y. Yu, and K. Zhang(2022)Invariant action effect model for reinforcement learning.InProceedings of the AAAI Conference on Artificial Intelligence,Cited by:§2.3.
- [22]Z. Zhu, H. Tian, X. Chen, K. Zhang, and Y. Yu(2025)Offline model-based reinforcement learning with causal structured world models.Frontiers of Computer Science19(4),pp. 194347.Cited by:§2.3.
Similar Articles
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
This paper introduces Agent-Authored World Modeling (AAWM), a training procedure that constructs world-model supervision based on the policy's own decision needs rather than next-observation prediction, aligning the learning objective with the dynamics required for effective decision-making.
World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning
This paper proposes Privileged-Future On-Policy Self-Distillation (PF-OPSD) for controlled concrete reasoning, combining world models' visual simulation with language models' abstract reasoning to improve prediction accuracy and robustness on two new benchmarks.
World Model for Robot Learning: A Comprehensive Survey
This comprehensive survey reviews the literature on world models for robot learning, covering their roles in policy learning, planning, and simulation. It highlights key paradigms, benchmarks, and future directions for predictive modeling in embodied agents.
Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
This paper proposes a strategic robustness objective for learning simulators in model-based reinforcement learning, formulated as a minimax game between a model player and an adversarial policy player. Theoretical guarantees and a provably convergent algorithm are provided, with experiments showing reduced prediction error and improved real-world policy transfer.
Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
This paper introduces Counterfactual Latent World Models (CLWM) to address counterfactual collapse in world models, improving planning success in embodied reasoning under partial observability through contrastive objectives.