Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

arXiv cs.LG Papers

Summary

This paper identifies two distinct failures in learned physical simulators—instability in long rollouts and inability to adapt to changed laws—and proposes separate structural solutions: symplectic integration for stability and explicit factorization for counterfactual generalization.

arXiv:2609.19674v1 Announce Type: new Abstract: A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to $100\times$ the training horizon, while equal-capacity predictors, an energy-regularized predictor, and a tuned neural ODE diverge. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law. Crucially, the two mechanisms are separable: removing the structure responsible for long-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability. This double dissociation, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three-body system and remains visible when the physical state must be inferred from pixels rather than provided directly. The result is a concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.
Original Article
View Cached Full Text

Cached at: 09/18/26, 09:02 AM

# Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models
Source: [https://arxiv.org/html/2609.19674](https://arxiv.org/html/2609.19674)
Parivesh PriyeAffiliation:Georgia Institute of TechnologyLu WeiAffiliation:Stony Brook UniversityHaibin LingAffiliation:Westlake University

###### Abstract

A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change\. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one\. We show that these two failures require*different structural remedies*\. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to100×100\\timesthe training horizon, while equal\-capacity predictors, an energy\-regularized predictor, and a tuned neural ODE diverge\. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never\-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law\. Crucially, the two mechanisms are separable: removing the structure responsible for long\-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability\. This*double dissociation*, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three\-body system and remains visible when the physical state must be inferred from pixels rather than provided directly\. The result is a concrete design principle for physical world models: long\-horizon stability and changed\-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other\.

## 1Introduction

A learned simulator predicts the next state of a physical system and iterates that prediction into a trajectory\. Modern models can perform well by conventional measures such as rollout accuracy or visual realism, yet they often fail in two distinct ways once the governing conditions move beyond training\. First, prediction errors accumulate until long rollouts drift or become unstable\. Second, when the physical law itself changes, for example through reversed gravity, a flipped coupling, or reversed time, the model often continues to follow the law seen during training\. Both failures remain visible in current systems\. The best reported video generator reaches only22%22\\%joint accuracy on the hard split of VideoPhy\-2, no model exceeds3\.3/53\.3/5on PhyGround, and PhyWorldBench includes a dedicated “anti\-physics” category for changed law failures\([Yi et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib21);[Bear and others, 2021](https://arxiv.org/html/2609.19674#bib.bib22);[Gu et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib24);[Bansal et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib26);[Nam et al\., 2026](https://arxiv.org/html/2609.19674#bib.bib27)\)\. Related limitations also appear in learned simulators across scales\([Sanchez\-Gonzalez et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib15);[Zhong et al\., 2021a](https://arxiv.org/html/2609.19674#bib.bib8)\)\. This paper asks:*which piece of physical structure fixes which failure?*

We address this question by encoding physical invariances directly rather than asking the model to infer them from finite data\. The ingredients are individually established, including Hamiltonian and Lagrangian neural networks, generalized coordinate models, constrained dynamics, and symplectic structure preserving parameterizations\([Greydanus et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib1);[Cranmer et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib2);[Lutter et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib3);[Finzi et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib7);[Zhong et al\., 2020b](https://arxiv.org/html/2609.19674#bib.bib4)\)\. What remains unclear is which structural commitment produces which behavior and where each commitment stops helping\. We show a*double dissociation*\. Conserving a learned energy through a symplectic update keeps long rollouts physically bounded but does not determine how the system should respond to a changed law\. Conversely, making an interaction force enter linearly through its physical coupling enables transfer to a never seen sign of that coupling but does not prevent long horizon drift\. Matched controls that remove or replace one ingredient at a time show that stability depends on geometric and energetic consistency, whereas counterfactual transfer depends on how the intervened parameter enters the dynamics\.

This paper makes four contributions\. First, we establish this double dissociation with matched controls, showing that conservation produces long horizon stability and a factored coupling produces counterfactual inversion, while neither produces the other and neither is explained by capacity, integrator choice, or formalism alone\. Second, we formalize the mechanism behind each behavior and map the boundary at which each prior stops helping\. Third, we show that the dissociation extends beyond smooth few body motion to generalized coordinates, contact, many body systems, mixed force families, and pixel based settings with grounded perception, including a learned object binder and frozen public video encoders\. Fourth, we extend the same structural account beyond conservative dynamics, using a passive friction port that reaches the correct temperature and a factored drive that generalizes across unseen forcing strengths and signs\.

We study a learned transition map from physical state, including positions and momenta, to the next state and iterate it into a trajectory\. To separate dynamics from perception, we provide state in three ways: as oracle state; as state recovered from synthetic renders by an encoder whose position channel is anchored by either a fixed renderer or a label free centroid; and as state discovered from pixels with no anchor\. The first two settings recover the target behaviors, while the third exposes a boundary of the current approach\. Accordingly, this study concerns physical law fidelity rather than visual fidelity and does not claim improvements in real video generation, image quality, or leaderboard performance\.

## 2Related work

Structured dynamics and parameter conditioning\.Hamiltonian and Lagrangian networks encode conservation laws or variational structure\([Greydanus et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib1);[Cranmer et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib2);[Lutter et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib3)\)\. Generalised coordinate, constrained, and intrinsically symplectic variants extend these ideas to rigid and structured systems\([Finzi et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib7);[Zhong et al\., 2020b](https://arxiv.org/html/2609.19674#bib.bib4);[Jin et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib6)\), with continuous conservation statements for the learned energy and discrete behaviour determined in part by the numerical integrator\([Zhong et al\., 2021a](https://arxiv.org/html/2609.19674#bib.bib8)\)\. Conditioning dynamics on physical parameters, including coupling interactions through multiplicative coefficients, is also standard\([Greydanus et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib1);[Cranmer et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib2);[Lutter et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib3);[Zhong et al\., 2020b](https://arxiv.org/html/2609.19674#bib.bib4);[Jin et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib6);[Brunton et al\., 2016](https://arxiv.org/html/2609.19674#bib.bib29);[Han et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib11);[Raissi et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib18)\), while context adaptation extends prediction to new parameter draws within the training family\([Kirchmeyer et al\., 2022](https://arxiv.org/html/2609.19674#bib.bib31);[Yin et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib32);[Wang et al\., 2022](https://arxiv.org/html/2609.19674#bib.bib33)\)\. Our setting differs in two respects\. First, we intervene on a coupling whose sign lies outside the training support\. Second, we compare matched factored and unrestricted versions of HNN, DeLaN, and SymODEN to determine which structural commitment supports stability and which supports counterfactual inversion\.

Simulators, dissipation, contacts, and perception\.Graph network simulators are strong relational predictors\([Sanchez\-Gonzalez et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib15);[Battaglia et al\., 2016](https://arxiv.org/html/2609.19674#bib.bib16);[Satorras et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib17)\), and training noise is a common remedy for rollout instability; in the simulator studied here, it delays but does not prevent unbounded drift\. Dissipative, port Hamiltonian, and GENERIC or metriplectic models separate reversible from irreversible dynamics\([Sosanya and Greydanus, 2022](https://arxiv.org/html/2609.19674#bib.bib12);[Zhong et al\., 2020a](https://arxiv.org/html/2609.19674#bib.bib5);[Desai et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib37);[Lee et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib36)\)\. D\-HNN evaluates unseen friction coefficients, port Hamiltonian networks recover energy, forcing, and dissipation, and compositional port Hamiltonian learning assembles systems from learned components\([Neary and Topcu, 2023](https://arxiv.org/html/2609.19674#bib.bib38)\)\. Differentiable contact models further extend structured dynamics to collisions\([Zhong et al\., 2021b](https://arxiv.org/html/2609.19674#bib.bib13);[Hochlehnert et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib14)\)\. On the perception side, Hamiltonian generative models and unsupervised Lagrangian models learn dynamics directly from images\([Toth et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib10);[Zhong and Leonard, 2020](https://arxiv.org/html/2609.19674#bib.bib19)\), while[Botev et al\. \(2021\)](https://arxiv.org/html/2609.19674#bib.bib20)identify continuous time reversibility as the most broadly useful physical prior for pixel based learning\. We show that structure alone does not induce a canonical latent representation, and that reversibility helps only when the underlying dynamics are themselves reversible\. While recent video world models and physical reasoning benchmarks report persistent inconsistencies with physical laws\([Yi et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib21);[Bear and others, 2021](https://arxiv.org/html/2609.19674#bib.bib22);[Riochet et al\., 2018](https://arxiv.org/html/2609.19674#bib.bib23);[Gu et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib24);[Assran et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib25);[Bansal et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib26);[Nam et al\., 2026](https://arxiv.org/html/2609.19674#bib.bib27);[Lin et al\., 2026](https://arxiv.org/html/2609.19674#bib.bib28)\), our work isolates the specific structural commitments required to guarantee long\-horizon stability and counterfactual generalisation\.

## 3Methods

### 3\.1Setup and the two behaviours we measure

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/overview.png)Figure 1:Overview of the two pieces of physical structure we impose \(left\), the double dissociation they produce \(centre\), and how each behaviour is measured and bounded \(right and bottom\)\.Figure[1](https://arxiv.org/html/2609.19674#S3.F1)summarizes the study\. On the left are the two structural commitments we impose: energy conservation, implemented by evolving a learned energyHθH\_\{\\theta\}with a symplectic step, and a factored coupling, in which the intervened parameterGGenters the force linearly while all remaining content is learned from forward law data\. The centre shows the resulting double dissociation\. Conservation provides stability down the columns, whereas factoring provides the changed law response across the rows, so only the model containing both structures is stable*and*answers the counterfactual\. The right and bottom specify the corresponding measurements, the matched controls used to rule out capacity, integrator choice, and formalism, and the boundaries at which each structure stops helping\. We make each component precise below\. ForNNbodies, the state isx=\(q,p\)x=\(q,p\), whereqqdenotes positions andppmomenta, and the parameter vectorc=\(G,m1,…\)c=\(G,m\_\{1\},\\dots\)contains the couplingGGand the masses\. A dynamics model is a transition map iterated into a rollout,

xt\+1=fθ\(xt,c\),x0:T=\(x0,fθ\(x0,c\),…\),x\_\{t\+1\}=f\_\{\\theta\}\(x\_\{t\},c\),\\qquad x\_\{0:T\}=\\big\(x\_\{0\},f\_\{\\theta\}\(x\_\{0\},c\),\\dots\\big\),\(1\)and is trained only on one step prediction from forward law data with attractive gravity,G\>0G\>0\. Our headline system is softened three body gravity,

V\(q\)=−∑i<jG​mi​mj∥qi−qj∥2\+ϵ2,V\(q\)=\-\\sum\_\{i<j\}\\frac\{Gm\_\{i\}m\_\{j\}\}\{\\sqrt\{\\lVert q\_\{i\}\-q\_\{j\}\\rVert^\{2\}\+\\epsilon^\{2\}\}\},with evaluation rollouts ofT=2000T=2000steps, one hundred times the training horizon\. The remaining systems, including Coulomb charges, colliding disks, linear drag, a double pendulum in generalised coordinates, a confined thermal bath, and driven particles, together with the datasets, optimiser, and training settings, are specified in App\.[A](https://arxiv.org/html/2609.19674#A1)\.

We ask two questions of the trained transition map\. The first is long horizon stability, measured by the relative drift of the true energy along a finite rollout,

D⁡\(t\)=\|E⁡\(xt\)−E⁡\(x0\)\|\|E⁡\(x0\)\|,t≤T,D\(t\)=\\frac\{\\lvert E\(x\_\{t\}\)\-E\(x\_\{0\}\)\\rvert\}\{\\lvert E\(x\_\{0\}\)\\rvert\},\\qquad t\\leq T,\(2\)whereEEis fixed by the reference Hamiltonian\. We report both the per episode median and the worst episode, and count nonfinite trajectories explicitly rather than assigning them an arbitrary numerical value\. The second behaviour is counterfactual inversion\. We train only on attractive gravity,G\>0G\>0, and evaluate the unseen repulsive regime,G<0G<0\. Writingx^0:T\\hat\{x\}\_\{0:T\}for the model rollout andx−0:T,x\+0:Tx^\{\-\}\_\{0:T\},x^\{\+\}\_\{0:T\}for the flipped law and training law reference trajectories from the same initial condition, respectively, the regime match indicator records which law the learned rollout follows,

RM=\[∥x^0:T−x0:T−∥<∥x^0:T−x0:T\+∥\],\\mathrm\{RM\}=\\mathbb\{1\}\\\!\\left\[\\lVert\\hat\{x\}\_\{0:T\}\-x^\{\-\}\_\{0:T\}\\rVert<\\lVert\\hat\{x\}\_\{0:T\}\-x^\{\+\}\_\{0:T\}\\rVert\\right\],\(3\)and we report normalized trajectory error, defined as squared trajectory error divided by the reference variance, alongside it \(Def\.[1](https://arxiv.org/html/2609.19674#Thmdefinition1)\)\.

### 3\.2Two structural commitments, and what each guarantees

A symplectic model evolves a learned HamiltonianHθH\_\{\\theta\}through a symplectic update\. In continuous time, the corresponding flow

q˙=∂Hθ∂p,p˙=−∂Hθ∂q\\dot\{q\}=\\frac\{\\partial H\_\{\\theta\}\}\{\\partial p\},\\qquad\\dot\{p\}=\-\\frac\{\\partial H\_\{\\theta\}\}\{\\partial q\}\(4\)conservesHθH\_\{\\theta\}\. The discrete leapfrog update, however, conserves neitherHθH\_\{\\theta\}nor the reference energy exactly\.

###### Proposition 1\(Near conservation of a modified Hamiltonian; formal statement in App\.[B](https://arxiv.org/html/2609.19674#A2)\)\.

Under the regularity, bounded orbit, and small step assumptions of backward error analysis, leapfrog applied to Eq\. equation[4](https://arxiv.org/html/2609.19674#S3.E4)preserves a modified HamiltonianH~θ=Hθ\+O⁡\(d​t2\)\\widetilde\{H\}\_\{\\theta\}=H\_\{\\theta\}\+O\(dt^\{2\}\)up to exponentially small error over exponentially long times\([Hairer et al\., 2006](https://arxiv.org/html/2609.19674#bib.bib30)\)\.

This statement concernsH~θ\\widetilde\{H\}\_\{\\theta\}, not the physical reference energy\. Moreover, a generic MLP potential need not be coercive, so boundedness of the true energy remains an empirical finite horizon property\. We therefore report learned energy conservation, physical energy drift, state escape, and trajectory fidelity as separate outcomes\.

The second commitment is a coupling factored potential, which makes the interaction force exactly linear inGG:

Uθ\(q\)=∑i<jφθ\(qi−qj\),V\(q,c\)=GUθ\(q\),Fi=−G∂qiUθ\(q\),U\_\{\\theta\}\(q\)=\\sum\_\{i<j\}\\varphi\_\{\\theta\}\(q\_\{i\}\-q\_\{j\}\),\\qquad V\(q,c\)=G\\,U\_\{\\theta\}\(q\),\\qquad F\_\{i\}=\-G\\,\\partial\_\{q\_\{i\}\}U\_\{\\theta\}\(q\),\(5\)whereφθ\\varphi\_\{\\theta\}is learned using attractive data only\.

###### Proposition 2\(Sign inversion by construction\)\.

For any learnedφθ\\varphi\_\{\\theta\}and anyGGindependent kinetic energy, the force in Eq\. equation[5](https://arxiv.org/html/2609.19674#S3.E5)is odd inGG,Fi​\(q,−G\)=−Fi​\(q,G\)F\_\{i\}\(q,\-G\)=\-F\_\{i\}\(q,G\)for allqq\. The dynamics at−G\-Gtherefore follow the same learned interaction under the reversed coupling without requiring training examples from the opposite sign regime\.

The model family supplies the sign inversion, while the learned interactionφθ\\varphi\_\{\\theta\}must still recover the correct magnitude and state dependence\. Counterfactual accuracy therefore holds only where the learned content remains accurate on the states visited by the flipped rollout\. By contrast, a nonfactored model conditions a black box potentialVθ​\(q,G\)V\_\{\\theta\}\(q,G\)onGGwithout imposing any relation betweenG\>0G\>0andG<0G<0\. Section[4\.2](https://arxiv.org/html/2609.19674#S4.SS2)shows that such models can represent the flipped law but do not learn to select it from forward law data alone\. Proofs are given in App\.[B](https://arxiv.org/html/2609.19674#A2)\.

### 3\.3Controls, and how claims are measured

Every claim is read against controls that remove or replace one structural ingredient at a time, so each behaviour can be attributed to the component it depends on\. Against conservation we compare an unstructured predictor, a fixed physical\-energy penalty, and a tuned neural ODE, all capacity\-matched, testing whether stability instead reflects capacity, the objective, or the mere use of an integrator\. Against factoring we compare a black box that receives the coupling without the linear form and one that injects it through positive\-localized features, testing whether inversion follows from capacity or conditioning alone\. We repeat the comparison across Hamiltonian, Lagrangian, and symplectic\-ODE formalisms and a graph\-network simulator\. Energy drift \(Eq\. equation[2](https://arxiv.org/html/2609.19674#S3.E2)\) and regime match \(Eq\. equation[3](https://arxiv.org/html/2609.19674#S3.E3)\) are the primary measures in chaotic regimes, with trajectory error secondary and divergence kept separate from finite error\. Every experiment is repeated across independent initialisations, reporting uncertainty as error bars, standard deviations, orpp\-values, and near\-binary outcomes are tested at the level of complete runs\. Each result is assigned one of four pre\-registered evidence classes \(well\-supported, directional, supported, boundary\)\. The full model list and matching protocol, the statistics, and the per\-run results are in App\.[A](https://arxiv.org/html/2609.19674#A1), App\.[C](https://arxiv.org/html/2609.19674#A3), and App\.[E](https://arxiv.org/html/2609.19674#A5)\.

## 4Results

In this section we first show that energy conservation stabilises long rollouts but does not answer a changed law\. We then show that a factored coupling answers the changed law but does not provide stability\. We combine both effects in a single dissociation table, rule out the main confounds, map where each structural prior stops helping, and finally show that the pattern extends beyond the headline system\.

### 4\.1Conservation makes long rollouts stable, but cannot answer a changed law

Table 1:Relative true\-energy driftD⁡\(t\)D\(t\)over the rollout\.The first question is whether conserving a learned energy stabilises long rollouts where alternative approaches do not\. We train the five models in Table[1](https://arxiv.org/html/2609.19674#S4.T1)using one\-step prediction on attractive\-gravity trajectories and roll each model out for20002000steps\. Every reported value comes from a single raw record that retains the finite status of each episode \(App\.[C\.1](https://arxiv.org/html/2609.19674#A3.SS1)\)\. In the table,\>104\{\>\}10^\{4\}denotes a per\-episode median above the display cap, while NF counts trajectories that become non\-finite by step20002000, out of192192episodes\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/fig1_a.png)Figure 2:Long\-horizon energy drift \(log scale, capped at10410^\{4\}\)\.The symplectic model is the only model that remains finite and below the10410^\{4\}cap in every episode \(Fig\.[2](https://arxiv.org/html/2609.19674#S4.F2)\)\. Remaining finite, however, is not the same as remaining accurate\. Its true\-energy drift grows from0\.400\.40at100100steps to7\.17\.1at20002000, corresponding to a710%710\\%error, because the near\-conservation result in Prop\.[1](https://arxiv.org/html/2609.19674#Thmproposition1)applies to a modified learned Hamiltonian rather than to the physical energy\. The one\-step MLP exceeds the cap in the median episode by step500500, and6161of192192trajectories become non\-finite between steps10001000and20002000\. An energy penalty withλ=1\\lambda=1does not improve this behavior, with123123of192192episodes becoming non\-finite\. Since a non\-finite trajectory has undefined rather than infinite error, we report non\-finite counts, energy drift, and rollout error separately instead of collapsing them into one ratio\. The defensible comparison is therefore finite rollout with large drift versus divergence, and the display cap should not be interpreted as a measured lower bound\.

The neural ODE is the stronger control: at matched step size, objective, capacity, and learning\-rate sweep, every trajectory stays finite and its100100\-step drift \(0\.080\.08\) is five times smaller than the symplectic model’s, yet by500500steps it reaches3×1023\\times 10^\{2\}–4×1034\\times 10^\{3\}and exceeds the cap by10001000\. The short horizon favours the neural ODE, the long horizon favours Hamiltonian structure, and since equal step size does not imply equal integration error for RK4 and leapfrog, the comparison is between complete updates rather than integrators in isolation\.

The most informative control is the energy\-penalised model, whose target is the analytic physical energy: its short\-horizon drift is far smaller than the symplectic model’s, yet grows without bound by20002000steps \(with trajectory error following\), so penalising the true energy delays drift without bounding it, and the two are always reported together\. Two further probes on the factored model sharpen this picture: a short\-rollout loss more than halves the20002000\-step drift while preserving inversion, and over10,00010\{,\}000steps the factored model’s drift saturates from1\.91\.9toward0\.890\.89while the non\-factored symplectic model keeps drifting\.

Stability is nevertheless all that conservation provides\. Trained only on attractive gravity, the same symplectic model does not follow the unseen repulsive law\. Its rollout remains closer to the attractive world, with regime match00\. We turn to that second behavior next\.

### 4\.2A factored coupling answers a changed law, but cannot make rollouts stable

Table 2:Following a never\-trained sign flip of the coupling toG∈\{−0\.5,−1\.0,−1\.5\}G\\in\\\{\-0\.5,\-1\.0,\-1\.5\\\}\.![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/fig1_b.png)Figure 3:Inversion error atG=−1G=\-1; only the factored model falls below nMSE11\.We next ask whether a model trained only on attractive gravity can follow the unseen repulsive law \(Table[2](https://arxiv.org/html/2609.19674#S4.T2)\)\. Regime match measures the fraction of rollouts that follow the repulsive rather than the attractive reference, while rollout error is measured against the true repulsive trajectory\. The factored model follows the flipped\-law regime at every tested coupling, with regime match11and nMSE0\.030\.03–0\.180\.18forG∈\{−0\.5,−1\.0,−1\.5\}G\\in\\\{\-0\.5,\-1\.0,\-1\.5\\\}\. In contrast, the non\-factored symplectic model remains bounded but never inverts, with regime match00at every coupling and nMSE near2\.72\.7\. Its rollout remains closer to the attractive reference\. AtG=−1G=\-1, only the factored model falls below nMSE11, by four orders of magnitude \(Fig\.[3](https://arxiv.org/html/2609.19674#S4.F3)\)\. The mechanism is Prop\.[2](https://arxiv.org/html/2609.19674#Thmproposition2): because the force is linear inGG, the sign reversal is supplied by construction, while the learned interaction shapeφθ\\varphi\_\{\\theta\}determines the magnitude of the repulsive trajectory\. The non\-factored model fails not because it lacks capacity, but because its black\-box potential imposes no relation between the unseenG<0G<0regime and the observedG\>0G\>0regime\.

The converse also holds: factoring provides the counterfactual but not stability\. Placed on a non\-conserving base, the factored coupling still inverts the regime but diverges over long rollouts \(§[4\.3](https://arxiv.org/html/2609.19674#S4.SS3)\)\. On smooth parameter shifts the distinction becomes quantitative rather than binary: every model’s one\-step error rises smoothly, and the factored model achieves the lowest extrapolation error \(1010to30×30\\timesbelow the training band\) for the coupling it explicitly factors, but gains no advantage for the mass parameter it does not\.

### 4\.3The dissociation, and the controls that rule out confounds

Table 3:The double dissociation, from one matched comparison on oracle three\-body gravity\.The two effects meet in Table[3](https://arxiv.org/html/2609.19674#S4.T3), which crosses energy conservation with coupling factorization in one matched comparison\. The stability column reports true\-energy drift at20002000steps, where bounded models remain below22and divergent models reach the10410^\{4\}display cap\. The inversion column reports regime match together with rollout nMSE under the sign flip\. The pattern is a clean double dissociation\. Stability appears when the update is symplectic, while inversion appears when the coupling is factored, and the two effects vary independently\. The factored symplectic model is the only model that both remains stable and answers the changed law; the unstructured model does neither\. We next remove the most plausible confounds, namely capacity, integrator choice, and the choice of formalism\.

The first concern is that the counterfactual benefit may come from some feature that accompanies the factored potential rather than from factorization itself\. We test this by varying the state representation, parameter map, and integrator one at a time using an exact linear diagnostic\. Each candidate fits a stiffness matrix by ridge regression to force labels containing a fixed small non\-gradient curl component and observation noise\. Hamiltonian candidates are projected onto the symmetric cone, while the conditional parameter map is constructed from positive\-localised features\. The design crosses leapfrog, RK4, and implicit midpoint at three step sizes in two coercive coupled\-oscillator systems\. Every candidate fits the training regime comparably, so the comparison is not driven by state escape \(paired runs, Holmp≈0\.003p\\approx 0\.003\)\. At matched in\-distribution fit, the factored map reaches error0\.00350\.0035at the unseen coupling, while an unrestricted map trained on the same data reaches0\.9850\.985\. The Hamiltonian representation keeps true\-energy error at55to8×10−48\\times 10^\{\-4\}, compared with0\.310\.31to0\.330\.33for an unrestricted vector field\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_transformer_dr.png)Figure 4:Stronger baselines on the sign flip:20002000\-step latent\-norm \(stability, log\), with each model’s inversion outcome marked\.Two controls clarify this result\. First, the energy advantage tracks the injected curl: at zero curl both representations drift little \(0\.00070\.0007and0\.0020\.002\), but as contamination rises to curl fractions0\.030\.03,0\.060\.06,0\.100\.10the unrestricted model climbs to0\.350\.35,0\.940\.94,4\.304\.30while the Hamiltonian stays near0\.00080\.0008, and since no configuration escapes, the factorial isolates fidelity and inversion rather than mere prevention of divergence\. Second, the inversion gap reflects the extrapolation prior, not capacity: under the same objective a multiplicative map transfers whether it has eight scalars or4\.54\.5k parameters, a flexible network receiving the coupling directly only partially extrapolates \(0\.0710\.071–0\.0750\.075\), and one receiving it through positive\-localised features stays near1\.01\.0despite the most parameters\. The transferable ingredient is the multiplicative structure, supplied rather than identified from positive\-support data\.

The separation persists across formalisms: on the engine gravity flip both factored models invert, while the three non\-factored structured models fail at comparable stability\. Capacity is ruled out directly: across a width sweep the unstructured model diverges at every width and the factored model inverts at every width, while adding capacity to the non\-factored conserving model only worsens its drift\. Graph\-network simulators give the complementary control: factoring their messages restores the short\-horizon response, but both variants fail the20002000\-step test, so parameter transport alone does not provide long\-horizon prediction \(App\.[C\.2](https://arxiv.org/html/2609.19674#A3.SS2)\)\.

A stronger architecture and a direct data remedy leave the dissociation unchanged \(Fig\.[4](https://arxiv.org/html/2609.19674#S4.F4)\)\. A permutation\-equivariant set\-transformer, used as a stronger relational black box, still diverges over20002000steps and fails to invert the unseen sign flip\. Domain randomisation trains a black box on both signs of the coupling and therefore recovers inversion because the counterfactual is now part of its training distribution, but it provides no stability and diverges like the ordinary black box\. Only the factored symplectic model both remains bounded and inverts, showing that neither attention nor data coverage substitutes for the relevant structural commitment\.

### 4\.4Where conservation and factoring each stop helping

Both structural priors have clear limits\. Under linear drag, a port\-Hamiltonian model dissipates energy with the correct sign and order of magnitude, at roughly half the true rate \(−0\.014/−0\.063\-0\.014/\-0\.063compared with−0\.033/−0\.091\-0\.033/\-0\.091\)\. A conserving model cannot dissipate at all\. Conservation is therefore the wrong prior once the physical system genuinely loses energy\.

A factored law with the wrong parameter dependence gives a subtler failure\. A linear\-GGprior trained on aG2G^\{2\}law fits the observed regime well \(0\.014≈0\.0100\.014\\approx 0\.010\) but predicts a spurious sign flip with error0\.3430\.343\. Only the correctly specifiedG2G^\{2\}prior restores the counterfactual \(0\.0220\.022\)\. A wrong interaction structure, such as a pairwise prior fitted to a three\-body interaction, instead produces poor in\-distribution fit \(0\.430\.43\), so coupling\-power misspecification is the more dangerous case, failing without an in\-distribution warning\. The ambiguity nonetheless has a constructive resolution: when linear \(GG\) and even \(\|G\|\|G\|\) couplings fit the positive\-support data equally well, positive\-only queries reach unique selection only half the time \(0\.500\.50across four candidate maps\), whereas a single fixed query atG=−1G=\-1, using no ground truth, reaches1\.001\.00, and a non\-adaptive query over the full domain is intermediate \(0\.7750\.775; App\.[C\.7](https://arxiv.org/html/2609.19674#A3.SS7)\)\.

Perception introduces a third boundary: a single frame recovers position and fields but not velocity, so the inferred state falls off the correct energy shell, and the limiting variable is momentum rather than perception as a whole \(App\.[C\.4](https://arxiv.org/html/2609.19674#A3.SS4)\)\. One further diagnostic sharpens the map\. Time\-reversal consistency helps exactly when the underlying dynamics are reversible, giving retrace match11under conservative contact and only partial recovery under dissipation, which refines the conclusion of[Botev et al\. \(2021\)](https://arxiv.org/html/2609.19674#bib.bib20)\. A truth\-assisted transport score from the bound in Prop\.[3](https://arxiv.org/html/2609.19674#Thmproposition3)ranks which representation and map roll out worst under the flip \(pooled Spearman0\.870\.87versus0\.090\.09for the initial\-state term alone\), a model\-selection diagnostic rather than a certificate\.

### 4\.5The dissociation is not an artefact of the headline system

Table 4:Sign\-flip inversion error \(nMSE atG<0G<0\) across state groundings\.The same two mechanisms extend well beyond smooth three\-body gravity with oracle state\. In generalized coordinates, a Hamiltonian model with a learned mass metric on a MuJoCo double pendulum keeps its energy drift near0\.10\.1while unstructured baselines diverge, and the gravity flip still inverts only when the coupling is factored, with regime match11versus00\. Under contact, the symplectic model remains bounded as collisions become stiffer while the unstructured model diverges \(App\.[C\.3](https://arxiv.org/html/2609.19674#A3.SS3)\)\. The factored coupling also transfers across object count and force family\. A model trained on three bodies still inverts with thirty\-two; it transfers to Coulomb charges under a held\-out sign combination, with error0\.520\.52compared with1\.421\.42for a black box; and a single factored model trained on a mixed gravity\-plus\-Coulomb generator inverts both laws about120×120\\timesmore accurately than a black box receiving the same couplings\.

The inversion comes from the factored channel itself rather than from additional data or an unrestricted shortcut\. Increasing the training set does not improve the non\-factored model and slightly worsens it: its inversion error rises from nMSE0\.760\.76with3232trajectories to2\.672\.67with512512, because additional attractive trajectories reinforce a continuation that becomes incorrect after the sign flip\. The structural channel is also necessary rather than incidental\. When a free black\-box shortcut is added alongside the factored pathway, it absorbs about42%42\\%of the force and degrades inversion\. Penalising the shortcut reduces its contribution to10−410^\{\-4\}and restores fidelity, improving nMSE from0\.430\.43to0\.090\.09\(App\.[C\.5](https://arxiv.org/html/2609.19674#A3.SS5)\)\.

Because neither target behavior should depend on a hand\-built state representation, we also ground the counterfactual in pixels \(Table[4](https://arxiv.org/html/2609.19674#S4.T4)\)\. A latent representation learned from images recovers inversion once a fixed canonical kinetic form is imposed, reaching approximately the oracle\-state error\. A decode\-free objective inverts once a label\-free image centroid anchors the latent gauge\. A learned object binder over a fixed detector restores inversion when removing object identity had destroyed it, succeeding on10/1010/10seeds compared with0/100/10for the unstructured twin \(exact binomialp≈10−3p\\approx 10^\{\-3\}\)\. The same structure transfers to frozen public video encoders that we did not train, including VideoMAE and V\-JEPA2, where the factored readout inverts more reliably than its unstructured counterpart\.

Increasing encoder scale alone does not recover the physical readout: across V\-JEPA2 from roughly300300M to11B parameters the inversion error is essentially flat \(1\.02×1\.02\\times\) and the dissociation persists at the largest scale, so perception\-side scale does not supply the counterfactual\. The split appears even within one public model: a readout recovers position and momentum from frozen V\-JEPA2 tokens at nMSE0\.280\.28–0\.430\.43, so the encoder carries the physical state, yet its action\-conditioned predictor under no\-op actions gives no improvement over repeating the final frame\. What the representation must expose is momentum: a per\-frame encoder captures position but not velocity, and inversion returns only when momentum is estimated by differencing features across frames, which lowers inversion error from0\.620\.62to0\.0960\.096and matches the video encoders\. The momentum boundary of §[4\.4](https://arxiv.org/html/2609.19674#S4.SS4)therefore reappears one level higher, in the learned representation \(App\.[C\.8](https://arxiv.org/html/2609.19674#A3.SS8)\)\.

The structural account also extends beyond conservative dynamics\. In a confined thermal bath, a positive\-semidefinite friction port with fluctuation\-dissipation noise reaches the correct stationary temperature, whereas a conserving model does not thermalise\. The same passive port cannot represent its anti\-damped counterpart \(Prop\.[4](https://arxiv.org/html/2609.19674#Thmproposition4)\), so energy injection must enter through a separate factored drive\. That drive recovers the correct driven steady state across both trained and held\-out strengths, including a never\-seen cooling sign \(Prop\.[5](https://arxiv.org/html/2609.19674#Thmproposition5)\)\. When the system is observed through video, a grounded energy\-balance estimator recovers the damping rate\. Applied to time\-reversed footage, it recovers a negative rate that the sign\-constrained port refuses to represent, providing an arrow\-of\-time measurement paired with an architectural refusal \(App\.[C\.6](https://arxiv.org/html/2609.19674#A3.SS6),[C\.8](https://arxiv.org/html/2609.19674#A3.SS8)\)\.

## 5Conclusion and Limitations

In this work we investigated which structural commitments make a learned world model obey physical laws, evaluating each through long\-horizon stability and counterfactual generalisation\. First, energy conservation through a symplectic update provides long\-horizon stability where equal\-capacity predictors, tuned neural ordinary differential equations, and fixed energy penalties fail, yet it does not improve counterfactual generalisation\. Second, making the interaction force linear in the physical coupling enables extrapolation to never\-seen signs of that coupling, which an unrestricted parameter map fails to recover, yet it does not improve stability\. Third, we map the boundary of each mechanism and identify where its structural prior ceases to help\. Fourth, we extend the same account beyond conservative dynamics, using a passive friction port that thermalises correctly and a factored drive that generalises across unseen forcing strengths and signs\. The broader lesson is that physical structure does not make a model uniformly more physical\. Specific commitments produce specific behaviours, allowing a modeller to combine structural components according to the capabilities a task requires\.

However some limitations remains\. Conservation becomes the wrong prior when the true system genuinely dissipates energy, while factoring helps only when the intervened parameter follows the supplied functional form\. A misspecified coupling power can therefore fit the training regime well while failing under intervention\. A fixed kinetic term and a label\-free anchor reduce the ambiguity, but joint learned detection and binding remain open\. Finally, all evidence in this study comes from synthetic simulators and rendered observations rather than real video, the general deployment guarantee of our work is the next step we aim to ensure\.

## AI use statement

A large language model assisted the authors in polishing the language of the manuscript\. The authors verified that all claims, proofs, mathematical formulations, and reported values are valid, checked them against the implementation and results, and take full responsibility for the manuscript\.

## Ethics statement

This work uses only synthetic physics simulations, so no human subjects, personal data, or scraped datasets are involved\. We do not foresee direct harmful applications of the findings, which concern the architectural conditions for physical generalisation in learned simulators\. The main integrity\-relevant practices, namely prospective committed protocols with metrics fixed in advance, equal\-capacity baselines, and the reporting of negative and boundary results, are described in the paper and its record appendices\.

## Reproducibility statement

App\.[A](https://arxiv.org/html/2609.19674#A1)specifies all systems, models, integrators, metrics, and the perception pipeline\. App\.[B](https://arxiv.org/html/2609.19674#A2)gives the derivations and proofs\. App\.[C](https://arxiv.org/html/2609.19674#A3)gives the per\-experiment protocols, seeds, and per\-seed results\. App\.[F](https://arxiv.org/html/2609.19674#A6)gives the pre\-registration protocol \(metrics fixed in advance, the equal\-capacity rule, and non\-finite capping\), the statistics conventions, and the compute and hyperparameter defaults \(Table[23](https://arxiv.org/html/2609.19674#A6.T23)\)\. Our implementation and the pre\-registration documents are provided as supplementary material\.

## References

- Assranet al\.\(2025\)M\. Assran, A\. Bardes, D\. Fan, Q\. Garrido, Y\. LeCun, M\. Rabbat, N\. Ballas,et al\.V\-jepa 2: self\-supervised video models enable understanding, prediction and planning\.arXiv preprint\.External Links:2506\.09985Cited by:[§C\.8](https://arxiv.org/html/2609.19674#A3.SS8.p8.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Bansalet al\.\(2025\)H\. Bansal, C\. Peng, Y\. Bitton, R\. Goldenberg, A\. Grover, and K\. ChangVideoPhy\-2: a challenging action\-centric physical commonsense evaluation in video generation\.arXiv preprint\.External Links:2503\.06800Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Battagliaet al\.\(2016\)P\. W\. Battaglia, R\. Pascanu, M\. Lai, D\. Rezende, and K\. KavukcuogluInteraction networks for learning about objects, relations and physics\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:1612\.00222Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p3.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Bearet al\.\(2021\)D\. M\. Bearet al\.Physion: evaluating physical prediction from vision in humans and machines\.InNeurIPS Datasets and Benchmarks,External Links:2106\.08261Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Botevet al\.\(2021\)A\. Botev, A\. Jaegle, P\. Wirnsberger, D\. Hennes, and I\. HigginsWhich priors matter? benchmarking models for learning latent dynamics\.InNeurIPS Datasets and Benchmarks,External Links:2111\.05458Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1),[§4\.4](https://arxiv.org/html/2609.19674#S4.SS4.p3.1)\.
- Bruntonet al\.\(2016\)S\. L\. Brunton, J\. L\. Proctor, and J\. N\. KutzDiscovering governing equations from data by sparse identification of nonlinear dynamical systems\.Proceedings of the National Academy of Sciences \(PNAS\)113\(15\),pp\. 3932–3937\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1517384113)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Cranmeret al\.\(2020\)M\. Cranmer, S\. Greydanus, S\. Hoyer, P\. Battaglia, D\. Spergel, and S\. HoLagrangian neural networks\.ICLR 2020 Deep Differential Equations Workshop\.External Links:2003\.04630Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p2.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Desaiet al\.\(2021\)S\. A\. Desai, M\. Mattheakis, D\. Sondak, P\. Protopapas, and S\. J\. RobertsPort\-Hamiltonian neural networks for learning explicit time\-dependent dynamical systems\.InPhysical Review E 104, 034312,External Links:[Document](https://dx.doi.org/10.1103/PhysRevE.104.034312)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Finziet al\.\(2020\)M\. Finzi, K\. A\. Wang, and A\. G\. WilsonSimplifying hamiltonian and lagrangian neural networks via explicit constraints\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2010\.13581Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p2.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Greydanuset al\.\(2019\)S\. Greydanus, M\. Dzamba, and J\. YosinskiHamiltonian neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:1906\.01563Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p2.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Guet al\.\(2025\)J\. Gu, X\. Liu, Y\. Zeng, A\. Nagarajan, F\. Zhu, D\. Hong, Y\. Fan, Q\. Yan, K\. Zhou, M\. Liu, and X\. E\. WangPhyWorldBench: a comprehensive evaluation of physical realism in text\-to\-video models\.arXiv preprint arXiv:2507\.13428\.External Links:2507\.13428Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Haireret al\.\(2006\)E\. Hairer, C\. Lubich, and G\. WannerGeometric numerical integration: structure\-preserving algorithms for ordinary differential equations\.2nd edition,Springer Series in Computational Mathematics,Springer\.External Links:[Document](https://dx.doi.org/10.1007/978-3-662-05018-7)Cited by:[Appendix B](https://arxiv.org/html/2609.19674#A2.p3.1),[Proposition 1](https://arxiv.org/html/2609.19674#Thmproposition1.p1.1.1)\.
- Hanet al\.\(2021\)C\. Han, B\. Glaz, M\. Haile, and Y\. LaiAdaptable hamiltonian neural networks\.Physical Review Research3\(2\),pp\. 023156\.External Links:2102\.13235,[Document](https://dx.doi.org/10.1103/PhysRevResearch.3.023156)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Hochlehnertet al\.\(2021\)A\. Hochlehnert, A\. Terenin, S\. Sæmundsson, and M\. P\. DeisenrothLearning contact dynamics using physically structured neural networks\.InArtificial Intelligence and Statistics \(AISTATS\),External Links:2102\.11206Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Jinet al\.\(2020\)P\. Jin, Z\. Zhang, A\. Zhu, Y\. Tang, and G\. E\. KarniadakisSympNets: intrinsic structure\-preserving symplectic networks for identifying hamiltonian systems\.Neural Networks\.External Links:2001\.03750Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Kirchmeyeret al\.\(2022\)M\. Kirchmeyer, Y\. Yin, J\. Donà, N\. Baskiotis, A\. Rakotomamonjy, and P\. GallinariGeneralizing to new physical systems via context\-informed dynamics adaptation\.InInternational Conference on Machine Learning \(ICML\),External Links:2202\.01889Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Leeet al\.\(2021\)K\. Lee, N\. Trask, and P\. StinisMachine learning structure preserving brackets for forecasting irreversible processes \(GFINNs\)\.InNeurIPS,External Links:2106\.12619Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p6.2),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Linet al\.\(2026\)J\. Lin, A\. Akbari, Y\. He, L\. Zhao, P\. Zhao, Y\. Wang,et al\.PhyGround: benchmarking physical reasoning in generative world models\.arXiv preprint\.External Links:2605\.10806Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Lutteret al\.\(2019\)M\. Lutter, C\. Ritter, and J\. PetersDeep lagrangian networks: using physics as model prior for deep learning\.InInternational Conference on Learning Representations \(ICLR\),External Links:1907\.04490Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p3.1),[§1](https://arxiv.org/html/2609.19674#S1.p2.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Namet al\.\(2026\)H\. Nam, Q\. Le Lidec, L\. Maes, Y\. LeCun, and R\. BalestrieroCausal\-jepa: learning world models through object\-level latent masking\.arXiv preprint\.External Links:2602\.11389Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Neary and Topcu \(2023\)C\. Neary and U\. TopcuCompositional learning of dynamical system models using port\-Hamiltonian neural networks\.InLearning for Dynamics and Control Conference \(L4DC\),Proceedings of Machine Learning Research, Vol\.211\.External Links:2212\.00893Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Oquabet al\.\(2024\)M\. Oquab, T\. Darcet, T\. Moutakanni, H\. Vo, M\. Szafraniec, V\. Khalidov, P\. Fernandez, D\. Haziza, F\. Massa, A\. El\-Nouby,et al\.DINOv2: learning robust visual features without supervision\.Transactions on Machine Learning Research\.External Links:2304\.07193Cited by:[§C\.8](https://arxiv.org/html/2609.19674#A3.SS8.p2.1)\.
- Pickupet al\.\(2014\)L\. C\. Pickup, Z\. Pan, D\. Wei, Y\. Shih, C\. Zhang, A\. Zisserman, B\. Schölkopf, and W\. T\. FreemanSeeing the arrow of time\.InIEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),External Links:[Document](https://dx.doi.org/10.1109/CVPR.2014.262)Cited by:[§C\.8](https://arxiv.org/html/2609.19674#A3.SS8.p14.1)\.
- Raissiet al\.\(2019\)M\. Raissi, P\. Perdikaris, and G\. E\. KarniadakisPhysics\-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations\.Journal of Computational Physics378,pp\. 686–707\.External Links:[Document](https://dx.doi.org/10.1016/j.jcp.2018.10.045)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Riochetet al\.\(2018\)R\. Riochet, M\. Y\. Castro, M\. Bernard, A\. Lerer, R\. Fergus, V\. Izard, and E\. DupouxIntPhys: a framework and benchmark for visual intuitive physics reasoning\.arXiv preprint\.External Links:1803\.07616,[Document](https://dx.doi.org/10.1109/TPAMI.2021.3083839)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Sanchez\-Gonzalezet al\.\(2020\)A\. Sanchez\-Gonzalez, J\. Godwin, T\. Pfaff, R\. Ying, J\. Leskovec, and P\. W\. BattagliaLearning to simulate complex physics with graph networks\.InInternational Conference on Machine Learning \(ICML\),External Links:2002\.09405Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p3.1),[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Satorraset al\.\(2021\)V\. G\. Satorras, E\. Hoogeboom, and M\. WellingE\(n\) equivariant graph neural networks\.InInternational Conference on Machine Learning \(ICML\),External Links:2102\.09844Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Sosanya and Greydanus \(2022\)A\. Sosanya and S\. GreydanusDissipative hamiltonian neural networks: learning dissipative and conservative dynamics separately\.arXiv preprint\.External Links:2201\.10085Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Tao \(2016\)M\. TaoExplicit symplectic approximation of nonseparable hamiltonians: algorithm and long time performance\.Physical Review E94\(4\),pp\. 043303\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevE.94.043303)Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p4.1),[Table 23](https://arxiv.org/html/2609.19674#A6.T23.2.11.2.1.1)\.
- Tonget al\.\(2022\)Z\. Tong, Y\. Song, J\. Wang, and L\. WangVideoMAE: masked autoencoders are data\-efficient learners for self\-supervised video pre\-training\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2203\.12602Cited by:[§C\.8](https://arxiv.org/html/2609.19674#A3.SS8.p8.1)\.
- Tothet al\.\(2020\)P\. Toth, D\. J\. Rezende, A\. Jaegle, S\. Racanière, A\. Botev, and I\. HigginsHamiltonian generative networks\.InInternational Conference on Learning Representations \(ICLR\),External Links:1909\.13789Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Wanget al\.\(2022\)R\. Wang, R\. Walters, and R\. YuMeta\-learning dynamics forecasting using task inference\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2102\.10271,[Document](https://dx.doi.org/10.52202/068431-1573)Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Weiet al\.\(2018\)D\. Wei, J\. J\. Lim, A\. Zisserman, and W\. T\. FreemanLearning and using the arrow of time\.InIEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),External Links:[Document](https://dx.doi.org/10.1109/CVPR.2018.00840)Cited by:[§C\.8](https://arxiv.org/html/2609.19674#A3.SS8.p14.1)\.
- Yiet al\.\(2020\)K\. Yi, C\. Gan, Y\. Li, P\. Kohli, J\. Wu, A\. Torralba, and J\. B\. TenenbaumCLEVRER: collision events for video representation and reasoning\.International Conference on Learning Representations \(ICLR\)\.External Links:1910\.01442Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Yinet al\.\(2021\)Y\. Yin, I\. Ayed, E\. de Bézenac, N\. Baskiotis, and P\. GallinariLEADS: learning dynamical systems that generalize across environments\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2106\.04546Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Zhonget al\.\(2020a\)Y\. D\. Zhong, B\. Dey, and A\. ChakrabortyDissipative symoden: encoding hamiltonian dynamics with dissipation and control into deep learning\.ICLR 2020 Workshop DeepDiffEq\.External Links:2002\.08860Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Zhonget al\.\(2020b\)Y\. D\. Zhong, B\. Dey, and A\. ChakrabortySymplectic ode\-net: learning hamiltonian dynamics with control\.InInternational Conference on Learning Representations \(ICLR\),External Links:1909\.12077Cited by:[Appendix A](https://arxiv.org/html/2609.19674#A1.p3.1),[§1](https://arxiv.org/html/2609.19674#S1.p2.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Zhonget al\.\(2021a\)Y\. D\. Zhong, B\. Dey, and A\. ChakrabortyBenchmarking energy\-conserving neural networks for learning dynamics from data\.InLearning for Dynamics and Control \(L4DC\),External Links:2012\.02334Cited by:[§1](https://arxiv.org/html/2609.19674#S1.p1.1),[§2](https://arxiv.org/html/2609.19674#S2.p1.1)\.
- Zhonget al\.\(2021b\)Y\. D\. Zhong, B\. Dey, and A\. ChakrabortyExtending lagrangian and hamiltonian neural networks with differentiable contact models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2102\.06794Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.
- Zhong and Leonard \(2020\)Y\. D\. Zhong and N\. LeonardUnsupervised learning of lagrangian dynamics from images for prediction and control\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2007\.01926Cited by:[§2](https://arxiv.org/html/2609.19674#S2.p2.1)\.

## Appendix AFormal setup and definitions

A system ofNNbodies has statex=\(q,p\)x=\(q,p\), with positionsq∈ℝN​dq\\in\\mathbb\{R\}^\{Nd\}and conjugate momentap∈ℝN​dp\\in\\mathbb\{R\}^\{Nd\}\(ddis the spatial dimension, equal to22except for the pendulum, whereqqare joint angles\)\. A parameter vectorc=\(G,\{mi\}\)c=\(G,\\\{m\_\{i\}\\\}\)collects the coupling constant and masses\. A dynamics model is a mapfθ:\(xt,c\)↦xt\+1f\_\{\\theta\}:\(x\_\{t\},c\)\\mapsto x\_\{t\+1\}, iterated to a rolloutx0:Tx\_\{0:T\}\. Ground truth comes from a reference integratorΦ\\Phi\. Training minimises one\-step squared error on forward\-law data only\. Thekk\-step variant instead sums this error over akk\-step rollout\.

The evaluated systems are as follows\. \(i\)Conservative gravity\(oracle state\):H=∑i∥pi∥2/2​mi\+V⁡\(q\)H=\\sum\_\{i\}\\lVert p\_\{i\}\\rVert^\{2\}/2m\_\{i\}\+V\(q\)with softened\-core potential

V\(q\)=−∑i<jG​mi​mj∥qi−qj∥2\+ϵ2,ϵ=0\.1,dt=0\.002\.V\(q\)=\-\\sum\_\{i<j\}\\frac\{G\\,m\_\{i\}m\_\{j\}\}\{\\sqrt\{\\lVert q\_\{i\}\-q\_\{j\}\\rVert^\{2\}\+\\epsilon^\{2\}\}\},\\qquad\\epsilon=0\.1,\\;dt=0\.002\.\(6\)The softening keeps the ground\-truth integrator near\-conservative \(drift∼0\.01\{\\sim\}0\.01;ϵ=10−2\\epsilon=10^\{\-2\}increases the reference drift by110×110\\times\)\. This softening and the associated reference\-integrator drift are reported with all gravity results\. \(ii\)Coulomb: the same form withG​mi​mjGm\_\{i\}m\_\{j\}replaced by a charge productk​zi​zjk\\,z\_\{i\}z\_\{j\},zi∈\{±1\}z\_\{i\}\\in\\\{\\pm 1\\\}\. \(iii\)Springs\-with\-cutoff\(necessity screen\)\. \(iv\)Soft\-sphere diskson a periodic torus \(minimum\-image\), contact stiffness series set by the stiffnessk∈\{30,300,∞\}k\\in\\\{30,300,\\infty\\\}\. \(v\)Dissipative gravity: Eq\. equation[6](https://arxiv.org/html/2609.19674#A1.E6)plus linear drag,p˙i=−∂V/∂qi−γpi/mi\\dot\{p\}\_\{i\}=\-\\partial V/\\partial q\_\{i\}\-\\gamma\\,p\_\{i\}/m\_\{i\}\. \(vi\)Double pendulum\(MuJoCo\): joint anglesqq, configuration\-dependentM⁡\(q\)M\(q\), generalised momentump=M⁡\(q\)​q˙p=M\(q\)\\dot\{q\}, RK4 engine ground truth; conservative and joint\-damped variants\.

Parameter counts are approximately matched in the five\-model gravity comparison\. Other comparisons use the experiment\-specific budgets reported below\. Equal parameter count does not imply equal functional capacity\. The unstructured MLP is a residual MLP,xt\+1=xt\+gθ​\(xt,c\)x\_\{t\+1\}=x\_\{t\}\+g\_\{\\theta\}\(x\_\{t\},c\)\. The energy\-penalised model is the unstructured MLP plus an energy\-drift penaltyλ​\|Eθ​\(xt\+1\)−Eθ​\(xt\)\|\\lambda\\,\\lvert E\_\{\\theta\}\(x\_\{t\+1\}\)\-E\_\{\\theta\}\(x\_\{t\}\)\\rvert\. The interpretation of this baseline depends on howEθE\_\{\\theta\}is obtained: a jointly learned unconstrained energy can become constant and make the penalty vacuous\. Its failure alone therefore cannot establish the inadequacy of a penalty evaluated using a fixed physical energy\. The penalty energy is*fixed and analytic*, not learned or pretrained: the true physical energyE=∑i∥pi∥2/2​mi−G​∑i<jmi​mj/ri​j2\+ϵ2E=\\sum\_\{i\}\\lVert p\_\{i\}\\rVert^\{2\}/2m\_\{i\}\-G\\sum\_\{i<j\}m\_\{i\}m\_\{j\}/\\sqrt\{r\_\{ij\}^\{2\}\+\\epsilon^\{2\}\}with the true masses andGGand the same softeningϵ=0\.1\\epsilon=0\.1as the generator, applied asλ​\(E⁡\(x^t\+1\)−E⁡\(xt\)\)2\\lambda\\,\(E\(\\hat\{x\}\_\{t\+1\}\)\-E\(x\_\{t\}\)\)^\{2\}on the de\-standardised prediction\. The penalty does not by itself exclude the constant\-energy solution \(a wrong trajectory that holdsEEnear its initial value\), which is why rollout error is always co\-reported with energy drift\. In the raw record theλ=100\\lambda=100arm does not take that route: its energy error and its rollout error grow together \(drift115115–\>104\{\>\}10^\{4\}with normalised rollout error8282–\>104\{\>\}10^\{4\}at20002000steps\), after an early phase in which its energy error is below the symplectic model’s \(Table[6](https://arxiv.org/html/2609.19674#A3.T6)\)\. The symplectic model is a symplectic HNN,Hθ=Tθ​\(p\)\+Vθ​\(q,c\)H\_\{\\theta\}=T\_\{\\theta\}\(p\)\+V\_\{\\theta\}\(q,c\), one leapfrog step\. The factored symplectic model is the factored HNN, potential Eq\. equation[5](https://arxiv.org/html/2609.19674#S3.E5); the canonical*fixed\-kinetic*variant setsT=∑i∥pi∥2/2​miT=\\sum\_\{i\}\\lVert p\_\{i\}\\rVert^\{2\}/2m\_\{i\}\. DeLaN\([Lutter et al\., 2019](https://arxiv.org/html/2609.19674#bib.bib3)\): LagrangianLθ=12​q˙⊤​Mθ​\(q\)​q˙−V⁡\(q,c\)L\_\{\\theta\}=\\tfrac\{1\}\{2\}\\dot\{q\}^\{\\top\}M\_\{\\theta\}\(q\)\\dot\{q\}\-V\(q,c\), Euler–Lagrange accelerations, factored/ non\-factoredVV\. SymODEN\([Zhong et al\., 2020b](https://arxiv.org/html/2609.19674#bib.bib4)\): control\-Hamiltonian, learnedMθ​\(q\)−1M\_\{\\theta\}\(q\)^\{\-1\}andVθV\_\{\\theta\}, RK4\. Generalised\-coordinate HNN:H=12​p⊤​Mθ​\(q\)−1​p\+g​Uθ​\(q\)H=\\tfrac\{1\}\{2\}p^\{\\top\}M\_\{\\theta\}\(q\)^\{\-1\}p\+g\\,U\_\{\\theta\}\(q\), Cholesky\-SPD metric, factored potential\. Port\-Hamiltonian: conservative flow plus a learned dissipative port−γ​Rθ​\(p\)\-\\gamma\\,R\_\{\\theta\}\(p\)\. The graph\-network simulator and its factored variant\([Sanchez\-Gonzalez et al\., 2020](https://arxiv.org/html/2609.19674#bib.bib15);[Battaglia et al\., 2016](https://arxiv.org/html/2609.19674#bib.bib16)\): an interaction network; the factored form multiplies the aggregated message byGG\. Hybrid \(necessity\): factored pathway plus a free black\-box shortcut with anL2L\_\{2\}penaltyλ\\lambdaon the shortcut force\.

Separable Hamiltonians use symplectic leapfrog\. The non\-separable generalised\-coordinate Hamiltonian uses Tao’s explicit symplectic scheme\([Tao, 2016](https://arxiv.org/html/2609.19674#bib.bib9)\), admitted only after a calibration gate \(a known Hamiltonian must stay within10−410^\{\-4\}relative energy drift over 2000 steps\)\.

The relative energy drift is

D⁡\(t\)=\|E⁡\(xt\)−E⁡\(x0\)\|\|E⁡\(x0\)\|,D\(t\)=\\frac\{\\lvert E\(x\_\{t\}\)\-E\(x\_\{0\}\)\\rvert\}\{\\lvert E\(x\_\{0\}\)\\rvert\},\(7\)withEEthe true energy of the predicted state; regime\-match and the affine identifiability residual are Def\.[1](https://arxiv.org/html/2609.19674#Thmdefinition1)\. Energy\-shed rate \(dissipation\) and time\-reversal match \(retrace a reversed\-momentum end state\) are defined analogously\.

For the stochastic dissipative results \(§[4\.5](https://arxiv.org/html/2609.19674#S4.SS5), App\.[C\.6](https://arxiv.org/html/2609.19674#A3.SS6)\), we separate mechanical friction from a closed\-system GENERIC formulation\. For canonical kinetic energy and momentum\-space friction, the continuous model is

d​q=∇pH​d​t,d​p=\[−∇qH−M​∇pH\+fext\]​d​t\+Σ​d​W,Σ​Σ⊤=2​kB​T​M,M⪰0\.dq=\\nabla\_\{p\}H\\,dt,\\qquad dp=\\big\[\-\\nabla\_\{q\}H\-M\\nabla\_\{p\}H\+f\_\{\\rm ext\}\\big\]dt\+\\Sigma\\,dW,\\qquad\\Sigma\\Sigma^\{\\top\}=2k\_\{B\}T\\,M,\\quad M\\succeq 0\.\(8\)The deterministic friction contribution satisfiesH˙fric=−\(∇pH\)⊤​M​∇pH≤0\\dot\{H\}\_\{\\rm fric\}=\-\(\\nabla\_\{p\}H\)^\{\\top\}M\\nabla\_\{p\}H\\leq 0\. This is a passive mechanical channel, not a claim that subsystem energy cannot rise through work or thermal fluctuations\. State\-dependent mobilities require a separately specified stochastic drift and convention; the isolated OU result below concerns constant coefficients\. A closed\-system GENERIC model additionally requires degeneracy conditionsL∇S=0L\\nabla S=0andM∇E=0M\\nabla E=0, beyond antisymmetry and positive semidefiniteness\([Lee et al\., 2021](https://arxiv.org/html/2609.19674#bib.bib36)\)\. Those closed\-system laws are not established by the passive\-channel tests here\. The consequences of the imposed relation are reported\. \(vii\)Confined LangevinNN\-body: softened gravity plus a harmonic trap and an exact Ornstein–Uhlenbeck thermostat at bath temperatureTT; the diagnostic observable is the kinetic temperatureT^=⟨p2/m⟩\\hat\{T\}=\\langle p^\{2\}/m\\rangle\(equal toTTat equilibrium by equipartition,kB=1k\_\{B\}\{=\}1\)\. \(viii\)Damped double pendulumwith linear and nonlinear \(v2v^\{2\}\) joint friction \(the engine of §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)\)\. \(ix\)Driven\-dissipative particles: system \(vii\) plus a linear active drive of strengthAAopposing frictionγ\\gamma, reaching a*non\-equilibrium steady state*atTeff=T​γ/\(γ−A\)T\_\{\\text\{eff\}\}=T\\gamma/\(\\gamma\-A\)\(App\.[B](https://arxiv.org/html/2609.19674#A2)\)\. Models add a dissipation channel to the symplectic backbone: a diagonal positive\-semidefinite friction with fluctuation–dissipation noise \(the*guarantee*side, “M⪰0M\\succeq 0”\); the full PSD friction operatorM=L​L⊤M=LL^\{\\top\}\(Cholesky\) with correlated noise; a*factored*frictionM=γ​M0M=\\gamma M\_\{0\}whose sign flips withγ\\gammaand permits inversion with respect toγ\\gamma; and, for system \(ix\), a factored driveA​wθA\\,w\_\{\\theta\}with net dampingγf−A​w\\gamma\_\{f\}\-A\\,wand fluctuation–dissipation noise tied toγf\\gamma\_\{f\}\. Primary metric is the temperature relative error\|T^−T∗\|/T∗\|\\hat\{T\}\-T^\{\\ast\}\|/T^\{\\ast\}\(against the correct equilibrium or steady\-state temperatureT∗T^\{\\ast\}, the true simulator’s measured value which targets agreement with the discrete reference rather than cancellation of model\-dependent integration bias\); secondary is the energy\-shed rate and the injection rate atγ<0\\gamma<0\.

A scene is rendered to synthetic frames by a fixed differentiable rendererRR\(Gaussian or coloured\-disk splats at body positions\); an encoderEϕE\_\{\\phi\}maps a frame stack to a latent\(q^,p^\)\(\\hat\{q\},\\hat\{p\}\)\. In the supervised setting,EϕE\_\{\\phi\}is trained against the oracle state\. In the discovered setting, it is trained only through future\-frame prediction,∥R⁡\(rollout of​Eϕ​\(⋅\)\)−frames∥2\\lVert R\(\\text\{rollout of \}E\_\{\\phi\}\(\\cdot\)\)\-\\text\{frames\}\\rVert^\{2\}, without oracle\-state supervision\. The symplectic\-consistency prior adds∥p^/m−vobs∥2\\lVert\\hat\{p\}/m\-v\_\{\\text\{obs\}\}\\rVert^\{2\},vobsv\_\{\\text\{obs\}\}a finite difference of soft\-argmax centroids across frames \(an observed velocity, not oracle state\)\. In the end\-to\-end pixel comparison every model sharesEϕE\_\{\\phi\}andRR; a separate branch removesRRto test unsupervised discovery\.

## Appendix BDerivations

###### Proof of Proposition[2](https://arxiv.org/html/2609.19674#Thmproposition2)\.

Writefi\(q,G\)=−∂V/∂qi=−G∂qiUθ\(q\)f\_\{i\}\(q,G\)=\-\\partial V/\\partial q\_\{i\}=\-G\\,\\partial\_\{q\_\{i\}\}U\_\{\\theta\}\(q\)\.VVis linear inGGwithGG\-independent coefficient∑i<jφθ\\sum\_\{i<j\}\\varphi\_\{\\theta\}, so∂V/∂qi\\partial V/\\partial q\_\{i\}is linear inGGandfi​\(q,−G\)=−fi​\(q,G\)f\_\{i\}\(q,\-G\)=\-f\_\{i\}\(q,G\)\. The Hamiltonian vector field’s momentum componentp˙=f\\dot\{p\}=ftherefore changes sign withGG, whereasq˙=∂T/∂p\\dot\{q\}=\\partial T/\\partial pis unchanged\. The resulting trajectory is the flow generated by the same learned potential under coupling−G\-G\. ∎

*Comparison with the non\-factored model\.*A black\-box conditional potentialVθ​\(q,G\)V\_\{\\theta\}\(q,G\)has∂Vθ/∂q\\partial V\_\{\\theta\}/\\partial qwith no constraint linkingG\>0G\>0toG<0G<0, sofθ​\(q,−G\)f\_\{\\theta\}\(q,\-G\)is unconstrained byG\>0G\>0training and inversion is not guaranteed; empirically it fails \(App\.[C\.2](https://arxiv.org/html/2609.19674#A3.SS2)\)\. Thus, the dissociation isolates*factoring*, rather than conservation, as the inversion ingredient\.

*Proposition[1](https://arxiv.org/html/2609.19674#Thmproposition1), formal statement\.*LetHθH\_\{\\theta\}be analytic on a neighbourhood of a compact set that contains the numerical orbit, and let the stepd​tdtbe below the threshold of backward\-error analysis for that set\. Then the leapfrog map applied to equation[4](https://arxiv.org/html/2609.19674#S3.E4)is the exact time\-d​tdtflow of a modified HamiltonianH~θ=Hθ\+O⁡\(d​t2\)\\widetilde\{H\}\_\{\\theta\}=H\_\{\\theta\}\+O\(dt^\{2\}\)up to a remainder of sizeO\(e−c/dt\)O\(e^\{\-c/dt\}\)per step, soH~θ\\widetilde\{H\}\_\{\\theta\}is conserved up toO\(e−c/dt\)O\(e^\{\-c/dt\}\)over times of orderec/d​te^\{c/dt\}and the numericalHθH\_\{\\theta\}stays withinO⁡\(d​t2\)O\(dt^\{2\}\)of its initial value over that interval\([Hairer et al\., 2006](https://arxiv.org/html/2609.19674#bib.bib30)\)\. The statement is aboutH~θ\\widetilde\{H\}\_\{\\theta\}and the learnedHθH\_\{\\theta\}; it says nothing about the physical energy unlessHθH\_\{\\theta\}equals it\.*Empirical regularity \(not a theorem\)\.*A generic explicit mapxt\+1=xt\+gθ​\(xt\)x\_\{t\+1\}=x\_\{t\}\+g\_\{\\theta\}\(x\_\{t\}\)has no corresponding conservation guarantee\. Its energy behaviour therefore depends on the learned map; a contraction, for example, may remain bounded\. In these experiments, the unstructured and graph\-network models exhibit secular drift and reach the reporting cap\. The proposition supplies a sufficient condition for long\-time near\-conservation, whereas divergence of the unstructured models is an empirical observation\.*Remark\.*Training\-noise injection reduces the per\-step prediction error but does not itself provide an energy\-boundedness guarantee\. In the graph\-network simulator it delays, but does not prevent, the observed drift \(App\.[C\.2](https://arxiv.org/html/2609.19674#A3.SS2)\)\. The learnedH~θ≠Htrue\\widetilde\{H\}\_\{\\theta\}\\neq H\_\{\\text\{true\}\}, which is why the symplectic model is*bounded but not flat*; multi\-step training and longer horizons shrink and saturate the gap\.

###### Proposition 3\(Finite\-time transport bound\)\.

LetF⋆F^\{\\star\}be the reference field,FθtrueF\_\{\\theta\}^\{\\rm true\}the learned content evaluated with the true parameter map, andFθdeclF\_\{\\theta\}^\{\\rm decl\}the deployed field using the declared map, all at the tested coupling\. Suppose the reference and deployed trajectories remain in a compact regionΩ\\Omegaon\[0,T\]\[0,T\], andFθdeclF\_\{\\theta\}^\{\\rm decl\}is Lipschitz there with constantL≥0L\\geq 0\. Define

ϵc=supΩ‖Fθtrue−F⋆‖,ϵm=supΩ‖Fθdecl−Fθtrue‖\.\\epsilon\_\{c\}=\\sup\_\{\\Omega\}\\\|F\_\{\\theta\}^\{\\rm true\}\-F^\{\\star\}\\\|,\\qquad\\epsilon\_\{m\}=\\sup\_\{\\Omega\}\\\|F\_\{\\theta\}^\{\\rm decl\}\-F\_\{\\theta\}^\{\\rm true\}\\\|\.Then

‖x^t−xt‖≤eL​t​\(‖x^0−x0‖\+t⁡\(ϵc\+ϵm\)\),0≤t≤T\.\\\|\\hat\{x\}\_\{t\}\-x\_\{t\}\\\|\\leq e^\{Lt\}\\big\(\\\|\\hat\{x\}\_\{0\}\-x\_\{0\}\\\|\+t\(\\epsilon\_\{c\}\+\\epsilon\_\{m\}\)\\big\),\\qquad 0\\leq t\\leq T\.

###### Proof\.

Adding and subtracting the deployed field at the reference state bounds the trajectory\-error growth byL​‖x^−x‖\+ϵc\+ϵmL\\\|\\hat\{x\}\-x\\\|\+\\epsilon\_\{c\}\+\\epsilon\_\{m\}\. Integration and Gronwall’s inequality give the displayed bound\. ∎

The auxiliary true\-map field and both suprema are truth\-assisted quantities\. They are not generally available at deployment\. The bound assumes trajectories remain inΩ\\Omegaand therefore does not establish global stability\. The reported pooled Spearman association \(0\.870\.87versus0\.090\.09using only initial\-state error\) describes a ranking on the evaluated models \(§[4\.4](https://arxiv.org/html/2609.19674#S4.SS4)\)\. It is not coverage calibration or a certified prediction of error on new model families\.

###### Proposition 4\(Passive friction and sign\-reversed anti\-damping\)\.

Consider the deterministic momentum\-friction channelp˙\|fric=−M​∇pH\\dot\{p\}\|\_\{\\rm fric\}=\-M\\nabla\_\{p\}HwithM⪰0M\\succeq 0\. Its contribution toH˙\\dot\{H\}is nonpositive for every represented state\. IfM=γ​M0M=\\gamma M\_\{0\}, withM0≻0M\_\{0\}\\succ 0, changingγ\\gammato a negative value makes this contribution positive whenever∇pH≠0\\nabla\_\{p\}H\\neq 0\. Thus a globally PSD parameterisation cannot represent that sign\-reversed channel\. For a bath withT\>0T\>0, negativeγ\\gammaalso preventsΣ​Σ⊤=2​kB​T​γ​M0\\Sigma\\Sigma^\{\\top\}=2k\_\{B\}T\\gamma M\_\{0\}from being a real noise covariance\.

###### Proof\.

The chain rule givesH˙fric=−\(∇pH\)⊤​M​∇pH\\dot\{H\}\_\{\\rm fric\}=\-\(\\nabla\_\{p\}H\)^\{\\top\}M\\nabla\_\{p\}H\. Its sign follows from positive semidefiniteness; for negativeγ\\gammaand nonzero velocity the sign reverses\. A covariance is positive semidefinite, whereas2​kB​T​γ​M02k\_\{B\}T\\gamma M\_\{0\}is negative definite\. ∎

This restriction concerns the same passive channel\. It neither forbids external work or thermal energy injection nor makes every anti\-damped stochastic model ill\-posed\. It also does not establish the invariant law of a coupled numerical simulator \(App\.[C\.6](https://arxiv.org/html/2609.19674#A3.SS6)\)\.

###### Proposition 5\(Effective stationary temperature for linear drive\)\.

LetH⁡\(q,p\)=U⁡\(q\)\+‖p‖2/\(2​m\)H\(q,p\)=U\(q\)\+\\\|p\\\|^\{2\}/\(2m\),m\>0m\>0,γ\>0\\gamma\>0,T\>0T\>0, andA<γA<\\gamma\. Consider the continuous underdamped dynamics

d​q=\(p/m\)​d​t,d​p=−∇U​\(q\)​d​t−\(γ−A\)​\(p/m\)​d​t\+2​γ​kB​T​d​W\.dq=\(p/m\)\\,dt,\\qquad dp=\-\\nabla U\(q\)\\,dt\-\(\\gamma\-A\)\(p/m\)\\,dt\+\\sqrt\{2\\gamma k\_\{B\}T\}\\,dW\.If the density is normalizable and boundary terms vanish, thenρ∞\(q,p\)∝exp\[−H\(q,p\)/\(kBTeff\)\]\\rho\_\{\\infty\}\(q,p\)\\propto\\exp\[\-H\(q,p\)/\(k\_\{B\}T\_\{\\rm eff\}\)\]is stationary, whereTeff=T​γ/\(γ−A\)T\_\{\\rm eff\}=T\\gamma/\(\\gamma\-A\)\. Each momentum component has variancem​kB​Teffmk\_\{B\}T\_\{\\rm eff\}\.

###### Proof\.

Hamiltonian transport annihilates any differentiable density depending only onHH\. The momentum probability current from friction and noise is

\(γ−A\)​\(p/m\)​ρ∞\+γ​kB​T​∇pρ∞=\(\(γ−A\)−γ​T/Teff\)​\(p/m\)​ρ∞=0\.\(\\gamma\-A\)\(p/m\)\\rho\_\{\\infty\}\+\\gamma k\_\{B\}T\\nabla\_\{p\}\\rho\_\{\\infty\}=\\big\(\(\\gamma\-A\)\-\\gamma T/T\_\{\\rm eff\}\\big\)\(p/m\)\\rho\_\{\\infty\}=0\.Thus the Fokker–Planck operator annihilates the stated density\. Its momentum marginal is Gaussian with the given variance\. ∎

The isolated OU momentum process has no stationary probability law whenA≥γA\\geq\\gammaandT\>0T\>0\. The proposition makes no corresponding unqualified assertion for arbitrary coupled nonlinear models, and proves neither uniqueness nor exact preservation by a numerical integrator\. The externally driven bath interpretation is distinct from the mathematical effective Gibbs form\. A factored drive supplies the coefficient dependence; predictive accuracy still depends on the remaining learned dynamics\.

###### Definition 1\(Regime\-match and affine recovery\)\.

*Regime\-match\.*Letx^\\hat\{x\}be the model’s rollout under the flipped law, andx−,x\+x^\{\-\},x^\{\+\}the ground\-truth flipped\-law and training\-law \(default\) rollouts, both from the*true*initial state \(so initialisation error is counted, not cancelled\)\. The regime is matched iff∥x^−x−∥<∥x^−x\+∥\\lVert\\hat\{x\}\-x^\{\-\}\\rVert<\\lVert\\hat\{x\}\-x^\{\+\}\\rVert; we report the matched fraction and the normalised MSE∥x^−x−∥2/Var⁡\(x−\)\\lVert\\hat\{x\}\-x^\{\-\}\\rVert^\{2\}/\\mathrm\{Var\}\(x^\{\-\}\), for nonzero reference variance\. Ties do not count as matches\. A non\-finite rollout is reported as a failure rather than assigned a nearest\-reference label\. Two conventions are made explicit here\. First, the distance is the batch\-pooled normalised MSE over the whole rollout, which gives one regime\-match per run rather than per episode\. Second, the sign convention is fixed: the correct reference is rolled at the tested couplingGtestG\_\{\\rm test\}and the default reference at\|Gtest\|\|G\_\{\\rm test\}\|, and the regime is scored as𝟙\[nMSE\(x^,x−\)<nMSE\(x^,x\+\)\]\\mathbb\{1\}\[\\mathrm\{nMSE\}\(\\hat\{x\},x^\{\-\}\)<\\mathrm\{nMSE\}\(\\hat\{x\},x^\{\+\}\)\]\.*Affine recovery residual*of a discovered latentz^\\hat\{z\}with respect to the true\(q,p\)\(q,p\):minA,b⁡∥A​z^\+b−\(q,p\)∥2/Var⁡\(q,p\)\\min\_\{A,b\}\\lVert A\\hat\{z\}\+b\-\(q,p\)\\rVert^\{2\}/\\mathrm\{Var\}\(q,p\)\. This measures linear recoverability under a fitted affine map, not uniqueness of the latent coordinates\. Values below0\.50\.5are a descriptive threshold used in this study, not an identifiability theorem\.

## Appendix CExtended results and findings

This appendix reports the findings summarised in the main text, together with their figures, multi\-seed estimates, experiment\-specific methods, and stated limitations\. The unified notation of App\.[A](https://arxiv.org/html/2609.19674#A1)is used: statex=\(q,p\)x=\(q,p\), couplingGG, relative energy driftD⁡\(t\)D\(t\)\(Eq\.[2](https://arxiv.org/html/2609.19674#S3.E2)\), regime\-match and affine identifiability residual \(Def\.[1](https://arxiv.org/html/2609.19674#Thmdefinition1)\)\. Each finding is tagged*\[well\-supported\]*,*\[directional\]*,*\[supported\]*, or*\[boundary\]*\(criteria in App\.[E](https://arxiv.org/html/2609.19674#A5)\)\.

Aggregates are the mean over at least three seeds, and at least five seeds at scale, with the available aggregate summaries in App\.[C](https://arxiv.org/html/2609.19674#A3)\. Where a bracket is labelled a normal\-approximation 95% interval, it retains that meaning even for small seed counts and is not an observed range\. Such small\-sample intervals are descriptive and do not justify significance from non\-overlap\. Individual seed values are not tabulated for every comparison\.

Display\-clipped finite errors and non\-finite numerical failures are distinct\. A sentinel substituted for a failure is not a measured error and cannot establish a quantitative error ratio\. These figures retain their plotting thresholds, and the text distinguishes what the two mean\. A pre\-registered absolute\-error*backstop*marks a perception result invalid when even its in\-distribution fidelity is worse than the threshold \(such runs are reported as invalid, not compared\)\. Regime\-match is meaningful only for a*bounded*rollout\. For a diverged rollout, such as that of the graph network, comparison with the nearest reference degenerates and is not used as evidence\. Such cases are identified explicitly\. Table[5](https://arxiv.org/html/2609.19674#A3.T5)summarises all the findings; in it, “match” is the regime match \(1 = follows the unseen flipped law\), “nMSE” the normalised inversion error, and “drift” the relative true\-energy drift of Eq\.[2](https://arxiv.org/html/2609.19674#S3.E2)\.

Table 5:Summary of findings and evidence classes\.FindingSystemSummary statisticClassStabilitygravitysymplectic model drift→7\.10\.40\\\!\\to\\\!7\.1vs cap10410^\{4\}well\-supportedTuning\-robustgravityevery unstructured arm passes the cap;λ=100\\lambda\{=\}100delays but does not bound driftwell\-supportedMulti\-stepgravityfactored symplectic model 2000\-step drift→1\.94\.4\\\!\\to\\\!1\.9directional10k\-stepgravityfactored symplectic model drift saturates→0\.891\.9\\\!\\to\\\!0\.89well\-supportedNon\-factoredgravitysymplectic model match00\(no inversion\)boundaryFactoredgravityfactored symplectic model match11, nMSE0\.030\.03–0\.180\.18well\-supportedParam\. extrap\.gravitysmooth:2×2\\timesedge, no abrupt degradationdirectionalDeLaN/SymODENenginefactored match11vs non\-factored00well\-supportedGraph\-netgravityreported response labels11vs00; both fail long\-horizon testboundaryContactsdisksbounded stiffness series; sym\-inversion11directionalEnginependulumfactored symplectic model match11, nMSE8\.7​e−48\.7e\{\-\}4well\-supportedContact aug\.disksno better than smooth modelboundaryDissipationdragconservation sheds00; port shedsboundarySupervised perc\.pixelsinversion retained; error increases with speeddirectionalDiscovered perc\.pixelsfixed\-kinetic nMSE≈0\.0670\.075\\\!\\approx\\\!0\.067well\-supportedMomentum recov\.pixelspp\-resid→0\.350\.72\\\!\\to\\\!0\.35; pos\. opendirectionalPixel comparisonpixelsfactored fixed\-kinetic model nMSE0\.120\.12; others failwell\-supportedNecessitygravityshortcut42%→10−442\\%\\\!\\to\\\!10^\{\-4\}; nMSE→0\.090\.43\\\!\\to\\\!0\.09directionalComp\. & scalegravitymatch11toN=32N\{=\}32; drift→1020\.2\\\!\\to\\\!102directionalCoulomb, held\-out signCoulombfactored0\.520\.52vs black\-box1\.421\.42;2\.7×2\.7\\timeswell\-supportedMixed two\-family inversionCoulombfactored0\.00350\.0035vs black\-box0\.4250\.425;120×120\\timeswell\-supportedData\-efficiencygravityinverts with 32 traj\.; symplectic model error rises with datawell\-supportedCapacitygravityunstructured MLP diverges at every widthwell\-supportedCompositiondragthree ingredients composedirectionalControlgravitystructured plans transfer; unstructured MLP misestimatesdirectionalRenderingpixelsstructured video consistent; unstructured jumpssupportedNoise screengravityrobust to30%30\\%state noisesupportedTemperatureLangevinstructuredT^\\hat\{T\}has low error; unstruct\. divergeswell\-supportedTensionLangevinpassivity precludes channel anti\-damping; factored form permits itwell\-supportedFull operatorLangevinfullMMhas higher error than diagonalMMsupportedEngine dissipationpendulumport sheds energy, correct sign,∼\{\\sim\}half ratedirectionalNoise robustnessLangevincorrectT^\\hat\{T\}to∼10%\{\\sim\}10\\%meas\. noisedirectionalDriven drivedrivenfactored drive inverts strength\+sign, err≤0\.09\\leq 0\.09well\-supportedMisspecificationgravitylow in\-dist\. error, incorrect counterfactualwell\-supportedTransformer/dom\-randgravitytransformer diverges and does not invertwell\-supportedGrounded JEPApixelsanchored factored hybrid match11, nMSE0\.400\.40vs0\.680\.68well\-supportedAnonymous objectpixelsidentity removed: both collapse to match0\.330\.33boundaryLearned bindingpixelslearned tracker\+\+factored inverts \(10/1010/10seeds, nMSE0\.150\.15; control0/100/10\)well\-supportedDetection\+\+bindingpixelsslot\-attention does not reach it \(match00\)boundaryExternal encoderpixelsfrozen VideoMAE/V\-JEPA2 invert \(0\.130\.13/0\.590\.59, both 10\-seed, factored10/1010/10p=0\.001p\{=\}0\.001; readout\-matched2\.6×2\.6\\timesover our CNN; scale\-flat300300M→\\to11B, gain1\.02×1\.02\\times\)well\-supportedMomentum axispixelsdifferencing DINOv2 tokens closes the image\-SSL gap \(→0\.0960\.62\\\!\\to\\\!0\.096\)well\-supportedEncoder compositionpixelsVideoMAE\+\+learned detect/bind\+\+prior does not compose \(unstable, no clean inversion\)boundarySecond law from videopixelsreported damping diagnostic orders damped\>1\.25\\\!\>\\\!cons0\.840\.84\(5/55/5seeds\); passive friction by constructionsupported \(ordering only\)Floor located\+\+refusalpixelsoracle control rules out grounding noise as the sole explanation; training\-level refusal directionally consistent5/55/5but below the pre\-registered barboundaryEnergy\-aux unificationpixelsthe auxiliary\-loss lever fails all three pre\-registered predictions; with a learned meter it backfires \(σdamped\\sigma\_\{\\text\{damped\}\}→0\.621\.25\\\!\\to\\\!0\.62\)pre\-registered negativeMeasure, don’t fitpixelsthe energy\-balance estimator removes the residual bias \(γ^cons\\hat\{\\gamma\}\_\{\\text\{cons\}\}→−0\.0060\.84\\\!\\to\\\!\-0\.006\); arrow\-of\-time rate read withσ≥0\\sigma\\\!\\geq\\\!0refusal; composition instability cured, grounding boundary standswell\-supported \(one gate marginal\)### C\.1A\. Stability from conservation \(extends §[4\.1](https://arxiv.org/html/2609.19674#S4.SS1)\)

Conservation→\\tostability\.Five approximately capacity\-matched models \(unstructured MLP, energy\-penalised MLPs, and neural ODE at∼139\{\\sim\}139k parameters; symplectic model at∼138\{\\sim\}138k\) predict one step of conservative gravity and roll toT=2000T=2000\(100×100\\timesthe 20\-step training horizon\)\. The unstructured MLP and theλ=1\\lambda=1penalised MLP pass the10410^\{4\}cap in the median episode by step500500and lose6161and123123of192192episodes to non\-finite states between steps10001000and20002000\. Theλ=100\\lambda=100penalised MLP and the neural ODE stay finite in every episode but pass the cap\. The symplectic model stays finite and below the cap in every episode \(Table[6](https://arxiv.org/html/2609.19674#A3.T6)\)\. It remains bounded but has non\-zero drift because near\-conservation concerns a modified learned Hamiltonian under the assumptions of Prop\.[1](https://arxiv.org/html/2609.19674#Thmproposition1), not the true energy\. Multi\-step and longer\-horizon training reduce this discrepancy\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_a1.png)Figure 5:Relative energy driftD⁡\(t\)D\(t\)vs rollout step \(logyy, 3 seeds; the shaded band is the plotting normal\-approximation interval, descriptive at three seeds\) in the original three\-model display\. The unstructured MLP and energy\-penalised model reach the10410^\{4\}display cap by∼\{\\sim\}step 300; the symplectic model plateaus\. Numerical summaries are taken from the raw record \(Table[6](https://arxiv.org/html/2609.19674#A3.T6)\), not from this capped display\.Table 6:Raw stability record: per\-episode median relative energy drift over6464held\-out episodes, mean \[min, max\] over three runs;\>104\{\>\}10^\{4\}marks a median above the cap\. NF is the number of episodes non\-finite by step20002000\(of192192; none is non\-finite by step10001000\);err2000\\mathrm\{err\}\_\{2000\}is the median normalised rollout error at20002000steps, range over seeds\. Learning rates were selected per seed on held\-out one\-step error from\{10−3,3×10−4,10−4\}\\\{10^\{\-3\},3\\times 10^\{\-4\},10^\{\-4\}\\\}\.#### The neural\-ODE and energy\-penalty controls\.

The neural vector field with RK4 uses matched capacity, the same step size and one\-step objective, and the same learning\-rate sweep\. All three seeds remain finite in every episode\. Its median energy drift is0\.080\.08at100100steps,3×1023\\times 10^\{2\}–4×1034\\times 10^\{3\}at500500, and above the cap by10001000\. Equal step size does not imply equal integration error across RK4 and leapfrog\. Theλ=100\\lambda=100energy\-penalised MLP has the lowest energy error of all the models at100100and500500steps, after which the drift grows without bound, with the rollout error growing alongside it\.

All models share the evaluation episodes within a seed, so the record supports paired comparisons\. The summaries here support a long\-horizon energy\-error contrast, not a claim that Hamiltonian structure is necessary for finiteness\. The comparison is generated from one raw record rather than from capped plots\. For each model and run, that record stores the learning\-rate grid and the selected rate, the held\-out one\-step error used for selection, the per\-episode finite and non\-finite status and counts at each checkpoint, the uncapped relative energy drift on finite episodes \(median, minimum, maximum, and fraction above11\), and the rollout error, together with the selected checkpoint for each model and seed\. Seed\-level summaries are the mean, minimum, and maximum over seeds of per\-episode medians\.

### C\.2B\. Law\-inversion and the double dissociation \(extends §[4\.1](https://arxiv.org/html/2609.19674#S4.SS1)–[4\.2](https://arxiv.org/html/2609.19674#S4.SS2)\)

Conservation alone does not provide inversion\.Under never\-trained repulsiveG<0G<0, the symplectic model is stable but its rollout stays closer to the training\-law trajectory \(regime\-match00at everyGGand seed\)\. It learned a black\-boxVθ​\(q,G\)V\_\{\\theta\}\(q,G\)onG\>0G\>0only \(Prop\.[2](https://arxiv.org/html/2609.19674#Thmproposition2)contrast\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_a3.png)Figure 6:Non\-factored HNN: stable but does not invert; repulsive rollout matches the attractive default\.The factored coupling yields inversion\.Trained on attraction only, the factored symplectic model \(Eq\.[5](https://arxiv.org/html/2609.19674#S3.E5)\) tracks unseen repulsion at nMSE0\.030\.03–0\.180\.18acrossG=−0\.5G=\-0\.5to−1\.5\-1\.5\(regime\-match11everywhere\), more than an order of magnitude below the non\-factored symplectic model and far below the \(capped\) unstructured baselines, while attaining the lowest in\-distribution error with fewer parameters\. Together with the non\-factored result, this is the double dissociation\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_a3v2_a.png)\(a\)inversion error
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_a3v2_b.png)\(b\)regime\-match

Figure 7:Factored HNN tracks the true reversed \(G<0G<0\) reference trajectory; the other evaluated models diverge or remain closer to the training\-law trajectory\.DeLaN/SymODEN: comparison across formalisms and integrators\.On the engine gravity\-flip, three non\-factored structured models fail and both factored models invert \(Table[7](https://arxiv.org/html/2609.19674#A3.T7)\), isolating the axis to the factoring\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b7delan.png)Figure 8:DeLaN/SymODEN\.Inversion across five structured models: non\-factored fail \(match 0\), factored invert \(match 1\), all equally stable\.Table 7:DeLaN/SymODENengine gravity\-flip inversion\.Matched causal factorial\.A natural worry is that our stability and inversion contrasts change several things at once, namely the state representation, the update rule, and the numerical integrator\. We therefore run a controlled factorial that varies these three axes independently\. We cross a Hamiltonian representation against an unrestricted vector field, a coupling\-factored parameter map against an unrestricted conditional map, and three integrators \(leapfrog, RK4, and implicit midpoint\) at three step sizes, over twelve paired seeds in two coercive coupled\-oscillator systems, with every model fitted to the same observed regime\. The unrestricted vector field has more coefficients than the Hamiltonian model\. Parameter counts alone do not establish equal approximation capacity\. Two results follow\. First, at matched fit the Hamiltonian representation reaches a maximum true\-energy error of about5×10−45\\times 10^\{\-4\}to8×10−48\\times 10^\{\-4\}, against0\.310\.31to0\.330\.33for the unrestricted field, and a trajectory error of5×10−45\\times 10^\{\-4\}to1\.7×10−31\.7\\times 10^\{\-3\}against0\.0440\.044to0\.0590\.059\(Table[8](https://arxiv.org/html/2609.19674#A3.T8)\)\. Second, at the unseen negative coupling the factored map reaches an error of about0\.00350\.0035, while the unrestricted conditional map stays near0\.9850\.985, a gap that holds in both systems \(Holm\-adjustedp≈0\.003p\\approx 0\.003, twelve seeds; Table[9](https://arxiv.org/html/2609.19674#A3.T9)\)\. The differences persist across the evaluated representations, parameter maps, and solvers\. This does not establish equal functional capacity or exclude all optimisation and discretisation effects\. Every factorial cell stays bounded here \(escape fraction00\), so the factorial isolates a fidelity and inversion effect rather than a boundedness effect\. The boundedness contrast comes instead from the softened\-gravity system of §[4\.1](https://arxiv.org/html/2609.19674#S4.SS1), where the unrestricted andλ=1\\lambda=1energy\-penalised models pass the10410^\{4\}cap by step 500 and lose episodes to non\-finite states by step 2000, while the symplectic model stays finite and below the cap\.

Table 8:Controlled factorial at step size0\.050\.05over twelve training seeds\. Ranges span the two systems and the relevant model cells\.Table 9:Matched\-state response at the unseen negative coupling\.Graph\-network response and long\-horizon failure\.Both graph variants fail the 2000\-step test, so message factoring alone does not prevent their observed long\-horizon failure\. The source reports nearest\-reference labels of 1 for the factored variant and 0 for the other variant\. Without verified short\-horizon finiteness and fidelity, those labels are descriptive and are not counted here as validated counterfactual success\. Training\-noise injection does not remove the reported long\-horizon failure \(Table[10](https://arxiv.org/html/2609.19674#A3.T10)\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_gns_a.png)\(a\)2000\-step drift
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_gns_b.png)\(b\)inversion regime\-match

Figure 9:Graph\-network long\-horizon failure and reported response labels\. The labels require separate short\-horizon validity checks before being interpreted as counterfactual success\.Table 10:Graph\-network summaries\. Failure is separated from finite energy error; reported response labels are descriptive pending horizon\-validity checks\.
### C\.3C\. Contacts and the engine \(extends §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)\)

Contacts\.On colliding disks \(contact stiffness series: softk=30k\{=\}30, stiffk=300k\{=\}300, impulsive\), the factored symplectic model’s energy stays bounded across the whole series and time\-reversal \(symmetry\) inversion remains present \(Table[11](https://arxiv.org/html/2609.19674#A3.T11)\); parameter\-inversion regime\-matches \(match 1\) at soft/stiff, though its fidelity degrades \(nMSE​0\.76→2\.9\\text\{nMSE \}0\.76\\to 2\.9\), and has no referent at the impulsive level \(§[C\.4](https://arxiv.org/html/2609.19674#A3.SS4)\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b2_a.png)\(a\)energy drift across the stiffness series
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b2_b.png)\(b\)time\-reversal regime\-match

Figure 10:Contact stiffness series: bounded energy and surviving symmetry\-inversion where unconstrained models diverge\.Table 11:Across the stiffness series \(factored symplectic model\)\.
### C\.4D\. The boundary map \(extends §[4\.4](https://arxiv.org/html/2609.19674#S4.SS4)\)

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/fig4_a.png)\(a\)Dissipation: conservation is incompatible
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_misspec.png)\(b\)A misspecified factored law
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/fig4_b.png)\(c\)Perception: momentum is the open axis

Figure 11:The boundaries\. Conservation cannot shed energy under drag; a factored law of the wrong power fits in\-distribution yet gives a wrong counterfactual; and a single frame recovers position but not momentum\.The impulsive\-contact boundary\.The stiffness\-sign intervention used here has no specified counterpart in the impulsive hard\-contact model\. This does not rule out interventions on restitution, friction, or forcing in a separately defined hybrid law\. The tested differentiable\-restitution variant does not improve on the smooth model and loses time\-reversal performance at the impulsive level; the free\-learned impulse model diverges \(Table[12](https://arxiv.org/html/2609.19674#A3.T12)\)\. These results delimit the tested augmentation rather than all contact\-capable structured models\.

Table 12:Drift at the soft, stiff, and impulsive levels\.Conservation is incompatible with dissipative dynamics\.With linear drag the truth sheds energy \(Fig\.[11](https://arxiv.org/html/2609.19674#A3.F11)\); symplectic model/factored symplectic model shed∼0\{\\sim\}0, and only a port\-Hamiltonian variant more closely reproduces the reference shedding rate and inverts the drag sign \(Table[13](https://arxiv.org/html/2609.19674#A3.T13)\)\. Energy\-shed rate is primary here \(trajectory\-MSE does not separate the regimes short\-horizon\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b4.png)Figure 12:Conservative models have approximately zero energy\-shed rate; the port variant more closely reproduces the reference rate\.Table 13:Energy\-shed rate \(negative = losing energy\)\.Supervised perception\.With a soft\-argmax encoder supervised to the oracle state, position is near\-perfect and flat across a visual\-difficulty series; the limit is object*speed*\(per\-body momentum error0\.010\.01slow→0\.87\\to 0\.87fast\), and downstream inversion is retained \(factored symplectic model nMSE0\.062/0\.0640\.062/0\.064≈\\approxoracle0\.090\.09\)\. Reference trajectories are rolled from the true initial state so that perception error is counted rather than cancelled\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b1a.png)Figure 13:Supervised perception: position error remains stable; inversion is retained; momentum error increases with object speed\.Discovered perception leaves a large linear momentum\-recovery error; a canonical kinetic recovers it\.A latent learned end\-to\-end by prediction leaves a large linear momentum\-recovery error \(affinepp\-residual0\.720\.72\) with renderer\-pinned position \(qq\-residual0\.0040\.004\); a canonical*fixed*kinetic recovers inversion with error comparable to the oracle result \(3\-seed, 500\-epoch nMSE0\.0650\.065–0\.0920\.092≈\\approxoracle0\.0670\.067; backstop0\.0450\.045\)\. A learned decoder is invalid against the backstop\. The value is the 500\-epoch, three\-seed result; a 60\-epoch run is undertrained and disagrees\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b1b.png)Figure 14:Discovered latent: momentum poorly recovered by an affine readout \(highpp\-residual\); the fixed canonical kinetic reduces inversion nMSE to0\.0650\.065–0\.0920\.092, compared with the oracle value0\.0670\.067\.Momentum partially recoverable; position grounding open \[directional/boundary\]\.A symplectic\-consistency prior drops the momentum residual0\.72→0\.350\.72\\to 0\.35\(<0\.5<0\.5, improved linear recovery\) while keeping position \(qq\-residual0\.050\.05\) and low inversion error \(nMSE0\.100\.10\), at valid fidelity \(backstop0\.1490\.149\)\. Herevobsv\_\{\\text\{obs\}\}is frame\-derived, not oracle\. Removing the renderer to*discover*position increases the residual \(qq\-residual0\.890\.89; backstop1\.841\.84, invalid\): the residual boundary is decoder\-free position grounding\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_b1c.png)Figure 15:Symplectic\-consistency improves linear momentum recovery \(pp\-residual→0\.350\.72\\\!\\to\\\!0\.35\) while retaining position; removing renderer\-based grounding produces a high position residual\.Pixel\-based dynamics comparison under grounded perception \[well\-supported, scoped\]\.Trained end\-to\-end on synthetic renders with a shared fixed disk\-renderer \(only the dynamics differ\), symplectic models stay bounded while unstructured MLP/graph\-network simulator/factored graph\-network simulator/JEPA\-style predictor diverge, and only the factored fixed\-kinetic model approaches the oracle inversion error \(Table[14](https://arxiv.org/html/2609.19674#A3.T14)\)\. The factored symplectic model \(learned kinetic\) fails like the non\-factored models, reproducing the empirical recovery condition on pixels\. The comparison evaluates dynamics*on*pixels with renderer\-grounded position, not unsupervised inference*from*pixels\. The baselines are restricted to the evaluated architecture families and do not constitute a published leaderboard\. Oracle\-latent reconstruction is0\.220\.22–0\.530\.53\. The fully unsupervised branch is invalid \(backstops2\.82\.8–3\.33\.3\) even though the decoder renders from oracle latents, a position\-discovery failure that confirms, rather than resolves, the identifiability and position\-grounding boundaries above\. The decode\-free JEPA\-style predictor column here is the bare objective \(no grounding\)\. It is developed into a grounded hybrid, and into its own boundary, in App\.[C\.8](https://arxiv.org/html/2609.19674#A3.SS8)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_c4_a.png)\(a\)stability \(latent\-norm growth\)
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_c4_b.png)\(b\)anti\-physics inversion error

Figure 16:End\-to\-end on pixels: \(a\) latent\-norm growth \(log\) diverges for all but symplectic models; \(b\) anti\-physics inversion error is comparable to the oracle result only for the factored\-plus\-canonical\-kinetic model\. In\-distribution backstops are within the valid range throughout \(Table[14](https://arxiv.org/html/2609.19674#A3.T14)\)\.Table 14:End\-to\-end on pixels \(grounded perception\)\.
### C\.5E\. Necessity, compositionality, generality, and scale \(extends §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)\)

Necessity\.A first screen yields a null result because always\-active interactions provide no available shortcut\. A second screen makes a shortcut available: a hybrid \(factored pathway\+\+free black\-box term\) assigns∼42%\{\\sim\}42\\%of the force, with inversion degraded55–10×10\\times\. AnL2L\_\{2\}penalty drives the share to10−410^\{\-4\}and restores fidelity \(nMSE0\.43→0\.090\.43\\to 0\.09\)\. These results support the shortcut\-use interpretation within this physical system\.

Cross\-family\.On Coulomb with mixed signs, a charge\-product coupling has the lowest error on a held\-out all\-repulsive configuration \(error0\.520\.52vs black\-box1\.421\.42, a modest∼2\.7×\{\\sim\}2\.7\\timesmargin\)\. A single factored model spanning a mixed gravity\+\+Coulomb generator inverts both∼120×\{\\sim\}120\\timesbetter than a black\-box that reads the couplings \(nMSE0\.00350\.0035vs0\.4250\.425\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_coulomb.png)Figure 17:Held\-out all\-repulsive Coulomb: the charge\-product coupling generalises across sign\.Data\-efficiency\.The factored symplectic model inverts at every training\-set size down to 32 trajectories\. Giving the non\-factored symplectic model*more*data makes its inversion*worse*\(0\.76→2\.670\.76\\to 2\.67across32→51232\\to 512\)\. This is not undertraining at small data: the symplectic model’s*in\-distribution*\(G\>0G\{\>\}0\) rollout stays bounded with low drift at every size \(drift∼1\.4\{\\sim\}1\.4at 32 trajectories\), so it fits the training law well\. Within this experiment, the inversion failure is therefore not explained by insufficient training data\.

Capacity and composition \[well\-supported / directional\]\.Across a width sweep32→51232\\to 512, the unstructured MLP diverges at every width and the factored symplectic model inverts at every width\. Non\-factored conservation worsens with capacity\. Factored potential, dissipative port, and necessity penalty compose without interference under dissipation, with the necessity\-fidelity effect muted in\-regime \(shortcut13%13\\%vs the earlier42%42\\%\)\.

Calibration of in\-regime planning\.Planning an impulse through each model’s rollout and executing it in the true physics, structured models’ plans transfer \(executed≈\\approxpredicted\) while unstructured MLP misestimates its own plans\. The unstructured model reports the lowest predicted error but the highest realised error\. Because confidence intervals for realised error overlap, the supported result concerns prediction–execution calibration rather than a control\-performance advantage\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_control.png)Figure 18:Planned vs executed error: structured plans transfer; the unstructured model misestimates its own plans\.Rendering and noise screen\.A structured model’s rendered rollout is physically consistent \(re\-perceived energy conserved\) whereas the unstructured model’s bodies jump discontinuously, a re\-visualisation of the stability gap in pixels, not a benchmark\. Under i\.i\.d\. initial\-state noise to30%30\\%of the state s\.d\., the factored symplectic model keeps regime\-match 1 and inversion nMSE∼0\.1\{\\sim\}0\.1, a screen \(not perception error\) that motivated the move to pixels\.

### C\.6F\. Beyond conservation: stochastic dissipation and driven injection \(extends §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)\)

This subsection extends §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)using the passive\-friction formulation of Eq\. equation[8](https://arxiv.org/html/2609.19674#A1.E8)and the temperature metric of App\.[A](https://arxiv.org/html/2609.19674#A1)\. All aggregates are mean \[95% CI\] over≥3\\geq 3seeds; temperature error is capped at100100and compared against the true simulator’s*measured*equilibrium or steady\-state temperature which targets agreement with the discrete reference rather than cancellation of model\-dependent integration bias\.

Structured fluctuation–dissipation models approach the target temperature; guarantee–invertibility tension\.On the confined Langevin system \(three bodies, three seeds, four\(γ,T\)\(\\gamma,T\)settings\), a structured model with a positive\-semidefinite friction and fluctuation–dissipation noise reaches the target temperature with relative error0\.190\.19\(≤0\.10\\leq 0\.10at strong drag\), while an unstructured neural SDE*diverges*\(T^\\hat\{T\}error∼103\{\\sim\}10^\{3\}–10410^\{4\}\) and a reversible\-only model*never thermalises*\(11\.411\.4\): the conservation→\\tostability analogue for the thermal case \(Table[15](https://arxiv.org/html/2609.19674#A3.T15)\)\. The same table shows the tension: the guarantee\-side model has the lowest temperature error among the evaluated models but injects energy only at the reversible baseline \(injection0\.27≈0\.290\.27\\approx 0\.29, i\.e\. it*cannot*invert\), whereas the factored\-friction model injects∼7×\{\\sim\}7\\timesthe baseline \(1\.911\.91\) at the cost of being a worse thermostat without the passive\-channel sign guarantee \(Prop\.[4](https://arxiv.org/html/2609.19674#Thmproposition4)\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f29_a.png)\(a\)temperature error by model structure
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f29_b.png)\(b\)guarantee–invertibility trade\-off

Figure 19:The structured FDT model has the lowest temperature error; the unstructured model diverges and the reversible\-only model does not thermalise\. The positive\-semidefinite model cannot reverse its passive friction channel, whereas the signed factored model can\.Table 15:Temperature error and energy\-injection rate atγ<0\\gamma<0\(confined Langevin; injection near the reversible baseline0\.290\.29means “cannot invert”\)\.Full and diagonal friction parameterisations\.The full positive\-semidefinite operatorM=L​L⊤M=LL^\{\\top\}\(Cholesky\) with correlated fluctuation–dissipation noise retains the qualitative thermalisation result but has higher temperature error \(0\.460\.46versus0\.190\.19\) and is more difficult to train\. In these mechanical experiments, the diagonal parameterisation therefore performs better; evaluation in non\-mechanical dissipative domains remains future work\.

Dissipation in MuJoCo, with linear and nonlinear friction\.On the MuJoCo double pendulum in generalised coordinates, the structured dissipative port reproduces the true energy\-shed rate for both linear and nonlinear \(v2v^\{2\}\) friction, while an energy\-conserving control under\-sheds \(Table[16](https://arxiv.org/html/2609.19674#A3.T16)\)\. The unstructured model also fits smooth deterministic friction\. The comparison is therefore between the structured model and the conserving control\. The distinction from the unstructured model occurs in the stochastic temperature results above\. Under anti\-friction \(γ<0\\gamma<0\), the port injects energy at rate\+10\.1\+10\.1for linear friction but diverges under nonlinear friction in all seeds\. This instability is consistent with Prop\.[4](https://arxiv.org/html/2609.19674#Thmproposition4)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f31_a.png)\(a\)linear friction
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f31_b.png)\(b\)nonlinear \(v2v^\{2\}\) friction

Figure 20:The structured port sheds the true energy\-loss rate on the engine for linear andv2v^\{2\}friction; the conservation control under\-sheds\.Table 16:Energy\-shed rate \(negative==losing energy\), model vs true\.Energy injection as a factored drive generalises over strength and sign\.Modelling energy injection as an explicit factored*drive*AA\(Prop\.[5](https://arxiv.org/html/2609.19674#Thmproposition5)\), a structured model reaches the correct driven non\-equilibrium steady stateTeff=T​γ/\(γ−A\)T\_\{\\text\{eff\}\}=T\\gamma/\(\\gamma\-A\)across trained and held\-out drive strengths and the never\-seen cooling sign \(rel\-err≤0\.094\\leq 0\.094, Table[17](https://arxiv.org/html/2609.19674#A3.T17)\), while a control without drive input remains at the bath temperature \(rel\-err climbing to0\.680\.68at the extremes\) and the unstructured model diverges\. Beyond the friction \(A≥γA\\geq\\gamma\) there is no steady state and the driven models exhibit unbounded growth \(diverged fraction: factored drive67%67\\%, unstructured100%100\\%, control without drive input0%0\\%because it does not represent the drive\), consistent with the expected unbounded energy growth when net damping is non\-positive\. This parallels the gravity\-coupling result: a factored external drive generalises across sign by construction, whereas a friction\-sign flip is thermodynamically ill\-posed and should not generalise at all\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f33_a.png)\(a\)correct driven steady state
![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/bc_f33_b.png)\(b\)generalising over drive strength/sign

Figure 21:The factored\-drive model follows the true driven\-steady\-state curve across drive strengths and both signs; the control without drive input cannot track it\.Table 17:Driven\-steady\-state temperature rel\-err \(∗\*==trained drive; rest held\-out\)\.
### C\.7G\. Misspecification and stronger baselines \(extends §[4\.2](https://arxiv.org/html/2609.19674#S4.SS2)and §[4\.4](https://arxiv.org/html/2609.19674#S4.SS4)\)

Inaccurate counterfactuals under coupling\-power misspecification\.Every earlier inversion used a prior whose form matched the law\. Here we deliberately misspecify it \(softened 3\-body gravity, factored priorV=G​∑ϕθV=G\\sum\\phi\_\{\\theta\}, force linear inGG, 3 seeds\)\. When the true coupling is instead*even*inGG\(force∝G2\\propto G^\{2\}, so flippingGGleaves it unchanged\), the model fits in\-distribution as well as the matched case yet predicts a sign flip the truth never performs, with no in\-distribution warning\. The correctG2G^\{2\}structure restores the expected counterfactual, and a mismatched*interaction*structure \(a 3\-body term under a pairwise prior\) instead produces poor in\-distribution fit \(Table[18](https://arxiv.org/html/2609.19674#A3.T18)\)\. Parameter factoring therefore depends on the assumed functional form: coupling\-power misspecification is not detected by in\-distribution error here, whereas interaction\-structure misspecification produces elevated in\-distribution error\.

Table 18:Misspecification series \(nMSE; 3 seeds\)\. “true\-cf differs”≈0\{\\approx\}0means the true counterfactual does not change under the flip\.Transformer and domain\-randomisation comparisons\.A permutation\-equivariant set\-transformer \(bodies as tokens, self\-attention\), a relational unstructured model,*diverges*over 2000 steps \(∼2\{\\sim\}2orders more slowly than the MLP, but still unbounded\) and, evaluated at the same 150\-step inversion horizon, does not invert \(regime\-match00, counterfactual nMSE∼0\.3\{\\sim\}0\.3; Table[19](https://arxiv.org/html/2609.19674#A3.T19)\)\. Thus, stability is associated with symplectic structure and inversion with factoring\. The same distinction is observed with this additional baseline\. Regime\-match is primary here\. At 50 steps, the attractive and repulsive trajectories remain close, producing a small transformer nMSE despite regime\-match00\. This value reflects the short horizon rather than inversion\. Domain randomisation, in which the black\-box model is trained on both signs ofGG, inverts \(regime\-match11\) because the counterfactual is then in\-distribution, but it still diverges, so it gains no stability\. The factored model is the only evaluated model that both inverts \(zero\-shot\) and stays bounded\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_transformer_dr.png)Figure 22:The factored model is bounded and inverts; the transformer diverges and does not invert; domain randomisation inverts after training on both coupling signs but diverges\.Table 19:Transformer \+ domain\-randomisation \(3 seeds; inversion read at the 150\-step horizon\)\. Regime\-match is primary; nMSE secondary\.
### C\.8H\. Grounding a decode\-free objective, and the object\-binding boundary \(extends §[4\.5](https://arxiv.org/html/2609.19674#S4.SS5)\)

Encoder readout, scaling, and image\-SSL\.Is the advantage in the encoder or in the readout? VideoMAE’s0\.130\.13improves on our own CNN’s0\.450\.45, but that comparison changes two things at once, namely the encoder*and*the readout \(our CNN used the earlier pooled head, whereas the public encoders use the spatial cross\-attention\)\. We deconfound by running our*own*trained CNN, frozen, through the*same*spatial readout: it improves to0\.340\.34, so the localisation\-aware readout explains part of the gap, but it stays2\.6×2\.6\\timesworse than VideoMAE under the identical readout\. The “better than our own encoder” claim therefore softens from∼5×\\sim\\\!5\\timesto a readout\-matched∼2\.6×\\sim\\\!2\.6\\times, and the residual earns a mechanism: the public video\-SSL representation is richer than our task\-trained CNN’s, not merely better\-read\. The readout\-limited reading was then tested at scale: the same frozen\-encoder harness on V\-JEPA2 at ViT\-L \(300300M\), ViT\-H \(600600M\), and ViT\-g \(11B\), identical spatial readout and head budget\. The best\-rung gain over ViT\-L is1\.02×1\.02\\times\(factored inversion nMSE0\.585→0\.5730\.585\\to 0\.573, both rungs1010seeds; ViT\-H0\.600\.60at33\) against a pre\-registered1\.5×1\.5\\timesweak\-gain bar, and the escalation itself is instructive: a33\-seed ViT\-g interim read3×3\\timesbetter \(0\.1950\.195\) until the pre\-registered1010\-seed escalation surfaced the rung’s own heavy\-tail seed \(3\.63\.6\), collapsing the gain\. The full dissociation reproduces at11B \(factored drift∼3\{\\sim\}3bounded on10/1010/10seeds; unstructured0\.770\.77, diverges5/105/10\): perception\-side scale does not buy the physics readout here\.

Image\-SSL and the momentum axis\.A frozen*image*\-SSL encoder, DINOv2\([Oquab et al\., 2024](https://arxiv.org/html/2609.19674#bib.bib34)\), transfers only partially under the spatial readout: the stability dissociation holds \(factored drift2\.82\.8vs unstructured6\.86\.8\) but the inversion\-nMSE gap is small \(0\.620\.62vs0\.650\.65\), consistent with a per\-frame encoder exposing position but not momentum\. We tested that mechanism directly: a readout that reads a per\-frame position from DINOv2’s tokens and*differences*it across strided frames \(momentum lives in differences, at the representation level\) closes the gap to nMSE0\.0960\.096, matching VideoMAE, while its unstructured twin stays at0\.520\.52\. So image\-SSL encoders*do*transfer fully once momentum is differenced rather than read from a single frame\. The apparent “video\-encoders\-only” advantage resolves into one momentum\-visibility law\. \(A global pooled readout remains a confound throughout: VideoMAE\+\+pool is null, nMSE0\.760\.76\.\) The law was then*quantified*on the cited model itself \(a supervised\(q,p\)\(q,p\)\-readout probe at matched head budget,33seeds×\\times44temporal configurations\): position readout is configuration\-independent \(nMSE0\.040\.04–0\.090\.09everywhere\), while momentum readout splits exactly along the temporal\-difference channel: V\-JEPA2’s native multi\-frame clip reads momentum at0\.540\.54against1\.351\.35for the same clip with its centre frame repeated \(temporal information ablated at identical compute; a2\.5×2\.5\\timesadvantage, pre\-registered bar2×2\\times\), and DINOv2\-single\-frame sits at1\.29≈1\.29\\approxthe frame\-repeated video encoder \(ratio0\.950\.95\)\. A video encoder therefore buys exactly the temporal\-difference channel the momentum boundary demands, that is, the channel rather than the architecture \(Table[4](https://arxiv.org/html/2609.19674#S4.T4)\)\. This is encoder transfer on synthetic renders, not real\-video object physics \(joint learned detection\-and\-binding remains the open rung\)\.

This block develops the decode\-free JEPA\-style predictor column of the pixel comparison into a grounded family \(3 seeds, matched encoder/data/objective/anchor/budget, so the only variable within a pair is the dynamics module\)\. The unstructured dynamics head has36,74836\{,\}748parameters and the factored fixed\-kinetic head4,4174\{,\}417\(∼8×\{\\sim\}8\\timessmaller\), so every hybrid win is against a strictly larger unconstrained competitor\. CI lower bounds are clipped at the physical floor of each non\-negative metric for display\.

A label\-free image anchor lets the physics prior separate on a decode\-free objective\.A decode\-free JEPA\-style objective \(predict future*latents*against an exponential\-moving\-average target; no decoder\) with the physics prior alone does not invert: the fixed\-kinetic factored JEPA\-style predictor fails \(nMSE1\.461\.46\) and its latent grows \(∼22\{\\sim\}22\), reproducing the empirical recovery condition that the prior needs a fixed latent gauge to act on\. Supplying that gauge*without a renderer or state labels*\(a soft\-argmax image centroid for position and a central finite difference for momentum\) separates dynamics structure from grounding cleanly \(Table[20](https://arxiv.org/html/2609.19674#A3.T20), Fig\.[23](https://arxiv.org/html/2609.19674#A3.F23)\): the anchored unstructured model never inverts \(match00, nMSE0\.680\.68\), while the anchored factored hybrid inverts on every seed \(match11, nMSE0\.400\.40\) and stays bounded \(1\.391\.39\),∼41%\{\\sim\}41\\%lower error under matched everything\. This partially closes the residual position\-grounding boundary \(position grounding without a decoder\)\. Two limits: the anchor assumes calibrated per\-body channels, and its centroid/finite\-difference grounding is noisier than the renderer, so the hybrid \(0\.400\.40\) trails the renderer\-grounded factored fixed\-kinetic model \(0\.120\.12\): a reliable*regime*inversion decoder\-free, not near\-oracle fidelity\. A grounding\-quality control isolates which of the two differences drives that gap: replacing the noisy centroid with an exact\-position anchor \(same world coordinates and finite\-difference momentum, minus centroid and occlusion noise\), holding the JEPA objective fixed, does*not*move the number \(nMSE0\.4050\.405vs0\.4010\.401, match11every seed\), so the anchor is exonerated and the residual gap to0\.120\.12is a property of the decode\-free objective, not of grounding quality\. A composited\-RGB ladder confirms the operative ingredient is persistent temporal grounding, not per\-body channels: naive RGB differencing breaks both variants, but a seven\-frame robust tracker \(palette demixing, visibility\-weighted trajectory fitting; a controlled probe, not a general perception method\) restores the hybrid advantage \(nMSE0\.450\.45vs0\.820\.82, every seed\)\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_c4jepa.png)Figure 23:Decode\-free inversion nMSE across the grounding ladder: the image\-anchored factored hybrid inverts \(match11\) where its unstructured twin does not, on channels and on tracked RGB; with object identity removed the advantage disappears\.Table 20:Decode\-free JEPA family \(3 seeds\)\. “match” is anti\-physics regime\-match; nMSE the inversion error; growth the6060\-step latent\-norm ratio\.Removing object identity relocates the blocker to learned object binding\.When per\-body channels*and*colour identity are removed \(all bodies share one monochrome frame, bound across time by anonymous peak detection and matching\), the hybrid advantage vanishes: both the structured and unstructured variants match the flipped law on only1/31/3seeds with indistinguishable error \(nMSE∼0\.76\{\\sim\}0\.76\), because close encounters make the momentum/identity assignment ambiguous\. The physics dynamics are not the blocker; learned persistent*object binding*\(long\-context slots or tracks\) is\. The label\-free anchor moves the boundary from decoder\-free*position*grounding down to*object*binding, and this anonymous\-object result marks where the hand\-engineered binder fails; the next two results attack that blocker directly\.

A learned binder resolves the anonymous\-object boundary, and the physics prior disambiguates the learned binding\.The anonymous\-object collapse traces to the hand\-engineered binder: it detectsNNblobs and assigns identity across frames by a hard minimum\-acceleration match, which is ambiguous at a close encounter, so the prior is grounded on mis\-bound\(q,p\)\(q,p\)\. Replacing that match with a learned, differentiable, confidence\-aware soft\-assignment \(a “learned tracker”; the detector is unchanged, and the binder is the encoder, its grounded centroids rolled and matched under the decode\-free objective\) resolves it \(Table[21](https://arxiv.org/html/2609.19674#A3.T21), Fig\.[24](https://arxiv.org/html/2609.19674#A3.F24)\): the learned\-tracker factored hybrid inverts on every seed \(match11, nMSE0\.150\.15, bounded growth∼1\.4\{\\sim\}1\.4\), while its unstructured twin does not \(match00, nMSE0\.640\.64\) and the hand\-engineered floor does not \(match0\.330\.33, nMSE0\.760\.76\)\. Extended to the ten\-seed standard, the dissociation is exact: factored inverts on𝟏𝟎/𝟏𝟎\\mathbf\{10/10\}seeds and the unstructured control on𝟎/𝟏𝟎\\mathbf\{0/10\}\(one\-sided exact binomialp≈10−3p\\\!\\approx\\\!10^\{\-3\}each\)\. This is the same double dissociation once more, now on*learned*binding: with the same binder, factored inverts and unstructured fails\. It holds whether the binder is trained through the physics loss \(0\.1520\.152\) or self\-supervised on temporal smoothness and then frozen \(0\.1550\.155\), so physics need not train the binder end\-to\-end to obtain the separation; and0\.150\.15is below the channelised hybrid \(0\.400\.40\), approaching the renderer\-grounded0\.120\.12, because the compact anonymous disks detect cleanly and only the matching needed fixing\.

![Refer to caption](https://arxiv.org/html/2609.19674v1/figures/ap_c4bind.png)Figure 24:Inversion nMSE on the anonymous world across the binding ladder: only the learned\-tracker factored hybrid inverts \(below the0\.50\.5guide\), where the hand\-engineered floor and slot\-attention do not\.Table 21:Learned binding on the anonymous world \(3 seeds\)\. “match” is anti\-physics regime\-match; nMSE the inversion error\.Joint learned detection\-and\-binding is the next rung\.A slot\-attention binder, which must learn*detection*as well as assignment, does not reach the inversion \(match00, nMSE0\.90\.9–1\.01\.0across factored/unstructured and both couplings\)\. Learning the binding on top of a fixed detector is therefore solved; learning detection and binding jointly and decode\-free is not\. The perception frontier accordingly moves one rung further: from persistent object binding, now resolved by a learned soft\-assignment, to*joint*learned detection\-and\-binding\. The learned quantity in the tracker is only the assignment \(4,6094\{,\}609params, detector fixed\); the slot\-attention binder is105,728105\{,\}728params\.

The grounded approach transfers to a frozen public video encoder we did not build\.To test whether the pixel result depends on our specific encoder, we replace the trained CNN with a frozen public encoder \(weights not trained; only a small head and the dynamics train\) in the tracked setting, reading out\(q,p\)\(q,p\)withNNobject queries that cross\-attend over its patch tokens\. On two frozen*video*\-SSL encoders, VideoMAE\([Tong et al\., 2022](https://arxiv.org/html/2609.19674#bib.bib35)\)and the V\-JEPA2 model\([Assran et al\., 2025](https://arxiv.org/html/2609.19674#bib.bib25)\)this paper’s introduction cites \(both arms now at the full 10\-seed standard: factored inverts10/1010/10, exact binomialp=0\.001p=0\.001\), the factored approach inverts on every seed with bounded drift while its unstructured twin is far worse and unstable\. Over1010seeds VideoMAE inverts at nMSE0\.130\.13\(oracle0\.0670\.067\) against0\.650\.65unstructured \(which diverges on a seed\); V\-JEPA2 inverts at0\.590\.59against0\.820\.82unstructured \(the 10\-seed expansion widens the 3\-seed0\.250\.25with two heavy\-tail seeds; the unstructured twin diverges5/105/10, drift up to∼3100\{\\sim\}3100\)\. This is the same dissociation on representations we did not build \(Table[4](https://arxiv.org/html/2609.19674#S4.T4)\)\.

The world\-model half of V\-JEPA2, characterised, with the positive control as the finding\.Our external\-encoder result tested the public model’s*encoder*; here we test its*predictor*: the official action\-conditioned V\-JEPA2\-AC \(ViT\-g encoder \+ 300M frame\-causal predictor, frozen, official release weights\), rolled autoregressively on our rendered physics \(native 256px, 16\-frame context at the dynamics cadence\) with action tokens held at no\-op, in both the zero pose and the release example’s real end\-effector pose \(disclosed; a no\-op stream is itself part of the domain gap for a model trained on robot manipulation\)\. A pre\-registered positive control gates the physics probes: the predictor’s one\-step latent prediction on our renders must beat a copy\-last\-frame baseline by a stated margin before any probe is interpretable\. It does not: pred/copy=0\.978=0\.978–0\.9860\.986across33seeds and both no\-op variants, against a certified\-fair baseline \(consecutive\-frame representations sit at82%82\\%of the*unrelated*\-frame distance, since bodies traverse8%8\\%of the scene per frame, so there is real predictable signal that copying does not capture\)\. The same gate shows the*encoder*half does encode the state: our spatial readout recovers\(q,p\)\(q,p\)from its tokens at0\.280\.28–0\.430\.43nMSE\. Per the pre\-registered rule, the imagined\-rollout probes \(energy drift, counterfactual tracking, second\-law flags\) are VOID and reported as such, not as model failures\. The finding is that the representation transfers across the domain gap \(consistent with the encoder\-transfer result\) while the*predictive*half, under no\-op conditioning far from its training distribution, adds nothing over frame persistence\. This is a characterisation rather than a strawman: the domain gap \(Droid tabletop→\\torenderedNN\-body\) is the headline variable, the fairness artifact and both no\-op variants are in the record, and a same\-domain test would require fine\-tuning the predictor, which is out of scope for a frozen\-model claim\.

Composing a public encoder, learned binding, and the physics prior, the boundary that localises the residual\.The learned\-binder result and the frozen\-encoder result each remove one obstacle separately: the frozen\-encoder result keeps the hand\-engineered77\-frame tracker anchor, and the learned\-binder result’s detections are*real*image centroids\. We composed them so that no part of the perception stack is author\-built: a frozen VideoMAE reading the*anonymous*world, a*learned*soft\-argmax detector replacing the hand\-built centroids, and the learned tracker, under a factored prior\. It does*not*compose\. The factored cell is unstable across seeds \(inversion nMSE0\.80\.8–7\.57\.5, latent growth1\.21\.2–42×42\\times; the same seed even flips between bounded and divergent under GPU nondeterminism, i\.e\. it sits on the edge of stability\) and never reaches the clean inversion of the learned\-binder result \(0\.150\.15\) or the frozen\-encoder result \(0\.130\.13\)\. The unstructured control is worse \(nMSE up to315315\)\. The factored prior keeps only a weak stability edge\. The boundary is informative: this composition removes*both*physical\-grounding sources that made the two earlier results work \(real centroids; a built anchor\), and the physics prior, which can disambiguate binding and transfer across frozen encoders*given*grounding, does not bootstrap grounding in this frozen\-token configuration\. This is the joint learned detection\-and\-grounding residual at the representation level: the tested joint detection\-and\-grounding configuration remains unresolved\.

Damping\-rate ordering from video \[partial\]\.A grounded seven\-frame tracker supplies latent\(q,p\)\(q,p\)for a disk world with conservative or damped dynamics\. The model uses the positive portp←p​exp⁡\[−softplus⁡\(a\)​Δ​t\]p\\leftarrow p\\exp\[\-\\mathrm\{softplus\}\(a\)\\Delta t\]\. The effective rate isr=softplus⁡\(a\)\>0r=\\mathrm\{softplus\}\(a\)\>0; the raw parameteraais not constrained to be nonnegative\. This update contracts canonical quadratic kinetic energy during the isolated port, not arbitrary total energy\.

The source reports rate diagnostics1\.251\.25for damped video and0\.840\.84for conservative video, against reference rates1\.51\.5and00, across five seeds\. Their ordering is recovered, and a label\-free window head tracks a medium switch\. The parameter\-to\-rate reporting convention must be checked before interpreting these diagnostics as calibrated physical rates\. The reportedσ\\sigmavalues are transformed effective rates,σ=softplus⁡\(log⁡σ\)\\sigma=\\mathrm\{softplus\}\(\\log\\sigma\)read from the trained port \(not the raw parameter and not a fitted diagnostic\), in units of inverse physical time at the frame cadence; the estimator validation below gives the grounding comparison and the untested step\-size sensitivity\. The conservative diagnostic remains nonzero; per\-clip arrow\-of\-time detection is not demonstrated \(prediction nMSE approximately0\.90\.9\)\.

Oracle\-grounding and reversed\-video controls\.Replacing grounded pixels with true simulator states leaves the conservative diagnostic near0\.8850\.885, versus0\.840\.84from pixels \(five seeds\)\. Oracle grounding therefore does not eliminate the floor\. This directs attention to model, objective, optimisation, and temporal\-resolution limitations rather than grounding alone; it does not uniquely identify the cause\. On reversed damped clips, the constrained\-port momentum\-loss floor rises from1\.031\.03to1\.241\.24, while the sign\-free control changes from0\.840\.84to0\.870\.87\. The pooled1\.42×1\.42\\timescontrast misses the declared2×2\\timesthreshold\. The retained findings are ordering and a channel constraint, not accurate conservative\-rate recovery or demonstrated video\-level refusal\.

A grounded energy\-balance estimator \[pre\-registered; estimator amended before the run, disclosed\]\.The preceding experiments show that oracle grounding does not eliminate the learned\-model floor; they do not uniquely identify its cause\. A separate estimator uses grounded positions and the known interaction law\. For unit masses and linear velocity drag, it regresses the windowed energy balanceΔ​Ej≈−2​γ​K¯j​Δ​t\\Delta E\_\{j\}\\approx\-2\\gamma\\bar\{K\}\_\{j\}\\Delta t, withE=K\+V⁡\(q,G\)E=K\+V\(q;G\)\. For general masses underp˙i\|drag=−γpi/mi\\dot\{p\}\_\{i\}\|\_\{\\rm drag\}=\-\\gamma p\_\{i\}/m\_\{i\}, the dissipated power is instead−γ∑i∥pi∥2/mi2\-\\gamma\\sum\_\{i\}\\\|p\_\{i\}\\\|^\{2\}/m\_\{i\}^\{2\}\. A time\-constant additive measurement bias cancels in the energy difference; time\-varying tracking noise and discretisation error need not cancel\.*Estimator validation\.*The learned\-rate readout stores the transformed rateσ=softplus⁡\(log⁡σ\)\\sigma=\\mathrm\{softplus\}\(\\log\\sigma\), applied asp←p​e−σ​Δ​tp\\leftarrow p\\,e^\{\-\\sigma\\Delta t\}withΔ​t\\Delta tthe frame cadence \(MMsimulator steps\), soσ\\sigmais a rate per unit physical time; masses are11in this world, so thep/mp/mandp/m2p/m^\{2\}conventions coincide\. The∼0\.85\\sim 0\.85positive bias of the learned conservative rate is*not*tracking noise: the oracle\-grounded arm \(true positions\) givesσcons=0\.85\\sigma\_\{\\rm cons\}=0\.85–0\.930\.93and the pixel\-grounded arm0\.800\.80–0\.890\.89\(five seeds each\), so the bias is the co\-training transient described above, which the grounded energy\-balance estimator removes \(γ^cons=−0\.006\\hat\{\\gamma\}\_\{\\rm cons\}=\-0\.006\)\. The damped rate is recovered at1\.221\.22–1\.301\.30against a truth of1\.51\.5\(about15%15\\%low\) under both groundings\. Sensitivity to the frame cadence was not tested, so the video result is an ordering and an approximate rate, not a calibrated measurement\. One physics amendment was made to the pre\-registration*before*the full run, on smoke evidence: the originally specified kinetic\-envelope form fails on self\-gravitating worlds, whereKK*grows*even without drag \(virial infall, measured0\.066→1\.520\.066\\to 1\.52\); the energy\-balance form replaced it with all gates unchanged\. Result: pixel\-groundedγ^cons=−0\.006\\hat\{\\gamma\}\_\{\\text\{cons\}\}=\-0\.006andγ^damped=1\.63\\hat\{\\gamma\}\_\{\\text\{damped\}\}=1\.63\(truth1\.51\.5\), both gates pass at55seeds; the port carriesσ=max⁡\(γ^,0\)\\sigma=\\max\(\\hat\{\\gamma\},0\), so the estimator supplies the rate and the architecture the guarantee\. The estimator is label\-free by construction \(per\-episode, no medium index\), and a sliding version reads a mid\-clip medium switch cleanly \(γ^\\hat\{\\gamma\}median−0\.05→1\.62\-0\.05\\to 1\.62across a conservative→\\todamped concatenation\), a label\-free discovery pattern completed at the estimator level\. Scope: the meter here uses the known interaction lawV⁡\(q,G\)V\(q;G\); on real video, selecting the form that*becomes*the meter is an open form\-selection problem\. On time\-reversed footage the estimator readsγ^=−1\.66\\hat\{\\gamma\}=\-1\.66while theσ≥0\\sigma\\\!\\geq\\\!0port clamps its carried rate to zero, giving arrow\-of\-time*rate*measurement with architectural refusal, extending the classification\-only contrast of[Pickup et al\. \(2014\)](https://arxiv.org/html/2609.19674#bib.bib39)and[Wei et al\. \(2018\)](https://arxiv.org/html/2609.19674#bib.bib40)\. The training\-floor refusal channel was pursued through two pre\-registered designs and is now*retired as an instrument*: the original sign\-free control reached1\.93×1\.93\\timesagainst the2×2\\timesbar \(not demonstrated under the pre\-registered threshold\), and the exponential\-twin control \(the port’s exact unrestricted twin, the cleanest nested pair\) inverted the prediction for an identified reason: a fixed growth rate amplifies the core’s error bye\|γ^\|​Δ​te^\{\|\\hat\{\\gamma\}\|\\Delta t\}per step \(compounding≈6\.9×\\approx\\\!6\.9\\timesin squared error over the horizon; predicted floor6\.66\.6, measured5\.55\.5\), so under exponential growth a training floor measures*error amplification*, not rate\-fit, so both models degrade on reversed footage on10/1010/10seeds and the channel cannot discriminate\. This is a metric retired by its own control; the refusal claim rests on the estimator channel and the architectural guarantee\. Finally, a distill\-initialisation probe on the composition boundary \(detector distilled from the centroid detector, then released with EMA, gradient clipping, and0\.1×0\.1\\timesdetector learning rate\)*removes the instability*\(latent growth1\.21\.2–1\.3×1\.3\\timeson every seed, versus the bounded↔\\leftrightarrowdivergent flipping\) while inversion plateaus at nMSE0\.730\.73, and the residual is now*localised*: the cured detector’s median position error against the centroid teacher is0\.710\.71–0\.750\.75world units, larger than the disk radius \(0\.60\.6\), so frozen tokens fail at*localisation*on this anonymous world before binding even starts \(measured on the trained stack; the tokens’ information ceiling is not probed\)\. That number is the opening problem for future work\.

## Appendix DClaim–evidence map

This appendix maps each claim of the paper to the evidence that supports it and the class we assign that evidence\.

ClaimEvidenceClassIn the evaluated conservative systems, conservation by construction is associated with lower long\-horizon energy error than the evaluated unstructured baselines; neural\-ODE finiteness shows that structure is not necessary for bounded finite\-horizon states\.Bounded vs diverged energy and robustness to tuning; 10k\-step saturation\.Well\-supported\.Within the evaluated model families, law inversion occurs with factored coupling and not with conservation alone\.Non\-factored fails while factored inverts in the reported fidelity\-valid comparisons; the double\-pendulum engine and DeLaN/SymODEN comparison\. Graph\-network response labels are descriptive pending short\-horizon validity checks\.Well\-supported\.Under grounded pixel training, a canonical fixed kinetic is additionally required for inversion in the evaluated factored models\.The discovered\-perception and pixel\-comparison results \(learned\-kinetic factored fails like non\-factored\)\.Supported\.The results extend to contact dynamics and a generalised\-coordinate double\-pendulum engine\.Contacts; the engine with learnedM⁡\(q\)M\(q\)\.Supported\.The evaluated boundaries include hard contacts, dissipation, reversibility, and ungrounded visual\-state discovery\.Contacts; dissipation; the reversibility flip; perception; the unsupervised pixel branch \(invalid\)\.Supported as boundaries\.When an explicit shortcut is available, penalising its use increases reliance on the factored pathway and improves inversion fidelity\.The shortcut dose–response; the null result without an available shortcut\.Supported\.The factored model exhibits object\-level transfer and cross\-family generalisation within the tested ranges\.NN\-transfer with the regime persisting to 32; Coulomb and the mixed generator\.Supported \(degradation reported\)\.The dissociation is retained under end\-to\-end pixel training with grounded perception\.The end\-to\-end pixel comparison\.Supported, scoped\.Discovered\-from\-pixels perception does not provide position grounding through the objective alone; a label\-free image anchor supplies it without a decoder, and a learned soft\-assignment binder then resolves object binding, leaving joint learned detection\-and\-binding as the open rung\.The discovered\-perception results show the objective alone fails; a decoder\-free image anchor grounds position and the physics prior then separates; a hand\-engineered binder collapses on anonymous objects; a learned tracker\+\+the physics prior inverts \(match11, nMSE0\.150\.15\); slot\-attention, learning detection too, does not yet reach it\.Supported \(boundary moved two levels down\)\.The structured FDT model has the lowest stationary\-temperature error among the evaluated models\.The structured model is correct while the unstructured model diverges and the reversible\-only model never thermalises; this holds under noise\.Well\-supported\.Under the passive\-friction formulation of Prop\.[4](https://arxiv.org/html/2609.19674#Thmproposition4), a global positive\-semidefinite friction guarantee is incompatible with friction\-sign inversion\.Prop\.[4](https://arxiv.org/html/2609.19674#Thmproposition4); the passive channel cannot anti\-damp, whereas sign reversal relinquishes that channel guarantee; the double\-pendulum engine\.Well\-supported \(with proof\)\.Energy injection, as a factored drive, generalises over drive strength and sign\.The driven steady state at trained and held\-out drives, including the unseen sign \(A≥γA\\geq\\gammaproduces the expected divergence\)\.Well\-supported\.Accurate counterfactual inference from factoring depends on correct specification of the coupling’s functional form\.Wrong coupling power gives good fit but a wrong counterfactual with no warning, and the correct structure restores the result\.Well\-supported \(a boundary\)\.The dissociation is reproduced with the evaluated graph\-network and transformer baselines; domain randomisation inverts only when both coupling signs are included in training and does not provide bounded rollouts\.The transformer diverges and does not invert at the inversion horizon; domain randomisation includes both signs in training and does not provide stability\.Well\-supported\.On a decode\-free future\-latent objective, the physics prior separates once a label\-free image anchor fixes the latent gauge; grounding*quality*is not the limiter\.The anchored factored hybrid inverts every seed at nMSE0\.400\.40vs anchored unstructured0\.680\.68; an exact\-position anchor does not improve on the noisy centroid, so the residual gap to the renderer\-grounded0\.120\.12is a property of the objective\.Well\-supported\.The grounded factored approach does not depend on our own encoder: it transfers to frozen public video encoders, one within2×2\\timesof oracle fidelity \(VideoMAE0\.130\.13vs oracle0\.0670\.067\) and one at∼8\.8×\{\\sim\}8\.8\\times\(V\-JEPA20\.590\.59\), with a localisation\-aware readout\.Frozen VideoMAE and V\-JEPA2\+\+spatial readout\+\+factored invert at nMSE0\.130\.13and0\.590\.59vs unstructured0\.650\.65/0\.820\.82and diverging \[1010seeds throughout\]; a global pooled readout misses it; readout\-matched, VideoMAE stays2\.6×2\.6\\timesbetter than our own frozen CNN; scale\-flat across V\-JEPA2 ViT\-L/H/g, gain1\.02×1\.02\\times; a frozen image encoder transfers fully once momentum is differenced across frames\.Well\-supported \(encoder transfer; synthetic renders\)\.
## Appendix EEvidence classification

We classify evidence aswell\-supportedwhen effects are large, rest on an exact test or a clear per\-seed separation rather than on the non\-overlap of marginal intervals, and replicate across the stated architectures, seeds, or epochs\. Findings in this class are conservation\-to\-stability; factored\-coupling\-to\-inversion \(across MLP, Hamiltonian, Lagrangian, and graph\-net; pixel\-trained under grounded perception\); the canonical\-fixed\-kinetic condition; the capacity and data\-efficiency results; the misspecification boundary; the transformer / domain\-randomisation dissociation; the grounded decode\-free hybrid, where an image anchor lets the physics prior separate on a JEPA\-style objective; the learned\-binding crack, where a learned soft\-assignment binder\+\+the physics prior inverts the anonymous\-object world while its unstructured twin fails; the external\-encoder transfer, where the factored approach inverts on two frozen public video encoders we did not build, one within2×2\\timesof oracle fidelity; and, beyond conservation, structure\-to\-correct\-temperature, the guarantee/invertibility tension \(with proof\), and the factored\-drive generalisation\. We classify evidence asdirectionalwhen only a single configuration was evaluated or the per\-seed separation is not clear\. Findings in this class are the reversibility effect on the engine; the Coulomb margin \(3×3\\timesvs1515–40×40\\timesfor gravity\);NN\-scaling \(N=32N\{=\}32\); the necessity dose\-response; in\-regime control \(overlapping CIs\); the perception partial\-recovery; and the engine energy\-shed and noise robustness\. Two further tags complete the vocabulary used in the findings table and claim map:supportedmarks effects that replicate but with a scoped configuration, a qualified control, or only some of their pre\-registered gates passed \(stated in the row\); andboundarymarks a mapped limit, a place where the approach demonstrably stops working, reported as a result of equal standing\. Apre\-registered negativeis a pre\-registered hypothesis whose gates failed, kept at full strength because the failure is the finding \(the energy\-aux unification is the instance\)\.

## Appendix FPre\-registration and reproducibility

The protocol, including the metrics, the equal\-capacity rule, the non\-finite cap, and the permitted conclusions, was fixed in dated documents ahead of the corresponding runs, and serves as a prospective specification and as provenance for the results\. The following were recorded before the corresponding experiments: the fixed\-kinetic and decorrelation designs; the perception and canonical\-latent design; the pixel\-based comparison and its success criterion; the necessity analysis; and the stochastic\-dissipative design, including the initial go/no\-go experiment and temperature as the primary metric\. The guarantee–invertibility trade\-off and the anti\-friction boundary interpretation in the engine and driven experiments were specified in advance\. For the later arc, the specifications are dated ahead of their results: the oracle\-grounding and refusal arms of the video dissipation experiment \(the first of which*refuted our initial grounding\-noise attribution*\); the energy\-auxiliary unification \(all three predictions failed under the pre\-registered thresholds\); and the final estimator and its gates, including one physics amendment recorded*before*the full run \(the kinetic\-envelope form fails on self\-gravitating worlds, and the energy\-balance form replaced it with the gates unchanged\) and the exponential\-twin control whose outcome retired the training\-floor instrument\. The video world\-model characterisation was specified in advance, with its thresholds fixed before any run; its positive control was registered as the experiment’s power condition, and its firing is the reported finding\. The frozen public components used across the external\-encoder experiments are DINOv2\-small, VideoMAE\-base, V\-JEPA2 \(ViT\-L/H/g\), and V\-JEPA2\-AC \(ViT\-g\), each from its official public release pinned to a fixed version\.*Statistics\.*Aggregates are the mean over at least three seeds, and at least five seeds at scale\. When five or more seeds are available, the reported bracket\[\[lo, hi\]\]is a normal\-approximation 95% confidence interval,mean±1\.96​SE\\text\{mean\}\\pm 1\.96\\,\\text\{SE\}, a normal\-approximation interval that we label as such wherever it appears; it is not an observed range, and five seeds do not by themselves justify the approximation, so for the core comparisons we report the per\-seed values and, where the design is paired, the paired difference with its interval\. Significance statements rest on exact tests with a stated null: the binomial regime\-match test has nullPr⁡\(match\)=12\\Pr\(\\text\{match\}\)=\\tfrac\{1\}\{2\}per seed, and the paired sign\-flip tests of the factorial have the null of exchangeable signs of the per\-seed differences, Holm\-corrected across the declared family\. Display clipping must be distinguished from numerical failure\. A finite value above the temperature\-error threshold of100100or drift threshold of10410^\{4\}remains a finite observation\. A non\-finite state or diagnostic is a numerical failure and has no measured finite error\. We do not interpret ratios to substituted failure sentinels quantitatively\. Recomputing finite\-only aggregate statistics requires the underlying per\-run records\.*Horizons\.*Stability is evaluated at 2000 steps,100×100\\timesthe training horizon\. Inversion error is evaluated at 150 steps, within the pre\-Lyapunov window used for all models\. At longer horizons, trajectory MSE saturates even for accurate chaotic\-system models\.*Regime\-match*is therefore the primary inversion metric and nMSE the secondary metric\. The released reproducibility materials comprise per\-experiment model checkpoints, aggregate result tables, the scripts that regenerate the figures, and a fixed software environment\.

Table 23:Compute and hyperparameters \(defaults; per\-experiment overrides in the released configs\)\. All runs use Adam and one\-step prediction loss unless noted\. Each run uses one GPU, with runs distributed across eight GPUs\.

Similar Articles

Mind the Sim-to-Real Gap & Think Like a Scientist

arXiv cs.AI

This paper studies when and how a planner should supplement a pre-trained simulator with real experiments in sequential decision problems, proposing Fisher-SEP to minimize posterior variance of a target policy's value.