SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors

arXiv cs.LG Papers

Summary

SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.

arXiv:2608.04060v1 Announce Type: new Abstract: Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps. We introduce SJEPA, a reconstruction-free JEPA framework that learns predictive representations whose induced dynamics admit compact symbolic descriptions. Its hybrid transition combines a symbolic law with a regularised neural correction for dynamics outside the selected grammar. The central principle is to learn the simplest adequate dynamics: representation constraints preserve informative, non-collapsed predictive coordinates, while operator compression favours low-complexity symbolic-neural transitions that remain predictively adequate. We formalise this principle through induced-dynamics complexity, analyse predictive-coordinate non-identifiability, and show that unconstrained operator compression creates a direct shortcut to representation collapse. The framework supports both alternating representation-equation learning and symbolic dynamics fitted to fixed representations. In controlled pendulum experiments, joint learning discovers substantially simpler symbolic dynamics with lower long-horizon rollout error and divergence than post-hoc fitting, while an unconstrained one-step diagnostic realises the predicted collapse shortcut. Under grammar misspecification, correction regularisation preserves the representable symbolic mechanism and directs the neural component towards residual dynamics. The results expose a controllable trade-off among predictive fidelity, representation quality, symbolic parsimony, and symbolic-neural allocation.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:45 AM

# SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic–Neural Predictors
Source: [https://arxiv.org/html/2608.04060](https://arxiv.org/html/2608.04060)
\(23 July 2026\)

###### Abstract

Joint\-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps\. We introduceSJEPA, a reconstruction\-free JEPA framework that learns predictive representations whose induced dynamics admit compact symbolic descriptions\. Its hybrid transition combines a symbolic law with a regularised neural correction for dynamics that the selected grammar cannot express adequately\. The central principle is to learn the simplest adequate dynamics: representation constraints restrict learning to informative, non\-collapsed predictive coordinates, while operator compression favours the lowest\-complexity symbolic–neural transition that remains predictively adequate\. We formalise this principle through induced\-dynamics complexity, analyse the non\-identifiability of predictive coordinates, and show that unconstrained operator compression creates a direct shortcut to representation collapse\. The framework supports both alternating representation–equation learning and symbolic dynamics fitted to fixed representations\. In controlled pendulum experiments, joint representation–equation learning discovers substantially simpler symbolic dynamics with lower long\-horizon rollout error and divergence than post\-hoc fitting, while an unconstrained one\-step diagnostic realises the collapse shortcut predicted by the theory\. Under grammar misspecification, correction regularisation preserves the representable symbolic mechanism and encourages the neural component to focus on residual dynamics\. The results demonstrate a controllable trade\-off among predictive fidelity, representation quality, symbolic parsimony, and symbolic–neural allocation\. SJEPA provides a complementary direction within the JEPA family and may be useful in applications where compact latent dynamics, explicit structural constraints, or controlled residual modelling are desirable\.

## 1Introduction

Joint\-embedding predictive architectures \(JEPAs\) learn by predicting representations of missing or future observations rather than reconstructing all observation details\(LeCun,[2022](https://arxiv.org/html/2608.04060#bib.bib1); Assranet al\.,[2023](https://arxiv.org/html/2608.04060#bib.bib2); Bardeset al\.,[2024](https://arxiv.org/html/2608.04060#bib.bib3); Assranet al\.,[2025](https://arxiv.org/html/2608.04060#bib.bib4)\)\. This design separates predictive semantics from pixel\-level variability and provides a natural foundation for general\-purpose world models\. Yet the transition mechanism itself is usually represented by a neural predictor\. Such a predictor can be accurate without revealing which variables interact, how actions alter the future, or whether the learned coordinates support a concise dynamical description\.

This paper asks a different question from standard representation learning:*can a JEPA learn not only predictive states, but elegant dynamics over those states?*We use “elegant” in a precise operational sense: an elegant transition is a compact and parsimonious law that remains adequate for prediction\. The objective is therefore not the shortest possible equation\. An equation that is too simple underfits; a representation that is made trivial merely to simplify the equation is also unacceptable\. The target is the*simplest adequate governing law*for an informative predictive state\. The core principle is that operator compression should select predictive coordinates whose induced dynamics are simple yet adequate, while representation constraints prevent that simplicity from being achieved through collapse\.

We introduce Symbolic JEPA \(SJEPA111We use the unhyphenated abbreviationSJEPAbecause the framework is general\-purpose rather than task\-specific\.\), a JEPA whose latent transition is decomposed into a symbolic governing law and a neural correction\. Given a context embeddingZCZ\_\{C\}and target\-side informationε\\varepsilon, withZTZ\_\{T\}providing the prediction target, the hybrid predictor is

Z^T=fℰ,α​\(ZC,ε\)\+cϕ​\(ZC,ε\)\.\\widehat\{Z\}\_\{T\}=f\_\{\\mathcal\{E\},\\alpha\}\(Z\_\{C\},\\varepsilon\)\+c\_\{\\phi\}\(Z\_\{C\},\\varepsilon\)\.\(Eq\.[3](https://arxiv.org/html/2608.04060#S3.E3)\)Here,ε\\varepsilonidentifies the target or transition, for example through a target location, temporal offset, or action\. The symbolic structureℰ\\mathcal\{E\}and coefficientsα\\alphacapture the dominant reusable dynamics, whilecϕc\_\{\\phi\}corrects effects that the selected grammar does not represent adequately\. This decomposition acts on the transition operator rather than partitioning the latent representation\.

SJEPA distinguishes representation compression from operator compression\. The former determines which future\-relevant state is retained; the latter seeks a concise law governing that state\. These objectives must be coupled carefully because an unconstrained pressure for simple dynamics can encourage the encoder to erase information until the transition becomes trivial\. We therefore formulateSJEPAas constrained operator compression:

minθ,θ¯,ℰ,α,ϕ\\displaystyle\\min\_\{\\theta,\\bar\{\\theta\},\\mathcal\{E\},\\alpha,\\phi\}Ω​\(ℰ\)\+λc​ℛcorr​\(ϕ\)\\displaystyle\\Omega\(\\mathcal\{E\}\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)\(Eq\.[10](https://arxiv.org/html/2608.04060#S4.E10)\)subject toℒpred​\(θ,θ¯,ℰ,α,ϕ\)≤δpred,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\(\\theta,\\bar\{\\theta\},\\mathcal\{E\},\\alpha,\\phi\)\\leq\\delta\_\{\\mathrm\{pred\}\},\(θ,θ¯\)∈Θrepr\.\\displaystyle\(\\theta,\\bar\{\\theta\}\)\\in\\Theta\_\{\\mathrm\{repr\}\}\.The objective favours a compact symbolic law while discouraging unnecessary reliance on the correction\. The predictive constraint prevents underfitting, whereasΘrepr\\Theta\_\{\\mathrm\{repr\}\}restricts the context and target encoders to informative, predictively adequate, and non\-collapsed representations\. This formulation induces the dynamics\-complexity functional𝒞dyn​\(Eθ,Eθ¯\)\\mathcal\{C\}\_\{\\mathrm\{dyn\}\}\(E\_\{\\theta\},E\_\{\\bar\{\\theta\}\}\): the minimum symbolic\-plus\-correction complexity required for adequate prediction in the coordinates selected by the encoders\.

Our contributions are:

1. 1\.Symbolic and hybrid JEPA dynamics\.We replace the opaque neural predictor in JEPA with a compact symbolic law, together with an optional neural correction for dynamics that cannot be expressed adequately by the selected symbolic grammar\.
2. 2\.Learning the simplest adequate dynamics\.We introduce operator compression as a complement to representation learning: the representation must remain informative and non\-collapsed, while its transition should be as compact as possible without sacrificing predictive adequacy\.
3. 3\.A principled and modular learning framework\.We develop objectives that control symbolic complexity and discourage the neural correction from unnecessarily absorbing representable dynamics\. The framework supports both alternating representation–equation learning and symbolic dynamics fitted to frozen pretrained encoders\.
4. 4\.Theoretical and empirical validation\.We explain why predictive coordinates are not unique, show how dynamics compression can encourage representation collapse, and validate these effects experimentally\. Joint learning discovers substantially simpler latent dynamics with more accurate and less divergent symbolic rollouts than post\-hoc equation fitting, while correction regularisation preserves the dominant symbolic mechanism under grammar misspecification\.
5. 5\.Extensions to uncertainty and control\.We provide Bayesian and action\-conditioned formulations that support uncertainty\-aware latent dynamics, sensitivity analysis, local linearisation, and model\-based planning\.

## 2Related Works

##### Joint\-embedding predictive architectures\.

The JEPA family has expanded across modalities and learning objectives\. I\-JEPA introduced masked latent prediction for images\(Assranet al\.,[2023](https://arxiv.org/html/2608.04060#bib.bib2)\), while MC\-JEPA jointly learns motion and content features and V\-JEPA extends feature prediction to video\(Bardeset al\.,[2023](https://arxiv.org/html/2608.04060#bib.bib37);[2024](https://arxiv.org/html/2608.04060#bib.bib3)\)\. Audio\-JEPA, Point\-JEPA, and 3D\-JEPA adapt the framework to audio and three\-dimensional data, and VL\-JEPA extends latent prediction to vision–language learning\(Tuncayet al\.,[2025](https://arxiv.org/html/2608.04060#bib.bib38); Saitoet al\.,[2025](https://arxiv.org/html/2608.04060#bib.bib39); Huet al\.,[2024](https://arxiv.org/html/2608.04060#bib.bib40); Chenet al\.,[2026](https://arxiv.org/html/2608.04060#bib.bib41)\)\. Other variants move toward structured world modelling and decision making: ACT\-JEPA jointly predicts actions and latent observation sequences, V\-JEPA 2 develops action\-conditioned video world models for planning, and C\-JEPA introduces object\-level latent interventions for causal reasoning\(Vujinovic and Kovacevic,[2025](https://arxiv.org/html/2608.04060#bib.bib42); Assranet al\.,[2025](https://arxiv.org/html/2608.04060#bib.bib4); Posneret al\.,[2026a](https://arxiv.org/html/2608.04060#bib.bib43)\)\. LeJEPA provides a leaner and theoretically grounded training objective based on isotropic representation regularisation\(Balestriero and LeCun,[2025](https://arxiv.org/html/2608.04060#bib.bib32)\)\. Complementing these deterministic formulations, Huang’s variational JEPA \(VJEPA\) develops a probabilistic JEPA framework that learns predictive distributions over future latent states and connects JEPA with predictive\-state representations and Bayesian filtering\(Huang,[2026b](https://arxiv.org/html/2608.04060#bib.bib44)\)\. Despite this diversity, most variants retain flexible neural predictors or primarily modify the representation objective, modality, or conditioning structure\.SJEPAis complementary: it asks whether the induced latent transition itself can be compressed into a compact symbolic law, with a controlled neural correction for residual dynamics\.

##### Symbolic regression and governing\-equation discovery\.

Symbolic regression searches over analytical expressions rather than assuming a fixed parametric form\(Koza,[1992](https://arxiv.org/html/2608.04060#bib.bib5); Schmidt and Lipson,[2009](https://arxiv.org/html/2608.04060#bib.bib6)\)\. SINDy identifies sparse governing dynamics within a user\-specified library of candidate functions, while SINDYc extends this formulation to systems with external inputs and control\(Bruntonet al\.,[2016b](https://arxiv.org/html/2608.04060#bib.bib7); Kaiseret al\.,[2018](https://arxiv.org/html/2608.04060#bib.bib8)\)\. More recent methods distil symbolic relations from trained neural models, use neural generators or pretrained transformers to propose expressions, or employ scalable evolutionary search\(Cranmeret al\.,[2020](https://arxiv.org/html/2608.04060#bib.bib9); Petersenet al\.,[2021](https://arxiv.org/html/2608.04060#bib.bib10); Biggioet al\.,[2021](https://arxiv.org/html/2608.04060#bib.bib11); Cranmer,[2023](https://arxiv.org/html/2608.04060#bib.bib12)\)\. Most such methods operate on a supplied set of explanatory variables; in governing\-equation discovery, the state coordinates and either their derivatives or transition data are typically observed or selected beforehand\.SJEPAinstead couples symbolic equation discovery with learning the predictive coordinates in which those equations are expressed\.

##### Joint coordinate and equation discovery\.

Several important precedents learn coordinates together with structured dynamics\. SINDy autoencoders jointly learn a nonlinear coordinate transformation, a reconstruction map, and sparse latent governing equations\(Championet al\.,[2019](https://arxiv.org/html/2608.04060#bib.bib29)\)\. T\-SHRED combines transformer\-based temporal encoding with a SINDy\-attention mechanism that regularises and interprets latent dynamics when forecasting full systems from sparse sensor histories\(Yermakovet al\.,[2026](https://arxiv.org/html/2608.04060#bib.bib30)\)\. DYSCO uses multiple noisy views and temporal contrastive learning to jointly recover latent trajectories and structured dynamics, with identification guarantees up to an affine indeterminacy and symbolic recovery within that gauge\(Muratore and Mathis,[2026](https://arxiv.org/html/2608.04060#bib.bib31)\)\. These works establish that representation learning and equation discovery can be coupled\.SJEPAaddresses a complementary setting through a reconstruction\-free JEPA context–target interface, optional target\-side information, and an explicit symbolic law plus neural correction\. It further uses minimum adequate operator complexity to guide representation learning while imposing representation constraints to prevent trivial dynamics from being obtained through collapse\.

##### Learned coordinates for dynamics\.

Latent\-dynamics models combine compressed representations with learned transition models for prediction and simulation\. World Models, for example, learns compact visual states together with recurrent latent dynamics, while Neural ODEs parameterise continuous\-time hidden\-state evolution and can be used within latent\-variable models\(Ha and Schmidhuber,[2018](https://arxiv.org/html/2608.04060#bib.bib14); Chenet al\.,[2018](https://arxiv.org/html/2608.04060#bib.bib13)\)\. Koopman\-based methods pursue a more specific form of simplification by seeking observables or coordinates in which nonlinear dynamics evolve linearly, or approximately linearly, within a suitable invariant subspace\(Bruntonet al\.,[2016a](https://arxiv.org/html/2608.04060#bib.bib15); Luschet al\.,[2018](https://arxiv.org/html/2608.04060#bib.bib16)\)\.SJEPAadopts a different but compatible inductive bias: the latent transition need not be linear, but should admit a compact symbolic description, with a controlled neural correction where the symbolic grammar is insufficient\. Linear latent dynamics remain a possible special case of this broader symbolic operator class\.

##### Information bottlenecks and non\-collapse\.

The information bottleneck seeks a compressed representation that preserves information relevant to a target, with the variational information bottleneck providing a practical latent\-variable formulation\(Tishbyet al\.,[1999](https://arxiv.org/html/2608.04060#bib.bib17); Alemiet al\.,[2017](https://arxiv.org/html/2608.04060#bib.bib18)\)\. Recent analysis of V\-JEPA through the information bottleneck and predictive information bottleneck distinguishes the removal of non\-predictive nuisance information from the separate problem of avoiding representation collapse; predictable nuisance information may still be retained\(Huang,[2026a](https://arxiv.org/html/2608.04060#bib.bib34)\)\. Non\-contrastive representation\-learning methods address collapse and redundancy through mechanisms such as variance preservation, cross\-correlation control, architectural asymmetry, target networks, and distributional regularisation\(Bardeset al\.,[2022](https://arxiv.org/html/2608.04060#bib.bib19); Zbontaret al\.,[2021](https://arxiv.org/html/2608.04060#bib.bib20); Grillet al\.,[2020](https://arxiv.org/html/2608.04060#bib.bib21); Balestriero and LeCun,[2025](https://arxiv.org/html/2608.04060#bib.bib32)\)\. These methods are not equivalent to one another or to an information\-theoretic bottleneck\. InSJEPA,ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}is therefore an umbrella notation for the representation\-side requirement that the learned state remain predictive, informative, and non\-collapsed\. This requirement is particularly important because operator compression introduces a direct shortcut: the encoder can make the transition trivial by making the representation nearly constant\. Our experiments instantiateℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}with a VICReg\-style regulariser and use an unconstrained one\-step diagnostic to show that the predicted fixed\-point collapse shortcut can be reached in practice\.

##### Mechanistic and neuro\-symbolic world models\.

Mechanistic World Models propose organising learned world knowledge around reusable explanatory mechanisms rather than predictive mappings alone\(Posneret al\.,[2026b](https://arxiv.org/html/2608.04060#bib.bib33)\)\. Neuro\-symbolic JEPA explores a complementary connection between JEPA and symbolic reasoning: rule\-informed JEPA uses explicit logical rules to shape the latent energy landscape and uses the learned continuous space to support differentiable rule discovery\(Huang and Raza,[2026](https://arxiv.org/html/2608.04060#bib.bib35)\)\.SJEPAaddresses a different form of symbolic structure\. It does not primarily inject logical rules into the representation or discover rules relating symbolic concepts; instead, it learns a compact symbolic operator that describes how a predictive latent state evolves from context to target, optionally conditioned on time, actions, or other side information\. A regularised neural correction accounts for residual dynamics that the selected grammar cannot express adequately\. The resulting equations should be interpreted as compact predictive mechanisms rather than automatically as physical laws\. Predictive, representation, complexity, and correction constraints are required to prevent an apparently simple symbolic description from being obtained through underfitting, collapse, or delegation of the dynamics to an unrestricted neural component\.

## 3SJEPA: Symbolic and Hybrid Latent Dynamics

### 3\.1Problem formulation and notation

A training example is\(XC,XT,ε\)∼𝒟\(X\_\{C\},X\_\{T\},\\varepsilon\)\\sim\\mathcal\{D\}, whereXCX\_\{C\}is a context observation or history,XTX\_\{T\}is a target observation or future segment, andε\\varepsilondenotes \(possibly empty\) side information supplied to the predictor\.ε\\varepsilonmay encode target location, mask geometry, time offset, action, goal, intervention, or other conditioning variables\. The context and target encoders produce

ZC=Eθ​\(XC\),ZT=Eθ¯​\(XT\),Z\_\{C\}=E\_\{\\theta\}\(X\_\{C\}\),\\qquad Z\_\{T\}=E\_\{\\bar\{\\theta\}\}\(X\_\{T\}\),\(1\)where both embeddings lie in a common latent space222Although the context and target encoders have different parameter values during EMA training, they are treated as two parameterisations of the same ambient latent coordinate space𝒵\\mathcal\{Z\}\. The target encoder supplies slowly moving prediction targets rather than defining a separate latent coordinate chart\. For temporal prediction, the free\-rollout objective additionally trainsHℰ,α,ϕH\_\{\\mathcal\{E\},\\alpha,\\phi\}on its own recursively generated outputs\.\. The target encoder may be an exponential\-moving\-average \(EMA\) copy of the context encoder; in the frozen\-encoder mode, both may instead be fixed copies of a pretrained encoder\. A standard deterministic JEPA predictor can be written as

Z^T=gψ​\(ZC,ε\),\\widehat\{Z\}\_\{T\}=g\_\{\\psi\}\(Z\_\{C\},\\varepsilon\),\(2\)whereψ\\psidenotes the parameters of an unrestricted neural predictor\. InSJEPA, the predictor is333Equation \([3](https://arxiv.org/html/2608.04060#S3.E3)\) gives the general, direct\-transition realisation ofSJEPA, in which the symbolic and neural components directly predict the target embedding\. For continuous\-time temporal systems \(e\.g\. a dynamical system\), the same decomposition may instead parameterise a latent vector field \(e\.g\. gradient\)Vℰ,α,ϕ​\(z,ε\)=Fℰ,α​\(z,ε\)\+cϕ​\(z,ε\),V\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,\\varepsilon\)=F\_\{\\mathcal\{E\},\\alpha\}\(z,\\varepsilon\)\+c\_\{\\phi\}\(z,\\varepsilon\),from which the complete transition is constructed by an integration operator:Hℰ,α,ϕ,Δ​t​\(z,ε\)=ℐΔ​t​\(z,Vℰ,α,ϕ,ε\)\.H\_\{\\mathcal\{E\},\\alpha,\\phi,\\Delta t\}\(z,\\varepsilon\)=\\mathcal\{I\}\_\{\\Delta t\}\\left\(z,V\_\{\\mathcal\{E\},\\alpha,\\phi\},\\varepsilon\\right\)\.Our later experiments use the forward\-Euler realisationHℰ,α,ϕ,Δ​t​\(z\)=z\+Δ​t​\[Fℰ,α​\(z\)\+cϕ​\(z\)\]\.H\_\{\\mathcal\{E\},\\alpha,\\phi,\\Delta t\}\(z\)=z\+\\Delta t\\left\[F\_\{\\mathcal\{E\},\\alpha\}\(z\)\+c\_\{\\phi\}\(z\)\\right\]\.Thus, the symbolic–neural decomposition may apply either to a direct transition map or to a vector field from which the transition map is constructed\.

Hℰ,α,ϕ​\(ZC,ε\)=Fℰ,α​\(ZC,ε\)\+cϕ​\(ZC,ε\),H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(Z\_\{C\},\\varepsilon\)=F\_\{\\mathcal\{E\},\\alpha\}\(Z\_\{C\},\\varepsilon\)\+c\_\{\\phi\}\(Z\_\{C\},\\varepsilon\),\(3\)whereFℰ,α≡Fℰ,αF\_\{\\mathcal\{E\},\\alpha\}\\equiv F\_\{\\mathcal\{E\},\\alpha\}is a symbolic expression with structureℰ\\mathcal\{E\}and coefficientsα\\alpha, andcϕc\_\{\\phi\}is a neural corrector for dynamics not captured adequately by the symbolic grammar\. Settingcϕ≡0c\_\{\\phi\}\\equiv 0recovers the purely symbolic model\.

The structureℰ\\mathcal\{E\}is a typed syntax tree drawn from a grammar𝔈\\mathfrak\{E\}of arithmetic, low\-order polynomial, transcendental, and domain\-specific primitives\. Its structural complexity is

Ω​\(ℰ\)=∑v∈nodes⁡\(ℰ\)wtype⁡\(v\),\\Omega\(\\mathcal\{E\}\)=\\sum\_\{v\\in\\operatorname\{nodes\}\(\\mathcal\{E\}\)\}w\_\{\\operatorname\{type\}\(v\)\},\(4\)wherewtype⁡\(v\)≥0w\_\{\\operatorname\{type\}\(v\)\}\\geq 0assigns a complexity cost to each operator or primitive\. This weighted score favours concise expressions while allowing more expressive or structurally complex primitives to receive larger penalties\.

In the differentiable sparse\-library implementation used in our experiments,ℰ\\mathcal\{E\}is represented by the active library support andα\\alphaby its coefficients\. In that case, the same notationΩ​\(ℰ\)\\Omega\(\\mathcal\{E\}\)denotes the sum of prespecified complexity weights over active output–term pairs; a smooth coefficient\-dependent surrogate is used during optimisation and the thresholded support is used for final reporting\. Alternatively, the structural score may be replaced by a minimum\-description\-length criterion444The minimum description length \(MDL\) principle selects the model that minimises the total code length required to describe both the model and the data not explained by it\. In our setting, this balances the description length of the symbolic structure and its coefficients against the remaining predictive error\. MDL therefore does not favour the shortest expression unconditionally; it favours a concise expression that accounts adequately for the observed latent transitions\.\.

The preceding definitions specify the representation and hybrid transition model\. Section[4](https://arxiv.org/html/2608.04060#S4)introduces predictive adequacy, symbolic complexity, and correction control, and combines them to formulate elegant\-dynamics learning as constrained operator compression\.

XCX\_\{C\}EθE\_\{\\theta\}ZCZ\_\{C\}ε\\varepsilonsymbolic lawFℰ,αF\_\{\\mathcal\{E\},\\alpha\}neural correctioncϕc\_\{\\phi\}\+\+Z^T\\widehat\{Z\}\_\{T\}XTX\_\{T\}Eθ¯E\_\{\\bar\{\\theta\}\}ZTZ\_\{T\}prediction lossFigure 1:SJEPAdecomposes the latent transition operator into a compact symbolic governing law and a neural correction for dynamics\. The decomposition acts on the transition operator rather than partitioning the latent representation\.
### 3\.2Why a hybrid predictor?

A purely symbolic law is directly inspectable but may fail when the chosen grammar omits relevant interactions555Discovering interactions, especially higher\-order ones, is a challenging problem in statistical learning and can be particularly difficult in physical and dynamical systems\.or when the latent coordinates retain unresolved variability\. A purely neural predictor is flexible but may absorb predictive regularities without exposing them\. The hybrid model therefore seeks a disciplined division of labour: the symbolic component represents the dominant reusable mechanism, while the regularised correction accounts for effects that a compact law does not capture adequately\. In a physical analogy,Fℰ,αF\_\{\\mathcal\{E\},\\alpha\}may represent a governing relation, whilecϕc\_\{\\phi\}may capture friction, unresolved interactions, approximation error, observation effects, or representation mismatch\.

This interpretation does not imply that the two components form a unique or intrinsically physical decomposition\. Prediction loss alone cannot determine how a transition should be divided betweenFℰ,αF\_\{\\mathcal\{E\},\\alpha\}andcϕc\_\{\\phi\}\. The symbolic\-complexity and correction penalties instead encode a modelling preference: use a concise symbolic explanation where adequate, and reserve the correction for the remaining predictive structure\. Section[4](https://arxiv.org/html/2608.04060#S4)formalises this allocation and its limitations\.

### 3\.3The encoder as a learned transform

In the end\-to\-end mode, the encoder pair can be interpreted as learning coordinates in which the context–target relation admits a simple symbolic description\. In the frozen\-encoder mode, the same interpretation instead asks whether the pretrained coordinates already expose such a relation\. Let𝒯ε\\mathcal\{T\}\_\{\\varepsilon\}denote the observation\-space target mechanism associated with side informationε\\varepsilon, written schematically as

XT≈𝒯ε​\(XC\)\.X\_\{T\}\\approx\\mathcal\{T\}\_\{\\varepsilon\}\(X\_\{C\}\)\.\(5\)For temporal tasks,𝒯ε\\mathcal\{T\}\_\{\\varepsilon\}is a dynamical transition; for masked or spatial prediction, it denotes the corresponding target\-generating relation\. In stochastic settings, it represents the relevant conditional target mechanism rather than a deterministic map\.

There are then two routes from the context observation to the target latent representation\. The first applies the observation\-space target mechanism and subsequently encodes the resulting target; the second encodes the context and predicts the target embedding from the context embedding together with the same side information\.SJEPAseeks consistency between these routes:

Eθ¯​\(𝒯ε​\(XC\)\)≈Hℰ,α,ϕ​\(Eθ​\(XC\),ε\)\.E\_\{\\bar\{\\theta\}\}\\left\(\\mathcal\{T\}\_\{\\varepsilon\}\(X\_\{C\}\)\\right\)\\approx H\_\{\\mathcal\{E\},\\alpha,\\phi\}\\left\(E\_\{\\theta\}\(X\_\{C\}\),\\varepsilon\\right\)\.\(6\)Thus,ε\\varepsilonidentifies the relevant target relation and is supplied explicitly, together with the context embeddingZC=Eθ​\(XC\)Z\_\{C\}=E\_\{\\theta\}\(X\_\{C\}\), to the latent predictorHℰ,α,ϕH\_\{\\mathcal\{E\},\\alpha,\\phi\}\.

The analogy is to selecting Laplace, Fourier, canonical, or other transformed coordinates that simplify a governing relation\. The claim is not that every system becomes linear or that the resulting coordinates must coincide with physical state variables\. Rather,the representation should expose informative, non\-collapsed coordinates in which an adequate symbolic context–target operator has low description complexity\. For temporal tasks, this operator is the latent transition law\. Such coordinates may be non\-unique and need not remain affinely aligned with the original state, as observed later in Experiment 1\.

## 4Learning Elegant Dynamics by Operator Compression

### 4\.1Predictive adequacy and operator complexity

For a pointwise latent\-space discrepancydd, the one\-step predictive risk is

ℒpred=𝔼𝒟​\[d​\(Hℰ,α,ϕ​\(ZC,ε\),sg⁡\(ZT\)\)\],\\mathcal\{L\}\_\{\\mathrm\{pred\}\}=\\mathbb\{E\}\_\{\\mathcal\{D\}\}\\left\[d\\left\(H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(Z\_\{C\},\\varepsilon\),\\operatorname\{sg\}\(Z\_\{T\}\)\\right\)\\right\],\(7\)wheresg⁡\(⋅\)\\operatorname\{sg\}\(\\cdot\)stops gradients through the target embedding\. The discrepancyddmay, for example, be a squared Euclidean or cosine distance\. For temporal prediction, one\-step accuracy may be supplemented by the multi\-step rollout lossℒroll\\mathcal\{L\}\_\{\\mathrm\{roll\}\}defined in Equation \([61](https://arxiv.org/html/2608.04060#A1.E61)\) of Appendix[A](https://arxiv.org/html/2608.04060#A1)\.

The neural correction should improve predictive fidelity without absorbing reusable structure that can be represented adequately by the symbolic law\. We therefore regularise both its output magnitude and parameter capacity:

ℛcorr​\(ϕ\)=𝔼𝒟​\[‖cϕ​\(ZC,ε\)‖22\]\+τϕ​‖ϕ‖22,\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)=\\mathbb\{E\}\_\{\\mathcal\{D\}\}\\left\[\\left\\\|c\_\{\\phi\}\(Z\_\{C\},\\varepsilon\)\\right\\\|\_\{2\}^\{2\}\\right\]\+\\tau\_\{\\phi\}\\\|\\phi\\\|\_\{2\}^\{2\},\(8\)whereτϕ≥0\\tau\_\{\\phi\}\\geq 0controls parameter regularisation\. The first term discourages large corrections at the prediction level, while the second limits unnecessary complexity in the correction model\.

###### Definition 4\.1\(Operator\-complexity score\)\.

For a fixed latent normalisation, symbolic grammar, and correction class, define the operator\-complexity score of the hybrid transitionHℰ,α,ϕ=Fℰ,α\+cϕH\_\{\\mathcal\{E\},\\alpha,\\phi\}=F\_\{\\mathcal\{E\},\\alpha\}\+c\_\{\\phi\}in Equation \([3](https://arxiv.org/html/2608.04060#S3.E3)\) by

Γ​\(ℰ,ϕ\)=Ω​\(ℰ\)\+λc​ℛcorr​\(ϕ\),\\Gamma\(\\mathcal\{E\},\\phi\)=\\Omega\(\\mathcal\{E\}\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\),\(9\)whereλc≥0\\lambda\_\{\\mathrm\{c\}\}\\geq 0controls the trade\-off between symbolic parsimony and reliance on the neural correction\.

###### Definition 4\.2\(Elegant dynamics\)\.

Fix an admissible prediction thresholdδpred\\delta\_\{\\mathrm\{pred\}\}and representation classΘrepr\\Theta\_\{\\mathrm\{repr\}\}\. An*elegant*SJEPAsolution is a feasible encoder–predictor tuple that minimisesΓ\\Gamma\. It therefore realises the simplest adequate transition within the chosen symbolic grammar and correction class, rather than the simplest transition without qualification\.

This definition balances predictive fidelity, symbolic parsimony, and reliance on the correction\. Predictive adequacy prevents a short but inaccurate law from being considered elegant\.

### 4\.2Representation compression versus operator compression

Representation learning and dynamics learning play complementary roles, but they answer different questions:

What future\-relevant state should be retained?and

What is the simplest adequate law governing that state?
Representation learning must avoid collapse, redundancy, and the removal of variables required for future prediction, while discarding observation\-specific nuisance information where possible\. The representation module may therefore implement a full information bottleneck\(Huang,[2026a](https://arxiv.org/html/2608.04060#bib.bib34)\)or a practical surrogate that preserves an informative, non\-collapsed predictive state\.

Dynamics learning faces a different trade\-off\. If the transition is insufficiently compressed, a flexible symbolic expression or neural correction may reproduce the data accurately while remaining opaque\. If it is compressed too aggressively, the resulting law may underfit, extrapolate poorly, or pressure the encoder towards artificially simple coordinates\. A misspecified grammar may omit mechanisms required for prediction, while an overly expressive correction may absorb the structure that the symbolic component is intended to reveal\. The objective is therefore not the simplest possible dynamics, but the simplest law that remains predictively adequate\.

Given an admissible representation, the dynamics module controls the induced transition through the symbolic\-complexity termΩ​\(ℰ\)\\Omega\(\\mathcal\{E\}\)and correction penaltyℛcorr​\(ϕ\)\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)\. Thus,SJEPAdoes not seek simplicity by discarding the state itself; it seeks an informative representation whose evolution admits a concise and adequate description\.

Table[1](https://arxiv.org/html/2608.04060#S4.T1)summarises the distinct failure regimes associated with the two forms of compression\. The slogan “compress the dynamics, not the representation” should therefore be interpreted precisely: the operator\-complexity objective must not achieve simplicity by erasing the state666For example, a collapsedSJEPAmodel may map every context and target observation to the same arbitrary vectorz0z\_\{0\}, so thatEθ​\(XC\)=Eθ¯​\(XT\)=z0E\_\{\\theta\}\(X\_\{C\}\)=E\_\{\\bar\{\\theta\}\}\(X\_\{T\}\)=z\_\{0\}\. An identity predictorHℰ,α,ϕ​\(z,ε\)=zH\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,\\varepsilon\)=z, or any transition havingz0z\_\{0\}as a fixed point, then predicts every target embedding perfectly\. Collapse is characterised by the absence of variation or information in the representation, not by whether its constant value is zero\.\. Representation compression may discard nuisance information only while preserving predictive sufficiency and non\-collapse; operator compression may simplify the transition only while preserving predictive adequacy\.

Table 1:Representation and operator compression address different objectives and failure modes\. Operator compression is meaningful only within an admissible predictive representation class\.
### 4\.3Constrained formulation and induced\-dynamics complexity

LetΘrepr\\Theta\_\{\\mathrm\{repr\}\}denote the admissible set of encoder pairs satisfying the selected predictive\-information and anti\-collapse requirements\. The idealSJEPAobjective is

minθ,θ¯,ℰ,α,ϕ\\displaystyle\\min\_\{\\theta,\\bar\{\\theta\},\\mathcal\{E\},\\alpha,\\phi\}Γ​\(ℰ,ϕ\)\\displaystyle\\Gamma\(\\mathcal\{E\},\\phi\)\(10\)s\.t\.\\displaystyle\\mathrm\{s\.t\.\}ℒpred≤δpred,\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\\leq\\delta\_\{\\mathrm\{pred\}\},\(θ,θ¯\)∈Θrepr\.\\displaystyle\(\\theta,\\bar\{\\theta\}\)\\in\\Theta\_\{\\mathrm\{repr\}\}\.The predictive constraint prevents operator simplicity from being obtained through underfitting, while the representation constraint prevents it from being obtained through collapsed or uninformative coordinates\.

For fixed encoders, define the induced\-dynamics complexity

𝒞dyn​\(Eθ,Eθ¯\)=infℰ,α,ϕ\{Γ​\(ℰ,ϕ\):ℒpred≤δpred\}\.\\mathcal\{C\}\_\{\\mathrm\{dyn\}\}\(E\_\{\\theta\},E\_\{\\bar\{\\theta\}\}\)=\\inf\_\{\\mathcal\{E\},\\alpha,\\phi\}\\left\\\{\\Gamma\(\\mathcal\{E\},\\phi\):\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\\leq\\delta\_\{\\mathrm\{pred\}\}\\right\\\}\.\(11\)This quantity asks how simple the latent transition can be once the representation coordinates have been fixed\. It searches over symbolic structuresℰ\\mathcal\{E\}, coefficientsα\\alpha, and neural correctionsϕ\\phithat satisfy the predictive requirement, and returns the smallest symbolic\-plus\-correction complexity among them\. A low value indicates that adequate prediction requires a compact symbolic expression and little reliance on the correction\. If no predictor satisfies the constraint, then𝒞dyn​\(Eθ,Eθ¯\)=\+∞\.\\mathcal\{C\}\_\{\\mathrm\{dyn\}\}\(E\_\{\\theta\},E\_\{\\bar\{\\theta\}\}\)=\+\\infty\.For temporal applications in which recursive accuracy is part of the required notion of adequacy, the feasible set in Equation \([11](https://arxiv.org/html/2608.04060#S4.E11)\) may additionally impose777We retain the one\-step constraint in the task\-general definition because recursive rollout is not defined for every JEPA context–target relation\.ℒroll≤δroll\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\\leq\\delta\_\{\\mathrm\{roll\}\}\.

Representation selection may then be written as

min\(θ,θ¯\)∈Θrepr⁡𝒞dyn​\(Eθ,Eθ¯\)\.\\min\_\{\(\\theta,\\bar\{\\theta\}\)\\in\\Theta\_\{\\mathrm\{repr\}\}\}\\mathcal\{C\}\_\{\\mathrm\{dyn\}\}\(E\_\{\\theta\},E\_\{\\bar\{\\theta\}\}\)\.\(12\)The inner optimisation finds the simplest adequate transition for each fixed representation, while the outer optimisation selects, among informative and non\-collapsed representations, coordinates whose induced dynamics are easier to describe\. Equation \([12](https://arxiv.org/html/2608.04060#S4.E12)\) formalises the central principle ofSJEPA: a representation is valuable not only because it supports prediction, but also because its induced dynamics admit a concise and adequate governing law\.

### 4\.4Predictive non\-identifiability and coordinate selection

###### Proposition 4\.4\(Predictive coordinates are non\-identifiable\)\.

LetEEmap context and target observations into a common latent space𝒵\\mathcal\{Z\}, and suppose that

E​\(XT\)=F​\(E​\(XC\),ε\)E\(X\_\{T\}\)=F\(E\(X\_\{C\}\),\\varepsilon\)\(13\)almost surely, whereF:𝒵×𝒮ε→𝒵F:\\mathcal\{Z\}\\times\\mathcal\{S\}\_\{\\varepsilon\}\\rightarrow\\mathcal\{Z\}\. For any bijectionh:𝒵→𝒵~h:\\mathcal\{Z\}\\rightarrow\\widetilde\{\\mathcal\{Z\}\}, define

E~=h∘E\\widetilde\{E\}=h\\circ E\(14\)and

F~​\(z~,ε\)=h​\(F​\(h−1​\(z~\),ε\)\)\.\\widetilde\{F\}\(\\widetilde\{z\},\\varepsilon\)=h\\left\(F\(h^\{\-1\}\(\\widetilde\{z\}\),\\varepsilon\)\\right\)\.\(15\)Then

E~​\(XT\)=F~​\(E~​\(XC\),ε\)\\widetilde\{E\}\(X\_\{T\}\)=\\widetilde\{F\}\(\\widetilde\{E\}\(X\_\{C\}\),\\varepsilon\)\(16\)almost surely\. However, the symbolic description complexities ofFFandF~\\widetilde\{F\}need not agree\.

Intuitively, the proposition says that a predictive latent representation can be rewritten in any invertible coordinate system without changing what the model predicts\. The encoder first expresses the state in the new coordinates throughhh, while the transformed predictor undoes this change, applies the original transition, and maps the result back\. Prediction alone therefore cannot distinguish between these coordinate systems\. Their symbolic descriptions, however, may differ substantially: a transition that is simple in one coordinate system may become complicated in another\.

The transformed transition follows the coordinate route

z~C→h−1zC→F​\(⋅,ε\)zT→ℎz~T\.\\widetilde\{z\}\_\{C\}\\xrightarrow\{\\,h^\{\-1\}\\,\}z\_\{C\}\\xrightarrow\{\\,F\(\\cdot,\\varepsilon\)\\,\}z\_\{T\}\\xrightarrow\{\\,h\\,\}\\widetilde\{z\}\_\{T\}\.It maps the transformed context state back to the original coordinates, applies the original transition, and maps the predicted target state into the transformed coordinates\.

The proof is given in Appendix[B\.1\.1](https://arxiv.org/html/2608.04060#A2.SS1.SSS1)\. Prediction alone therefore does not determine a preferred coordinate system\. Operator complexity supplies an additional inductive criterion: among predictively adequate coordinates, prefer those whose transition is easier to describe\. This criterion does not guarantee that the selected coordinates coincide with physically meaningful variables\.

### 4\.5Dynamics compression creates a collapse shortcut

###### Proposition 4\.5\(Degenerate minimum without representation constraints\)\.

Assume that the encoder classes contain a constant mapE0​\(x\)=z0E\_\{0\}\(x\)=z\_\{0\}and that the predictor class contains a transitionH0H\_\{0\}satisfying

H0​\(z0,ε\)=z0H\_\{0\}\(z\_\{0\},\\varepsilon\)=z\_\{0\}\(17\)for every admissibleε\\varepsilon\. Suppose further thatH0H\_\{0\}attains the minimum value ofλs​Ω​\(ℰ\)\+λc​ℛcorr​\(ϕ\)\\lambda\_\{\\mathrm\{s\}\}\\Omega\(\\mathcal\{E\}\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)within the predictor class and that all objective terms are nonnegative\. Then the unconstrained objective

ℒpred\+λs​Ω​\(ℰ\)\+λc​ℛcorr​\(ϕ\)\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega\(\\mathcal\{E\}\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)\(18\)admits the collapsed solution\(E0,E0,H0\)\(E\_\{0\},E\_\{0\},H\_\{0\}\)as a global minimiser wheneverd​\(z0,z0\)=0d\(z\_\{0\},z\_\{0\}\)=0\.

This proposition says that, without explicit representation constraints, the easiest way to obtain simple and perfectly predictable latent dynamics may be to remove all variation from the representation\. Once both encoders map every observation to the same point, any predictor for whichz0z\_\{0\}is a fixed point achieves zero latent prediction error\. For example, the identity transition

H0​\(z,ε\)=zH\_\{0\}\(z,\\varepsilon\)=z\(19\)satisfiesH0​\(z0,ε\)=z0H\_\{0\}\(z\_\{0\},\\varepsilon\)=z\_\{0\}for everyz0z\_\{0\}\. The model may therefore attain zero predictive loss and minimal operator complexity while retaining no information about the observations or their dynamics\.

Appendix[B\.1\.2](https://arxiv.org/html/2608.04060#A2.SS1.SSS2)gives the proof\. This failure mode is more specific than the general possibility of collapse in non\-contrastive representation learning: operator compression creates a direct incentive to make the transition trivial by first making the representation trivial\. Representation constraints are therefore essential to ensure that the method simplifies the governing law rather than erasing the state on which it acts\.

### 4\.6Controlling allocation to the correction

###### Proposition 4\.6\(Non\-identifiability of the hybrid decomposition\)\.

Letrrbe any function for whichFℰ,α\+rF\_\{\\mathcal\{E\},\\alpha\}\+rremains in the symbolic function class andcϕ−rc\_\{\\phi\}\-rremains in the correction class\. Then

Fℰ,α\+cϕ=\(Fℰ,α\+r\)\+\(cϕ−r\),F\_\{\\mathcal\{E\},\\alpha\}\+c\_\{\\phi\}=\(F\_\{\\mathcal\{E\},\\alpha\}\+r\)\+\(c\_\{\\phi\}\-r\),\(20\)so predictive loss alone cannot determine how structure is allocated between the two components\.

Intuitively, the same overall predictor can be obtained by transferring part of the transition function between the symbolic law and the neural correction\. This fact, proved in Appendix[B\.2\.1](https://arxiv.org/html/2608.04060#A2.SS2.SSS1), motivates two forms of control\. The symbolic\-complexity penalty prevents uncontrolled symbolic growth, whileℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}prevents a powerful correction network from explaining predictable structure unnecessarily\. The resulting decomposition represents an inductive preference; stronger identifiability would require additional assumptions on the grammar, correction class, data support, or relationship between the two function spaces\.

###### Proposition 4\.7\(Regularised allocation to the correction\)\.

Fix the encoders and a symbolic lawFF\. Let

m​\(z,ε\)=𝔼​\[ZT∣ZC=z,ε\],m\(z,\\varepsilon\)=\\mathbb\{E\}\[Z\_\{T\}\\mid Z\_\{C\}=z,\\varepsilon\],\(21\)and suppose the correction ranges over all square\-integrable functions\. Forλ\>0\\lambda\>0, the minimiser of

𝔼​\[‖ZT−F​\(ZC,ε\)−c​\(ZC,ε\)‖22\+λ​‖c​\(ZC,ε\)‖22\]\\mathbb\{E\}\\left\[\\left\\\|Z\_\{T\}\-F\(Z\_\{C\},\\varepsilon\)\-c\(Z\_\{C\},\\varepsilon\)\\right\\\|\_\{2\}^\{2\}\+\\lambda\\left\\\|c\(Z\_\{C\},\\varepsilon\)\\right\\\|\_\{2\}^\{2\}\\right\]\(22\)is

c⋆​\(z,ε\)=m​\(z,ε\)−F​\(z,ε\)1\+λc^\{\\star\}\(z,\\varepsilon\)=\\frac\{m\(z,\\varepsilon\)\-F\(z,\\varepsilon\)\}\{1\+\\lambda\}\(23\)almost surely\.

The proof is given in Appendix[B\.2\.2](https://arxiv.org/html/2608.04060#A2.SS2.SSS2)\. Intuitively, the regulariser prevents the correction from fully absorbing the discrepancy left by the symbolic law, with largerλ\\lambdaassigning less of that discrepancy to the correction\. Whenλ=0\\lambda=0, an unrestricted correction absorbs the complete conditional\-mean discrepancym−Fm\-F\. Forλ\>0\\lambda\>0, this discrepancy is shrunk by the factor\(1\+λ\)−1\(1\+\\lambda\)^\{\-1\}, limiting how much predictive structure is assigned to the correction\. Because Proposition[4\.7](https://arxiv.org/html/2608.04060#S4.Thmdefinition7)holdsFFfixed, stronger regularisation leaves more discrepancy unexplained in this analytical setting\. In the full joint optimisation, however, it encourages structure expressible by the symbolic grammar to be retained in the symbolic law\. A finite neural correction generally learns a regularised projection of the residual onto its chosen function class\.

### 4\.7Long\-horizon prediction is a separate requirement

One\-step adequacy does not imply accurate or non\-divergent recursive prediction\. LetTTdenote the true latent transition andHHthe learned hybrid transition\.

###### Proposition 4\.8\(Rollout error under uniform one\-step adequacy\)\.

Assume thatH​\(⋅,u\)H\(\\cdot,u\)isLL\-Lipschitz for every relevant action or conditioning valueuu, and that

supz,u‖T​\(z,u\)−H​\(z,u\)‖2≤δ\.\\sup\_\{z,u\}\\left\\\|T\(z,u\)\-H\(z,u\)\\right\\\|\_\{2\}\\leq\\delta\.\(24\)If

eh=‖zt\+h−z^t\+h‖2,e\_\{h\}=\\left\\\|z\_\{t\+h\}\-\\widehat\{z\}\_\{t\+h\}\\right\\\|\_\{2\},\(25\)then

eh≤Lh​e0\+\{δ​Lh−1L−1,L≠1,h​δ,L=1\.e\_\{h\}\\leq L^\{h\}e\_\{0\}\+\\begin\{cases\}\\displaystyle\\delta\\frac\{L^\{h\}\-1\}\{L\-1\},&L\\neq 1,\\\\\[6\.0pt\] h\\delta,&L=1\.\\end\{cases\}\(26\)

Intuitively, long\-horizon error depends not only on the one\-step approximation errorδ\\delta, but also on how strongly the learned transition amplifies existing errors through its Lipschitz constantLL\. Appendix[B\.3\.1](https://arxiv.org/html/2608.04060#A2.SS3.SSS1)gives the proof\. The bound clarifies both the promise and the limitation of operator compression: a compact law may reduce local approximation error and expose properties relevant to stability analysis, but compactness alone does not guaranteeL<1L<1or prevent errors from accumulating\. Long\-horizon prediction must therefore be assessed separately through recursive rollout metrics, divergence rates, and, where appropriate, invariant or attractor statistics\.

## 5Representation and Dynamics Learning

This section turns the constrained principle of Section[4](https://arxiv.org/html/2608.04060#S4)into a practical learning procedure\. We first describe the representation regulariser that preserves informative, non\-collapsed latent states\. We then introduce a unified practical objective combining prediction, rollout, representation quality, symbolic complexity, and correction control\. Finally, we present two deployment modes: end\-to\-end alternating learning, which jointly adapts the coordinates and their dynamics, and frozen\-encoder learning, which discovers dynamics in an existing representation space\. Exact experimental instantiations are given in Section[8](https://arxiv.org/html/2608.04060#S8), while stage\-specific objectives and optimisation details are deferred to Appendix[C](https://arxiv.org/html/2608.04060#A3)\.

### 5\.1A modular representation regulariser

The representation module should preserve information needed for prediction while preventing the collapse shortcut identified in Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5)\. We write its training objective abstractly as

ℒrepr=ℒpred\+λr​ℒroll\+λIB​ℛIB,\\mathcal\{L\}\_\{\\mathrm\{repr\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{IB\}\}\\mathcal\{R\}\_\{\\mathrm\{IB\}\},\(27\)whereλr=0\\lambda\_\{\\mathrm\{r\}\}=0when multi\-step prediction is not applicable\.

One information\-theoretic choice is the predictive bottleneck

ℛIBMI=I​\(XC;ZC\)−βrel​I​\(ZC;ZT\),\\mathcal\{R\}\_\{\\mathrm\{IB\}\}^\{\\mathrm\{MI\}\}=\\mathrm\{I\}\(X\_\{C\};Z\_\{C\}\)\-\\beta\_\{\\mathrm\{rel\}\}\\mathrm\{I\}\(Z\_\{C\};Z\_\{T\}\),\(28\)which compresses the context representation while preserving information relevant to the target\(Huang,[2026a](https://arxiv.org/html/2608.04060#bib.bib34)\)\. In practice,ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}may instead be implemented using a variational bottleneck, VICReg\-style variance and covariance constraints, Barlow Twins, latent noise, dimensional restrictions, or related non\-collapse mechanisms\. These alternatives are not mathematically equivalent; they are modular ways of enforcing an informative predictive representation\. Appendix[A\.3](https://arxiv.org/html/2608.04060#A1.SS3)describes representative choices\.

Later, our experiments use a VICReg\-style regulariser because its variance term directly prevents constant embeddings, while its covariance and invariance terms discourage redundant coordinates and encourage consistency under the selected observation augmentations\. Its exact form is specified with the experimental objectives in Section[8](https://arxiv.org/html/2608.04060#S8)\.

### 5\.2Unified practical objective

A practical scalar relaxation of the constrained objective in Equation \([10](https://arxiv.org/html/2608.04060#S4.E10)\) is

ℒSJEPAtrain=ℒpred\+λr​ℒroll\+λIB​ℛIB\+λs​Ωtrain​\(ℰ,α\)\+λc​ℛcorr​\(ϕ\),\\mathcal\{L\}\_\{\\textsc\{SJEPA\}\}^\{\\mathrm\{train\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{IB\}\}\\mathcal\{R\}\_\{\\mathrm\{IB\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\}\(\\mathcal\{E\},\\alpha\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\),\(29\)where all weights are nonnegative\. The five terms respectively control one\-step prediction, recursive prediction, representation quality, symbolic parsimony, and reliance on the neural correction\.

The optimisation\-compatible complexityΩtrain\\Omega^\{\\mathrm\{train\}\}may be the discrete structural scoreΩ​\(ℰ\)\\Omega\(\\mathcal\{E\}\)or a differentiable surrogate when symbolic coefficients and support are learned continuously\. The final discovered law is evaluated using the thresholded structural complexity defined in Equation \([4](https://arxiv.org/html/2608.04060#S3.E4)\)\. Exact forms are given in Appendix[C](https://arxiv.org/html/2608.04060#A3)\.

Different learning stages use restrictions of Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\): terms that are constant or inapplicable in a particular stage are omitted\. Because encoder learning and symbolic search are nonconvex, Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) should be understood as a controllable scalarisation of the constrained principle rather than a guaranteed equivalent formulation\. Bayesian fitting criteria and downstream planning objectives are separate from this deterministic training loss\.

### 5\.3End\-to\-end alternating learning

End\-to\-endSJEPAjointly searches for an informative representation and a compact transition law\. Direct joint optimisation is difficult because symbolic structure search is discrete or sparsity\-driven, whereas encoder, coefficient, and correction learning are continuous\. We therefore alternate between two complementary phases:

1. 1\.Dynamics search:hold the encoders fixed and find a compact symbolic law, together with a correction where required, in the current latent coordinates\.
2. 2\.Space search:hold the symbolic structure fixed or locally relaxed and update the encoder and continuous dynamics parameters so that the representation remains predictive and non\-collapsed while supporting a simpler transition\.

Algorithm 1Alternating training for end\-to\-endSJEPA0:Dataset

𝒟\\mathcal\{D\}, context encoder

EθE\_\{\\theta\}, target encoder

Eθ¯E\_\{\\bar\{\\theta\}\}, grammar

𝔈\\mathfrak\{E\}, number of cycles

CC
1:Warm\-start

\(θ,θ¯\)\(\\theta,\\bar\{\\theta\}\)using a neural JEPA predictor and the representation objective in Equation \([27](https://arxiv.org/html/2608.04060#S5.E27)\)

2:Initialise symbolic structure

ℰ\\mathcal\{E\}, coefficients

α\\alpha, and correction

cϕc\_\{\\phi\}
3:forouter cycle

c=1,…,Cc=1,\\ldots,Cdo

4:Encode latent tuples

\{\(ZC\(i\),ZT\(i\),ε\(i\)\)\}\\\{\(Z\_\{C\}^\{\(i\)\},Z\_\{T\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\)\\\}
5:Dynamics search:fix the encoders and minimise Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) over

\(ℰ,α,ϕ\)\(\\mathcal\{E\},\\alpha,\\phi\), omitting

ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}
6:Space search:fix or locally relax

ℰ\\mathcal\{E\}and minimise Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) over

\(θ,α,ϕ\)\(\\theta,\\alpha,\\phi\)
7:Update

θ¯\\bar\{\\theta\}by exponential moving average of

θ\\theta
8:endfor

9:return

\(Eθ,Eθ¯,Fℰ,α,cϕ\)\(E\_\{\\theta\},E\_\{\\bar\{\\theta\}\},F\_\{\\mathcal\{E\},\\alpha\},c\_\{\\phi\}\)

Alternation implements the learned\-transform interpretation of Section[4](https://arxiv.org/html/2608.04060#S4): improved coordinates permit simpler equations, while the symbolic objective provides pressure towards coordinates with lower induced\-dynamics complexity\. The representation regulariser prevents this pressure from being satisfied through collapse\. Detailed phase\-specific objectives, differentiable symbolic relaxations, and alternative optimisation schemes are given in Appendix[C](https://arxiv.org/html/2608.04060#A3)\.

### 5\.4Frozen pretrained encoders

Representation learning isoptional\. Given a frozen pretrained encoderEpreE\_\{\\mathrm\{pre\}\}, such as a ViT, DINO, MAE, I\-JEPA, or V\-JEPA encoder\(Dosovitskiyet al\.,[2021](https://arxiv.org/html/2608.04060#bib.bib22); Caronet al\.,[2021](https://arxiv.org/html/2608.04060#bib.bib23); Heet al\.,[2022](https://arxiv.org/html/2608.04060#bib.bib24); Assranet al\.,[2023](https://arxiv.org/html/2608.04060#bib.bib2); Bardeset al\.,[2024](https://arxiv.org/html/2608.04060#bib.bib3)\), define

ZC=Epre​\(XC\),ZT=Epre​\(XT\),Z\_\{C\}=E\_\{\\mathrm\{pre\}\}\(X\_\{C\}\),\\qquad Z\_\{T\}=E\_\{\\mathrm\{pre\}\}\(X\_\{T\}\),and optimise only\(ℰ,α,ϕ\)\(\\mathcal\{E\},\\alpha,\\phi\)\. Because the representation is fixed,ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}and the EMA update are omitted, while the symbolic and correction modules remain unchanged\. This mode turns𝒞dyn​\(Epre,Epre\)\\mathcal\{C\}\_\{\\mathrm\{dyn\}\}\(E\_\{\\mathrm\{pre\}\},E\_\{\\mathrm\{pre\}\}\)into a diagnostic of whether the pretrained representation already exposes compact dynamics\. It is computationally simpler than end\-to\-end learning, but it cannot reshape coordinates that hide an otherwise simple transition\.

Table 2:The two training modes share the same symbolic–correction transition model\.

## 6BayesianSJEPA: Uncertainty over Latent Dynamics

We further introduce uncertainty into dynamics learning through a Bayesian formulation ofSJEPA\. The encoders remain deterministic and may be trained jointly beforehand or kept frozen, while Bayesian inference is applied only to the latent transition model\. In particular, uncertainty is represented over the symbolic structure, its coefficients, the transition covariance, and an optional Gaussian\-process correction\.

### 6\.1Bayesian symbolic regression in latent space

The deterministic SJEPA formulation selects a single symbolic transition law\. With finite data, however, several expressions may explain the observed latent transitions nearly equally well, leaving uncertainty about the symbolic structure, its coefficients, and residual transition variability\. BayesianSJEPArepresents this uncertainty explicitly rather than committing immediately to one equation\. Throughout this section, the encoders are treated as fixed, so uncertainty is introducedonlyover the latent dynamics\.

Given latent transition data

𝒟Z=\{\(zC\(i\),zT\(i\),ε\(i\)\)\}i=1n,\\mathcal\{D\}\_\{Z\}=\\left\\\{\\left\(z\_\{C\}^\{\(i\)\},z\_\{T\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\\right\\\}\_\{i=1\}^\{n\},we model each target embedding as

zT\(i\)∣zC\(i\),ε\(i\),ℰ,α,Σ∼𝒩​\(Fℰ,α​\(zC\(i\),ε\(i\)\),Σ\)\.z\_\{T\}^\{\(i\)\}\\mid z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\},\\mathcal\{E\},\\alpha,\\Sigma\\sim\\mathcal\{N\}\\left\(F\_\{\\mathcal\{E\},\\alpha\}\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\),\\Sigma\\right\)\.\(30\)Here,ℰ\\mathcal\{E\}specifies the symbolic structure,α\\alphacontains its numerical coefficients, andΣ\\Sigmarepresents latent transition variability not explained by the symbolic mean\.

To favour parsimonious laws, we place the complexity prior

p​\(ℰ\)=exp⁡\{−γ​Ω​\(ℰ\)\}Zγ,Zγ=∑ℰ∈𝔈exp⁡\{−γ​Ω​\(ℰ\)\},p\(\\mathcal\{E\}\)=\\frac\{\\exp\\left\\\{\-\\gamma\\Omega\(\\mathcal\{E\}\)\\right\\\}\}\{Z\_\{\\gamma\}\},\\qquad Z\_\{\\gamma\}=\\sum\_\{\\mathcal\{E\}\\in\\mathfrak\{E\}\}\\exp\\left\\\{\-\\gamma\\Omega\(\\mathcal\{E\}\)\\right\\\},\(31\)whereγ≥0\\gamma\\geq 0controls the preference for simpler expressions\. We assume that the candidate structure space is finite or countable and thatZγ<∞Z\_\{\\gamma\}<\\infty\. Together with priorsp​\(α∣ℰ\)p\(\\alpha\\mid\\mathcal\{E\}\)andp​\(Σ\)p\(\\Sigma\), Bayes’ rule gives

p​\(ℰ,α,Σ∣𝒟Z\)∝p​\(𝒟Z∣ℰ,α,Σ\)​p​\(α∣ℰ\)​p​\(Σ\)​p​\(ℰ\)\.p\(\\mathcal\{E\},\\alpha,\\Sigma\\mid\\mathcal\{D\}\_\{Z\}\)\\propto p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,\\Sigma\)p\(\\alpha\\mid\\mathcal\{E\}\)p\(\\Sigma\)p\(\\mathcal\{E\}\)\.\(32\)The posterior balances predictive fit, symbolic complexity, and prior plausibility of the coefficients and residual covariance\.

For a new context embeddingzCz\_\{C\}and side informationε\\varepsilon, the posterior predictive distribution averages over plausible symbolic structures and parameter values:

p​\(zT∣zC,ε,𝒟Z\)=∑ℰ∫p​\(zT∣zC,ε,ℰ,α,Σ\)​p​\(ℰ,α,Σ∣𝒟Z\)​dα​dΣ\.p\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{D\}\_\{Z\}\)=\\sum\_\{\\mathcal\{E\}\}\\int p\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{E\},\\alpha,\\Sigma\)p\(\\mathcal\{E\},\\alpha,\\Sigma\\mid\\mathcal\{D\}\_\{Z\}\)\\mathrm\{d\}\\alpha\\,\\mathrm\{d\}\\Sigma\.\(33\)This distribution propagates uncertainty about the symbolic law, its coefficients, and residual transition variability into the predicted target embedding\. BayesianSJEPAcan therefore retain several competing explanations when the available data do not identify one law decisively\.

###### Theorem 6\.1\(MAP equivalence for penalised symbolic regression\)\.

Assume that the latent transitions are conditionally independent under Equation \([30](https://arxiv.org/html/2608.04060#S6.E30)\), with fixed isotropic covarianceΣ=σ2​I\\Sigma=\\sigma^\{2\}Iforσ2\>0\\sigma^\{2\}\>0\. Let

p​\(ℰ,α\)=p​\(α∣ℰ\)​p​\(ℰ\),p\(\\mathcal\{E\},\\alpha\)=p\(\\alpha\\mid\\mathcal\{E\}\)p\(\\mathcal\{E\}\),wherep​\(ℰ\)p\(\\mathcal\{E\}\)is the proper complexity prior in Equation \([31](https://arxiv.org/html/2608.04060#S6.E31)\)\. Then any maximum\-a\-posteriori estimator of\(ℰ,α\)\(\\mathcal\{E\},\\alpha\)minimises

12​σ2​∑i=1n‖zT\(i\)−Fℰ,α​\(zC\(i\),ε\(i\)\)‖22\+γ​Ω​\(ℰ\)−log⁡p​\(α∣ℰ\),\\frac\{1\}\{2\\sigma^\{2\}\}\\sum\_\{i=1\}^\{n\}\\left\\\|z\_\{T\}^\{\(i\)\}\-F\_\{\\mathcal\{E\},\\alpha\}\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\\right\\\|\_\{2\}^\{2\}\+\\gamma\\Omega\(\\mathcal\{E\}\)\-\\log p\(\\alpha\\mid\\mathcal\{E\}\),\(34\)and any minimiser of Equation \([34](https://arxiv.org/html/2608.04060#S6.E34)\) is a MAP estimator, provided the extrema exist\.

The proof is given in Appendix[B\.3\.2](https://arxiv.org/html/2608.04060#A2.SS3.SSS2)\. Intuitively, the theorem shows that deterministic symbolic penalties can be interpreted as negative log\-priors: penalising complex expressions corresponds to assigning them lower prior probability\. The symbolic\-complexity penalty is induced by the structure prior, while coefficient regularisation is determined byp​\(α∣ℰ\)p\(\\alpha\\mid\\mathcal\{E\}\)\. MAP estimation returns one most probable symbolic law and coefficient vector\. By contrast, Equation \([33](https://arxiv.org/html/2608.04060#S6.E33)\) averages over competing structures and parameter values, retaining uncertainty when the data do not identify one law decisively\.

### 6\.2Bayesian hybrid predictor with a Gaussian\-process correction

To represent systematic residual dynamics not captured by the symbolic grammar, we augment the symbolic mean with a Gaussian\-process correction888Equation \([35](https://arxiv.org/html/2608.04060#S6.E35)\) is the direct\-transition formulation\. For a continuous\-time vector\-field model observed at step sizeΔ​tt\\Delta t\_\{t\}, the corresponding one\-step likelihood may instead usezt\+1∣zt,Δ​tt,ℰ,α,g,ΣΔ​tt∼𝒩​\(zt\+Δ​tt​\[Fℰ,α​\(zt\)\+g​\(zt\)\],ΣΔ​tt\)\.z\_\{t\+1\}\\mid z\_\{t\},\\Delta t\_\{t\},\\mathcal\{E\},\\alpha,g,\\Sigma\_\{\\Delta t\_\{t\}\}\\sim\\mathcal\{N\}\\left\(z\_\{t\}\+\\Delta t\_\{t\}\\left\[F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{t\}\)\+g\(z\_\{t\}\)\\right\],\\Sigma\_\{\\Delta t\_\{t\}\}\\right\)\.Our later experiments use the deterministic counterpart of this residual formulation, whereas the task\-general Bayesian development is written using the direct\-transition notation\.:

zT=Fℰ,α​\(zC,ε\)\+g​\(zC,ε\)\+η,g∼𝒢​𝒫​\(0,k\),η∼𝒩​\(0,Σ\)\.z\_\{T\}=F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{C\},\\varepsilon\)\+g\(z\_\{C\},\\varepsilon\)\+\\eta,\\qquad g\\sim\\mathcal\{GP\}\(0,k\),\\qquad\\eta\\sim\\mathcal\{N\}\(0,\\Sigma\)\.\(35\)The Gaussian process represents input\-dependent residual structure, whereasη\\etarepresents transition variability remaining after conditioning on the symbolic law and correction\. As in the deterministic hybrid model, the symbolic–GP decomposition is not identifiable from predictive fit alone\. Its allocation depends on the symbolic\-structure prior, coefficient priors, GP kernel family, kernel\-amplitude prior, and transition\-noise prior\. These priors provide the Bayesian analogue of deterministic correction control: they encode a preference for a compact symbolic explanation with a modest residual correction, but they do not guarantee a unique decomposition\.

For fixed\(ℰ,α\)\(\\mathcal\{E\},\\alpha\), define the GP input

si=\(zC\(i\),ε\(i\)\)\.s\_\{i\}=\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\., and the symbolic residualri=zT\(i\)−Fℰ,α​\(zC\(i\),ε\(i\)\)\.r\_\{i\}=z\_\{T\}^\{\(i\)\}\-F\_\{\\mathcal\{E\},\\alpha\}\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\.Stacking the residuals gives

𝐫=vec⁡\(\[r1,…,rn\]⊤\)∈ℝn​dz,\\mathbf\{r\}=\\operatorname\{vec\}\\left\(\[r\_\{1\},\\ldots,r\_\{n\}\]^\{\\top\}\\right\)\\in\\mathbb\{R\}^\{nd\_\{z\}\},\(36\)wheredzd\_\{z\}is the latent dimension\. Let𝐊g\\mathbf\{K\}\_\{g\}denote the covariance matrix induced by the vector\-valued GP over the stacked inputs and outputs\. After integrating outgg, the residual vector has the marginal distribution

𝐫∣ℰ,α∼𝒩​\(0,𝐂\),𝐂=𝐊g\+In⊗Σ\.\\mathbf\{r\}\\mid\\mathcal\{E\},\\alpha\\sim\\mathcal\{N\}\\left\(0,\\mathbf\{C\}\\right\),\\qquad\\mathbf\{C\}=\\mathbf\{K\}\_\{g\}\+I\_\{n\}\\otimes\\Sigma\.\(37\)Letϑg\\vartheta\_\{g\}denote the GP kernel hyperparameters\. The corresponding negative log marginal likelihood is\(Rasmussen and Williams,[2005](https://arxiv.org/html/2608.04060#bib.bib36)\)

−log⁡p​\(𝒟Z∣ℰ,α,ϑg,Σ\)=12​𝐫⊤​𝐂−1​𝐫⏟residual data fit\+12​log⁡\|𝐂\|⏟complexity penalty\+n​dz2​log⁡\(2​π\)\.\-\\log p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,\\vartheta\_\{g\},\\Sigma\)=\\underbrace\{\\frac\{1\}\{2\}\\mathbf\{r\}^\{\\top\}\\mathbf\{C\}^\{\-1\}\\mathbf\{r\}\}\_\{\\text\{residual data fit\}\}\+\\underbrace\{\\frac\{1\}\{2\}\\log\|\\mathbf\{C\}\|\}\_\{\\text\{complexity penalty\}\}\+\\frac\{nd\_\{z\}\}\{2\}\\log\(2\\pi\)\.\(38\)The quadratic term measures the symbolic residual after accounting for the covariance structure permitted by the GP and transition noise\. The log\-determinant term is an Occam factor: it penalises covariance structures that can explain a broad range of residual functions, unless their additional flexibility is supported by improved fit\. The final term is constant with respect to the model parameters\.

When independent GPs are used for the latent coordinates, Equation \([38](https://arxiv.org/html/2608.04060#S6.E38)\) decomposes across coordinates\. Let𝐫j∈ℝn\\mathbf\{r\}\_\{j\}\\in\\mathbb\{R\}^\{n\}be the residuals for coordinatejjand let

𝐂j=𝐊j\+σj2​In\.\\mathbf\{C\}\_\{j\}=\\mathbf\{K\}\_\{j\}\+\\sigma\_\{j\}^\{2\}I\_\{n\}\.\([37](https://arxiv.org/html/2608.04060#S6.E37)b\)Then

−log⁡p​\(𝒟Z∣ℰ,α,ϑg,Σ\)=∑j=1dz\[12​𝐫j⊤​𝐂j−1​𝐫j\+12​log⁡\|𝐂j\|\+n2​log⁡\(2​π\)\]\.\-\\log p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,\\vartheta\_\{g\},\\Sigma\)=\\sum\_\{j=1\}^\{d\_\{z\}\}\\left\[\\frac\{1\}\{2\}\\mathbf\{r\}\_\{j\}^\{\\top\}\\mathbf\{C\}\_\{j\}^\{\-1\}\\mathbf\{r\}\_\{j\}\+\\frac\{1\}\{2\}\\log\|\\mathbf\{C\}\_\{j\}\|\+\\frac\{n\}\{2\}\\log\(2\\pi\)\\right\]\.\([38](https://arxiv.org/html/2608.04060#S6.E38)b\)The Bayesian hybrid predictor provides a probabilistic analogue of the symbolic–correction decomposition in deterministicSJEPA\. The resulting allocation remains prior\-dependent, while the marginal likelihood balances residual fit against the covariance flexibility permitted by the selected GP kernel and quantifies uncertainty about unexplained dynamics\.

## 7Explicit Action Dependence and Planning

In control tasks, the side informationε\\varepsilonmay contain an actionutu\_\{t\}\. For a discrete\-time direct\-transition model, the action\-conditioned predictor is999For a continuous\-time vector\-field realisation, the planner instead uses the integrated transitionz^t\+1=Hℰ,α,ϕ,Δ​tt​\(zt,ut\)=ℐΔ​tt​\(zt,Fℰ,α\+cϕ,ut\)\.\\widehat\{z\}\_\{t\+1\}=H\_\{\\mathcal\{E\},\\alpha,\\phi,\\Delta t\_\{t\}\}\(z\_\{t\},u\_\{t\}\)=\\mathcal\{I\}\_\{\\Delta t\_\{t\}\}\\left\(z\_\{t\},F\_\{\\mathcal\{E\},\\alpha\}\+c\_\{\\phi\},u\_\{t\}\\right\)\.

z^t\+1=Hℰ,α,ϕ​\(zt,ut\)=Fℰ,α​\(zt,ut\)\+cϕ​\(zt,ut\)\.\\widehat\{z\}\_\{t\+1\}=H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z\_\{t\},u\_\{t\}\)=F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{t\},u\_\{t\}\)\+c\_\{\\phi\}\(z\_\{t\},u\_\{t\}\)\.\(39\)The symbolic component makes the role of the action explicit\. For example, the discovered law may reveal whether an action affects a particular latent coordinate linearly, through an interaction with the current state, or only within a particular operating regime\. This provides information that is difficult to extract directly from an opaque neural predictor\.

WhenHℰ,α,ϕH\_\{\\mathcal\{E\},\\alpha,\\phi\}is differentiable, its behaviour around the current \(latent\) state–action pair\(zt,ut\)\(z\_\{t\},u\_\{t\}\)can be approximated by a first\-order local model\. Define101010For the forward\-Euler realisationHℰ,α,ϕ,Δ​t​\(z\)=z\+Δ​t​\[Fℰ,α​\(z\)\+cϕ​\(z\)\]H\_\{\\mathcal\{E\},\\alpha,\\phi,\\Delta t\}\(z\)=z\+\\Delta t\\left\[F\_\{\\mathcal\{E\},\\alpha\}\(z\)\+c\_\{\\phi\}\(z\)\\right\], the state Jacobian of the complete transition isAt=I\+Δ​tt​\[∂Fℰ,α∂z\+∂cϕ∂z\]\(zt,ut\),A\_\{t\}=I\+\\Delta t\_\{t\}\\left\[\\frac\{\\partial F\_\{\\mathcal\{E\},\\alpha\}\}\{\\partial z\}\+\\frac\{\\partial c\_\{\\phi\}\}\{\\partial z\}\\right\]\_\{\(z\_\{t\},u\_\{t\}\)\},whereas the corresponding continuous\-time local analysis uses the Jacobian of the vector field itself\. These two stability criteria should not be conflated\.

At=∂Hℰ,α,ϕ​\(z,u\)∂z\|\(zt,ut\),Bt=∂Hℰ,α,ϕ​\(z,u\)∂u\|\(zt,ut\)\.A\_\{t\}=\\left\.\\frac\{\\partial H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,u\)\}\{\\partial z\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\},\\qquad B\_\{t\}=\\left\.\\frac\{\\partial H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,u\)\}\{\\partial u\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\}\.\(40\)For small perturbationsΔ​zt\\Delta z\_\{t\}andΔ​ut\\Delta u\_\{t\}, these matrices give

Δ​zt\+1≈At​Δ​zt\+Bt​Δ​ut\.\\Delta z\_\{t\+1\}\\approx A\_\{t\}\\Delta z\_\{t\}\+B\_\{t\}\\Delta u\_\{t\}\.\(41\)Intuitively,AtA\_\{t\}describes how a small change in the current latent state propagates to the next state, whileBtB\_\{t\}describes how a small change in the action affects the next state\. Thus,AtA\_\{t\}captures local state sensitivity andBtB\_\{t\}captures local action sensitivity\.

These matrices provide a direct interface to classical local analysis\. Around an equilibrium\(z⋆,u⋆\)\(z^\{\\star\},u^\{\\star\}\)satisfying

z⋆=Hℰ,α,ϕ​\(z⋆,u⋆\),z^\{\\star\}=H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z^\{\\star\},u^\{\\star\}\),\(42\)the eigenvalues ofA⋆A^\{\\star\}can be used to assess local stability, while the pair\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)can be examined to determine which latent directions are locally influenced by the available actions\. The same local model may also be used by linear or locally linear control procedures, or as an approximation inside iterative trajectory\-optimization and model\-predictive\-control methods\. These linearised analyses do not guarantee global stability or controllability; they describe the model only in a neighbourhood of the selected operating point\.

For polynomial or elementary symbolic expressions, the derivatives ofFℰ,αF\_\{\\mathcal\{E\},\\alpha\}can often be obtained analytically\. Automatic differentiation can be used for the neural correctioncϕc\_\{\\phi\}, so that

At=∂Fℰ,α∂z\|\(zt,ut\)\+∂cϕ∂z\|\(zt,ut\),Bt=∂Fℰ,α∂u\|\(zt,ut\)\+∂cϕ∂u\|\(zt,ut\)\.A\_\{t\}=\\left\.\\frac\{\\partial F\_\{\\mathcal\{E\},\\alpha\}\}\{\\partial z\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\}\+\\left\.\\frac\{\\partial c\_\{\\phi\}\}\{\\partial z\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\},\\qquad B\_\{t\}=\\left\.\\frac\{\\partial F\_\{\\mathcal\{E\},\\alpha\}\}\{\\partial u\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\}\+\\left\.\\frac\{\\partial c\_\{\\phi\}\}\{\\partial u\}\\right\|\_\{\(z\_\{t\},u\_\{t\}\)\}\.\(43\)This decomposition further shows whether the dominant action sensitivity is explained by the symbolic law or delegated to the neural correction\.

Beyond local analysis, a generic external planner can use the full nonlinear hybrid predictor to simulate candidate action sequences\. LetKKdenote the planning horizon\. Starting from the current latent statez^t=zt\\widehat\{z\}\_\{t\}=z\_\{t\}, the planner recursively computes

z^t\+h\+1=Hℰ,α,ϕ​\(z^t\+h,ut\+h\),h=0,…,K−1\.\\widehat\{z\}\_\{t\+h\+1\}=H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(\\widehat\{z\}\_\{t\+h\},u\_\{t\+h\}\),\\qquad h=0,\\ldots,K\-1\.\(44\)It then selects an action sequence by solving

ut:t\+K−1⋆=arg​minut:t\+K−1⁡\[∑h=0K−1ℓ​\(z^t\+h,ut\+h\)\+ℓT​\(z^t\+K\)\],u^\{\\star\}\_\{t:t\+K\-1\}=\\operatorname\*\{arg\\,min\}\_\{u\_\{t:t\+K\-1\}\}\\left\[\\sum\_\{h=0\}^\{K\-1\}\\ell\(\\widehat\{z\}\_\{t\+h\},u\_\{t\+h\}\)\+\\ell\_\{T\}\(\\widehat\{z\}\_\{t\+K\}\)\\right\],\(45\)whereℓ\\ellis a stage cost andℓT\\ell\_\{T\}is a terminal cost\.

The planner itself is not the methodological contribution ofSJEPA\. The contribution is an action\-conditioned predictive model whose governing structure can be inspected, differentiated, locally linearised, and used by a range of external planning or control methods\. In BayesianSJEPA, the deterministic rollout can be replaced by a posterior predictive rollout, and the planner may minimise posterior expected cost or a risk\-sensitive objective that accounts for uncertainty\.

Compact action\-conditioned laws may also improve sample efficiency and extrapolation by reusing the same structural mechanism across states and action sequences\. This remains an empirical hypothesis rather than an unconditional guarantee: a misspecified symbolic grammar, an inaccurate latent representation, or a dominant neural correction can reduce both structural clarity and planning performance\.

## 8Experiments

We conduct two controlled experiments targeting the two claims that defineSJEPA\. Experiment 1 askswhether joint representation and operator learning discovers predictive coordinates with simpler induced dynamics than post\-hoc equation discovery\. Experiment 2 removes representation learning and testswhether correction regularisation preserves a compact symbolic mechanism under controlled grammar misspecification\. Detailed data generation, architectures, optimisation settings, checkpoint selection, per\-seed equations, and additional metrics are given in Appendix[D](https://arxiv.org/html/2608.04060#A4)\. Unless stated otherwise, reported values are means and standard deviations over three optimisation seeds\.

##### Residual dynamics and evaluation protocol\.

Both experiments instantiate Equation \([3](https://arxiv.org/html/2608.04060#S3.E3)\) using residual vector\-field predictors,

z^t\+1=zt\+Δ​tt​\[Fℰ,α​\(zt\)\+cϕ​\(zt\)\],\\widehat\{z\}\_\{t\+1\}=z\_\{t\}\+\\Delta t\_\{t\}\\left\[F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{t\}\)\+c\_\{\\phi\}\(z\_\{t\}\)\\right\],\(46\)withcϕ≡0c\_\{\\phi\}\\equiv 0for symbolic\-only models\. This parameterisation treatsFℰ,α​\(zt\)\+cϕ​\(zt\)F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{t\}\)\+c\_\{\\phi\}\(z\_\{t\}\)as an estimate of the continuous\-time latent derivativez˙t\\dot\{z\}\_\{t\}and obtains the next state through a forward\-Euler step, while explicitly accounting for variable step sizesΔ​tt\\Delta t\_\{t\}\. Training combines one\-step prediction with a ten\-step free rollout\. Training and model selection use disjoint trajectory splits\. Model parameters and symbolic coefficients are learned from the training trajectories, while separate validation trajectories are used to select checkpoints and, in Experiment 1, the best completed alternating cycle\. When validation risks are nearly tied, the lower\-complexity thresholded symbolic model is preferred\. Test and OOD trajectories are used only for final evaluation after all model selection is complete\. Experiment 1 evaluates prediction directly in the latent space and evaluates physical\-state rollouts after fitting an affine map from training latents to the physical state111111The affine map converts learned latent coordinates into physical\-state coordinates so that rollout errors can be measured meaningfully\.; Experiment 2 directly uses the physical state as the latent coordinate,zt=\(qt,pt\)z\_\{t\}=\(q\_\{t\},p\_\{t\}\), so predicted rollouts can be evaluated in physical\-state space without fitting an additional affine alignment map\.

##### Experimental objectives and model selection\.

Each experimental condition uses the applicable terms from the unified objective in Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\); Table[3](https://arxiv.org/html/2608.04060#S8.T3)summarises the active objective terms in each condition\. Experiment 1 jointly studies representation and operator learning, whereas Experiment 2 fixes the state coordinates and isolates the allocation between the symbolic law and neural correction\.

In Experiment 1, the representation regulariser is

ℛIBexp=λinv​ℒinv\+λvar​ℒvar\+λcov​ℒcov\+λmean​ℒmean\.\\mathcal\{R\}\_\{\\mathrm\{IB\}\}^\{\\mathrm\{exp\}\}=\\lambda\_\{\\mathrm\{inv\}\}\\mathcal\{L\}\_\{\\mathrm\{inv\}\}\+\\lambda\_\{\\mathrm\{var\}\}\\mathcal\{L\}\_\{\\mathrm\{var\}\}\+\\lambda\_\{\\mathrm\{cov\}\}\\mathcal\{L\}\_\{\\mathrm\{cov\}\}\+\\lambda\_\{\\mathrm\{mean\}\}\\mathcal\{L\}\_\{\\mathrm\{mean\}\}\.\(47\)Here,ℒinv\\mathcal\{L\}\_\{\\mathrm\{inv\}\}encourages consistency under observation augmentation,ℒvar\\mathcal\{L\}\_\{\\mathrm\{var\}\}imposes a per\-coordinate variance floor,ℒcov\\mathcal\{L\}\_\{\\mathrm\{cov\}\}penalises off\-diagonal latent covariance, andℒmean\\mathcal\{L\}\_\{\\mathrm\{mean\}\}provides weak centering\. The variance term provides the principal protection against the constant\-representation shortcut in Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5)\. The weights in Equation \([47](https://arxiv.org/html/2608.04060#S8.E47)\) control the relative contributions withinℛIBexp\\mathcal\{R\}\_\{\\mathrm\{IB\}\}^\{\\mathrm\{exp\}\}, whileλIB\\lambda\_\{\\mathrm\{IB\}\}in Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) controls its overall strength\.

Table 3:Active terms from Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) in each experimental condition\. A dash indicates that the corresponding term is absent\. The unregularised hybrid retains a neural correction but setsλc=0\\lambda\_\{\\mathrm\{c\}\}=0, so the correction penalty is inactive\.Model parameters are learned from the training trajectories\. Checkpoints and, forSJEPA, completed alternating cycles are selected using the held\-out validation risk

ℛval=ℒpredval\+λr​ℒrollval,\\mathcal\{R\}\_\{\\mathrm\{val\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}^\{\\mathrm\{val\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}^\{\\mathrm\{val\}\},\(48\)whereλr\\lambda\_\{\\mathrm\{r\}\}is the rollout weight used by the corresponding condition\. For Experiment 1 conditions that update the encoder, candidates must satisfy the prescribed non\-collapse diagnostics; the no\-ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}and fixed\-point diagnostic conditions are exempt from this requirement\. Among candidates whose validation risks lie within a predefined relative tolerance, the lower\-complexity symbolic law is selected\. Full loss definitions, weights, diagnostics, and selection tolerances are given in Appendix[D](https://arxiv.org/html/2608.04060#A4)\.

### 8\.1Experiment 1: Learning coordinates with elegant dynamics

##### Design\.

We simulate the continuous\-time pendulum

q˙=p,p˙=−sin⁡\(q\),\\dot\{q\}=p,\\qquad\\dot\{p\}=\-\\sin\(q\),\(49\)using fourth\-order Runge–Kutta integration\. At each transition, the time step is sampled independently from

Δ​tt∈\{0\.025,0\.04,0\.06\}\.\\Delta t\_\{t\}\\in\\\{0\.025,0\.04,0\.06\\\}\.
The two\-dimensional physical state is embedded in a noisy3232\-dimensional observation by combining raw and nonlinear state features and applying fixed orthogonal mixing\. Thus,qqandppare not supplied as named input coordinates, although they remain linearly recoverable from the noiseless full observation because the raw state features are included\. The experiment therefore tests predictive\-coordinate selection under high\-dimensional mixing and nonlinear distractor features rather than recovery from a genuinely nonlinear observation inverse\. The context encoder and its EMA target copy map these observations to a common two\-dimensional latent space\.

We compare:

1. 1\.Neural JEPA, which jointly learns the encoder and a neural residual vector field;
2. 2\.post\-hoc symbolic, which freezes the Neural JEPA coordinates and subsequently fits a symbolic vector field, as detailed in Appendix[D\.4](https://arxiv.org/html/2608.04060#A4.SS4);
3. 3\.SJEPAsymbolic, which alternates symbolic\-dynamics search and representation\-space search withcϕ≡0c\_\{\\phi\}\\equiv 0;
4. 4\.anunconstrained one\-step collapse diagnostic, which removes representation regularisation and rollout supervision and trains the symbolic model using only the one\-step objective;
5. 5\.acollapsed fixed\-point control, in which both encoders are initialised as the same nonzero constant map\.

The unconstrained one\-step diagnostic sets both the representation\-regularisation weight and rollout\-loss weight to zero in order to expose the one\-step fixed\-point shortcut directly\. It therefore demonstrates that the unconstrained one\-step objective can reach the collapse mechanism identified in Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5); it is not a matched ablation isolating the effect of removingℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}alone\. The collapsed fixed\-point control uses the same one\-step objective and constructs the degenerate solution explicitly\. Because a constant representation is preserved under repeated identity transitions, the existence of the shortcut itself is not specific to one\-step training\.

The symbolic library is

Θ​\(z\)=\{1,z1,z2,sin⁡\(z1\),sin⁡\(z2\),cos⁡\(z1\),cos⁡\(z2\),z12,z22,z1​z2\}\.\\Theta\(z\)=\\left\\\{1,z\_\{1\},z\_\{2\},\\sin\(z\_\{1\}\),\\sin\(z\_\{2\}\),\\cos\(z\_\{1\}\),\\cos\(z\_\{2\}\),z\_\{1\}^\{2\},z\_\{2\}^\{2\},z\_\{1\}z\_\{2\}\\right\\\}\.\(50\)The reported symbolic complexity is the weighted complexity of the terms whose fitted coefficients exceed the fixed reporting threshold specified in Appendix[D](https://arxiv.org/html/2608.04060#A4)\.

##### Joint coordinate search exposes a substantially simpler law\.

Table[4](https://arxiv.org/html/2608.04060#S8.T4)compares the 3 predictive models\. Neural JEPA achieves the lowest raw prediction and rollout errors, as expected from its flexible neural vector field\. The primary operator\-compression comparison is therefore between post\-hoc symbolic regression and joint symbolicSJEPA\.

Joint coordinate search reduces mean symbolic complexity from26\.026\.0to4\.674\.67, a factor of approximately5\.65\.6\. Within the respective learned latent spaces, it also reduces one\-step latent MSE from2\.09×10−32\.09\\times 10^\{\-3\}to5\.81×10−45\.81\\times 10^\{\-4\}, a factor of approximately3\.63\.6\. Because latent MSE depends on the scale and coordinate system of the learned representation, this comparison is descriptive rather than coordinate\-invariant\. The physical\-state rollout results provide a complementary cross\-model comparison: joint symbolicSJEPAreduces test rollout MSE from1\.2291\.229to0\.4670\.467, OOD rollout MSE from9\.2629\.262to3\.4353\.435, and OOD divergence rate from0\.1780\.178to0\.0170\.017\. Figure[2](https://arxiv.org/html/2608.04060#S8.F2)\(a\) shows the corresponding reduction in recursively accumulated error\.

Table 4:Experiment 1 results\. Latent MSE is evaluated within each model’s learned embedding space and is therefore coordinate\-dependent\. Rollout columns report clipped physical\-state MSE after applying an affine map fitted only on training latents; divergence thresholds are defined in Appendix[D](https://arxiv.org/html/2608.04060#A4)\. A dash denotes a metric that is not applicable, and bold identifies the better of the two symbolic methods\. Values are mean±\\pmstandard deviation over 3 optimisation seeds\.The dominantSJEPAstructure is consistent across seeds: every discovered law containsz2z\_\{2\}in the equation forz˙1\\dot\{z\}\_\{1\}andz1z\_\{1\}in the equation forz˙2\\dot\{z\}\_\{2\}\. Up to a reversal of latent orientation, the learned dynamics are therefore oscillator\-like:

z˙1≈−a​z2,z˙2≈b​z1\+lower\-magnitude terms,a,b\>0\.\\dot\{z\}\_\{1\}\\approx\-az\_\{2\},\\qquad\\dot\{z\}\_\{2\}\\approx bz\_\{1\}\+\\text\{lower\-magnitude terms\},\\qquad a,b\>0\.\(51\)
To illustrate the difference between the three predictors, consider the selected models for the same seed\. Neural JEPA represents the vector field by an unrestricted neural network,

z˙=fψ​\(z\),\\dot\{z\}=f\_\{\\psi\}\(z\),\(52\)and therefore does not produce a closed\-form symbolic equation\. Symbolic regression fitted post hoc to the Neural JEPA coordinates gives

z˙1\\displaystyle\\dot\{z\}\_\{1\}=−1\.189​sin⁡\(z2\)−0\.602\+0\.414​cos⁡\(z2\)\+0\.244​sin⁡\(z1\)\+0\.238​z22\\displaystyle=\-1\.189\\sin\(z\_\{2\}\)\-0\.602\+0\.414\\cos\(z\_\{2\}\)\+0\.244\\sin\(z\_\{1\}\)\+0\.238z\_\{2\}^\{2\}−0\.188​z1\+0\.074​cos⁡\(z1\)\+0\.069​z12,\\displaystyle\\hskip 34\.14322pt\-0\.188z\_\{1\}\+0\.074\\cos\(z\_\{1\}\)\+0\.069z\_\{1\}^\{2\},z˙2\\displaystyle\\dot\{z\}\_\{2\}=1\.888​sin⁡\(z1\)\+1\.133​cos⁡\(z1\)−0\.967−0\.386​z1\+0\.320​z12−0\.245​z1​z2\.\\displaystyle=1\.888\\sin\(z\_\{1\}\)\+1\.133\\cos\(z\_\{1\}\)\-0\.967\-0\.386z\_\{1\}\+0\.320z\_\{1\}^\{2\}\-0\.245z\_\{1\}z\_\{2\}\.By contrast, joint representation and equation learning produces

z˙1=−0\.806​z2,z˙2=0\.822​z1\+0\.177−0\.143​z12\.\\dot\{z\}\_\{1\}=\-0\.806z\_\{2\},\\qquad\\dot\{z\}\_\{2\}=0\.822z\_\{1\}\+0\.177\-0\.143z\_\{1\}^\{2\}\.The complete per\-seedSJEPAequations are reported in Appendix[D\.5\.1](https://arxiv.org/html/2608.04060#A4.SS5.SSS1)\.

The post\-hoc and Neural JEPA models share the same prediction\-only coordinates, whereasSJEPAlearns a different latent coordinate system\. Their coefficients should therefore not be compared term by term\. The relevant comparison is whether each representation admits an accurate and compact induced transition\. In this example, the prediction\-only coordinates require a mixture of polynomial and trigonometric terms, while joint representation and equation learning exposes a much more concise cross\-coupled system\.

Together with Table[4](https://arxiv.org/html/2608.04060#S8.T4), these results support our operator\-compression claim:SJEPAdiscovers substantially simpler symbolic dynamics and achieves lower physical\-state rollout error than symbolic regression fitted post hoc\. They do not imply that the symbolic model universally exceeds Neural JEPA in predictive accuracy\. Rather,SJEPAtrades some predictive flexibility for a considerably more concise and inspectable transition\.

##### Representation constraints prevent trivial operator compression\.

The unconstrained one\-step diagnostic121212As this diagnostic removes bothℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}and rollout supervision, it demonstrates that the unconstrained one\-step objective can reach the theoretical fixed\-point shortcut; it is not a matched ablation isolating the effect of removingℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}alone\.exhibits the degenerate behaviour predicted by Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5)\. In all 3 seeds, the learned vector field becomes

z˙1=0,z˙2=0\.\\dot\{z\}\_\{1\}=0,\\qquad\\dot\{z\}\_\{2\}=0\.Under the residual parameterisation in Equation \([46](https://arxiv.org/html/2608.04060#S8.E46)\), this zero vector field implements the identity transitionz^t\+1=zt\\widehat\{z\}\_\{t\+1\}=z\_\{t\}\. Its latent prediction error is smaller than that of every non\-collapsed model, but Table[5](https://arxiv.org/html/2608.04060#S8.T5)shows that this apparent success is obtained by reducing the representation to a region with almost no variation\. The constructive fixed\-point control reaches essentially exact collapse at a nonzero constant latent state\.

Table[4](https://arxiv.org/html/2608.04060#S8.T4)evaluates the predictive and rollout performance of the three non\-collapsed models\. Table[5](https://arxiv.org/html/2608.04060#S8.T5)instead isolates the collapse mechanism by comparing regularisedSJEPAwith two diagnostic conditions\. The no\-ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}model tests whether collapse emerges when representation control is removed, while the fixed\-point control constructs the constant\-representation solution directly\. We report absolute latent\-variation statistics because a low latent prediction error can be obtained trivially when all observations are encoded near the same point\.

Table 5:Experiment 1 representation\-collapse diagnostics\. Latent MSE is evaluated on test transitions, whileminj⁡Std⁡\(Zj\)\\min\_\{j\}\\operatorname\{Std\}\(Z\_\{j\}\)andTr⁡\(Cov⁡Z\)\\operatorname\{Tr\}\(\\operatorname\{Cov\}Z\)are computed from context embeddings of the training trajectories\. RegularisedSJEPAis included as a non\-collapsed reference; the no\-ℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}and fixed\-point conditions are collapse diagnostics rather than matched predictive baselines\. Very low latent prediction error is uninformative when absolute latent variation vanishes\. Values are mean±\\pmstandard deviation over 3 optimisation seeds\.Figure[2](https://arxiv.org/html/2608.04060#S8.F2)\(b\) visualises the corresponding loss of latent variation\. Relative to the regularised model, the one\-step collapse diagnostic reduces the minimum coordinate standard deviation by approximately two orders of magnitude, while the fixed\-point control is effectively constant\. The experiment therefore distinguishes genuine operator compression from representation collapse: the regularised model retains substantial two\-dimensional variation, whereas the unconstrained objective obtains a trivial identity transition by making successive embeddings almost constant\. Effective rank alone is insufficient for this diagnosis because it measures relative dimensional usage but is insensitive to the absolute scale of latent variation\.

![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment1_rollout_curve_v4.png)

\(a\) Physical\-state rollout error

![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment1_collapse_min_std_v4.png)

\(b\) Minimum latent coordinate standard deviation

Figure 2:Experiment 1 evaluates operator compression and representation collapse\.\(a\)Joint symbolicSJEPAaccumulates substantially less physical\-state rollout error than symbolic regression fitted post hoc, although the flexible Neural JEPA predictor remains the most accurate\.\(b\)RemovingℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}causes absolute latent variation to nearly vanish, while the fixed\-point control is effectively constant\. Curves and error bars show means and one standard deviation over 3 seeds\.As reported in Table[12](https://arxiv.org/html/2608.04060#A4.T12)\(Appendix\.[D](https://arxiv.org/html/2608.04060#A4)\), the learnedSJEPAcoordinates are less affinely aligned with the original physical state than the prediction\-only coordinates: test affine\-probeR2R^\{2\}decreases from0\.814±0\.0630\.814\\pm 0\.063to0\.293±0\.1540\.293\\pm 0\.154\. We therefore interpret Experiment 1 as evidence that joint learning can produce simpler, non\-collapsed latent dynamics, rather than as evidence that operator compression preserves a simple affine correspondence with the original state\. An illustrative example of the nonlinear latent geometry learned in Experiment 1 is shown in Figure[4](https://arxiv.org/html/2608.04060#A4.F4)\(a\)\.

### 8\.2Experiment 2: Hybrid dynamics under controlled grammar mismatch

##### Design\.

Experiment 2 isolates the transition decomposition by using state\-aligned coordinates, namely the physical state,

and learning no encoder\. The true dynamics include quadratic drag:

q˙=p,p˙=−sin⁡\(q\)−0\.4​p​\|p\|\.\\dot\{q\}=p,\\qquad\\dot\{p\}=\-\\sin\(q\)\-0\.4p\|p\|\.\(54\)The incomplete output\-specific grammar permits

q˙:\{p\},p˙:\{sin⁡\(q\)\},\\dot\{q\}:\\\{p\\\},\\qquad\\dot\{p\}:\\\{\\sin\(q\)\\\},\(55\)whereas the complete grammar additionally permitsp​\|p\|p\|p\|in the second equation\. Linearppis deliberately excluded fromp˙\\dot\{p\}so that it cannot act as a symbolic surrogate for the omitted quadratic drag\.

We compare symbolic\-only dynamics under the incomplete and complete grammars, a neural\-only vector field, and hybrid models with and without the correction penalty\. The two hybrid models use the incomplete grammar and the same bounded neural correction architecture\. Thus, their only methodological difference is whetherℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}is active\.

We measure reliance on the neural correction using the normalised correction\-energy ratio

ρcorr=𝔼​\[‖cϕ​\(z\)‖22\]𝔼​\[‖Fℰ,α​\(z\)\+cϕ​\(z\)‖22\]\+η0,η0=10−12\.\\rho\_\{\\mathrm\{corr\}\}=\\frac\{\\mathbb\{E\}\\left\[\\\|c\_\{\\phi\}\(z\)\\\|\_\{2\}^\{2\}\\right\]\}\{\\mathbb\{E\}\\left\[\\\|F\_\{\\mathcal\{E\},\\alpha\}\(z\)\+c\_\{\\phi\}\(z\)\\\|\_\{2\}^\{2\}\\right\]\+\\eta\_\{0\}\},\\qquad\\eta\_\{0\}=10^\{\-12\}\.\(Eq\.[64](https://arxiv.org/html/2608.04060#A1.E64)\)A small value indicates limited reliance on the correction, but is meaningful only when considered jointly with predictive error and symbolic complexity\. For symbolic\-only models,cϕ≡0c\_\{\\phi\}\\equiv 0and henceρcorr=0\\rho\_\{\\mathrm\{corr\}\}=0; the ratio is not applicable to the neural\-only model because it has no symbolic–correction decomposition\.

##### A complete grammar recovers the governing law\.

Table[6](https://arxiv.org/html/2608.04060#S8.T6)shows that the incomplete symbolic model underfits the drag mechanism, whereas restoring the omitted primitive reduces one\-step state MSE by approximately two orders of magnitude and test rollout MSE by approximately277277times\. Across all 3 seeds, the complete symbolic model selects exactly the intended three terms and recovers

q˙=\(0\.9830±0\.0005\)​p,p˙=−\(0\.9553±0\.0010\)​sin⁡\(q\)−\(0\.3724±0\.0007\)​p​\|p\|\.\\dot\{q\}=\(0\.9830\\pm 0\.0005\)p,\\qquad\\dot\{p\}=\-\(0\.9553\\pm 0\.0010\)\\sin\(q\)\-\(0\.3724\\pm 0\.0007\)p\|p\|\.The true coefficients are\(1,−1,−0\.4\)\(1,\-1,\-0\.4\), corresponding to relative errors of approximately1\.7%1\.7\\%,4\.5%4\.5\\%, and6\.9%6\.9\\%\. Figure[3](https://arxiv.org/html/2608.04060#S8.F3)\(a\) shows the corresponding long\-horizon behaviour: the incomplete symbolic law accumulates substantial error, whereas restoring the omitted primitive produces rollouts approaching those of the neural predictor while retaining an explicit governing equation\.

Table 6:Experiment 2 results under controlled grammar mismatch\. The normalised correction\-energy ratioρcorr\\rho\_\{\\mathrm\{corr\}\}measures reliance on the neural correction and should be interpreted jointly with predictive error and symbolic complexity\. Values are mean±\\pmstandard deviation over three optimisation seeds\.
##### Correction regularisation preserves the symbolic mechanism\.

Under the incomplete grammar, the regularised hybrid retains

q˙=\(0\.9581±0\.0008\)​p,p˙=−\(0\.7942±0\.0024\)​sin⁡\(q\),\\dot\{q\}=\(0\.9581\\pm 0\.0008\)p,\\qquad\\dot\{p\}=\-\(0\.7942\\pm 0\.0024\)\\sin\(q\),and has a normalised correction\-energy ratio of0\.06270\.0627\. Withoutℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}, the symbolic coefficients shrink to

q˙=\(0\.3324±0\.0065\)​p,p˙=−\(0\.1718±0\.0746\)​sin⁡\(q\),\\dot\{q\}=\(0\.3324\\pm 0\.0065\)p,\\qquad\\dot\{p\}=\-\(0\.1718\\pm 0\.0746\)\\sin\(q\),while the normalised correction\-energy ratio rises to0\.5550\.555\. Thus, an unrestricted correction largely takes over dynamics that the symbolic component could otherwise explain\.

To evaluate whether the correction captures the mechanism omitted by the symbolic law, let

rp​\(q,p\)=ftrue,p​\(q,p\)−Fℰ,α,p​\(q,p\)r\_\{p\}\(q,p\)=f\_\{\\mathrm\{true\},p\}\(q,p\)\-F\_\{\\mathcal\{E\},\\alpha,p\}\(q,p\)\(Eq\.[81](https://arxiv.org/html/2608.04060#A4.E81)\)be the complete residual left in theppvector field\. A scalar calibration is fitted using training data and then evaluated on held\-out test transitions\. Table[7](https://arxiv.org/html/2608.04060#S8.T7)shows that the regularised correction closely recovers both this total residual and the physical drag term\. Figure[3](https://arxiv.org/html/2608.04060#S8.F3)\(b\) visualises the near one\-to\-one relationship between the calibrated correction and the complete symbolic residual\. The drag\-specific comparison is reported as a secondary diagnostic in Figure[4](https://arxiv.org/html/2608.04060#A4.F4)\(b\)\. The unregularised model remains predictive, but its correction is much less specifically aligned with the omitted mechanism\.

Table 7:Correction\-allocation diagnostics in Experiment 2\. Residual denotes the complete test\-setpp\-field residual after the fitted symbolic law; drag denotes the isolated contribution−0\.4​p​\|p\|\-0\.4p\|p\|\. The reportedR2R^\{2\}values use a scalar calibration fitted on the training set\.![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment2_rollout_curve_v4.png)

\(a\) State\-space rollout error

![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment2_correction_vs_total_residual_v4.png)

\(b\) Recovery of the symbolic residual

Figure 3:Experiment 2 evaluates controlled grammar mismatch and symbolic–neural allocation\.\(a\)The incomplete symbolic grammar accumulates substantial rollout error, whereas the complete symbolic grammar approaches the neural predictor\. Both hybrid models substantially improve over the incomplete symbolic model\.\(b\)For the regularised hybrid using theincompletegrammar, the neural correction closely tracks the held\-outpp\-field residual left by the fitted symbolic component after scalar calibration on the training set\. The dashed line denotes exact recovery\.The unregularised hybrid attains a numerically lower one\-step state MSE,\(1\.888±0\.038\)×10−5\(1\.888\\pm 0\.038\)\\times 10^\{\-5\}compared with\(2\.474±0\.083\)×10−5\(2\.474\\pm 0\.083\)\\times 10^\{\-5\}for the regularised hybrid, and a lower mean OOD rollout MSE,0\.0778±0\.00810\.0778\\pm 0\.0081compared with0\.1045±0\.00310\.1045\\pm 0\.0031, as shown in Table[6](https://arxiv.org/html/2608.04060#S8.T6)\. Conversely, the regularised hybrid achieves the lower test rollout MSE,0\.0671±0\.00340\.0671\\pm 0\.0034compared with0\.0879±0\.01410\.0879\\pm 0\.0141\. Consequently,ℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}should not be interpreted as providing an unconditional accuracy improvement\. Its role is to control symbolic–neural allocation: the symbolic component retains the dominant reusable mechanism, while the correction is reserved for residual dynamics that the grammar does not express adequately\.

### 8\.3Experiments summary and scope

The two experiments support complementary parts of the proposed framework\.

Experiment 1 shows that operator complexity can guide predictive coordinate selection\. Relative to symbolic regression fitted post hoc to frozen Neural JEPA coordinates, joint symbolicSJEPAreduces mean weighted symbolic complexity from26\.026\.0to4\.674\.67, a factor of approximately5\.65\.6\(Table[4](https://arxiv.org/html/2608.04060#S8.T4)\)\. It also reduces physical\-state test rollout MSE from1\.2291\.229to0\.4670\.467, OOD rollout MSE from9\.2629\.262to3\.4353\.435, and OOD divergence rate from0\.1780\.178to0\.0170\.017\. Across all three seeds, the discovered laws retain the same dominant oscillator\-like cross\-coupling betweenz1z\_\{1\}andz2z\_\{2\}\. The within\-representation one\-step latent MSE is also approximately3\.63\.6times lower, although this comparison is descriptive because latent MSE depends on the scale and coordinate system of each learned representation\.

The collapse diagnostics establish the complementary necessity of representation constraints\. RemovingℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}produces nearly constant embeddings and a zero vector field, which implements an identity transition under the residual parameterisation\. The constructive fixed\-point control reaches the same degenerate solution directly\. Extremely low latent prediction error and zero symbolic complexity can therefore indicate representation collapse rather than successful dynamics discovery, confirming the shortcut identified in Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5)\.

These findings do not imply thatSJEPAuniversally outperforms a neural predictor\. Neural JEPA remains the most accurate model in raw one\-step and rollout prediction\. Moreover, the jointly learned symbolic coordinates are less affinely aligned with the original physical state than the prediction\-only coordinates\. Experiment 1 therefore demonstrates a trade\-off:SJEPAsacrifices some predictive flexibility and simple physical\-state alignment in exchange for a substantially more concise symbolic transition with lower recursive rollout error and divergence than the post\-hoc symbolic baseline\.

Experiment 2 separates grammar adequacy and symbolic–neural allocation from representation learning\. When the grammar contains the required quadratic\-drag primitive, the symbolic model selects the intended governing structure in every seed and recovers all coefficients to within approximately7%7\\%of their true values\. Its rollout error approaches that of the neural predictor while retaining an explicit equation, showing that a sufficiently expressive grammar can remove the need for a residual correction in this controlled setting\.

Under the incomplete grammar, the regularised hybrid retains the representable pendulum terms and has a normalised correction\-energy ratio of0\.06270\.0627, compared with0\.5550\.555for the unregularised hybrid\. After scalar calibration on the training set, the regularised correction predicts the held\-out complete symbolic residual withR2=0\.994R^\{2\}=0\.994and is strongly aligned with the deliberately omitted quadratic\-drag mechanism, with drag\-specificR2=0\.901R^\{2\}=0\.901\. Withoutℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}, the symbolic coefficients shrink substantially and the correction absorbs much of the dynamics that the grammar can already represent\.

Correction regularisation is nevertheless not an unconditional accuracy improvement\. The unregularised hybrid attains a numerically lower one\-step state MSE and lower mean OOD rollout MSE in this experiment, whereas the regularised hybrid attains the lower test rollout MSE\. The role ofℛcorr\\mathcal\{R\}\_\{\\mathrm\{corr\}\}is therefore explanatory allocation: it encourages the symbolic component to retain compact representable structure and reserves the neural component for the remaining residual, rather than guaranteeing the smallest prediction error under every evaluation condition\.

These experiments are controlled diagnostics rather than evidence of large\-scale generality\. They test coordinate selection, representation collapse, grammar adequacy, rollout behaviour, and symbolic–neural allocation in systems where the underlying state and omitted mechanism are known\. Evaluation with visual observations at scale, genuinely pretrained foundation encoders, action\-conditioned control tasks, real scientific datasets, and Bayesian posterior calibration remains future work\.

## 9Discussion

##### Key innovation\.

The contribution ofSJEPAis not merely to insert symbolic regression into a JEPA predictor, and joint coordinate–equation learning has important precedents\. Its distinctive contribution is to make the complexity of the induced transition operator an explicit criterion for learning a reconstruction\-free predictive representation\. The symbolic law is therefore not fitted only as a post\-hoc explanation of fixed coordinates; in the end\-to\-end mode, it directly influences which predictive coordinates are selected\. The constrained formulation, induced\-dynamics complexity, and collapse analysis formalise this principle as learning the simplest adequate dynamics over an informative, non\-collapsed state\.

A second innovation is the controlled symbolic–neural decompositionHℰ,α,ϕ=Fℰ,α\+cϕH\_\{\\mathcal\{E\},\\alpha,\\phi\}=F\_\{\\mathcal\{E\},\\alpha\}\+c\_\{\\phi\}\. The symbolic component represents the dominant reusable mechanism, while the correction accounts for predictive structure that the selected grammar cannot express adequately\. Because this decomposition is not identifiable from prediction loss alone, the symbolic\-complexity and correction penalties encode an explicit allocation preference\. The framework thereby addresses two coupled questions: which coordinates expose compact dynamics, and how predictable structure should be allocated between an explicit symbolic law and a flexible residual model\.131313At a broader level, this raises a more general question that we leave to separate work: given a task and dataset, which representation geometry or latent space is most suitable for JEPA learning? Different modelling objectives and selection criteria implicitly favour different latent geometries\.?

##### Empirical implications\.

The controlled experiments validate the two main deterministic consequences of the framework\. First, allowing operator complexity to influence representation learning produces substantially more concise symbolic dynamics with lower rollout error and divergence than fitting equations post hoc, provided that explicit representation constraints prevent collapse\. Second, correction regularisation controls how predictable structure is allocated between the symbolic law and neural residual under grammar misspecification\. These findings establish a controllable trade\-off among predictive fidelity, operator simplicity, representation quality, and symbolic–neural allocation rather than a universal improvement in raw predictive accuracy\.

##### Interpretation of elegant dynamics\.

A compact symbolic transition can expose signs, interactions, symmetries, action couplings, and candidate invariants that are difficult to inspect in an unrestricted neural predictor\. It can also be differentiated, simplified, checked against known constraints, and reused across trajectories\. These advantages do not make a learned latent equation automatically physical, causal, or valid outside the observed regime\. Its interpretation remains conditional on the learned coordinates, symbolic grammar, data support, predictive tolerance, and adequacy of the residual model\.

The oscillator\-like coordinates discovered in Experiment 1 illustrate both the value and the ambiguity of operator compression\. The learned transition is concise and produces substantially lower rollout error than the post\-hoc symbolic baseline, but the coordinates are nonlinearly related to the original physical state\. Operator compression selects coordinates with simple evolution, not necessarily coordinates with direct physical semantics\. Likewise, symbolic simplicity and linear simplicity are not mutually exclusive: the approximately linear oscillator discovered bySJEPAis a special case of compact symbolic dynamics\. Comparisons with Koopman\-style methods should therefore jointly evaluate predictive risk, operator complexity, coordinate quality, rollout behaviour, and downstream control utility under matched budgets\.

These results motivate reporting prediction error, rollout stability, symbolic complexity, representation variation, coordinate alignment, and correction reliance together\. No single metric is sufficient: low predictive error may accompany collapse, low symbolic complexity may reflect underfitting, and a small correction may simply indicate that an incomplete symbolic model has been left uncorrected\.

##### Scope, modularity, and limitations\.

SJEPAis modular with respect to the encoder, representation regulariser, symbolic search procedure, correction architecture, and downstream planner\. Experiment 1 evaluates the end\-to\-end mode in which representation and symbolic dynamics are learned jointly\. Experiment 2 uses fixed, state\-aligned coordinates and isolates the corresponding fixed\-coordinate question: whether the supplied representation admits an adequate compact law and how unexplained structure is allocated to a correction\. Direct experiments with genuinely pretrained JEPA or foundation encoders remain to be conducted\.

The empirical validation is deliberately narrow\. Both experiments use variants of one controlled pendulum system, two\-dimensional states, prescribed symbolic libraries, and 3 optimisation seeds\. They do not establish that compact symbolic latent laws will emerge in high\-dimensional video, partially observed environments, scientific datasets, or embodied control tasks\. The present differentiable sparse\-library implementation also explores a more restricted model class than unrestricted expression\-tree symbolic regression\.

Simplicity is relative to the coordinate system, grammar, primitive weights, coefficient threshold, predictive tolerance, and observed regime\. Alternating representation–operator optimisation is more expensive than fitting a single neural predictor, is susceptible to local optima, and does not guarantee recovery of a globally simplest adequate law\. A correction with excessive capacity may conceal grammar misspecification, whereas an overly restricted correction may force residual effects into distorted symbolic coefficients\. Moreover, discontinuous, stochastic, multiscale, or regime\-switching systems may require local, piecewise, hierarchical, context\-dependent, or distributional symbolic mechanisms rather than one compact global equation\.

##### Bayesian and decision\-making extensions\.

The present experiments validate only the deterministic framework\. The Bayesian formulation provides uncertainty over symbolic structures, coefficients, residual covariance, and an optional Gaussian\-process correction, but posterior calibration and decision\-theoretic utility have not yet been evaluated\. Similarly, the action\-conditioned formulation exposes sensitivities, local linearisations, and an interface to external planning, but the paper reports no control experiment\. Claims concerning posterior calibration, sample efficiency, control performance, or uncertainty\-aware planning therefore remain hypotheses for future empirical study\.

## 10Conclusion

We introducedSJEPA, a reconstruction\-free joint\-embedding predictive framework that learns latent transitions as compact symbolic laws with optional neural corrections\. Its central principle is to complement representation learning with operator compression: the representation must preserve an informative, non\-collapsed predictive state, while the transition model should realise the simplest adequate law over that state\. We formalised this principle through constrained operator compression and induced\-dynamics complexity, characterised the coordinate non\-identifiability of predictive representations, identified the collapse shortcut introduced by unconstrained dynamics simplification, and analysed how correction regularisation controls the allocation between symbolic and neural dynamics\. The framework supports both alternating representation–equation learning and symbolic dynamics over fixed representations, with Bayesian and action\-conditioned formulations providing extensions to uncertainty and decision making\.

The controlled experiments support the two principal deterministic claims\. Joint representation and symbolic\-dynamics learning reduces symbolic complexity by approximately5\.65\.6times and produces substantially lower physical\-state rollout error and OOD divergence than symbolic regression fitted post hoc to prediction\-only coordinates\. The unconstrained one\-step diagnostic yields near\-constant embeddings and trivial identity dynamics, demonstrating that the predicted collapse shortcut is attainable and that operator simplicity is meaningful only for an admissible representation\. Under controlled grammar misspecification, a complete grammar recovers the governing structure and coefficients consistently, while correction regularisation reduces the normalised correction\-energy ratio from0\.5550\.555to0\.06270\.0627and yields a residual correction with calibrated testR2=0\.994R^\{2\}=0\.994, while limiting unnecessary reliance on the correction for representable dynamics\.

These findings expose a controllable trade\-off rather than a universal accuracy advantage\. Flexible neural predictors remain more accurate, learned symbolic coordinates need not be simply aligned with physical variables, and correction regularisation may exchange some predictive flexibility for more controlled symbolic–neural allocation\. The central conclusion is therefore that operator compression can select predictive coordinates whose induced dynamics are simple yet adequate, while representation constraints and correction control ensure that this simplicity is achieved by compressing the transition law rather than collapsing the representation or delegating the dynamics to the neural correction\.

## References

- A\. A\. Alemi, I\. Fischer, J\. V\. Dillon, and K\. Murphy \(2017\)Deep variational information bottleneck\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HyxQzBceg)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.
- M\. Assran, Q\. Duval, I\. Misra, P\. Bojanowski, P\. Vincent, M\. Rabbat, Y\. LeCun, and N\. Ballas \(2023\)Self\-supervised learning from images with a joint\-embedding predictive architecture\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 15619–15629\.Cited by:[§1](https://arxiv.org/html/2608.04060#S1.p1.1),[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1),[§5\.4](https://arxiv.org/html/2608.04060#S5.SS4.p1.1)\.
- M\. Assran, A\. Bardes, D\. Fan, Q\. Garrido, R\. Howes, Mojtaba, Komeili, M\. Muckley, A\. Rizvi, C\. Roberts, K\. Sinha, A\. Zholus, S\. Arnaud, A\. Gejji, A\. Martin, F\. R\. Hogan, D\. Dugas, P\. Bojanowski, V\. Khalidov, P\. Labatut, F\. Massa, M\. Szafraniec, K\. Krishnakumar, Y\. Li, X\. Ma, S\. Chandar, F\. Meier, Y\. LeCun, M\. Rabbat, and N\. Ballas \(2025\)V\-jepa 2: self\-supervised video models enable understanding, prediction and planning\.External Links:2506\.09985,[Link](https://arxiv.org/abs/2506.09985)Cited by:[§1](https://arxiv.org/html/2608.04060#S1.p1.1),[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Balestriero and Y\. LeCun \(2025\)LeJEPA: provable and scalable self\-supervised learning without the heuristics\.External Links:2511\.08544,[Link](https://arxiv.org/abs/2511.08544)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.
- A\. Bardes, Q\. Garrido, J\. Ponce, X\. Chen, M\. Rabbat, Y\. LeCun, M\. Assran, and N\. Ballas \(2024\)Revisiting feature prediction for learning visual representations from video\.Transactions on Machine Learning Research\.External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=QaCCuDfBk2)Cited by:[§1](https://arxiv.org/html/2608.04060#S1.p1.1),[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1),[§5\.4](https://arxiv.org/html/2608.04060#S5.SS4.p1.1)\.
- A\. Bardes, J\. Ponce, and Y\. LeCun \(2022\)VICReg: variance\-invariance\-covariance regularization for self\-supervised learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=xm6YD62D1Ub)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.
- A\. Bardes, J\. Ponce, and Y\. LeCun \(2023\)MC\-jepa: a joint\-embedding predictive architecture for self\-supervised learning of motion and content features\.External Links:2307\.12698,[Link](https://arxiv.org/abs/2307.12698)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Biggio, T\. Bendinelli, A\. Neitz, A\. Lucchi, and G\. Parascandolo \(2021\)Neural symbolic regression that scales\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139\.Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- S\. L\. Brunton, B\. W\. Brunton, J\. L\. Proctor, and J\. N\. Kutz \(2016a\)Koopman invariant subspaces and finite linear representations of nonlinear dynamical systems for control\.PLOS ONE11\(2\),pp\. 1–19\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0150171),[Link](https://doi.org/10.1371/journal.pone.0150171)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px4.p1.1)\.
- S\. L\. Brunton, J\. L\. Proctor, and J\. N\. Kutz \(2016b\)Discovering governing equations from data by sparse identification of nonlinear dynamical systems\.Proceedings of the National Academy of Sciences113\(15\),pp\. 3932–3937\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1517384113),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.1517384113),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.1517384113Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Caron, H\. Touvron, I\. Misra, H\. Jégou, J\. Mairal, P\. Bojanowski, and A\. Joulin \(2021\)Emerging properties in self\-supervised vision transformers\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 9650–9660\.Cited by:[§5\.4](https://arxiv.org/html/2608.04060#S5.SS4.p1.1)\.
- K\. Champion, B\. Lusch, J\. N\. Kutz, and S\. L\. Brunton \(2019\)Data\-driven discovery of coordinates and governing equations\.Proceedings of the National Academy of Sciences116\(45\),pp\. 22445–22451\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1906995116),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.1906995116),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.1906995116Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Chen, M\. Shukor, T\. Moutakanni, W\. Chung, J\. Yu, T\. Kasarla, Y\. Bang, A\. Bolourchi, Y\. LeCun, and P\. Fung \(2026\)VL\-jepa: joint embedding predictive architecture for vision\-language\.External Links:2512\.10942,[Link](https://arxiv.org/abs/2512.10942)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- R\. T\. Q\. Chen, Y\. Rubanova, J\. Bettencourt, and D\. K\. Duvenaud \(2018\)Neural ordinary differential equations\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/69386f6bb1dfed68692a24c8686939b9-Paper.pdf)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px4.p1.1)\.
- M\. Cranmer, A\. Sanchez\-Gonzalez, P\. Battaglia, R\. Xu, K\. Cranmer, D\. Spergel, and S\. Ho \(2020\)Discovering symbolic models from deep learning with inductive biases\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Cranmer \(2023\)Interpretable machine learning for science with pysr and symbolicregression\.jl\.External Links:2305\.01582,[Link](https://arxiv.org/abs/2305.01582)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. Houlsby \(2021\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=YicbFdNTTy)Cited by:[§5\.4](https://arxiv.org/html/2608.04060#S5.SS4.p1.1)\.
- J\. Grill, F\. Strub, F\. Altché, C\. Tallec, P\. H\. Richemond, E\. Buchatskaya, C\. Doersch, B\. A\. Pires, Z\. D\. Guo, M\. G\. Azar, B\. Piot, K\. Kavukcuoglu, R\. Munos, and M\. Valko \(2020\)Bootstrap your own latent a new approach to self\-supervised learning\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.
- D\. Ha and J\. Schmidhuber \(2018\)Recurrent world models facilitate policy evolution\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/2de5d16682c3c35007e4e92982f1a2ba-Paper.pdf)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px4.p1.1)\.
- K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollár, and R\. Girshick \(2022\)Masked autoencoders are scalable vision learners\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 16000–16009\.Cited by:[§5\.4](https://arxiv.org/html/2608.04060#S5.SS4.p1.1)\.
- N\. Hu, H\. Cheng, Y\. Xie, S\. Li, and J\. Zhu \(2024\)3D\-jepa: a joint embedding predictive architecture for 3d self\-supervised representation learning\.External Links:2409\.15803,[Link](https://arxiv.org/abs/2409.15803)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Huang and H\. Raza \(2026\)Knowledge, rules and their embeddings: two paths towards neuro\-symbolic JEPA\.arXiv preprint arXiv:2603\.13265\.Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px6.p1.1)\.
- Y\. Huang \(2026a\)On the information bottleneck of VJEPA\.Note:HAL preprintHAL: hal\-05622405External Links:[Link](https://hal.science/hal-05622405)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2),[§4\.2](https://arxiv.org/html/2608.04060#S4.SS2.p2.1),[§5\.1](https://arxiv.org/html/2608.04060#S5.SS1.p2.1)\.
- Y\. Huang \(2026b\)VJEPA: variational joint embedding predictive architectures as probabilistic world models\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=omJqMJt2fC)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Kaiser, J\. N\. Kutz, and S\. L\. Brunton \(2018\)Sparse identification of nonlinear dynamics for model predictive control in the low\-data limit\.Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences474\(2219\),pp\. 20180335\.External Links:ISSN 1364\-5021,[Document](https://dx.doi.org/10.1098/rspa.2018.0335),[Link](https://doi.org/10.1098/rspa.2018.0335),https://royalsocietypublishing\.org/rspa/article\-pdf/doi/10\.1098/rspa\.2018\.0335/632569/rspa\.2018\.0335\.pdfCited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- J\. R\. Koza \(1992\)Genetic programming: on the programming of computers by means of natural selection\.MIT Press,Cambridge, MA\.External Links:ISBN 9780262111706Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. LeCun \(2022\)A path towards autonomous machine intelligence\.Note:OpenReview preprintExternal Links:[Link](https://openreview.net/forum?id=BZ5a1r-kVsf)Cited by:[§1](https://arxiv.org/html/2608.04060#S1.p1.1)\.
- B\. Lusch, J\. N\. Kutz, and S\. L\. Brunton \(2018\)Deep learning for universal linear embeddings of nonlinear dynamics\.Nature Communications9,pp\. 4950\.External Links:[Document](https://dx.doi.org/10.1038/s41467-018-07210-0),[Link](https://doi.org/10.1038/s41467-018-07210-0)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px4.p1.1)\.
- P\. Muratore and M\. W\. Mathis \(2026\)Extracting governing equations from latent dynamics via multi\-view contrastive learning\.External Links:2606\.13260,[Link](https://arxiv.org/abs/2606.13260)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px3.p1.1)\.
- B\. K\. Petersen, M\. L\. Larma, T\. N\. Mundhenk, C\. P\. Santiago, S\. K\. Kim, and J\. T\. Kim \(2021\)Deep symbolic regression: recovering mathematical expressions from data via risk\-seeking policy gradients\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=m5Qsh0kBQG)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- I\. Posner, A\. Lei, and B\. Schölkopf \(2026a\)From observation to insight: mechanistic world models and the quest for autonomous discovery\.External Links:2607\.12474,[Link](https://arxiv.org/abs/2607.12474)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- I\. Posner, A\. Lei, and B\. Schölkopf \(2026b\)From observation to insight: mechanistic world models and the quest for autonomous discovery\.External Links:2607\.12474,[Link](https://arxiv.org/abs/2607.12474)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px6.p1.1)\.
- C\. E\. Rasmussen and C\. K\. I\. Williams \(2005\)Gaussian processes for machine learning\.The MIT Press\.External Links:ISBN 9780262256834,[Document](https://dx.doi.org/10.7551/mitpress/3206.001.0001),[Link](https://doi.org/10.7551/mitpress/3206.001.0001),https://direct\.mit\.edu/book\-pdf/2514321/book\_9780262256834\.pdfCited by:[§6\.2](https://arxiv.org/html/2608.04060#S6.SS2.p2.6)\.
- A\. Saito, P\. Kudeshia, and J\. Poovvancheri \(2025\)Point\-jepa: joint embedding predictive architecture for 3d point cloud self\-supervised learning\.InIEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\),Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Schmidt and H\. Lipson \(2009\)Distilling free\-form natural laws from experimental data\.Science324\(5923\),pp\. 81–85\.External Links:[Document](https://dx.doi.org/10.1126/science.1165893),[Link](https://www.science.org/doi/abs/10.1126/science.1165893),https://www\.science\.org/doi/pdf/10\.1126/science\.1165893Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Tishby, F\. C\. Pereira, and W\. Bialek \(1999\)The information bottleneck method\.InProceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing,pp\. 368–377\.Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.
- L\. Tuncay, E\. Labbé, E\. Benetos, and T\. Pellegrini \(2025\)Audio\-jepa: joint\-embedding predictive architecture for audio representation learning\.External Links:2507\.02915,[Link](https://arxiv.org/abs/2507.02915)Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Vujinovic and A\. Kovacevic \(2025\)ACT\-jepa: novel joint\-embedding predictive architecture for efficient policy representation learning\.arXiv preprint arXiv:2501\.14622\.Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Yermakov, D\. Zoro, L\. M\. Gao, and J\. N\. Kutz \(2026\)T\-shred: symbolic regression for regularization and model discovery with transformer shallow recurrent decoders\.Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences384\(2317\),pp\. 20240586\.External Links:ISSN 1364\-503X,[Document](https://dx.doi.org/10.1098/rsta.2024.0586),[Link](https://doi.org/10.1098/rsta.2024.0586),https://royalsocietypublishing\.org/rsta/article\-pdf/doi/10\.1098/rsta\.2024\.0586/6131778/rsta\.2024\.0586\.pdfCited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Zbontar, L\. Jing, I\. Misra, Y\. LeCun, and S\. Deny \(2021\)Barlow twins: self\-supervised learning via redundancy reduction\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 12310–12320\.Cited by:[§2](https://arxiv.org/html/2608.04060#S2.SS0.SSS0.Px5.p1.2)\.

## Appendix AObjectives and Representation Constraints

### A\.1Notation

Table 8:Main notation\. Operator\-complexity quantities are defined relative to the selected latent normalisation, symbolic grammar, correction class, and evaluation distribution\.
### A\.2Temporal Objectives and Correction Diagnostics

For temporal tuples

\(X≤t,Xt\+1:t\+K,εt:t\+K−1\),\\left\(X\_\{\\leq t\},X\_\{t\+1:t\+K\},\\varepsilon\_\{t:t\+K\-1\}\\right\),and, for continuous\-time models, step sizesΔ​tt:t\+K−1\\Delta t\_\{t:t\+K\-1\}, define

zt\\displaystyle z\_\{t\}=Eθ​\(X≤t\),\\displaystyle=E\_\{\\theta\}\(X\_\{\\leq t\}\),\(56\)zt\+htar\\displaystyle z\_\{t\+h\}^\{\\mathrm\{tar\}\}=Eθ¯​\(Xt\+h\),h=1,…,K\.\\displaystyle=E\_\{\\bar\{\\theta\}\}\(X\_\{t\+h\}\),\\qquad h=1,\\ldots,K\.\(57\)The target encoder supplies slowly moving prediction targets in the same ambient latent space as the context encoder\.

To cover both temporal realisations, define the complete one\-step transition used at rollout stephhby

ℋt\+h​\(z\)=\{Hℰ,α,ϕ​\(z,εt\+h\),direct\-transition realisation,ℐΔ​tt\+h​\(z,Vℰ,α,ϕ,εt\+h\),vector\-field realisation\.\\mathcal\{H\}\_\{t\+h\}\(z\)=\\begin\{cases\}H\_\{\\mathcal\{E\},\\alpha,\\phi\}\\left\(z,\\varepsilon\_\{t\+h\}\\right\),&\\text\{direct\-transition realisation\},\\\\\[4\.0pt\] \\mathcal\{I\}\_\{\\Delta t\_\{t\+h\}\}\\left\(z,V\_\{\\mathcal\{E\},\\alpha,\\phi\},\\varepsilon\_\{t\+h\}\\right\),&\\text\{vector\-field realisation\}\.\\end\{cases\}\(58\)The free recursive rollout is

z^t=zt,z^t\+h\+1=ℋt\+h​\(z^t\+h\),h=0,…,K−1\.\\widehat\{z\}\_\{t\}=z\_\{t\},\\qquad\\widehat\{z\}\_\{t\+h\+1\}=\\mathcal\{H\}\_\{t\+h\}\\left\(\\widehat\{z\}\_\{t\+h\}\\right\),\\qquad h=0,\\ldots,K\-1\.\(59\)
For the forward\-Euler realisation used in the reported experiments, Equation \([58](https://arxiv.org/html/2608.04060#A1.E58)\) becomes

ℋt\+h​\(z\)=z\+Δ​tt\+h​\[Fℰ,α​\(z,εt\+h\)\+cϕ​\(z,εt\+h\)\]\.\\mathcal\{H\}\_\{t\+h\}\(z\)=z\+\\Delta t\_\{t\+h\}\\left\[F\_\{\\mathcal\{E\},\\alpha\}\\left\(z,\\varepsilon\_\{t\+h\}\\right\)\+c\_\{\\phi\}\\left\(z,\\varepsilon\_\{t\+h\}\\right\)\\right\]\.\(60\)
The rollout loss is

ℒroll=∑h=1Kwh​d​\(z^t\+h,sg⁡\(zt\+htar\)\),wh≥0\.\\mathcal\{L\}\_\{\\mathrm\{roll\}\}=\\sum\_\{h=1\}^\{K\}w\_\{h\}d\\left\(\\widehat\{z\}\_\{t\+h\},\\operatorname\{sg\}\\left\(z\_\{t\+h\}^\{\\mathrm\{tar\}\}\\right\)\\right\),\\qquad w\_\{h\}\\geq 0\.\(61\)Teacher\-forced and free\-rollout terms may be combined\. For temporal applications in which recursive accuracy is part of the required notion of adequacy, the constrained formulation may additionally impose

ℒroll≤δroll\.\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\\leq\\delta\_\{\\mathrm\{roll\}\}\.The task\-general definition retains only the one\-step constraint because recursive rollout is not defined for every JEPA context–target relation\.

When correction control is applied during temporal rollout, the correction output may be penalised at every recursively visited state\. In the vector\-field realisation, the correction penalty and diagnostics are evaluated on the raw vector\-field correctioncϕc\_\{\\phi\}before numerical integration and before multiplication byΔ​t\\Delta t\. This prevents them from depending artificially on the integration step size\.

The correction\-energy ratio is evaluated on the decomposed predictive object itself: the complete predicted target in the direct\-transition realisation, or the latent vector field in the continuous\-time realisation\. Define

𝒢ℰ,α,ϕ​\(z,ε\)=\{Hℰ,α,ϕ​\(z,ε\),direct\-transition realisation,Vℰ,α,ϕ​\(z,ε\),vector\-field realisation\.\\mathcal\{G\}\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,\\varepsilon\)=\\begin\{cases\}H\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,\\varepsilon\),&\\text\{direct\-transition realisation\},\\\\\[3\.0pt\] V\_\{\\mathcal\{E\},\\alpha,\\phi\}\(z,\\varepsilon\),&\\text\{vector\-field realisation\}\.\\end\{cases\}\(62\)The normalised correction\-energy ratio is

ρcorr=𝔼​\[‖cϕ​\(ZC,ε\)‖22\]𝔼​\[‖𝒢ℰ,α,ϕ​\(ZC,ε\)‖22\]\+η0,η0\>0\.\\rho\_\{\\mathrm\{corr\}\}=\\frac\{\\mathbb\{E\}\\left\[\\left\\\|c\_\{\\phi\}\(Z\_\{C\},\\varepsilon\)\\right\\\|\_\{2\}^\{2\}\\right\]\}\{\\mathbb\{E\}\\left\[\\left\\\|\\mathcal\{G\}\_\{\\mathcal\{E\},\\alpha,\\phi\}\(Z\_\{C\},\\varepsilon\)\\right\\\|\_\{2\}^\{2\}\\right\]\+\\eta\_\{0\}\},\\qquad\\eta\_\{0\}\>0\.\(63\)For the vector\-field experiments, this reduces to

ρcorr=𝔼​\[‖cϕ​\(z\)‖22\]𝔼​\[‖Fℰ,α​\(z\)\+cϕ​\(z\)‖22\]\+η0\.\\rho\_\{\\mathrm\{corr\}\}=\\frac\{\\mathbb\{E\}\\left\[\\left\\\|c\_\{\\phi\}\(z\)\\right\\\|\_\{2\}^\{2\}\\right\]\}\{\\mathbb\{E\}\\left\[\\left\\\|F\_\{\\mathcal\{E\},\\alpha\}\(z\)\+c\_\{\\phi\}\(z\)\\right\\\|\_\{2\}^\{2\}\\right\]\+\\eta\_\{0\}\}\.\(64\)The residual identity contributionzzin Equation \([60](https://arxiv.org/html/2608.04060#A1.E60)\) and the factorΔ​t\\Delta tare excluded because they are not part of the symbolic–neural decomposition of the vector field\.

A low value ofρcorr\\rho\_\{\\mathrm\{corr\}\}indicates limited reliance on the correction, but is not automatically preferable: a near\-zero correction paired with high predictive error may indicate underfitting\. The ratio is therefore interpreted jointly with predictive risk and symbolic complexity\. Because the symbolic and correction components need not be orthogonal,ρcorr\\rho\_\{\\mathrm\{corr\}\}is a normalised energy ratio rather than an additive fraction of the total dynamics\.

### A\.3Representation Regularisation

The theory requires an admissible, non\-collapsed predictive representation\. Different implementations impose different assumptions\.

Table 9:Representative implementations ofℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}\. VICReg and Barlow Twins are practical surrogates for representation quality and non\-collapse, not exact estimators of the mutual\-information bottleneck\.A concrete VICReg\-style instance is

ℛIBVIC=λinv​ℒinv\+λvar​ℒvar\+λcov​ℒcov\.\\mathcal\{R\}\_\{\\mathrm\{IB\}\}^\{\\mathrm\{VIC\}\}=\\lambda\_\{\\mathrm\{inv\}\}\\mathcal\{L\}\_\{\\mathrm\{inv\}\}\+\\lambda\_\{\\mathrm\{var\}\}\\mathcal\{L\}\_\{\\mathrm\{var\}\}\+\\lambda\_\{\\mathrm\{cov\}\}\\mathcal\{L\}\_\{\\mathrm\{cov\}\}\.\(65\)The three terms are the usual invariance, per\-coordinate variance\-floor, and off\-diagonal covariance penalties\. InSJEPA, the variance term blocks the constant\-coordinate shortcut of Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5); the covariance term discourages redundant linear dependence between coordinates; and the invariance term encourages consistency under the specified observation augmentations\. The experimental instantiation in Appendix[D](https://arxiv.org/html/2608.04060#A4)additionally includes a weak mean\-centering termλmean​ℒmean\\lambda\_\{\\mathrm\{mean\}\}\\mathcal\{L\}\_\{\\mathrm\{mean\}\}\.

## Appendix BProofs

### B\.1Predictive Coordinates and Collapse

#### B\.1\.1Proof of Proposition[4\.4](https://arxiv.org/html/2608.04060#S4.Thmdefinition4)

For almost every training tuple,

F~​\(E~​\(XC\),ε\)\\displaystyle\\widetilde\{F\}\(\\widetilde\{E\}\(X\_\{C\}\),\\varepsilon\)=h​\(F​\(h−1​\(h​\(E​\(XC\)\)\),ε\)\)\\displaystyle=h\\left\(F\\left\(h^\{\-1\}\(h\(E\(X\_\{C\}\)\)\),\\varepsilon\\right\)\\right\)=h​\(F​\(E​\(XC\),ε\)\)\\displaystyle=h\(F\(E\(X\_\{C\}\),\\varepsilon\)\)=h​\(E​\(XT\)\)\\displaystyle=h\(E\(X\_\{T\}\)\)=E~​\(XT\)\.\\displaystyle=\\widetilde\{E\}\(X\_\{T\}\)\.Thus, the transformed representation is equally exact\. Symbolic complexity is not invariant under general conjugacies: a simple linear or polynomialFFcan become a complicated rational, transcendental, or non\-representableF~\\widetilde\{F\}, and conversely\. Therefore, predictive exactness alone does not identify coordinates, while an operator\-complexity criterion can distinguish them\.□\\square

#### B\.1\.2Proof of Proposition[4\.5](https://arxiv.org/html/2608.04060#S4.Thmdefinition5)

Under the constant encoders,

ZC=ZT=z0Z\_\{C\}=Z\_\{T\}=z\_\{0\}for every example\. By assumption,

H0​\(z0,ε\)=z0H\_\{0\}\(z\_\{0\},\\varepsilon\)=z\_\{0\}for every admissibleε\\varepsilon\. Hence

d​\(H0​\(ZC,ε\),ZT\)=d​\(z0,z0\)=0,d\\left\(H\_\{0\}\(Z\_\{C\},\\varepsilon\),Z\_\{T\}\\right\)=d\(z\_\{0\},z\_\{0\}\)=0,and thereforeℒpred=0\\mathcal\{L\}\_\{\\mathrm\{pred\}\}=0\. Every feasible model has nonnegative predictive loss, whileH0H\_\{0\}attains the minimum weighted operator penalty by assumption\. Consequently,\(E0,E0,H0\)\(E\_\{0\},E\_\{0\},H\_\{0\}\)attains the global lower bound of the unconstrained objective\.□\\square

### B\.2Symbolic–Neural Allocation

#### B\.2\.1Proof of Proposition[4\.6](https://arxiv.org/html/2608.04060#S4.Thmdefinition6)

Direct substitution gives

\(Fℰ,α\+r\)\+\(cϕ−r\)=Fℰ,α\+cϕ\.\(F\_\{\\mathcal\{E\},\\alpha\}\+r\)\+\(c\_\{\\phi\}\-r\)=F\_\{\\mathcal\{E\},\\alpha\}\+c\_\{\\phi\}\.Both decompositions therefore induce identical predictions on every input and have identical predictive risk\. Unless the penalties or function classes distinguish them, the allocation of structure is not identifiable\.□\\square

#### B\.2\.2Proof of Proposition[4\.7](https://arxiv.org/html/2608.04060#S4.Thmdefinition7)

Condition on\(ZC,ε\)=\(z,e\)\(Z\_\{C\},\\varepsilon\)=\(z,e\)and writem=m​\(z,e\)m=m\(z,e\)anda=F​\(z,e\)a=F\(z,e\)\. For any correction valuecc,

𝔼​\[‖ZT−a−c‖22∣z,e\]\\displaystyle\\mathbb\{E\}\[\\\|Z\_\{T\}\-a\-c\\\|\_\{2\}^\{2\}\\mid z,e\]=𝔼​\[‖ZT−m‖22∣z,e\]\+‖m−a−c‖22,\\displaystyle=\\mathbb\{E\}\[\\\|Z\_\{T\}\-m\\\|\_\{2\}^\{2\}\\mid z,e\]\+\\\|m\-a\-c\\\|\_\{2\}^\{2\},\(66\)because the conditional cross term vanishes\. The first term does not depend oncc\. Hence the conditional objective is

q​\(c\)=‖m−a−c‖22\+λ​‖c‖22\.q\(c\)=\\\|m\-a\-c\\\|\_\{2\}^\{2\}\+\\lambda\\\|c\\\|\_\{2\}^\{2\}\.It is strictly convex forλ\>0\\lambda\>0, and

∇cq​\(c\)=−2​\(m−a−c\)\+2​λ​c\.\\nabla\_\{c\}q\(c\)=\-2\(m\-a\-c\)\+2\\lambda c\.Setting the gradient to zero gives\(1\+λ\)​c=m−a\(1\+\\lambda\)c=m\-a, proving Equation \([23](https://arxiv.org/html/2608.04060#S4.E23)\)\. Integrating the pointwise minimiser over\(ZC,ε\)\(Z\_\{C\},\\varepsilon\)completes the proof\.□\\square

### B\.3Rollout and Bayesian Results

#### B\.3\.1Proof of Proposition[4\.8](https://arxiv.org/html/2608.04060#S4.Thmdefinition8)

For the same actionut\+hu\_\{t\+h\}in the true and learned rollouts,

eh\+1\\displaystyle e\_\{h\+1\}=‖T​\(zt\+h,ut\+h\)−H​\(z^t\+h,ut\+h\)‖2\\displaystyle=\\\|T\(z\_\{t\+h\},u\_\{t\+h\}\)\-H\(\\widehat\{z\}\_\{t\+h\},u\_\{t\+h\}\)\\\|\_\{2\}≤‖T​\(zt\+h,ut\+h\)−H​\(zt\+h,ut\+h\)‖2\\displaystyle\\leq\\\|T\(z\_\{t\+h\},u\_\{t\+h\}\)\-H\(z\_\{t\+h\},u\_\{t\+h\}\)\\\|\_\{2\}\+‖H​\(zt\+h,ut\+h\)−H​\(z^t\+h,ut\+h\)‖2\\displaystyle\\quad\+\\\|H\(z\_\{t\+h\},u\_\{t\+h\}\)\-H\(\\widehat\{z\}\_\{t\+h\},u\_\{t\+h\}\)\\\|\_\{2\}≤δ\+L​eh\.\\displaystyle\\leq\\delta\+Le\_\{h\}\.Iterating this recursion yields

eh≤Lh​e0\+δ​∑j=0h−1Lj\.e\_\{h\}\\leq L^\{h\}e\_\{0\}\+\\delta\\sum\_\{j=0\}^\{h\-1\}L^\{j\}\.Evaluating the geometric sum gives Equation \([26](https://arxiv.org/html/2608.04060#S4.E26)\)\.□\\square

#### B\.3\.2Proof of Theorem[6\.1](https://arxiv.org/html/2608.04060#S6.Thmdefinition1)

Under the conditionally independent Gaussian model in Equation \([30](https://arxiv.org/html/2608.04060#S6.E30)\) and the isotropic covariance assumptionΣ=σ2​I\\Sigma=\\sigma^\{2\}I, the negative log\-likelihood is

−log⁡p​\(𝒟Z∣ℰ,α,σ2​I\)\\displaystyle\-\\log p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,\\sigma^\{2\}I\)=n​dz2​log⁡\(2​π​σ2\)\\displaystyle=\\frac\{nd\_\{z\}\}\{2\}\\log\(2\\pi\\sigma^\{2\}\)\+12​σ2​∑i=1n‖zT\(i\)−Fℰ,α​\(zC\(i\),ε\(i\)\)‖22,\\displaystyle\\quad\+\\frac\{1\}\{2\\sigma^\{2\}\}\\sum\_\{i=1\}^\{n\}\\left\\\|z\_\{T\}^\{\(i\)\}\-F\_\{\\mathcal\{E\},\\alpha\}\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\\right\\\|\_\{2\}^\{2\},wheredzd\_\{z\}is the target\-embedding dimension\.

From the structure prior in Equation \([31](https://arxiv.org/html/2608.04060#S6.E31)\),

−log⁡p​\(ℰ\)=γ​Ω​\(ℰ\)\+const\.\-\\log p\(\\mathcal\{E\}\)=\\gamma\\Omega\(\\mathcal\{E\}\)\+\\mathrm\{const\}\.Bayes’ rule gives

−log⁡p​\(ℰ,α∣𝒟Z\)\\displaystyle\-\\log p\(\\mathcal\{E\},\\alpha\\mid\\mathcal\{D\}\_\{Z\}\)=−log⁡p​\(𝒟Z∣ℰ,α,σ2​I\)−log⁡p​\(α∣ℰ\)−log⁡p​\(ℰ\)\+const\\displaystyle=\-\\log p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,\\sigma^\{2\}I\)\-\\log p\(\\alpha\\mid\\mathcal\{E\}\)\-\\log p\(\\mathcal\{E\}\)\+\\mathrm\{const\}=12​σ2​∑i=1n‖zT\(i\)−Fℰ,α​\(zC\(i\),ε\(i\)\)‖22\\displaystyle=\\frac\{1\}\{2\\sigma^\{2\}\}\\sum\_\{i=1\}^\{n\}\\left\\\|z\_\{T\}^\{\(i\)\}\-F\_\{\\mathcal\{E\},\\alpha\}\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\)\\right\\\|\_\{2\}^\{2\}\+γ​Ω​\(ℰ\)−log⁡p​\(α∣ℰ\)\+const,\\displaystyle\\quad\+\\gamma\\Omega\(\\mathcal\{E\}\)\-\\log p\(\\alpha\\mid\\mathcal\{E\}\)\+\\mathrm\{const\},where the final constant is independent of\(ℰ,α\)\(\\mathcal\{E\},\\alpha\)\. Therefore, maximising the posterior is equivalent to minimising the objective in Equation \([34](https://arxiv.org/html/2608.04060#S6.E34)\)\.□\\square

## Appendix CSymbolic Model and Optimisation

### C\.1Grammar and Structural Complexity

This subsection describes a general expression\-tree realisation ofSJEPA\. The reported experiments instead use the fixed differentiable symbolic libraries specified in Appendix[D](https://arxiv.org/html/2608.04060#A4)\.

A grammar should be expressive enough to represent plausible mechanisms but small enough to make search and structural analysis meaningful\. A generic typed grammar is

ℰ::=\\displaystyle\\mathcal\{E\}::=\{\}α​∣zj∣​εk​∣ℰ\+ℰ∣​ℰ−ℰ∣ℰ×ℰ\\displaystyle\\alpha\\mid z\_\{j\}\\mid\\varepsilon\_\{k\}\\mid\\mathcal\{E\}\+\\mathcal\{E\}\\mid\\mathcal\{E\}\-\\mathcal\{E\}\\mid\\mathcal\{E\}\\times\\mathcal\{E\}∣safeDiv⁡\(ℰ,ℰ\)∣​sin⁡\(ℰ\)​∣cos⁡\(ℰ\)∣​expclip⁡\(ℰ\)∣logsafe⁡\(ℰ\)\.\\displaystyle\\mid\\operatorname\{safeDiv\}\(\\mathcal\{E\},\\mathcal\{E\}\)\\mid\\sin\(\\mathcal\{E\}\)\\mid\\cos\(\\mathcal\{E\}\)\\mid\\exp\_\{\\mathrm\{clip\}\}\(\\mathcal\{E\}\)\\mid\\log\_\{\\mathrm\{safe\}\}\(\\mathcal\{E\}\)\.Domain\-safe operators avoid undefined evaluations during search\. For vector dynamics, one expression is learned per target coordinate, with optional shared subexpressions\.

A weighted\-tree complexity may satisfy

wconstant<wlinear<wproduct<wdivision<wtranscendental,w\_\{\\mathrm\{constant\}\}<w\_\{\\mathrm\{linear\}\}<w\_\{\\mathrm\{product\}\}<w\_\{\\mathrm\{division\}\}<w\_\{\\mathrm\{transcendental\}\},or use a code length derived from operator frequencies\. Complexity weights must be fixed before test evaluation or tuned on validation data; otherwise, equation simplicity can be selected post hoc\.

For unrestricted expression\-tree search, equivalent expressions should be canonicalised before scoring\. Recommended steps include constant folding, commutative sorting, algebraic simplification, removal of neutral elements, coefficient thresholding, and numerical equivalence checks on held\-out points\. In the reported sparse\-library experiments, the feature ordering is fixed, and canonicalisation reduces primarily to coefficient thresholding and removal of inactive output–term pairs\.

### C\.2Differentiable Sparse\-Library Realisation

For discrete expression\-tree search,Ωtrain\\Omega^\{\\mathrm\{train\}\}may equal the structural complexityΩ​\(ℰ\)\\Omega\(\\mathcal\{E\}\)in Equation \([4](https://arxiv.org/html/2608.04060#S3.E4)\)\. For the differentiable sparse\-library implementation used in the experiments, we use

Ωtrain​\(ℰ,α\)=∑j=1dz∑k=1Pwk​αk​j2\+ξα,ξα\>0,\\Omega^\{\\mathrm\{train\}\}\(\\mathcal\{E\},\\alpha\)=\\sum\_\{j=1\}^\{d\_\{z\}\}\\sum\_\{k=1\}^\{P\}w\_\{k\}\\sqrt\{\\alpha\_\{kj\}^\{\\,2\}\+\\xi\_\{\\alpha\}\},\\qquad\\xi\_\{\\alpha\}\>0,\(67\)whereαk​j\\alpha\_\{kj\}is the coefficient of library termkkin output coordinatejj, andPPdenotes the number of candidate library terms\. This smooth quantity guides optimisation, while the thresholded support is used to compute the reported structural complexityΩ​\(ℰ^\)\\Omega\(\\widehat\{\\mathcal\{E\}\}\)\.

### C\.3Stage\-Specific Objectives

The practical training objective in Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) contains five common components\. Each optimisation stage restricts its trainable variables and omits terms that are constant or inapplicable\.

##### Representation warm start\.

Before symbolic search, the encoder and a neural predictor are trained using

ℒwarm=ℒpred\+λr​ℒroll\+λIB​ℛIB\.\\mathcal\{L\}\_\{\\mathrm\{warm\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{IB\}\}\\mathcal\{R\}\_\{\\mathrm\{IB\}\}\.\(68\)The target encoder is updated by exponential moving average\.

##### Dynamics search\.

With the encoders fixed, the dynamics phase minimises

ℒdyn=ℒpred\+λr​ℒroll\+λs​Ωtrain​\(ℰ,α\)\+λc​ℛcorr​\(ϕ\)\\mathcal\{L\}\_\{\\mathrm\{dyn\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\}\(\\mathcal\{E\},\\alpha\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)\(69\)over\(ℰ,α,ϕ\)\(\\mathcal\{E\},\\alpha,\\phi\)\.

##### Space search\.

With the discrete symbolic structure fixed or locally relaxed, the space phase minimises

ℒspace=ℒpred\+λr​ℒroll\+λIB​ℛIB\+λs​Ωtrain​\(ℰ,α\)\+λc​ℛcorr​\(ϕ\)\\mathcal\{L\}\_\{\\mathrm\{space\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{IB\}\}\\mathcal\{R\}\_\{\\mathrm\{IB\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\}\(\\mathcal\{E\},\\alpha\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\)\(70\)over the encoder and continuously optimised dynamics parameters\. When bothℰ\\mathcal\{E\}andα\\alphaare fixed, the symbolic\-complexity term is constant and may be omitted\. Whenα\\alphaor a relaxed support remains trainable, the term remains active\.

##### Frozen\-encoder search\.

For a permanently frozen encoder, the objective becomes

ℒfrozen=ℒpred\+λr​ℒroll\+λs​Ωtrain​\(ℰ,α\)\+λc​ℛcorr​\(ϕ\),\\mathcal\{L\}\_\{\\mathrm\{frozen\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\}\(\\mathcal\{E\},\\alpha\)\+\\lambda\_\{\\mathrm\{c\}\}\\mathcal\{R\}\_\{\\mathrm\{corr\}\}\(\\phi\),\(71\)optimised only over\(ℰ,α,ϕ\)\(\\mathcal\{E\},\\alpha,\\phi\)\.

### C\.4Relation Between Constrained and Penalised Formulations

For completeness, consider a finite\-dimensional convex relaxation with objectiveΓ​\(w\)\\Gamma\(w\)and convex constraintsgj​\(w\)≤0g\_\{j\}\(w\)\\leq 0\. If Slater’s condition holds, strong duality gives multipliersλj⋆≥0\\lambda\_\{j\}^\{\\star\}\\geq 0such that every primal optimum minimises

Γ​\(w\)\+∑jλj⋆​gj​\(w\)\.\\Gamma\(w\)\+\\sum\_\{j\}\\lambda\_\{j\}^\{\\star\}g\_\{j\}\(w\)\.This standard result motivates Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\)\. It does not establish global equivalence for neural encoders, discrete expression search, or nonconvex correction models\. In those settings, constraint\-aware search and Pareto reporting are safer than interpreting one penalty weight as canonical\.

Equation \([29](https://arxiv.org/html/2608.04060#S5.E29)\) should therefore be interpreted as a scalarisation of the constrained principle\. The regularisation weights control trade\-offs rather than recovering a unique constrained optimum\. Models should be assessed jointly using predictive error, rollout error, symbolic complexity, representation\-collapse diagnostics, and the normalised correction\-energy ratio\. Across multiple regularisation settings, this gives the empirical trade\-off

\(ℒpred,ℒroll,Ω​\(ℰ^\),ρcorr\),\\left\(\\mathcal\{L\}\_\{\\mathrm\{pred\}\},\\mathcal\{L\}\_\{\\mathrm\{roll\}\},\\Omega\(\\widehat\{\\mathcal\{E\}\}\),\\rho\_\{\\mathrm\{corr\}\}\\right\),subject to an admissible non\-collapsed representation\.

## Appendix DDetailed Experimental Setup and Additional Results

This section gives the complete implementation details for the two experiments in Section[8](https://arxiv.org/html/2608.04060#S8)\. Both experiments use the same continuous\-time simulator, residual\-predictor convention, trajectory counts, temporal windows, and validation\-based checkpoint procedure unless stated otherwise\.

### D\.1Experimental Protocol

##### Simulation and data splits\.

The physical state iss=\(q,p\)s=\(q,p\), with vector field

fκ​\(q,p\)=\[p−sin⁡\(q\)−κ​p​\|p\|\]\.f\_\{\\kappa\}\(q,p\)=\\begin\{bmatrix\}p\\\\ \-\\sin\(q\)\-\\kappa p\|p\|\\end\{bmatrix\}\.\(72\)Experiment 1 usesκ=0\\kappa=0, whereas Experiment 2 usesκ=0\.4\\kappa=0\.4\. Trajectories are generated using fourth\-order Runge–Kutta integration\. At every transition, the time step is sampled independently from

Δ​t∈\{0\.025,0\.04,0\.06\}\.\\Delta t\\in\\\{0\.025,0\.04,0\.06\\\}\.
Each experiment uses300300training trajectories,8080validation trajectories,8080test trajectories, and8080OOD trajectories, with100100transitions per trajectory\. The initial\-state ranges are listed in Table[10](https://arxiv.org/html/2608.04060#A4.T10)\.

Table 10:Initial\-state ranges used to generate trajectories\. Validation and test sets use independent trajectories from the same ranges\.The fixed dataset seeds are10011001,10041004,10021002, and10031003for the Experiment 1 training, validation, test, and OOD sets, respectively\. Experiment 2 uses20012001,20042004,20022002, and20032003\. Model optimisation is repeated with seeds

##### Experiment 1 observation map\.

Experiment 1 does not provide the physical state as named input coordinates\. Define the raw feature vector

ψ\(q,p\)=\[\\displaystyle\\psi\(q,p\)=\\big\[q,p,sin⁡q,cos⁡q,sin⁡\(2​q\),cos⁡\(2​q\),sin⁡p,cos⁡p,tanh⁡q,tanh⁡p,\\displaystyle q,p,\\sin q,\\cos q,\\sin\(2q\),\\cos\(2q\),\\sin p,\\cos p,\\tanh q,\\tanh p,qp,q2,p2,q3/6,p3/6,1,tanh\(Ws\+b\),sin\(Ws\+b\)\],\\displaystyle qp,q^\{2\},p^\{2\},q^\{3\}/6,p^\{3\}/6,1,\\tanh\(Ws\+b\),\\sin\(Ws\+b\)\\big\],whereW∈ℝ8×2W\\in\\mathbb\{R\}^\{8\\times 2\}andb∈ℝ8b\\in\\mathbb\{R\}^\{8\}are fixed random parameters\. The final observation is

X=Q​ψ​\(q,p\)\+ξ,ξ∼𝒩​\(0,0\.0032​I\),X=Q\\psi\(q,p\)\+\\xi,\\qquad\\xi\\sim\\mathcal\{N\}\(0,0\.003^\{2\}I\),\(73\)whereQ∈ℝ32×32Q\\in\\mathbb\{R\}^\{32\\times 32\}is a fixed orthogonal matrix\. The random observation map uses seed101101\. Every observation dimension is standardised using the mean and standard deviation of the training set only\.

Because the raw coordinatesqqandppare included inψ​\(q,p\)\\psi\(q,p\)andQQis orthogonal, the noiseless physical state is linearly recoverable from the complete observation\. The nonlinear terms act as mixed distractor and auxiliary features\. This controlled design is intended to study predictive\-coordinate and operator selection rather than nonlinear observability\.

### D\.2Model and Training Configuration

##### Architectures\.

The Experiment 1 encoder is

32⟶96⟶96⟶2,32\\longrightarrow 96\\longrightarrow 96\\longrightarrow 2,with SiLU activations in the hidden layers\. The target encoder is an exponential\-moving\-average copy with decay0\.990\.99\.

The neural vector field has two hidden layers of width9696with SiLU activations and bounded output

fψ​\(z\)=3​tanh⁡\(MLPψ⁡\(z\)\)\.f\_\{\\psi\}\(z\)=3\\tanh\\left\(\\operatorname\{MLP\}\_\{\\psi\}\(z\)\\right\)\.The correction network in Experiment 2 has two hidden layers of width4848, usestanh\\tanhactivations, and has bounded output

cϕ​\(z\)=tanh⁡\(MLPϕ⁡\(z\)\)\.c\_\{\\phi\}\(z\)=\\tanh\\left\(\\operatorname\{MLP\}\_\{\\phi\}\(z\)\\right\)\.
All predictors use the residual transition in Equation \([46](https://arxiv.org/html/2608.04060#S8.E46)\)\. Consequently, the zero vector field implements an identity transition:

z^t\+1=zt\.\\widehat\{z\}\_\{t\+1\}=z\_\{t\}\.If both encoders map every input to the same arbitrary constantz0z\_\{0\}, this identity transition predicts every target embedding perfectly\. The constant need not be the zero vector\.

##### Representation regularisation\.

The practical representation term is

ℛIBexp=2​ℒinv\+3​ℒvar\+0\.25​ℒcov\+0\.01​ℒmean\.\\mathcal\{R\}\_\{\\mathrm\{IB\}\}^\{\\mathrm\{exp\}\}=2\\mathcal\{L\}\_\{\\mathrm\{inv\}\}\+3\\mathcal\{L\}\_\{\\mathrm\{var\}\}\+0\.25\\mathcal\{L\}\_\{\\mathrm\{cov\}\}\+0\.01\\mathcal\{L\}\_\{\\mathrm\{mean\}\}\.The invariance term compares two observation\-noise augmentations with standard deviation0\.0120\.012\. The variance term imposes a unit\-scale coordinate\-variance target, the covariance term penalises off\-diagonal covariance, and the mean term weakly centres the representation\. The total representation term has unit weight in the practical objective\.

##### Symbolic libraries and coefficient initialisation\.

For Experiment 1, the symbolic feature library is

Θ​\(z\)=\[1,z1,z2,sin⁡z1,sin⁡z2,cos⁡z1,cos⁡z2,z12,z22,z1​z2\]\.\\Theta\(z\)=\\left\[1,z\_\{1\},z\_\{2\},\\sin z\_\{1\},\\sin z\_\{2\},\\cos z\_\{1\},\\cos z\_\{2\},z\_\{1\}^\{2\},z\_\{2\}^\{2\},z\_\{1\}z\_\{2\}\\right\]\.The corresponding complexity weights are

\(0\.5,1,1,2,2,2,2,1\.5,1\.5,2\)\.\(0\.5,1,1,2,2,2,2,1\.5,1\.5,2\)\.
Experiment 2 uses an output\-specific mask rather than permitting every feature in every equation\. The incomplete grammar is

q˙:\{p\},p˙:\{sin⁡\(q\)\},\\dot\{q\}:\\\{p\\\},\\qquad\\dot\{p\}:\\\{\\sin\(q\)\\\},whereas the complete grammar is

q˙:\{p\},p˙:\{sin⁡\(q\),p​\|p\|\}\.\\dot\{q\}:\\\{p\\\},\\qquad\\dot\{p\}:\\\{\\sin\(q\),p\|p\|\\\}\.The complexity weight ofp​\|p\|p\|p\|is33\. Linearppis not permitted in the equation forp˙\\dot\{p\}\.

Coefficients are initialised by sequentially thresholded ridge regression\. The ridge parameter is10−510^\{\-5\}and the STLSQ threshold is0\.0350\.035\. Continuous training uses

Ωtrain=∑j,kwk​αk​j2\+10−8\.\\Omega^\{\\mathrm\{train\}\}=\\sum\_\{j,k\}w\_\{k\}\\sqrt\{\\alpha\_\{kj\}^\{2\}\+10^\{\-8\}\}\.After each symbolic phase, coefficients with magnitude below0\.01750\.0175are set to zero\. Final equations and discrete complexity are reported using the stricter threshold

\|αk​j\|≥0\.05\.\|\\alpha\_\{kj\}\|\\geq 0\.05\.

##### Optimisation settings\.

Table 11:Optimisation settings used in the full three\-seed experiments\.The Experiment 1 warm neural JEPA is trained for at most130130epochs\. Neural JEPA is then continued for at most9090epochs\. Post\-hoc symbolic regression trains the symbolic dynamics for at most6565epochs with both encoders frozen\.

Joint symbolicSJEPAperforms four alternating cycles\. Each cycle consists of at most6565epochs of dynamics search with frozen encoders and at most8585epochs of representation\-space search\. The one\-step collapse diagnostic is trained for at most100100epochs using

ℒcollapse=ℒpred\+λs​Ωtrain,\\mathcal\{L\}\_\{\\mathrm\{collapse\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\},with neitherℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}norℒroll\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\. The constructive fixed\-point control is initialised at

z0=\(0\.65,−0\.35\)z\_\{0\}=\(0\.65,\-0\.35\)and trained for at most five epochs with the same one\-step objective\. Each Experiment 2 model is trained for at most260260epochs\.

### D\.3Validation and Evaluation

##### Checkpoint and cycle selection\.

Validation is performed every five epochs using at most sixteen validation batches\. Define

ℛval=ℒpredval\+0\.35​ℒrollval\.\\mathcal\{R\}\_\{\\mathrm\{val\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}^\{\\mathrm\{val\}\}\+0\.35\\mathcal\{L\}\_\{\\mathrm\{roll\}\}^\{\\mathrm\{val\}\}\.For Experiment 1 conditions that update the encoder, checkpoints failing the prescribed non\-collapse thresholds are ineligible whenever at least one eligible checkpoint is available\. The one\-step collapse and fixed\-point diagnostic conditions are exempt from this rule\. Eligible checkpoints are selected lexicographically:

1. 1\.lower validation predictive risk is preferred;
2. 2\.when two risks are within a relative tolerance of2%2\\%, lower discrete symbolic complexity is preferred;
3. 3\.the regularised validation objective and then the earlier epoch break any remaining tie\.

For checkpoint selection, a representation is considered non\-collapsed when

minj⁡Std⁡\(Zj\)≥0\.20andTr⁡\(Cov⁡Z\)≥0\.10\.\\min\_\{j\}\\operatorname\{Std\}\(Z\_\{j\}\)\\geq 0\.20\\qquad\\text\{and\}\\qquad\\operatorname\{Tr\}\(\\operatorname\{Cov\}Z\)\\geq 0\.10\.The one\-step collapse and explicit fixed\-point controls are selected without imposing this constraint because their purpose is to measure collapse\. Their observed collapse statistics are nevertheless recorded\.

The same validation rule is applied within every standard training phase, and the best completed alternating cycle is restored for final testing\. The selected Experiment 1 cycles are

c7⋆=3,c19⋆=1,c37⋆=4\.c^\{\\star\}\_\{7\}=3,\\qquad c^\{\\star\}\_\{19\}=1,\\qquad c^\{\\star\}\_\{37\}=4\.

##### Evaluation metrics\.

Experiment 1 reports:

- •latent one\-step MSE against the corresponding next\-state target embedding; end\-to\-end JEPA models use the EMA target embedding, whereas the post\-hoc symbolic baseline uses next\-state embeddings from the same frozen context encoder used to construct its transition dataset;
- •affine\-aligned physical\-state rollout MSE, using an affine map fitted only on training latents;
- •minimum latent coordinate standard deviation and covariance trace;
- •affine state\-probeR2R^\{2\};
- •thresholded symbolic complexity;
- •clipped rollout MSE and divergence rate\.

A rollout is marked divergent when the latent norm exceeds4040, the mapped state norm exceeds2525, or a non\-finite value occurs\. Diverged or missing errors are assigned the clipping value100100when computing the clipped mean\.

Experiment 2 reports the same state and rollout metrics directly in\(q,p\)\(q,p\)coordinates\. It uses the normalised correction\-energy ratio defined in Equation \([Eq\.64](https://arxiv.org/html/2608.04060#S8.Ex23)\), with expectations computed over independent test states\. The residual identity contributionztz\_\{t\}is excluded from the denominator because including it would make the diagnostic depend on the arbitrary scale of the state coordinates\.

For theppcoordinate, define the symbolic residual

rp​\(z\)=ftrue,p​\(z\)−Fℰ,α,p​\(z\)\.r\_\{p\}\(z\)=f\_\{\\mathrm\{true\},p\}\(z\)\-F\_\{\\mathcal\{E\},\\alpha,p\}\(z\)\.A scale\-only calibration

a⋆=arg​mina⁡𝔼train​\[\(rp​\(z\)−a​cϕ,p​\(z\)\)2\]a^\{\\star\}=\\operatorname\*\{arg\\,min\}\_\{a\}\\mathbb\{E\}\_\{\\mathrm\{train\}\}\\left\[\\left\(r\_\{p\}\(z\)\-ac\_\{\\phi,p\}\(z\)\\right\)^\{2\}\\right\]is fitted on training transitions\. Correlation and calibratedR2R^\{2\}are then evaluated on the independent test set\. The same procedure is applied to the isolated drag target

### D\.4Post\-Hoc Symbolic Baseline

The post\-hoc symbolic baseline separates representation learning from equation discovery\. It tests whether coordinates learned solely for neural prediction already admit a compact symbolic transition, without allowing symbolic simplicity to influence the representation itself\.

##### Stage 1: learning prediction\-oriented coordinates\.

We first train Neural JEPA by jointly optimising the context encoder and neural residual vector field\. Given an observationxtx\_\{t\}, the selected model produces

zt=Eθ⋆​\(xt\),z\_\{t\}=E\_\{\\theta^\{\\star\}\}\(x\_\{t\}\),\(74\)and predicts the next latent state according to

z^t\+1=zt\+Δ​tt​fψ⋆​\(zt\)\.\\widehat\{z\}\_\{t\+1\}=z\_\{t\}\+\\Delta t\_\{t\}f\_\{\\psi^\{\\star\}\}\(z\_\{t\}\)\.\(75\)The Neural JEPA checkpoint is selected using the held\-out validation risk defined in Equation \([48](https://arxiv.org/html/2608.04060#S8.E48)\)\. At this stage, the representation is optimised for predictive accuracy and representation quality, but not for the complexity of a subsequently fitted symbolic equation\.

##### Stage 2: freezing the learned coordinates\.

After model selection, the encoder parametersθ⋆\\theta^\{\\star\}are frozen\. The training trajectories are encoded to construct

𝒟Z=\{\(zt\(i\),zt\+1\(i\),Δ​tt\(i\)\)\}i=1N,\\mathcal\{D\}\_\{Z\}=\\left\\\{\\left\(z\_\{t\}^\{\(i\)\},z\_\{t\+1\}^\{\(i\)\},\\Delta t\_\{t\}^\{\(i\)\}\\right\)\\right\\\}\_\{i=1\}^\{N\},\(76\)where

zt\(i\)=Eθ⋆​\(xt\(i\)\),zt\+1\(i\)=Eθ⋆​\(xt\+1\(i\)\)\.z\_\{t\}^\{\(i\)\}=E\_\{\\theta^\{\\star\}\}\\left\(x\_\{t\}^\{\(i\)\}\\right\),\\qquad z\_\{t\+1\}^\{\(i\)\}=E\_\{\\theta^\{\\star\}\}\\left\(x\_\{t\+1\}^\{\(i\)\}\\right\)\.\(77\)Using the same frozen context encoder at both endpoints ensures that the fitted symbolic transition and its recursive rollouts remain in one latent coordinate system\. The EMA target encoder is used during Neural JEPA training but is not used to construct the post\-hoc symbolic transition dataset\. No gradients from the symbolic model are propagated into the frozen encoder\.

##### Stage 3: fitting symbolic dynamics\.

A symbolic vector field is fitted to the frozen latent transitions using the same candidate library and reporting threshold as the jointly learned symbolic model\. Its residual transition is

z^t\+1sym=zt\+Δ​tt​Fℰ,α​\(zt\)\.\\widehat\{z\}\_\{t\+1\}^\{\\mathrm\{sym\}\}=z\_\{t\}\+\\Delta t\_\{t\}F\_\{\\mathcal\{E\},\\alpha\}\(z\_\{t\}\)\.\(78\)The symbolic structureℰ\\mathcal\{E\}and coefficientsα\\alphaare learned by minimising

minℰ,αℒpred\+λr​ℒroll\+λs​Ωtrain\.\\min\_\{\\mathcal\{E\},\\alpha\}\\quad\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\lambda\_\{\\mathrm\{r\}\}\\mathcal\{L\}\_\{\\mathrm\{roll\}\}\+\\lambda\_\{\\mathrm\{s\}\}\\Omega^\{\\mathrm\{train\}\}\.\(79\)The representation regulariserℛIB\\mathcal\{R\}\_\{\\mathrm\{IB\}\}is absent because the encoder is frozen, and no neural correction is used\. Checkpoints are selected using the validation trajectories; among candidates with nearly tied validation risks, the lower\-complexity thresholded symbolic law is preferred\.

##### Evaluation\.

One\-step latent prediction error is evaluated directly in the frozen Neural JEPA coordinate system\. For physical\-state rollout evaluation, an affine map from the frozen training latents to\(q,p\)\(q,p\)is fitted using the training trajectories and then applied to test and OOD rollouts\. The physical\-state variables are used only for this evaluation alignment and are not provided to the symbolic regression procedure\.

The post\-hoc baseline therefore implements

learn predictive coordinates⟶freeze the coordinates⟶fit symbolic dynamics\.\\text\{learn predictive coordinates\}\\;\\longrightarrow\\;\\text\{freeze the coordinates\}\\;\\longrightarrow\\;\\text\{fit symbolic dynamics\}\.\(80\)By contrast,SJEPAalternates representation\-space and symbolic\-dynamics optimisation, allowing the latent coordinates to change so that the induced transition becomes simpler while maintaining predictive adequacy and the prescribed non\-collapse constraints\.

### D\.5Additional Experimental Results

#### D\.5\.1Per\-seed equations for Experiment 1

After applying the fixed reporting threshold, the selectedSJEPAequations are

Seed 7:z˙1\\displaystyle\\text\{Seed 7:\}\\qquad\\dot\{z\}\_\{1\}=−0\.806​z2,\\displaystyle=\-0\.806z\_\{2\},z˙2\\displaystyle\\dot\{z\}\_\{2\}=0\.822​z1\+0\.177−0\.143​z12,\\displaystyle=0\.822z\_\{1\}\+0\.177\-0\.143z\_\{1\}^\{2\},Seed 19:z˙1\\displaystyle\\text\{Seed 19:\}\\qquad\\dot\{z\}\_\{1\}=−0\.847​z2\+0\.146−0\.085​z12,\\displaystyle=\-0\.847z\_\{2\}\+0\.146\-0\.085z\_\{1\}^\{2\},z˙2\\displaystyle\\dot\{z\}\_\{2\}=0\.811​z1\+0\.185−0\.143​z12\+0\.116​z1​z2,\\displaystyle=0\.811z\_\{1\}\+0\.185\-0\.143z\_\{1\}^\{2\}\+0\.116z\_\{1\}z\_\{2\},Seed 37:z˙1\\displaystyle\\text\{Seed 37:\}\\qquad\\dot\{z\}\_\{1\}=0\.867​z2,\\displaystyle=0\.867z\_\{2\},z˙2\\displaystyle\\dot\{z\}\_\{2\}=−0\.835​z1\.\\displaystyle=\-0\.835z\_\{1\}\.
The sign reversal in seed3737reflects the non\-identifiability of latent orientation rather than a different qualitative mechanism\. The cross\-coordinate oscillator terms are selected in all three seeds\. By contrast, the post\-hoc symbolic models select most of the available polynomial and trigonometric library, with mean weighted complexity

#### D\.5\.2Representation and alignment metrics

Table 12:Additional Experiment 1 representation and alignment metrics\.The nonzero affine probe score of the one\-step collapse diagnostic does not contradict near\-collapse\. An affine map can amplify very small state\-correlated variations\. The decisive collapse diagnostics are the absolute coordinate standard deviation and covariance trace\.

#### D\.5\.3Coefficient and support stability in Experiment 2

Every permitted symbolic term is selected in all three Experiment 2 seeds\. The fitted coefficients have extremely small variation:

Incomplete symbolic:q˙\\displaystyle\\text\{Incomplete symbolic:\}\\qquad\\dot\{q\}=\(0\.9607±0\.0017\)​p,\\displaystyle=\(0\.9607\\pm 0\.0017\)p,p˙\\displaystyle\\dot\{p\}=−\(0\.8382±0\.0008\)​sin⁡\(q\),\\displaystyle=\-\(0\.8382\\pm 0\.0008\)\\sin\(q\),Regularised hybrid:q˙\\displaystyle\\text\{Regularised hybrid:\}\\qquad\\dot\{q\}=\(0\.9581±0\.0008\)​p,\\displaystyle=\(0\.9581\\pm 0\.0008\)p,p˙\\displaystyle\\dot\{p\}=−\(0\.7942±0\.0024\)​sin⁡\(q\),\\displaystyle=\-\(0\.7942\\pm 0\.0024\)\\sin\(q\),Unregularised hybrid:q˙\\displaystyle\\text\{Unregularised hybrid:\}\\qquad\\dot\{q\}=\(0\.3324±0\.0065\)​p,\\displaystyle=\(0\.3324\\pm 0\.0065\)p,p˙\\displaystyle\\dot\{p\}=−\(0\.1718±0\.0746\)​sin⁡\(q\),\\displaystyle=\-\(0\.1718\\pm 0\.0746\)\\sin\(q\),Complete symbolic:q˙\\displaystyle\\text\{Complete symbolic:\}\\qquad\\dot\{q\}=\(0\.9830±0\.0005\)​p,\\displaystyle=\(0\.9830\\pm 0\.0005\)p,p˙\\displaystyle\\dot\{p\}=−\(0\.9553±0\.0010\)​sin⁡\(q\)−\(0\.3724±0\.0007\)​p​\|p\|\.\\displaystyle=\-\(0\.9553\\pm 0\.0010\)\\sin\(q\)\-\(0\.3724\\pm 0\.0007\)p\|p\|\.
The complete grammar therefore recovers both the structure and coefficients of the true vector field consistently\. Under the incomplete grammar, correction regularisation preserves the symbolic pendulum mechanism; without it, the correction absorbs much of both equations\.

#### D\.5\.4Qualitative and mechanism\-specific diagnostics

Figure[4](https://arxiv.org/html/2608.04060#A4.F4)collects two secondary diagnostics\. Panel \(a\) illustrates the geometry of one selectedSJEPArepresentation in Experiment 1, while panel \(b\) evaluates the drag\-specific neural correction in Experiment 2\. The colour progression follows a curved latent manifold rather than either coordinate axis, suggesting that the true angleqqis encoded nonlinearly across both latent dimensions\. This visual evidence is qualitative; affine alignment is assessed quantitatively using the probeR2R^\{2\}\.

The correction is primarily evaluated against the complete residual

rp=ftrue,p−Fℰ,α,p,r\_\{p\}=f\_\{\\mathrm\{true\},p\}\-F\_\{\\mathcal\{E\},\\alpha,p\},\(81\)because this residual contains both the omitted drag and any coefficient error remaining in the fitted symbolic law\. Figure[4](https://arxiv.org/html/2608.04060#A4.F4)\(b\) additionally compares the calibrated correction with the isolated quadratic\-drag term\. This is therefore a secondary mechanism\-specific diagnostic rather than the primary correction target\.

![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment1_latent_coordinates_v4.png)

\(a\) ExampleSJEPAlatent geometry

![Refer to caption](https://arxiv.org/html/2608.04060v1/experiment2_correction_vs_drag_v4.png)

\(b\) Drag\-specific correction diagnostic

Figure 4:Additional qualitative and mechanism\-specific diagnostics\.\(a\)Nonlinear latent geometry in Experiment 1\. Each point is an SJEPA embedding\(z1,z2\)\(z\_\{1\},z\_\{2\}\)coloured by the true pendulum angleqq\. The smooth colour ordering indicates that information aboutqqis retained, while the curved geometry shows that the learned coordinates are not simply affinely aligned with the original state\.\(b\)Drag\-specific diagnostic for Experiment 2 using the regularised hybrid model with the incomplete grammar, whose symbolic component containsq˙∝p\\dot\{q\}\\propto pandp˙∝−sin⁡\(q\)\\dot\{p\}\\propto\-\\sin\(q\)\. The panel compares the calibratedpp\-coordinate correction on held\-out states with the omitted quadratic\-drag term−0\.4​p​\|p\|\-0\.4p\|p\|\. By contrast, Figure[3](https://arxiv.org/html/2608.04060#S8.F3)\(b\) compares the same correction with the complete residual left by the fitted symbolic law, which also includes error arising from imperfect estimation of thesin⁡\(q\)\\sin\(q\)coefficient\. The dashed line denotes exact recovery\.

## Appendix EBayesian Inference Details

This section provides additional details for the Bayesian formulation in the main text\. The encoders remain deterministic throughout: Bayesian uncertainty is introduced only over the latent transition law, its coefficients, transition covariance, and, in the hybrid model, the Gaussian\-process correction\.

### E\.1Finite\-Candidate Symbolic Inference

Exact inference over all symbolic expression trees is generally intractable\. Possible schemes include reversible\-jump Markov chain Monte Carlo, sequential Monte Carlo over expression trees, nested sampling, and Bayesian inference over a finite candidate set\. A practical implementation forSJEPAuses symbolic search as a proposal mechanism and subsequently performs Bayesian model comparison over the resulting candidate structures\.

Let

𝒮M=\{ℰm\}m=1M\\mathcal\{S\}\_\{M\}=\\left\\\{\\mathcal\{E\}\_\{m\}\\right\\\}\_\{m=1\}^\{M\}denote a finite candidate set generated using candidate\-generation data𝒟gen\\mathcal\{D\}\_\{\\mathrm\{gen\}\}\. Candidates may be obtained from sparse\-library search, genetic programming, beam search, or another symbolic\-discovery procedure\. Algebraically equivalent or numerically duplicate expressions are canonicalised and removed\. To preserve diversity, candidates may be retained from the predictive\-error–complexity Pareto set rather than selecting only the lowest\-error expressions\.

A disjoint evidence set𝒟evid\\mathcal\{D\}\_\{\\mathrm\{evid\}\}is used for Bayesian fitting and model comparison\. For each candidate,

p​\(𝒟evid∣ℰm\)=∫p​\(𝒟evid∣ℰm,α,Σ\)​p​\(α∣ℰm\)​p​\(Σ\)​dα​dΣ\.p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)=\\int p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\},\\alpha,\\Sigma\)p\(\\alpha\\mid\\mathcal\{E\}\_\{m\}\)p\(\\Sigma\)\\mathrm\{d\}\\alpha\\,\\mathrm\{d\}\\Sigma\.\(82\)The evidence integrates over symbolic coefficients and transition covariance rather than evaluating only one fitted parameter value\. It therefore accounts for parameter uncertainty and penalises structures whose apparent fit depends on a narrowly tuned or unnecessarily flexible parameterisation\.

The integral in Equation \([82](https://arxiv.org/html/2608.04060#A5.E82)\) is generally unavailable in closed form\. It may be estimated using posterior sampling, nested sampling, variational inference, or a Laplace approximation\. For the latter, let

ϑm=\(αm,λm\),Σm=L​\(λm\)​L​\(λm\)⊤,\\vartheta\_\{m\}=\\left\(\\alpha\_\{m\},\\lambda\_\{m\}\\right\),\\qquad\\Sigma\_\{m\}=L\(\\lambda\_\{m\}\)L\(\\lambda\_\{m\}\)^\{\\top\},\(83\)whereλm\\lambda\_\{m\}contains unconstrained parameters defining a lower\-triangular Cholesky factorL​\(λm\)L\(\\lambda\_\{m\}\)with positive diagonal entries\. Thus,Σm\\Sigma\_\{m\}remains positive definite throughout optimisation\.

Let

ϑ^m=arg​maxϑm⁡p​\(ϑm∣𝒟evid,ℰm\)\\widehat\{\\vartheta\}\_\{m\}=\\operatorname\*\{arg\\,max\}\_\{\\vartheta\_\{m\}\}p\(\\vartheta\_\{m\}\\mid\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{m\}\)\(84\)be the posterior mode, and let

Hm=−∇ϑm2log⁡p​\(𝒟evid,ϑm∣ℰm\)\|ϑm=ϑ^mH\_\{m\}=\-\\nabla\_\{\\vartheta\_\{m\}\}^\{2\}\\log p\\left\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\vartheta\_\{m\}\\mid\\mathcal\{E\}\_\{m\}\\right\)\\bigg\|\_\{\\vartheta\_\{m\}=\\widehat\{\\vartheta\}\_\{m\}\}\(85\)be the negative Hessian of the log joint density at that mode\. Ifdm=dim\(ϑm\)d\_\{m\}=\\dim\(\\vartheta\_\{m\}\)andHmH\_\{m\}is positive definite, the Laplace approximation gives

log⁡p​\(𝒟evid∣ℰm\)≈\\displaystyle\\log p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)\\approx\{\}log⁡p​\(𝒟evid∣ℰm,ϑ^m\)\+log⁡p​\(ϑ^m∣ℰm\)\\displaystyle\\log p\\left\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\},\\widehat\{\\vartheta\}\_\{m\}\\right\)\+\\log p\\left\(\\widehat\{\\vartheta\}\_\{m\}\\mid\\mathcal\{E\}\_\{m\}\\right\)\+dm2​log⁡\(2​π\)−12​log⁡\|Hm\|\.\\displaystyle\+\\frac\{d\_\{m\}\}\{2\}\\log\(2\\pi\)\-\\frac\{1\}\{2\}\\log\|H\_\{m\}\|\.\(86\)The log\-determinant term accounts for posterior concentration: a structure whose good fit is confined to a narrow parameter region need not receive the same evidence as one that explains the data over a larger plausible region\.

Within the selected candidate set, define

p​\(ℰm∣𝒮M\)∝exp⁡\{−γ​Ω​\(ℰm\)\}\.p\(\\mathcal\{E\}\_\{m\}\\mid\\mathcal\{S\}\_\{M\}\)\\propto\\exp\\left\\\{\-\\gamma\\Omega\(\\mathcal\{E\}\_\{m\}\)\\right\\\}\.\(87\)The posterior structure probabilities are

p​\(ℰm∣𝒟evid,𝒮M\)=p​\(𝒟evid∣ℰm\)​p​\(ℰm∣𝒮M\)∑r=1Mp​\(𝒟evid∣ℰr\)​p​\(ℰr∣𝒮M\)\.p\(\\mathcal\{E\}\_\{m\}\\mid\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{S\}\_\{M\}\)=\\frac\{p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)p\(\\mathcal\{E\}\_\{m\}\\mid\\mathcal\{S\}\_\{M\}\)\}\{\\displaystyle\\sum\_\{r=1\}^\{M\}p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{r\}\)p\(\\mathcal\{E\}\_\{r\}\\mid\\mathcal\{S\}\_\{M\}\)\}\.\(88\)For numerical stability, define

am=log⁡p​\(𝒟evid∣ℰm\)−γ​Ω​\(ℰm\)\.a\_\{m\}=\\log p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)\-\\gamma\\Omega\(\\mathcal\{E\}\_\{m\}\)\.\(89\)The normalised weight is

wm=exp⁡\(am−amax\)∑r=1Mexp⁡\(ar−amax\),amax=max1≤r≤M⁡ar\.w\_\{m\}=\\frac\{\\exp\(a\_\{m\}\-a\_\{\\max\}\)\}\{\\displaystyle\\sum\_\{r=1\}^\{M\}\\exp\(a\_\{r\}\-a\_\{\\max\}\)\},\\qquad a\_\{\\max\}=\\max\_\{1\\leq r\\leq M\}a\_\{r\}\.\(90\)
For each candidate, the structure\-conditional posterior predictive distribution is

p​\(zT∣zC,ε,𝒟evid,ℰm\)=∫p​\(zT∣zC,ε,ℰm,α,Σ\)​p​\(α,Σ∣𝒟evid,ℰm\)​dα​dΣ\.p\\left\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{m\}\\right\)=\\int p\\left\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{E\}\_\{m\},\\alpha,\\Sigma\\right\)p\\left\(\\alpha,\\Sigma\\mid\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{m\}\\right\)\\mathrm\{d\}\\alpha\\,\\mathrm\{d\}\\Sigma\.\(91\)Bayesian model averaging gives

p​\(zT∣zC,ε,𝒟evid,𝒮M\)=∑m=1Mwm​p​\(zT∣zC,ε,𝒟evid,ℰm\)\.p\\left\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{S\}\_\{M\}\\right\)=\\sum\_\{m=1\}^\{M\}w\_\{m\}p\\left\(z\_\{T\}\\mid z\_\{C\},\\varepsilon,\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{m\}\\right\)\.\(92\)
A Monte Carlo approximation first draws

J\(b\)∼Categorical⁡\(w1,…,wM\),J^\{\(b\)\}\\sim\\operatorname\{Categorical\}\(w\_\{1\},\\ldots,w\_\{M\}\),\(93\)then draws

\(α\(b\),Σ\(b\)\)∼p​\(α,Σ∣𝒟evid,ℰJ\(b\)\),\\left\(\\alpha^\{\(b\)\},\\Sigma^\{\(b\)\}\\right\)\\sim p\\left\(\\alpha,\\Sigma\\mid\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{J^\{\(b\)\}\}\\right\),\(94\)and finally samples or evaluates the corresponding latent transition\.

The symbolic\-search stage constructs a finite, data\-dependent support over which Bayesian model comparison is performed; its original search score is not itself treated as a posterior probability\. The resulting weights are approximate posterior probabilities conditional on the selected candidate set𝒮M\\mathcal\{S\}\_\{M\}, rather than the exact posterior over the full grammar\. Hyperparameters such asγ\\gamma, prior scales, andMMmay be chosen using validation data, while test and OOD data are reserved for final posterior\-predictive evaluation\. The encoders remain point\-estimated and are not included in the Bayesian averaging\.

Algorithm 2Finite\-candidate Bayesian symbolic inference0:Fixed encoders, candidate\-generation data

𝒟gen\\mathcal\{D\}\_\{\\mathrm\{gen\}\}, evidence data

𝒟evid\\mathcal\{D\}\_\{\\mathrm\{evid\}\}, symbolic grammar

𝔈\\mathfrak\{E\}
1:Generate candidate structures using

𝒟gen\\mathcal\{D\}\_\{\\mathrm\{gen\}\}
2:Canonicalise expressions and remove duplicate structures

3:Retain

MMcandidates from the predictive\-error–complexity Pareto set evaluated on

𝒟gen\\mathcal\{D\}\_\{\\mathrm\{gen\}\}
4:for

m=1,…,Mm=1,\\ldots,Mdo

5:Infer

p​\(α,Σ∣𝒟evid,ℰm\)p\(\\alpha,\\Sigma\\mid\\mathcal\{D\}\_\{\\mathrm\{evid\}\},\\mathcal\{E\}\_\{m\}\)
6:Estimate

log⁡p​\(𝒟evid∣ℰm\)\\log p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)
7:Set

am←log⁡p​\(𝒟evid∣ℰm\)−γ​Ω​\(ℰm\)a\_\{m\}\\leftarrow\\log p\(\\mathcal\{D\}\_\{\\mathrm\{evid\}\}\\mid\\mathcal\{E\}\_\{m\}\)\-\\gamma\\Omega\(\\mathcal\{E\}\_\{m\}\)
8:endfor

9:Normalise

\{am\}m=1M\\\{a\_\{m\}\\\}\_\{m=1\}^\{M\}using Equation \([90](https://arxiv.org/html/2608.04060#A5.E90)\)

10:Form the posterior predictive distribution using Equation \([92](https://arxiv.org/html/2608.04060#A5.E92)\)

11:returnCandidate structures, posterior weights, and posterior predictive model

### E\.2Gaussian\-Process Hybrid Inference

##### Output\-wise GP correction\.

For a fixed symbolic structureℰ\\mathcal\{E\}, coefficientsα\\alpha, and latent output coordinatejj, define

si=\(zC\(i\),ε\(i\)\),mi​j=Fℰ,α,j​\(si\),ri​j=zT,j\(i\)−mi​j\.s\_\{i\}=\\left\(z\_\{C\}^\{\(i\)\},\\varepsilon^\{\(i\)\}\\right\),\\qquad m\_\{ij\}=F\_\{\\mathcal\{E\},\\alpha,j\}\(s\_\{i\}\),\\qquad r\_\{ij\}=z\_\{T,j\}^\{\(i\)\}\-m\_\{ij\}\.Let

gj∼𝒢​𝒫​\(0,kϑg,j\),g\_\{j\}\\sim\\mathcal\{GP\}\\left\(0,k\_\{\\vartheta\_\{g,j\}\}\\right\),\(95\)and define

\[Kj\]a​b=kϑg,j​\(sa,sb\),Cj=Kj\+σj2​In\.\[K\_\{j\}\]\_\{ab\}=k\_\{\\vartheta\_\{g,j\}\}\(s\_\{a\},s\_\{b\}\),\\qquad C\_\{j\}=K\_\{j\}\+\\sigma\_\{j\}^\{2\}I\_\{n\}\.\(96\)The residual vector

𝒓j=\(r1​j,…,rn​j\)⊤\\bm\{r\}\_\{j\}=\(r\_\{1j\},\\ldots,r\_\{nj\}\)^\{\\top\}has the marginal distribution

𝒓j∣ℰ,α,ϑg,j,σj2∼𝒩​\(𝟎,Cj\)\.\\bm\{r\}\_\{j\}\\mid\\mathcal\{E\},\\alpha,\\vartheta\_\{g,j\},\\sigma\_\{j\}^\{2\}\\sim\\mathcal\{N\}\\left\(\\bm\{0\},C\_\{j\}\\right\)\.\(97\)Consequently, its negative log marginal likelihood is

−log⁡p​\(𝒓j∣ℰ,α,ϑg,j,σj2\)=12​𝒓j⊤​Cj−1​𝒓j\+12​log⁡\|Cj\|\+n2​log⁡\(2​π\)\.\-\\log p\\left\(\\bm\{r\}\_\{j\}\\mid\\mathcal\{E\},\\alpha,\\vartheta\_\{g,j\},\\sigma\_\{j\}^\{2\}\\right\)=\\frac\{1\}\{2\}\\bm\{r\}\_\{j\}^\{\\top\}C\_\{j\}^\{\-1\}\\bm\{r\}\_\{j\}\+\\frac\{1\}\{2\}\\log\|C\_\{j\}\|\+\\frac\{n\}\{2\}\\log\(2\\pi\)\.\(98\)The quadratic term measures how well the GP explains systematic residual structure left by the symbolic law\. The log\-determinant term controls the covariance flexibility used to explain the data\. Their balance concerns residual fit and GP complexity, but does not by itself identify a unique symbolic–GP allocation\.

##### Joint symbolic–GP posterior\.

Letg=\(g1,…,gdz\)g=\(g\_\{1\},\\ldots,g\_\{d\_\{z\}\}\)and letϑg\\vartheta\_\{g\}collect the GP hyperparameters\. The joint posterior has the schematic form

p​\(ℰ,α,g,ϑg,Σ∣𝒟Z\)∝p​\(𝒟Z∣ℰ,α,g,Σ\)​p​\(g∣ϑg\)​p​\(ϑg\)​p​\(α∣ℰ\)​p​\(Σ\)​p​\(ℰ\)\.p\(\\mathcal\{E\},\\alpha,g,\\vartheta\_\{g\},\\Sigma\\mid\\mathcal\{D\}\_\{Z\}\)\\propto p\(\\mathcal\{D\}\_\{Z\}\\mid\\mathcal\{E\},\\alpha,g,\\Sigma\)p\(g\\mid\\vartheta\_\{g\}\)p\(\\vartheta\_\{g\}\)p\(\\alpha\\mid\\mathcal\{E\}\)p\(\\Sigma\)p\(\\mathcal\{E\}\)\.\(99\)Uncertainty inℰ\\mathcal\{E\}concerns symbolic structure; uncertainty inα\\alphaconcerns its numerical coefficients; uncertainty inggconcerns systematic residual dynamics; andΣ\\Sigmarepresents transition variability remaining after conditioning on the symbolic law and correction\. The symbolic–GP allocation remains dependent on the structure prior, coefficient priors, GP kernel and amplitude priors, and transition\-noise prior\.

Integrating outgggives the GP marginal likelihood in Equation \([98](https://arxiv.org/html/2608.04060#A5.E98)\)\. One may then infer or average overℰ\\mathcal\{E\},α\\alpha, kernel hyperparameters, andΣ\\Sigmausing posterior sampling, variational approximations, Laplace approximations, or a finite candidate ensemble\.

##### Posterior correction at a new input\.

For

s⋆=\(zC,ε\),s\_\{\\star\}=\(z\_\{C\},\\varepsilon\),define

𝒌⋆,j=\(kϑg,j​\(s⋆,s1\),…,kϑg,j​\(s⋆,sn\)\)⊤,k⋆⋆,j=kϑg,j​\(s⋆,s⋆\)\.\\bm\{k\}\_\{\\star,j\}=\\left\(k\_\{\\vartheta\_\{g,j\}\}\(s\_\{\\star\},s\_\{1\}\),\\ldots,k\_\{\\vartheta\_\{g,j\}\}\(s\_\{\\star\},s\_\{n\}\)\\right\)^\{\\top\},\\qquad k\_\{\\star\\star,j\}=k\_\{\\vartheta\_\{g,j\}\}\(s\_\{\\star\},s\_\{\\star\}\)\.Conditional on fixed symbolic and GP parameters,

gj​\(s⋆\)∣𝒟Z,ℰ,α,ϑg,j,σj2∼𝒩​\(μgj​\(s⋆\),vgj​\(s⋆\)\),g\_\{j\}\(s\_\{\\star\}\)\\mid\\mathcal\{D\}\_\{Z\},\\mathcal\{E\},\\alpha,\\vartheta\_\{g,j\},\\sigma\_\{j\}^\{2\}\\sim\\mathcal\{N\}\\left\(\\mu\_\{g\_\{j\}\}\(s\_\{\\star\}\),v\_\{g\_\{j\}\}\(s\_\{\\star\}\)\\right\),\(100\)where

μgj​\(s⋆\)\\displaystyle\\mu\_\{g\_\{j\}\}\(s\_\{\\star\}\)=𝒌⋆,j⊤​Cj−1​𝒓j,\\displaystyle=\\bm\{k\}\_\{\\star,j\}^\{\\top\}C\_\{j\}^\{\-1\}\\bm\{r\}\_\{j\},\(101\)vgj​\(s⋆\)\\displaystyle v\_\{g\_\{j\}\}\(s\_\{\\star\}\)=k⋆⋆,j−𝒌⋆,j⊤​Cj−1​𝒌⋆,j\.\\displaystyle=k\_\{\\star\\star,j\}\-\\bm\{k\}\_\{\\star,j\}^\{\\top\}C\_\{j\}^\{\-1\}\\bm\{k\}\_\{\\star,j\}\.\(102\)The posterior mean supplies the data\-supported correction to the symbolic law, while the posterior variance typically becomes larger in regions weakly supported by the latent transition data, subject to the selected kernel and hyperparameters\.

For vector\-valued target embeddings, one may use independent output\-wise GPs, a shared kernel with output\-specific parameters, or a matrix\-valued kernel that models dependence between latent coordinates\. Independent output\-wise models are computationally simplest, whereas multi\-output kernels can represent correlated transition uncertainty\.

### E\.3Posterior Predictive Rollouts

Posterior predictive rollouts propagate uncertainty recursively through the learned latent dynamics\. For rollout samplebb, first draw one coherent transition model:

\(ℰ\(b\),α\(b\),g\(b\),ϑg\(b\),Σ\(b\)\)∼p​\(ℰ,α,g,ϑg,Σ∣𝒟Z\)\.\\left\(\\mathcal\{E\}^\{\(b\)\},\\alpha^\{\(b\)\},g^\{\(b\)\},\\vartheta\_\{g\}^\{\(b\)\},\\Sigma^\{\(b\)\}\\right\)\\sim p\\left\(\\mathcal\{E\},\\alpha,g,\\vartheta\_\{g\},\\Sigma\\mid\\mathcal\{D\}\_\{Z\}\\right\)\.\(103\)Starting fromzt\(b\)=ztz\_\{t\}^\{\(b\)\}=z\_\{t\}, recursively sample

zt\+h\+1\(b\)∼𝒩​\(Fℰ\(b\),α\(b\)​\(zt\+h\(b\),εt\+h\)\+g\(b\)​\(zt\+h\(b\),εt\+h\),Σ\(b\)\),z\_\{t\+h\+1\}^\{\(b\)\}\\sim\\mathcal\{N\}\\left\(F\_\{\\mathcal\{E\}^\{\(b\)\},\\alpha^\{\(b\)\}\}\\left\(z\_\{t\+h\}^\{\(b\)\},\\varepsilon\_\{t\+h\}\\right\)\+g^\{\(b\)\}\\left\(z\_\{t\+h\}^\{\(b\)\},\\varepsilon\_\{t\+h\}\\right\),\\Sigma^\{\(b\)\}\\right\),\(104\)forh=0,…,K−1h=0,\\ldots,K\-1\.

Equation \([104](https://arxiv.org/html/2608.04060#A5.E104)\) gives the direct\-transition formulation\. For a vector\-field realisation, its mean is replaced by the corresponding integrated transition, for example

zt\+h\(b\)\+Δ​tt\+h​\[Fℰ\(b\),α\(b\)​\(zt\+h\(b\),εt\+h\)\+g\(b\)​\(zt\+h\(b\),εt\+h\)\]z\_\{t\+h\}^\{\(b\)\}\+\\Delta t\_\{t\+h\}\\left\[F\_\{\\mathcal\{E\}^\{\(b\)\},\\alpha^\{\(b\)\}\}\\left\(z\_\{t\+h\}^\{\(b\)\},\\varepsilon\_\{t\+h\}\\right\)\+g^\{\(b\)\}\\left\(z\_\{t\+h\}^\{\(b\)\},\\varepsilon\_\{t\+h\}\\right\)\\right\]under forward Euler\.

The same sampled structure, coefficients, and GP function are retained throughout a rollout\. This preserves the interpretation of epistemic uncertainty as uncertainty about one underlying transition model\. Resampling the symbolic law or GP independently at every step would instead introduce artificial temporal variation in the model itself\. Transition noise may still be sampled separately at each step because it represents aleatoric variability conditional on the sampled dynamics\.

Repeating Equations \([103](https://arxiv.org/html/2608.04060#A5.E103)\)–\([104](https://arxiv.org/html/2608.04060#A5.E104)\) produces an empirical approximation to

p​\(zt\+1:t\+K∣zt,εt:t\+K−1,𝒟Z\)\.p\\left\(z\_\{t\+1:t\+K\}\\mid z\_\{t\},\\varepsilon\_\{t:t\+K\-1\},\\mathcal\{D\}\_\{Z\}\\right\)\.\(105\)The resulting distribution reflects uncertainty over symbolic structure, coefficients, systematic correction, transition covariance, and the uncertain states visited during recursion\.

Calibration should be evaluated separately for one\-step and multi\-step prediction\. Even a well\-calibrated one\-step model can become miscalibrated over long horizons because uncertain states are repeatedly fed back into the transition, model misspecification compounds, and trajectories may enter regions poorly represented in the training data\. Appropriate diagnostics include horizon\-wise log predictive density, empirical interval coverage, calibration error, and task\-level decision quality under posterior\-predictive planning\.

Similar Articles

The Annotated JEPA

Hacker News Top

A step-by-step annotated implementation and explanation of Joint Embedding Predictive Architectures (JEPA) for self-supervised learning, covering I-JEPA, V-JEPA, and LeJEPA.

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

arXiv cs.LG

This paper audits Joint-embedding predictive architectures (JEPA) for LLM fine-tuning on a natural-language-to-regex task, testing twenty-two auxiliary objectives. The results show that hidden-state representation improvements are only weakly coupled to decoded-task accuracy, with no auxiliary surviving family-wise correction.