Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
Summary
The paper establishes a theoretical connection between probabilistic Joint-Embedding Predictive Learning (JEPA) and Hidden Markov Models (HMMs), providing a state-space interpretation and introducing Markov-Chain JEPA for enhanced consistency.
View Cached Full Text
Cached at: 08/17/26, 09:46 AM
# Your Probabilistic JEPA Is Secretly a Hidden Markov Model A State-Space Interpretation of Joint-Embedding Predictive Learning
Source: [https://arxiv.org/html/2608.13621](https://arxiv.org/html/2608.13621)
Yongchao Huangyongchao\.huang@abdn\.ac\.uk
###### Abstract
A hidden Markov model \(HMM\) combines three roles: inference of a hidden\-state belief from observations, propagation through a Markov transition, and emission back to observation space\. We show that full, time\-indexed Predictive Information Bottleneck VJEPA \(PIB\-VJEPA\) exposes the same computational structure: a stochastic context encoder plays the role of an amortized filtering distribution, a probabilistic predictor defines latent\-state dynamics, and a decoder, inverse target encoder, or induced implicit conditional supplies the emission direction\. We distinguish 4 progressively stronger levels of correspondence and give sufficient conditions for exact sequence\-level HMM equivalence\. To make the connection concrete, we introduce Markov\-Chain JEPA \(MCJEPA\), which replaces the latent predictor by a learned transition matrix; in the finite time\-homogeneous case, matrix powers guarantee exact multi\-horizon Chapman–Kolmogorov consistency\. Conditioned discrete\-state transitions, continuous\-state Markov kernels, and continuous\-time dynamics extend this construction, while deterministic temporal JEPA appears as a degenerate Dirac\-kernel special case\. We further interpret predictive information\-bottleneck learning as seeking a compact predictive state: compression promotesminimality, while residual predictability testssufficiency\. Controlled experiments support transition composition, the filtering interpretation, predictive Markovization in a known synthetic process, and the distinction between JEPA latent prediction and HMM\-style sequence learning\. Together, these results give temporal JEPA a principled state\-space interpretation\.
###### Contents
1. [1Introduction](https://arxiv.org/html/2608.13621#S1)
2. [2From JEPA to Markov\-Chain JEPA](https://arxiv.org/html/2608.13621#S2)1. [2\.1Latent Markov dynamics](https://arxiv.org/html/2608.13621#S2.SS1) 2. [2\.2The HMM–PIB\-VJEPA analogy at a glance](https://arxiv.org/html/2608.13621#S2.SS2) 3. [2\.3Temporal JEPA](https://arxiv.org/html/2608.13621#S2.SS3) 4. [2\.4MCJEPA: replace the predictor by a Markov chain](https://arxiv.org/html/2608.13621#S2.SS4) 5. [2\.5Exact multi\-horizon path consistency](https://arxiv.org/html/2608.13621#S2.SS5) 6. [2\.6Avoiding discrete\-state collapse](https://arxiv.org/html/2608.13621#S2.SS6)
3. [3From a Transition Matrix to a Neural Markov Kernel](https://arxiv.org/html/2608.13621#S3)1. [3\.1Neural discrete\-state transitions](https://arxiv.org/html/2608.13621#S3.SS1) 2. [3\.2Continuous\-state stochastic transitions](https://arxiv.org/html/2608.13621#S3.SS2) 3. [3\.3Continuous\-time extensions](https://arxiv.org/html/2608.13621#S3.SS3) 4. [3\.4Deterministic temporal JEPA as a degenerate kernel](https://arxiv.org/html/2608.13621#S3.SS4) 5. [3\.5When a recurrent predictor is Markov](https://arxiv.org/html/2608.13621#S3.SS5)
4. [4A Probabilistic JEPA Is Secretly an HMM](https://arxiv.org/html/2608.13621#S4)1. [4\.1The three distributions in PIB\-VJEPA](https://arxiv.org/html/2608.13621#S4.SS1) 2. [4\.2The encode–transition–emit correspondence](https://arxiv.org/html/2608.13621#S4.SS2) 3. [4\.3When no decoder is present: an implicit emission](https://arxiv.org/html/2608.13621#S4.SS3) 4. [4\.4Four levels of correspondence](https://arxiv.org/html/2608.13621#S4.SS4)
5. [5Information Bottleneck Learning as Markovization](https://arxiv.org/html/2608.13621#S5)
6. [6Residual Predictability as a Diagnostic of Markov Sufficiency](https://arxiv.org/html/2608.13621#S6)1. [6\.1Categorical probability innovations](https://arxiv.org/html/2608.13621#S6.SS1) 2. [6\.2Testing incremental predictability from older history](https://arxiv.org/html/2608.13621#S6.SS2) 3. [6\.3Distinguishing state insufficiency from predictor misspecification](https://arxiv.org/html/2608.13621#S6.SS3) 4. [6\.4Continuous\-state diagnostics](https://arxiv.org/html/2608.13621#S6.SS4)
7. [7Experiments](https://arxiv.org/html/2608.13621#S7)1. [7\.1Experiment 1: finite\-HMM recovery and Markov composition](https://arxiv.org/html/2608.13621#S7.SS1) 2. [7\.2Experiment 2: filtering resolves emission ambiguity](https://arxiv.org/html/2608.13621#S7.SS2) 3. [7\.3Experiment 3: predictive compression and Markovization](https://arxiv.org/html/2608.13621#S7.SS3) 4. [7\.4Experiment 4: HMM\-style training of PIB\-VJEPA](https://arxiv.org/html/2608.13621#S7.SS4) 5. [7\.5Summary of Experiments](https://arxiv.org/html/2608.13621#S7.SS5)
8. [8Discussion](https://arxiv.org/html/2608.13621#S8)
9. [9Conclusion](https://arxiv.org/html/2608.13621#S9)
10. [References](https://arxiv.org/html/2608.13621#bib)
11. [ATraining Objectives and Minimal Algorithm](https://arxiv.org/html/2608.13621#A1)
12. [BProofs](https://arxiv.org/html/2608.13621#A2)1. [B\.1Proof of Proposition](https://arxiv.org/html/2608.13621#A2.SS1) 2. [B\.2Proof of Proposition](https://arxiv.org/html/2608.13621#A2.SS2) 3. [B\.3Proof of Proposition](https://arxiv.org/html/2608.13621#A2.SS3)
13. [CThree Realizations of the Emission Direction](https://arxiv.org/html/2608.13621#A3)1. [C\.1Explicit decoder](https://arxiv.org/html/2608.13621#A3.SS1) 2. [C\.2Invertible target encoder](https://arxiv.org/html/2608.13621#A3.SS2) 3. [C\.3Implicit emission](https://arxiv.org/html/2608.13621#A3.SS3)
14. [DHMM\-Style Training Objectives for PIB\-VJEPA](https://arxiv.org/html/2608.13621#A4)1. [D\.1Sequence likelihood](https://arxiv.org/html/2608.13621#A4.SS1) 2. [D\.2Exact likelihood and filtering for categorical MCJEPA](https://arxiv.org/html/2608.13621#A4.SS2) 3. [D\.3Continuous\-state extension](https://arxiv.org/html/2608.13621#A4.SS3) 4. [D\.4Relation to JEPA and hybrid training](https://arxiv.org/html/2608.13621#A4.SS4)
15. [EExact HMM Representation Conditions](https://arxiv.org/html/2608.13621#A5)1. [E\.1Marginal and filtering consistency](https://arxiv.org/html/2608.13621#A5.SS1) 2. [E\.2Explicit and invertible emissions](https://arxiv.org/html/2608.13621#A5.SS2) 3. [E\.3Implicit\-emission case](https://arxiv.org/html/2608.13621#A5.SS3) 4. [E\.4Completion of the proof for Theorem\.](https://arxiv.org/html/2608.13621#A5.SS4)
16. [FExperimental Details](https://arxiv.org/html/2608.13621#A6)1. [F\.1Evaluation metrics](https://arxiv.org/html/2608.13621#A6.SS1) 2. [F\.2Experiment 1: finite\-HMM recovery and Markov composition](https://arxiv.org/html/2608.13621#A6.SS2) 3. [F\.3Experiment 2: filtering under emission ambiguity](https://arxiv.org/html/2608.13621#A6.SS3) 4. [F\.4Experiment 3: predictive compression and Markovization](https://arxiv.org/html/2608.13621#A6.SS4) 5. [F\.5Experiment 4: HMM\-style training of PIB\-VJEPA](https://arxiv.org/html/2608.13621#A6.SS5) 6. [F\.6Reproducibility and computation](https://arxiv.org/html/2608.13621#A6.SS6)
## 1Introduction
Joint\-Embedding Predictive Architectures \(JEPAs\) learn by predicting a target representation from an observed context rather than reconstructing the target observation itself\([10](https://arxiv.org/html/2608.13621#bib.bib5);[2](https://arxiv.org/html/2608.13621#bib.bib6);[3](https://arxiv.org/html/2608.13621#bib.bib7)\)\. Variational JEPA \(VJEPA\) makes this prediction probabilistic, replacing a point predictor with a conditional distribution over future latent states\([9](https://arxiv.org/html/2608.13621#bib.bib9)\)\. A full Predictive Information Bottleneck \(PIB\) extension additionally makes the current representation stochastic and explicitly controls how much information it retains about the observed history\([8](https://arxiv.org/html/2608.13621#bib.bib10)\)\.
This progression creates a natural question:*what familiar probabilistic model is hidden inside a fully stochastic temporal JEPA?*The basic state\-space analogy is immediate\. An HMM follows
X≤t⟶p\(St∣X≤t\)⟶p\(St\+1∣St\)⟶p\(Xt\+1∣St\+1\),X\_\{\\leq t\}\\longrightarrow p\(S\_\{t\}\\mid X\_\{\\leq t\}\)\\longrightarrow p\(S\_\{t\+1\}\\mid S\_\{t\}\)\\longrightarrow p\(X\_\{t\+1\}\\mid S\_\{t\+1\}\),\(1\)whereas full PIB\-VJEPA, when equipped with an explicit observation model, has the corresponding pipeline
X≤t⟶qθ\(Zt∣X≤t\)⟶pϕ\(Zt\+1∣Zt,ξt\)⟶pψ\(Xt\+1∣Zt\+1\)\.X\_\{\\leq t\}\\longrightarrow q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\longrightarrow p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\\longrightarrow p\_\{\\psi\}\(X\_\{t\+1\}\\mid Z\_\{t\+1\}\)\.\(2\)Here,XtX\_\{t\}denotes observation\-level data such as pixels, video frames, or time\-series measurements;ZtZ\_\{t\}denotes a stochastic predictive representation; andξt\\xi\_\{t\}contains side information such as an action, elapsed time, target position, or exogenous covariates\([9](https://arxiv.org/html/2608.13621#bib.bib9)\)\.
The observation\-space path in[Eq\.2](https://arxiv.org/html/2608.13621#S1.E2)is optional for the core JEPA objective\. It may be implemented by an explicit probabilistic decoderpψ\(Xt\+1∣Zt\+1\)p\_\{\\psi\}\(X\_\{t\+1\}\\mid Z\_\{t\+1\}\)\. Alternatively, if the target encoderfθ¯f\_\{\\bar\{\\theta\}\}is invertible on the modeled data domain, its inverse provides a deterministic state\-to\-observation map,
X^t\+1=fθ¯−1\(Z^t\+1\)\.\\widehat\{X\}\_\{t\+1\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(\\widehat\{Z\}\_\{t\+1\}\)\.\(3\)A reliable approximate inverse may similarly support reconstruction or forecasting, although it does not by itself define the normalized emission likelihood required for exact HMM\-style likelihood training\. Under these interpretations, observation\-level data correspond to HMM observations, the stochastic context encoder plays the hidden\-state inference role, the predictor propagates the latent state, and the decoder or inverse target encoder realizes the state\-to\-observation direction\.
[Figure1](https://arxiv.org/html/2608.13621#S1.F1)combines two complementary views of this analogy, and[Table1](https://arxiv.org/html/2608.13621#S1.T1)summarizes the correspondence component by component\. The figure separates two directions that are often conflated\. In an HMM, the emission distribution maps a hidden state to an observation, whereas filtering111Here,*filtering*is used in the state\-space sense: it denotes inference of the current hidden\-state beliefp\(St∣X≤t\)p\(S\_\{t\}\\mid X\_\{\\leq t\}\)from the observations available up to timett\. This belief is obtained recursively by combining the transition\-based prediction from the previous state with the evidence provided by the current observation\. It should not be interpreted only as noise removal, although a learned JEPA encoder may also suppress observation\-level noise or other prediction\-irrelevant variation\.maps observations to an inferred hidden state\. Likewise, a PIB\-VJEPA encoder defines a recognition or state\-inference distribution rather than an emission model\. Under the filtering\-consistency conditions developed later, the history\-dependent context encoder coincides with the corresponding HMM filtering distribution\. The direct emission analogue is instead the decoderpψ\(Xt∣Zt\)p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)or, when available, the inverse target encoderfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}\. The target encoder itself maps observations to latent states; only its inverse has the state\-to\-observation direction of an HMM emission\.
Hidden Markov modelX1X\_\{1\}X2X\_\{2\}X3X\_\{3\}S1S\_\{1\}S2S\_\{2\}S3S\_\{3\}p\(S2∣S1\)p\(S\_\{2\}\\mid S\_\{1\}\)p\(S3∣S2\)p\(S\_\{3\}\\mid S\_\{2\}\)p\(X1∣S1\)p\(X\_\{1\}\\mid S\_\{1\}\)p\(X2∣S2\)p\(X\_\{2\}\\mid S\_\{2\}\)p\(X3∣S3\)p\(X\_\{3\}\\mid S\_\{3\}\)observationshidden MarkovchainX≤tX\_\{\\leq t\}p\(St∣X≤t\)p\(S\_\{t\}\\mid X\_\{\\leq t\}\)p\(St\+1∣St\)p\(S\_\{t\+1\}\\mid S\_\{t\}\)p\(Xt\+1∣St\+1\)p\(X\_\{t\+1\}\\mid S\_\{t\+1\}\)observation historystate inferencestate transitionemissionFull PIB\-VJEPAX1X\_\{1\}X2X\_\{2\}X3X\_\{3\}Z1Z\_\{1\}Z2Z\_\{2\}Z3Z\_\{3\}pϕ\(Z2∣Z1,ξ1\)p\_\{\\phi\}\(Z\_\{2\}\\mid Z\_\{1\},\\xi\_\{1\}\)pϕ\(Z3∣Z2,ξ2\)p\_\{\\phi\}\(Z\_\{3\}\\mid Z\_\{2\},\\xi\_\{2\}\)qθq\_\{\\theta\}qθ¯q\_\{\\bar\{\\theta\}\}qθ¯q\_\{\\bar\{\\theta\}\}pψp\_\{\\psi\}orfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}pψp\_\{\\psi\}orfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}pψp\_\{\\psi\}orfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}observation\-level datapredictive latentstate processX≤tX\_\{\\leq t\}qθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)pϕ\(Zt\+1∣Zt,ξt\)p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)pψ\(Xt\+1∣Zt\+1\)p\_\{\\psi\}\(X\_\{t\+1\}\\mid Z\_\{t\+1\}\)orfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}observation historystochastic encoderlatent predictorobservation map
Figure 1:Unified graphical and computational comparison of an HMM and full PIB\-VJEPA\.Top:the classical HMM view, in which an unobserved Markov chain emits observations\.Bottom:the corresponding PIB\-VJEPA view, in which stochastic encoders infer predictive latent states from observation\-level data and a probabilistic predictor propagates those states\. Solid downward arrows in the PIB\-VJEPA diagram denote observation\-to\-state encoding, while dashed upward arrows denote the optional state\-to\-observation map implemented by an explicit decoder or an inverse target encoder\. The boxed rows reproduce the aligned computational pipelines in[Eqs\.1](https://arxiv.org/html/2608.13621#S1.E1)and[2](https://arxiv.org/html/2608.13621#S1.E2)\.Table 1:Component\-wise analogy between an HMM and a full PIB\-VJEPA\.The subtlety is therefore not the high\-level architecture, but its precise probabilistic and training interpretation\. Standard JEPA training matches predicted future representations to stop\-gradient target\-encoder representations and need not optimize an observation likelihood\. An explicit decoder or inverse target encoder can complete the observation\-space prediction path, but the existence of such a map does not by itself make the JEPA objective identical to HMM maximum likelihood\. Conversely, when neither explicit observation map is present, a local stochastic encoder can still induce an implicit emission distribution, although that induced conditional need not be tractable for generation or likelihood evaluation\.
#### Scope and terminology\.
JEPA denotes the general joint\-embedding predictive framework, while VJEPA denotes its probabilistic latent\-prediction formulation\([9](https://arxiv.org/html/2608.13621#bib.bib9)\)\. Our main object of study is the*full, time\-indexed PIB\-VJEPA*\([8](https://arxiv.org/html/2608.13621#bib.bib10)\), in which the current representation, future target representation, and latent transition are all probabilistic\. This formulation makes the state\-space structure most explicit and therefore provides the cleanest setting in which to develop the HMM correspondence\. We use*PIB\-VJEPA*throughout the main probabilistic development, referring to JEPA or VJEPA when discussing the broader architectural family or relevant special cases\. The latent Markov perspective is not restricted to PIB\-VJEPA: other probabilistic temporal JEPA formulations inherit the same encode–transition structure when their context and target representations are connected by a probabilistic latent predictor, while classical deterministic temporal JEPA is recovered as the degenerate case in which the relevant latent distributions collapse to point masses and the predictor becomes a Dirac transition kernel\. The sufficient\-condition result developed later is stated for probabilistic temporal JEPA more generally, with full PIB\-VJEPA providing the principal concrete realization\. Exact sequence\-level HMM equivalence nevertheless requires the additional emission and consistency conditions developed later\.
Our main claim is therefore:
> Full, time\-indexed PIB\-VJEPA exposes the encode–transition–emit structure of a hidden Markov model\. Its stochastic context encoder plays the hidden\-state inference or filtering role, its probabilistic predictor defines the latent\-state transition, and an explicit decoder, an invertible target encoder, or an induced implicit conditional supplies the emission direction\. This structural correspondence does not by itself imply that PIB\-VJEPA defines the same sequence distribution or is trained by the same objective as an HMM\. We therefore distinguish four progressively stronger levels: computational correspondence, emission\-complete latent\-state representation, sequence\-level HMM equivalence, and model\-and\-objective equivalence\. Sequence\-level HMM equivalence additionally requires Markov, emission, marginal\-consistency, and filtering\-consistency conditions, while model\-and\-objective equivalence further requires the corresponding sequence\-level probabilistic objective to participate in training\.
We develop this claim in four progressive steps222The main text presents the construction and central claims directly; detailed derivations and proofs are deferred to appendices\.:
1. 1\.We introduce*Markov\-Chain JEPA*\(MCJEPA\), in which a learned transition matrix replaces the usual latent\-space predictor, and show that its matrix powers guarantee exact Chapman\-\-Kolmogorov consistency between direct and composed multi\-step predictions333The acronym MC\-JEPA has previously been used for Motion\-and\-Content JEPA\([4](https://arxiv.org/html/2608.13621#bib.bib8)\)\. We use MCJEPA here as shorthand for*Markov\-Chain JEPA*, an unhyphenated, general\-purpose, task\-independent JEPA variant\.\.
2. 2\.We generalize the transition matrix to conditioned discrete transitions, continuous\-state Markov kernels, and continuous\-time dynamics, with deterministic temporal JEPA recovered as a degenerate Dirac\-kernel special case\.
3. 3\.We formalize the HMM correspondence through explicit\-decoder, inverse\-target\-encoder, and implicit\-emission constructions; distinguish progressively stronger levels of correspondence; and give sufficient conditions for an exact sequence\-level HMM representation\.
4. 4\.We interpret predictive information\-bottleneck learning as Markov\-state construction, separating*minimality*through predictive compression from*sufficiency*through residual predictability, and examine how JEPA latent prediction, hybrid JEPA–HMM learning, and HMM\-style sequence learning impose different probabilistic semantics on the same latent\-state architecture\.
We evaluate these claims in 4 controlled experiments \(Section\.[7](https://arxiv.org/html/2608.13621#S7)\) that respectively examine finite\-state transition recovery and path consistency, the filtering interpretation of the context encoder, predictive compression and Markovization, and the objective\-level distinction between JEPA latent prediction and HMM\-style sequence learning\.
## 2From JEPA to Markov\-Chain JEPA
### 2\.1Latent Markov dynamics
Before constructing Markov\-Chain JEPA, it is useful to locate it within the broader family of Markov models\. Markov dynamics can be organized along two independent axes: whether time is discrete or continuous, and whether the state space is discrete or continuous\. These choices give four common cases:
1. 1\.*Discrete time and discrete state:*the transition law is represented by a row\-stochastic transition matrix\. This is the setting adopted by the basic MCJEPA construction developed below\.
2. 2\.*Discrete time and continuous state:*the transition law is represented by a conditional probability density or, more generally, a Markov kernel over continuous latent representations\.
3. 3\.*Continuous time and discrete state:*the dynamics form a continuous\-time Markov chain specified by a transition\-rate generator\.
4. 4\.*Continuous time and continuous state:*the dynamics may be represented by a stochastic differential equation or another continuous\-time Markov process\.
We use*latent Markov dynamics*as an umbrella term for these 4 cases\. More specifically,*Markov chain*commonly refers to a discrete\-state process, whereas*Markov process*or*Markov kernel*also covers continuous\-state models\. Our main development focuses on discrete\-time prediction because temporal JEPA training is typically organized around frames, tokens, or measurements indexed by discrete steps\. We begin with the discrete\-time, discrete\-state case because it yields the most transparent connection to a classical HMM transition matrix\. We then relax the fixed\-matrix and discrete\-state assumptions using conditioned transition matrices and continuous\-state kernels\. Continuous\-time variants are included later to situate the framework more broadly and to accommodate irregularly sampled systems\.
### 2\.2The HMM–PIB\-VJEPA analogy at a glance
The construction is easiest to understand through the common encode–transition–emit pipeline summarized in[Table2](https://arxiv.org/html/2608.13621#S2.T2)\. In both an HMM and full PIB\-VJEPA, observations are used to infer a belief over a latent state, the state is advanced by a transition model, and the predicted state may be mapped back to observation space\. Compared with a classical HMM, PIB\-VJEPA primarily changes how these roles are parameterized and trained: state inference is amortized by an encoder, future\-state supervision is supplied by a target encoder, and the principal predictive objective is imposed in representation space rather than through an observation\-sequence likelihood\.
Table 2:Direct correspondence between an HMM and full PIB\-VJEPA\. Observation\-level inputs correspond to HMM observations, the stochastic context encoder plays the hidden\-state inference role, the predictor defines the latent transition, and an explicit decoder, inverse target encoder, or induced implicit conditional can supply the emission direction\.We now instantiate this correspondence for full PIB\-VJEPA\([8](https://arxiv.org/html/2608.13621#bib.bib10)\)\. Its stochastic context encoder
qθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)represents the current latent\-state belief, while its probabilistic predictor
pϕ\(Zt\+1∣Zt,ξt\)p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)defines the latent\-state transition\. The target encoder
qθ¯\(Zt\+1∣Xt\+1\)q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+1\}\\mid X\_\{t\+1\}\)provides the future latent distribution against which the prediction is trained\. The identification of the history\-dependent context encoder with an HMM filtering distribution becomes exact only under the filtering\-consistency conditions developed later\.
To complete the encode–transition–emit path explicitly, PIB\-VJEPA may additionally use a probabilistic observation model
pψ\(Xt\+1∣Zt\+1\)\.p\_\{\\psi\}\(X\_\{t\+1\}\\mid Z\_\{t\+1\}\)\.The resulting one\-step predictive observation distribution is
pΘ\(Xt\+1∣X≤t\)=∫∫\\displaystyle p\_\{\\Theta\}\(X\_\{t\+1\}\\mid X\_\{\\leq t\}\)=\\int\\\!\\\!\\intpψ\(Xt\+1∣Zt\+1\)pϕ\(Zt\+1∣Zt,ξt\)\\displaystyle p\_\{\\psi\}\(X\_\{t\+1\}\\mid Z\_\{t\+1\}\)\\,p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\(4\)×qθ\(Zt∣X≤t\)dZtdZt\+1,\\displaystyle\\times q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\,dZ\_\{t\}\\,dZ\_\{t\+1\},whereΘ=\(θ,ϕ,ψ\)\\Theta=\(\\theta,\\phi,\\psi\)\. This conditional implements the same operational sequence as HMM prediction: infer a current latent\-state belief from the observation history, propagate it through the transition model, and map the predicted state to the next observation\. When the context encoder coincides with the corresponding Bayesian filter, this becomes the usual HMM predictive construction\.
As an alternative to introducing a separate decoder, PIB\-VJEPA may use the inverse target\-encoder construction introduced in[Eq\.3](https://arxiv.org/html/2608.13621#S1.E3)\. If
fθ¯:𝒳→𝒮f\_\{\\bar\{\\theta\}\}:\\mathcal\{X\}\\rightarrow\\mathcal\{S\}is bijective on the modeled data domain, then
X^t\+1=fθ¯−1\(Z^t\+1\)\\widehat\{X\}\_\{t\+1\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(\\widehat\{Z\}\_\{t\+1\}\)provides a deterministic state\-to\-observation map without requiring a separate decoder\. A normalized emission likelihood additionally requires a tractable density, for example through a change\-of\-variables model or an explicit observation\-noise distribution\. Compatible dimensions and invertibility are strong conditions that ordinary compressed JEPA encoders generally do not satisfy, so we treat inverse target encoding as an alternative realization of the emission direction rather than a universal requirement of PIB\-VJEPA\.
### 2\.3Temporal JEPA
LetX≤tX\_\{\\leq t\}denote observations up to timett, and letXt\+hX\_\{t\+h\}denote a future observation or target segment\. A temporal JEPA uses an online encoder, an EMA target encoder, and a latent predictor\([9](https://arxiv.org/html/2608.13621#bib.bib9)\):
Zt=fθ\(X≤t\),Zt\+hT=fθ¯\(Xt\+h\),Z^t\+h=Pϕ\(Zt,h,ξt:t\+h\),Z\_\{t\}=f\_\{\\theta\}\(X\_\{\\leq t\}\),\\qquad Z^\{\\mathrm\{T\}\}\_\{t\+h\}=f\_\{\\bar\{\\theta\}\}\(X\_\{t\+h\}\),\\qquad\\widehat\{Z\}\_\{t\+h\}=P\_\{\\phi\}\(Z\_\{t\},h,\\xi\_\{t:t\+h\}\),whereξ\\xicontains side information such as elapsed time, action, target position, or known covariates\. Training matchesZ^t\+h\\widehat\{Z\}\_\{t\+h\}to the stop\-gradient targetZt\+hTZ^\{\\mathrm\{T\}\}\_\{t\+h\}\.
### 2\.4MCJEPA: replace the predictor by a Markov chain
We now specialize the JEPA latent stateZtZ\_\{t\}, previously allowed to be a general continuous representation, to a categorical predictive state,
Zt∈\{1,…,K\}\.Z\_\{t\}\\in\\\{1,\\ldots,K\\\}\.We retainZtZ\_\{t\}for the JEPA state and reserveStS\_\{t\}for the corresponding hidden state in the HMM notation\.
The online encoder returns a soft state distribution
qt=qθ\(Zt∣X≤t\)∈ΔK−1,q\_\{t\}=q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\in\\Delta^\{K\-1\},and the EMA target encoder returns
q¯t\+h=qθ¯\(Zt\+h∣Xt\+h\)∈ΔK−1\.\\bar\{q\}\_\{t\+h\}=q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\)\\in\\Delta^\{K\-1\}\.
The predictor is a constant, time\-homogeneous, row\-stochastic transition matrix
A∈\[0,1\]K×K,∑j=1KAij=1\.A\\in\[0,1\]^\{K\\times K\},\\qquad\\sum\_\{j=1\}^\{K\}A\_\{ij\}=1\.\(5\)Its entries represent the categorical JEPA transition probabilities
Aij=pϕ\(Zt\+1=j∣Zt=i\)\.A\_\{ij\}=p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i\)\.\(6\)
The one\-step andhh\-step predictive state distributions are444The linear\-algebra interpretation is straightforward\. The current state beliefqt∈ℝ1×Kq\_\{t\}\\in\\mathbb\{R\}^\{1\\times K\}is a row vector whoseiith entry,\(qt\)i=qθ\(Zt=i∣X≤t\),\(q\_\{t\}\)\_\{i\}=q\_\{\\theta\}\(Z\_\{t\}=i\\mid X\_\{\\leq t\}\),is the probability assigned to current stateii\. The transition matrixA∈ℝK×KA\\in\\mathbb\{R\}^\{K\\times K\}is row stochastic, withAij=pϕ\(Zt\+1=j∣Zt=i\),A\_\{ij\}=p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i\),so itsiith row is the next\-state distribution conditional on currently occupying stateii\. Their productq^t\+1=qtA∈ℝ1×K\\widehat\{q\}\_\{t\+1\}=q\_\{t\}A\\in\\mathbb\{R\}^\{1\\times K\}is again a probability row vector, and itsjjth entry is\(q^t\+1\)j=\(qtA\)j=∑i=1K\(qt\)iAij\.\(\\widehat\{q\}\_\{t\+1\}\)\_\{j\}=\(q\_\{t\}A\)\_\{j\}=\\sum\_\{i=1\}^\{K\}\(q\_\{t\}\)\_\{i\}A\_\{ij\}\.Thus, the predicted probability of being in statejjat timet\+1t\+1is obtained by summing, over all possible current statesii, the probability of currently being in stateiimultiplied by the probability of transitioning fromiitojj\. Equivalently,qtAq\_\{t\}Ais a convex combination of the rows ofAA\([14](https://arxiv.org/html/2608.13621#bib.bib12);[6](https://arxiv.org/html/2608.13621#bib.bib13)\), weighted by the current state beliefqtq\_\{t\}\. Becauseqtq\_\{t\}is normalized andAAis row stochastic, the resulting vector remains normalized\. The same interpretation applies toqtAhq\_\{t\}A^\{h\}, where\(Ah\)ij\(A^\{h\}\)\_\{ij\}is the total probability of reaching statejjafterhhsteps when starting from stateii, obtained by summing the probabilities of all length\-hhpaths through the possible intermediate states\. Consequently, the resulting predictive distribution depends only on the total horizonhh, not on how that horizon is decomposed into successive transition steps\.
q^t\+1=qtA,q^t\+h=qtAh\.\\widehat\{q\}\_\{t\+1\}=q\_\{t\}A,\\qquad\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\}\.\(7\)
We train the transition matrix and online encoder by matching each predicted distribution to the corresponding target\-encoder distribution:
ℒMC=𝔼\[∑h∈ℋwhKL\(sg\(q¯t\+h\)∥qtAh\)\],\\mathcal\{L\}\_\{\\mathrm\{MC\}\}=\\mathbb\{E\}\\left\[\\sum\_\{h\\in\\mathcal\{H\}\}w\_\{h\}\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+h\}\)\\,\\middle\\\|\\,q\_\{t\}A^\{h\}\\right\)\\right\],\(8\)whereℋ\\mathcal\{H\}is the set of prediction horizons andwh≥0w\_\{h\}\\geq 0controls the relative importance of horizonhh\. For example, one may use uniform weights,
wh=1\|ℋ\|,w\_\{h\}=\\frac\{1\}\{\|\\mathcal\{H\}\|\},or exponentially discounted weights,
wh=γh−1∑r∈ℋγr−1,γ∈\(0,1\],w\_\{h\}=\\frac\{\\gamma^\{h\-1\}\}\{\\sum\_\{r\\in\\mathcal\{H\}\}\\gamma^\{r\-1\}\},\\qquad\\gamma\\in\(0,1\],which place greater emphasis on near\-term predictions;γ=1\\gamma=1recovers uniform weighting\.
The target branch remains a slowly moving representation target\. The core JEPA objective does not require observation reconstruction, although an explicit decoder or inverse target\-encoder map may be used to map the predicted state back toXt\+hX\_\{t\+h\}for observation forecasting or HMM\-style likelihood modeling\.
X≤tX\_\{\\leq t\}stochasticencoderqtq\_\{t\}transitionAAqtAhq\_\{t\}A^\{h\}optional decoderorfθ¯−1f\_\{\\bar\{\\theta\}\}^\{\-1\}X^t\+h\\widehat\{X\}\_\{t\+h\}q¯t\+h\\bar\{q\}\_\{t\+h\}EMA targetencoderXt\+hX\_\{t\+h\}hhstepslatent loss
Figure 2:Markov\-Chain JEPA\. The context encoder infers a predictive\-state distribution, the transition matrix propagates it to future latent states, and training matches the prediction to an EMA target distribution\. The dashed observation branch is optional: a decoder or inverse target encoder can map the predicted latent state back to observation space when observation forecasting or an explicit emission model is required\.
### 2\.5Exact multi\-horizon path consistency
For temporal JEPA, adirectprediction fromtttot\+h1\+h2t\+h\_\{1\}\+h\_\{2\}agrees exactly with a prediction composed through theintermediatehorizon:
qtAh1\+h2=\(qtAh1\)Ah2\.q\_\{t\}A^\{h\_\{1\}\+h\_\{2\}\}=\(q\_\{t\}A^\{h\_\{1\}\}\)A^\{h\_\{2\}\}\.\(9\)This is the Chapman\-\-Kolmogorov law555The Chapman–Kolmogorov equation states that a transition across two consecutive intervals is obtained by marginalizing over every possible intermediate state:Ph1\+h2\(i,k\)=∑j=1KPh1\(i,j\)Ph2\(j,k\)\.P\_\{h\_\{1\}\+h\_\{2\}\}\(i,k\)=\\sum\_\{j=1\}^\{K\}P\_\{h\_\{1\}\}\(i,j\)P\_\{h\_\{2\}\}\(j,k\)\.For a time\-homogeneous chain,Ph=AhP\_\{h\}=A^\{h\}, soAh1\+h2=Ah1Ah2A^\{h\_\{1\}\+h\_\{2\}\}=A^\{h\_\{1\}\}A^\{h\_\{2\}\}; left\-multiplying by the current state distributionqtq\_\{t\}gives[Eq\.9](https://arxiv.org/html/2608.13621#S2.E9)\.for a time\-homogeneous finite\-state chain\.
###### Proposition 1\(Exact path consistency\)\.
Assume that the latent dynamics form a time\-homogeneous finite\-state Markov chain with a fixed row\-stochastic transition matrixAA\. Then, for any state distributionqtq\_\{t\}and nonnegative integersh1,h2h\_\{1\},h\_\{2\},
qtAh1\+h2=\(qtAh1\)Ah2\.q\_\{t\}A^\{h\_\{1\}\+h\_\{2\}\}=\\left\(q\_\{t\}A^\{h\_\{1\}\}\\right\)A^\{h\_\{2\}\}\.Consequently, all prediction paths whose transition lengths sum to the same total horizon produce the same predictive distribution\.
A proof is provided in[SectionB\.1](https://arxiv.org/html/2608.13621#A2.SS1)\. Proposition[1](https://arxiv.org/html/2608.13621#Thmproposition1)is useful for combining short\- and long\-horizon planning because a long\-horizon prediction can be computed either directly or by composing shorter transitions without introducing path\-dependent discrepancies\. This supports hierarchical planning and temporal abstraction while ensuring that every decomposition of the same total horizon yields a consistent predictive distribution\.
A softer, less constrained alternative learns a separate matrixAhA\_\{h\}for each horizon and penalizes
ℒCK=∑h1,h2‖Ah1\+h2−Ah1Ah2‖F2\.\\mathcal\{L\}\_\{\\mathrm\{CK\}\}=\\sum\_\{h\_\{1\},h\_\{2\}\}\\left\\\|A\_\{h\_\{1\}\+h\_\{2\}\}\-A\_\{h\_\{1\}\}A\_\{h\_\{2\}\}\\right\\\|\_\{F\}^\{2\}\.This variant tests whether one homogeneous chain is adequate or whether the data require horizon\-dependent dynamics\.
### 2\.6Avoiding discrete\-state collapse
The Markov\-chain prediction objective in[Eq\.8](https://arxiv.org/html/2608.13621#S2.E8)does not by itself guarantee that the categorical latent states contain useful information\. Because the online encoder, target encoder, and transition matrix are learned jointly, they may agree through a degenerate representation\. This is the discrete\-state analogue of representation collapse in deterministic self\-supervised learning\.
For each time indextt, the online encoder produces
qt=\(qt1,…,qtK\),qtk=qθ\(Zt=k∣X≤t\),q\_\{t\}=\\bigl\(q\_\{t1\},\\ldots,q\_\{tK\}\\bigr\),\\qquad q\_\{tk\}=q\_\{\\theta\}\(Z\_\{t\}=k\\mid X\_\{\\leq t\}\),whereqtkq\_\{tk\}is the probability that the current observation history is represented by latent statekk\.
Two simple degenerate solutions are particularly important\. First, all observations may be assigned to the same state:
qt≈ejfor everyt,q\_\{t\}\\approx e\_\{j\}\\qquad\\text\{for every \}t,whereej∈\{0,1\}Ke\_\{j\}\\in\\\{0,1\\\}^\{K\}is thejjth standard basis vector, with\(ej\)j=1\(e\_\{j\}\)\_\{j\}=1and\(ej\)k=0\(e\_\{j\}\)\_\{k\}=0for allk≠jk\\neq j\. The transition matrix can then place nearly all probability on the self\-transitionAjjA\_\{jj\}, allowing the online and target branches to agree without learning meaningful temporal structure\. We refer to this failure mode as*single\-state collapse*\.
Second, every observation may receive the same uniform assignment:
qt≈Unif\(K\)=\(1K,…,1K\)\.q\_\{t\}\\approx\\operatorname\{Unif\}\(K\)=\\left\(\\frac\{1\}\{K\},\\ldots,\\frac\{1\}\{K\}\\right\)\.A transition matrix that preserves the uniform distribution can again produce consistent predictions even though the latent state contains no information about the observation\. We refer to this as*uniform\-assignment collapse*\.
To measure how the available states are used across a minibatchℬ\\mathcal\{B\}, define the average state occupancy
q¯ℬ=1\|ℬ\|∑t∈ℬqt\.\\bar\{q\}\_\{\\mathcal\{B\}\}=\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{t\\in\\mathcal\{B\}\}q\_\{t\}\.We encourage aggregate use of the state space through
KL\(q¯ℬ∥Unif\(K\)\)\.\\mathrm\{KL\}\\left\(\\bar\{q\}\_\{\\mathcal\{B\}\}\\,\\middle\\\|\\,\\operatorname\{Unif\}\(K\)\\right\)\.This term is zero when aggregate occupancy is uniform and increases when most probability mass is concentrated on only a few states, thereby discouraging single\-state collapse and unused states\.
Aggregate diversity alone is not sufficient\. If every observation receives the uniform distribution, then
q¯ℬ=Unif\(K\),\\bar\{q\}\_\{\\mathcal\{B\}\}=\\operatorname\{Unif\}\(K\),so the occupancy penalty is again zero\. We therefore also control the entropy of each assignment,
H\(qt\)=−∑k=1Kqtklogqtk\.H\(q\_\{t\}\)=\-\\sum\_\{k=1\}^\{K\}q\_\{tk\}\\log q\_\{tk\}\.Entropy is maximal atlogK\\log Kfor a uniform assignment and minimal at zero for a one\-hot assignment\. Because this entropy appears with a positive weight in a minimized loss, it encourages comparatively confident, low\-entropy assignments\.
Combining the two effects gives
ℒstate=λoccKL\(q¯ℬ∥Unif\(K\)\)⏟occupancy\+λent1\|ℬ\|∑t∈ℬH\(qt\)⏟entropy,\\mathcal\{L\}\_\{\\mathrm\{state\}\}=\\underbrace\{\\lambda\_\{\\mathrm\{occ\}\}\\mathrm\{KL\}\\left\(\\bar\{q\}\_\{\\mathcal\{B\}\}\\,\\middle\\\|\\,\\operatorname\{Unif\}\(K\)\\right\)\}\_\{\\text\{occupancy\}\}\+\\underbrace\{\\lambda\_\{\\mathrm\{ent\}\}\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{t\\in\\mathcal\{B\}\}H\(q\_\{t\}\)\}\_\{\\text\{entropy\}\},\(10\)whereλocc,λent≥0\\lambda\_\{\\mathrm\{occ\}\},\\lambda\_\{\\mathrm\{ent\}\}\\geq 0\. The complete basic MCJEPA objective is therefore
ℒ=ℒMC\+ℒstate\.\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{state\}\}\.
The two regularizers inℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}play complementary roles\. The occupancy term encourages*diversity across observations*, while the entropy term encourages*confidence within each observation*\. Their combination therefore favors balanced but informative assignments: different observations can occupy different states while each individual observation receives a comparatively concentrated state distribution\.
The relative weights must nevertheless be selected carefully\. Ifλocc\\lambda\_\{\\mathrm\{occ\}\}is too large, the model may artificially force every minibatch to use all states even when the underlying state distribution is imbalanced\. Ifλent\\lambda\_\{\\mathrm\{ent\}\}is too large, the encoder may make prematurely hard and unstable assignments\. In practice, the entropy weight may be introduced gradually, and the uniform occupancy target may be replaced by a nonuniform priorπ\\piwhen unequal state frequencies are expected:
KL\(q¯ℬ∥π\)\.\\mathrm\{KL\}\\left\(\\bar\{q\}\_\{\\mathcal\{B\}\}\\,\\middle\\\|\\,\\pi\\right\)\.
## 3From a Transition Matrix to a Neural Markov Kernel
The taxonomy in[Section2\.1](https://arxiv.org/html/2608.13621#S2.SS1)places the basic MCJEPA construction in the*discrete\-time, discrete\-state*class\. Its transition matrix is the simplest realization of latent Markov dynamics: the predictive state is categorical, time advances in discrete steps, and, in the time\-homogeneous case, the same transition matrix is applied at every step\. As shown in[Eq\.7](https://arxiv.org/html/2608.13621#S2.E7), this gives the particularly simple multi\-step prediction
q^t\+h=qtAh\.\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\}\.
The Markov principle itself is more general\. It requires that the next\-state distribution depend on the past only through the current predictive state and the transition\-relevant side information666See sufficiency of predictive state\([9](https://arxiv.org/html/2608.13621#bib.bib9)\)\.:
p\(Zt\+1∣Z≤t,ξ≤t\)=p\(Zt\+1∣Zt,ξt\)\.p\(Z\_\{t\+1\}\\mid Z\_\{\\leq t\},\\xi\_\{\\leq t\}\)=p\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\.\(11\)The transition may therefore be fixed or conditioned, and the latent state may be discrete or continuous\. We now generalize the fixed transition matrix progressively while preserving this conditional\-independence structure\.
### 3\.1Neural discrete\-state transitions
The fixed matrixAAassumes that the transition law is time homogeneous and independent of external information\. A more expressive model allows a neural network to produce a row\-stochastic transition matrix conditioned on side information:
At=Aϕ\(ξt\),At∈\[0,1\]K×K,∑j=1K\(At\)ij=1\.A\_\{t\}=A\_\{\\phi\}\(\\xi\_\{t\}\),\\qquad A\_\{t\}\\in\[0,1\]^\{K\\times K\},\\qquad\\sum\_\{j=1\}^\{K\}\(A\_\{t\}\)\_\{ij\}=1\.\(12\)The side informationξt\\xi\_\{t\}may contain an action, elapsed time, goal, control input, regime indicator, or exogenous covariates\. Its entries have the interpretation
\(At\)ij=pϕ\(Zt\+1=j∣Zt=i,ξt\)\.\(A\_\{t\}\)\_\{ij\}=p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i,\\xi\_\{t\}\)\.
For a sequence of conditioned transitions, thehh\-step predictive distribution is
q^t\+h\\displaystyle\\widehat\{q\}\_\{t\+h\}=qt∏j=0h−1Aϕ\(ξt\+j\)\\displaystyle=q\_\{t\}\\prod\_\{j=0\}^\{h\-1\}A\_\{\\phi\}\(\\xi\_\{t\+j\}\)\(13\)=qtAϕ\(ξt\)Aϕ\(ξt\+1\)⋯Aϕ\(ξt\+h−1\),\\displaystyle=q\_\{t\}A\_\{\\phi\}\(\\xi\_\{t\}\)A\_\{\\phi\}\(\\xi\_\{t\+1\}\)\\cdots A\_\{\\phi\}\(\\xi\_\{t\+h\-1\}\),where the product is ordered chronologically from left to right\. The transition matrices at different steps need not commute, so this ordering matters\. In the time\-homogeneous, unconditioned case,
Aϕ\(ξt\+j\)=Afor allj,A\_\{\\phi\}\(\\xi\_\{t\+j\}\)=A\\qquad\\text\{for all \}j,and[Eq\.13](https://arxiv.org/html/2608.13621#S3.E13)reduces to theqtAhq\_\{t\}A^\{h\}construction in[Eq\.7](https://arxiv.org/html/2608.13621#S2.E7)\.
This construction supports action\-conditioned dynamics, changing goals or regimes, exogenous covariates, and irregular temporal intervals when elapsed time is included inξt\\xi\_\{t\}, while retaining a discrete and interpretable latent state space\. It also preserves exact path composition whenever the same chronologically ordered transition sequence is used along both prediction paths\.
### 3\.2Continuous\-state stochastic transitions
A categorical state may be too restrictive when the predictive representation varies continuously\. In this case, the transition matrix is replaced by a Markov kernel
pϕ\(Zt\+1∣Zt,ξt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\),which assigns a probability distribution over the next continuous latent state for each current state and side\-information value\.
A simple example is a Gaussian transition,
pϕ\(Zt\+1∣Zt,ξt\)=𝒩\(Zt\+1,μϕ\(Zt,ξt\),Σϕ\(Zt,ξt\)\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)=\\mathcal\{N\}\\\!\\left\(Z\_\{t\+1\};\\mu\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\),\\Sigma\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\\right\),\(14\)where the neural network outputs a conditional meanμϕ\\mu\_\{\\phi\}and a valid covariance matrixΣϕ\\Sigma\_\{\\phi\}\. The mean describes the expected latent evolution, while the covariance represents stochastic uncertainty and unresolved variation around that mean\.
The two\-step transition is obtained by marginalizing over the intermediate latent state:
pϕ\(2\)\(Zt\+2∣Zt,ξt:t\+1\)\\displaystyle p\_\{\\phi\}^\{\(2\)\}\\left\(Z\_\{t\+2\}\\mid Z\_\{t\},\\xi\_\{t:t\+1\}\\right\)=∫pϕ\(Zt\+2∣Zt\+1,ξt\+1\)pϕ\(Zt\+1∣Zt,ξt\)dZt\+1\.\\displaystyle=\\int p\_\{\\phi\}\(Z\_\{t\+2\}\\mid Z\_\{t\+1\},\\xi\_\{t\+1\}\)p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\\,dZ\_\{t\+1\}\.More generally,
pϕ\(h\)\(Zt\+h∣Zt,ξt:t\+h−1\)\\displaystyle p\_\{\\phi\}^\{\(h\)\}\\left\(Z\_\{t\+h\}\\mid Z\_\{t\},\\xi\_\{t:t\+h\-1\}\\right\)\(15\)=∫∏j=0h−1pϕ\(Zt\+j\+1∣Zt\+j,ξt\+j\)dZt\+1⋯dZt\+h−1\.\\displaystyle=\\int\\prod\_\{j=0\}^\{h\-1\}p\_\{\\phi\}\(Z\_\{t\+j\+1\}\\mid Z\_\{t\+j\},\\xi\_\{t\+j\}\)\\,dZ\_\{t\+1\}\\cdots dZ\_\{t\+h\-1\}\.This is the continuous\-state analogue of multiplying transition matrices in[Eqs\.7](https://arxiv.org/html/2608.13621#S2.E7)and[13](https://arxiv.org/html/2608.13621#S3.E13): all possible intermediate latent states are marginalized out\. Closed\-form composition is available only for restricted transition families, so neural models may instead use sampling, Monte Carlo integration, moment propagation, or learned approximations to multi\-step prediction\.
A single Gaussian kernel cannot represent a genuinely multimodal conditional distribution\. When multiple distinct futures are important, the same construction can instead use mixtures\([7](https://arxiv.org/html/2608.13621#bib.bib14)\), normalizing flows, diffusion\-based transitions, or other expressive conditional distributions\. The central requirement is not Gaussianity, but the Markov factorization in[Eq\.11](https://arxiv.org/html/2608.13621#S3.E11)\.
### 3\.3Continuous\-time extensions
Suppose that observations are recorded at physical times
τ1<τ2<⋯,\\tau\_\{1\}<\\tau\_\{2\}<\\cdots,with elapsed interval
Δtt=τt\+1−τt\.\\Delta t\_\{t\}=\\tau\_\{t\+1\}\-\\tau\_\{t\}\.Irregular sampling can already be handled within a discrete\-time model by including the observed interval in the side information,
ξt=\(ξ~t,Δtt\),\\xi\_\{t\}=\(\\widetilde\{\\xi\}\_\{t\},\\Delta t\_\{t\}\),whereξ~t\\widetilde\{\\xi\}\_\{t\}contains the remaining actions, controls, goals, or exogenous covariates\. The resulting discrete\-time transition model then learns how state evolution changes with the supplied propagation interval\.
A more explicit alternative is to model the latent dynamics directly in continuous time\. For a discrete latent state, a continuous\-time Markov chain is specified by a generator matrix
Qϕ\(ξ~t\),Q\_\{\\phi\}\(\\widetilde\{\\xi\}\_\{t\}\),whose off\-diagonal entries are nonnegative transition rates and whose rows sum to zero\. If the generator is held fixed over a propagation interval of durationΔt≥0\\Delta t\\geq 0, the corresponding transition matrix is
Aϕ\(Δt,ξ~t\)=exp\(ΔtQϕ\(ξ~t\)\)\.A\_\{\\phi\}\(\\Delta t,\\widetilde\{\\xi\}\_\{t\}\)=\\exp\\\!\\left\(\\Delta t\\,Q\_\{\\phi\}\(\\widetilde\{\\xi\}\_\{t\}\)\\right\)\.\(16\)Here,Δt\\Delta tdenotes a generic continuous\-time propagation duration, whereasΔtt=τt\+1−τt\\Delta t\_\{t\}=\\tau\_\{t\+1\}\-\\tau\_\{t\}denotes the particular interval between thettth and\(t\+1\)\(t\+1\)th observations\. Under a piecewise\-constant conditioning assumption over this interval,
At=Aϕ\(Δtt,ξ~t\)=exp\(ΔttQϕ\(ξ~t\)\)\.A\_\{t\}=A\_\{\\phi\}\(\\Delta t\_\{t\},\\widetilde\{\\xi\}\_\{t\}\)=\\exp\\\!\\left\(\\Delta t\_\{t\}\\,Q\_\{\\phi\}\(\\widetilde\{\\xi\}\_\{t\}\)\\right\)\.The same generator can therefore propagate the latent state across observation intervals of different lengths or across arbitrary forecasting horizons\. If the conditioning variables vary continuously within an interval, the corresponding transition requires the appropriate time\-varying generator composition rather than a single matrix exponential\.
For a continuous latent state, a continuous\-time stochastic transition may instead be represented by a stochastic differential equation\. Usingτ\\taufor continuous physical time to distinguish it from the discrete observation index,
dZ\(τ\)=bϕ\(Z\(τ\),ξ\(τ\)\)dτ\+Gϕ\(Z\(τ\),ξ\(τ\)\)dW\(τ\),dZ\(\\tau\)=b\_\{\\phi\}\(Z\(\\tau\),\\xi\(\\tau\)\)\\,d\\tau\+G\_\{\\phi\}\(Z\(\\tau\),\\xi\(\\tau\)\)\\,dW\(\\tau\),\(17\)wherebϕb\_\{\\phi\}is the drift,GϕG\_\{\\phi\}controls the diffusion, andW\(τ\)W\(\\tau\)is a Wiener process\. The drift describes systematic latent evolution, while the diffusion represents stochastic transition uncertainty\. SettingGϕ=0G\_\{\\phi\}=0recovers deterministic continuous\-time dynamics such as a neural ordinary differential equation\.
These continuous\-time constructions are not required for our main MCJEPA development, but they show that the transition\-matrix formulation belongs to a broader family of latent Markov models\.
### 3\.4Deterministic temporal JEPA as a degenerate kernel
A deterministic temporal JEPA uses a predictor
Z^t\+1=gϕ\(Zt,ξt\)\.\\widehat\{Z\}\_\{t\+1\}=g\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\.Probabilistically, this is the Dirac transition kernel
pϕ\(Zt\+1∣Zt,ξt\)=δ\(Zt\+1−gϕ\(Zt,ξt\)\)\.p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)=\\delta\\\!\\left\(Z\_\{t\+1\}\-g\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\\right\)\.\(18\)All conditional probability mass is concentrated at the predictor output, so the transition contains no intrinsic stochastic uncertainty\. For example, if the Gaussian mean in[Eq\.14](https://arxiv.org/html/2608.13621#S3.E14)satisfies
μϕ\(Zt,ξt\)=gϕ\(Zt,ξt\),\\mu\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)=g\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\),then the Gaussian transition approaches this deterministic case as
Σϕ\(Zt,ξt\)→0\.\\Sigma\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\\rightarrow 0\.
Classical deterministic temporal JEPA is therefore not separate from the latent Markov\-kernel perspective\. A deterministic encoder may be represented probabilistically by a point\-mass state distribution, while a deterministic predictor is represented by the Dirac transition in[Eq\.18](https://arxiv.org/html/2608.13621#S3.E18)\. MCJEPA, conditioned discrete\-state JEPA, continuous probabilistic VJEPA, and deterministic JEPA can thus be viewed within the same state\-space framework, differing primarily in their state representation and transition family\.
Table 3:Representative hierarchy of latent Markov transitions for temporal JEPA\. The main development focuses on discrete\-time models; continuous\-time variants situate the transition\-matrix formulation within the broader state\-space family\.
### 3\.5When a recurrent predictor is Markov
A recurrent predictor may appear to violate the first\-order Markov assumption because its prediction can depend on a summary of the entire preceding latent history\. LetMtM\_\{t\}denote a recurrent memory state and suppose that the next latent state is predicted according to
pϕ\(Zt\+1∣Zt,Mt,ξt\)\.p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},M\_\{t\},\\xi\_\{t\}\)\.\(19\)After sampling or predictingZt\+1Z\_\{t\+1\}, the recurrent memory may be updated deterministically as
Mt\+1=rϕ\(Mt,Zt\+1,ξt\)\.M\_\{t\+1\}=r\_\{\\phi\}\(M\_\{t\},Z\_\{t\+1\},\\xi\_\{t\}\)\.
The process need not be first\-order Markov inZtZ\_\{t\}alone, because two histories yielding the sameZtZ\_\{t\}but different memoriesMtM\_\{t\}may induce different next\-state distributions\. However, defining the augmented state
Z~t=\(Zt,Mt\)\\widetilde\{Z\}\_\{t\}=\(Z\_\{t\},M\_\{t\}\)restores a first\-order representation:
pϕ\(Z~t\+1∣Z~≤t,ξ≤t\)=pϕ\(Z~t\+1∣Z~t,ξt\)\.p\_\{\\phi\}\(\\widetilde\{Z\}\_\{t\+1\}\\mid\\widetilde\{Z\}\_\{\\leq t\},\\xi\_\{\\leq t\}\)=p\_\{\\phi\}\(\\widetilde\{Z\}\_\{t\+1\}\\mid\\widetilde\{Z\}\_\{t\},\\xi\_\{t\}\)\.
This distinction motivates a central representation\-learning objective of PIB\-VJEPA: ideally, the learned predictive stateZtZ\_\{t\}itself should summarize the information from the observation history that is relevant to future prediction\. When this succeeds, a simple first\-order transition inZtZ\_\{t\}is sufficient\. When substantial predictive information remains outsideZtZ\_\{t\}, the model must either enlarge the state, augment it with memory, use higher\-order dynamics, or accept that the latent process is not first\-order Markov in the chosen representation\.
## 4A Probabilistic JEPA Is Secretly an HMM
### 4\.1The three distributions in PIB\-VJEPA
The full, time\-indexed PIB\-VJEPA considered here contains three conditional distributions\([8](https://arxiv.org/html/2608.13621#bib.bib10)\):
qθ\(Zt∣X≤t\)\\displaystyle q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)current\-state encoder,\\displaystyle\\quad\\text\{current\-state encoder\},pϕ\(Zt\+1∣Zt,ξt\)\\displaystyle p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)latent transition predictor,\\displaystyle\\quad\\text\{latent transition predictor\},qθ¯\(Zt\+1∣Xt\+1\)\\displaystyle q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+1\}\\mid X\_\{t\+1\}\)future target encoder\.\\displaystyle\\quad\\text\{future target encoder\}\.The online encoderqθq\_\{\\theta\}maps the observation historyX≤tX\_\{\\leq t\}to a distribution over the current predictive stateZtZ\_\{t\}\. The transition modelpϕp\_\{\\phi\}propagates that state to a distribution overZt\+1Z\_\{t\+1\}, possibly conditioned on side informationξt\\xi\_\{t\}such as an action, elapsed time, or exogenous covariates\. The target encoderqθ¯q\_\{\\bar\{\\theta\}\}provides the future latent distribution against which the prediction is trained\. Here,θ¯\\bar\{\\theta\}denotes the slowly updated target\-encoder parameters, typically obtained as an exponential moving average of the online parametersθ\\theta\.
A compact form of the PIB\-VJEPA objective is\([8](https://arxiv.org/html/2608.13621#bib.bib10)\):
ℒPIB\-VJEPA=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{PIB\\text\{\-\}VJEPA\}\}=\{\}𝔼\[−logpϕ\(Zt\+1∣Zt,ξt\)\]\\displaystyle\\mathbb\{E\}\\left\[\-\\log p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\\right\]\(20\)\+γS𝔼KL\(qθ\(Zt∣X≤t\)∥prefS\(Zt\)\)\\displaystyle\+\\gamma\_\{\\mathrm\{S\}\}\\mathbb\{E\}\\mathrm\{KL\}\\left\(q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\,\\middle\\\|\\,p\_\{\\mathrm\{ref\}\}^\{\\mathrm\{S\}\}\(Z\_\{t\}\)\\right\)\+βT𝔼KL\(qθ¯\(Zt\+1∣Xt\+1\)∥prefT\(Zt\+1\)\)\.\\displaystyle\+\\beta\_\{\\mathrm\{T\}\}\\mathbb\{E\}\\mathrm\{KL\}\\left\(q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+1\}\\mid X\_\{t\+1\}\)\\,\\middle\\\|\\,p\_\{\\mathrm\{ref\}\}^\{\\mathrm\{T\}\}\(Z\_\{t\+1\}\)\\right\)\.The expectation is taken over training sequences, side information, and latent samples from the online and target encoders\. The superscriptsS\\mathrm\{S\}andT\\mathrm\{T\}label the*source\-side*current state and*target\-side*future state, respectively\. Accordingly,prefSp\_\{\\mathrm\{ref\}\}^\{\\mathrm\{S\}\}andprefTp\_\{\\mathrm\{ref\}\}^\{\\mathrm\{T\}\}are reference prior distributions for the current and future latent states\. Depending on the latent family, these may be standard Gaussian, uniform categorical, or other suitably chosen simple distributions\.
The nonnegative coefficients
γS≥0,βT≥0\\gamma\_\{\\mathrm\{S\}\}\\geq 0,\\qquad\\beta\_\{\\mathrm\{T\}\}\\geq 0control the strengths of the source\- and target\-side regularization\. IncreasingγS\\gamma\_\{\\mathrm\{S\}\}places greater pressure on the current representation to compress the observation history, while increasingβT\\beta\_\{\\mathrm\{T\}\}more strongly regularizes the future target representation\. These coefficients therefore trade predictive accuracy against latent compression and regularity\.
The first term in[Eq\.20](https://arxiv.org/html/2608.13621#S4.E20)encourages the current state to preserve information needed to predict the future target state\. The second term promotes compression of the observation history by regularizingqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)toward the reference priorprefSp\_\{\\mathrm\{ref\}\}^\{\\mathrm\{S\}\}, while the third regularizes the future target distribution towardprefTp\_\{\\mathrm\{ref\}\}^\{\\mathrm\{T\}\}\. Together, the three terms encourage a predictive latent state while controlling the information retained in its stochastic representation\.
### 4\.2The encode–transition–emit correspondence
A conventional HMM factorizes as\([13](https://arxiv.org/html/2608.13621#bib.bib1)\)
p\(s1:T,x1:T\)=p\(s1\)∏t=1T−1p\(st\+1∣st\)∏t=1Tp\(xt∣st\)\.p\(s\_\{1:T\},x\_\{1:T\}\)=p\(s\_\{1\}\)\\prod\_\{t=1\}^\{T\-1\}p\(s\_\{t\+1\}\\mid s\_\{t\}\)\\prod\_\{t=1\}^\{T\}p\(x\_\{t\}\\mid s\_\{t\}\)\.\(21\)The direct correspondence777For notational simplicity,[Eq\.21](https://arxiv.org/html/2608.13621#S4.E21)shows the unconditioned case\. Corresponding to JEPA, when observed side information is present, replacep\(st\+1∣st\)p\(s\_\{t\+1\}\\mid s\_\{t\}\)byp\(st\+1∣st,ξt\)p\(s\_\{t\+1\}\\mid s\_\{t\},\\xi\_\{t\}\)and interpret the sequence factorization conditional onξ1:T−1\\xi\_\{1:T\-1\}\. To keep the encoder notation compact, we usually suppress past side information inqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\); when relevant, it should be read asqθ\(Zt∣X≤t,ξ<t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\},\\xi\_\{<t\}\)\.was summarized in[Table2](https://arxiv.org/html/2608.13621#S2.T2)\. Observation\-level data play the role of HMM observations,ZtZ\_\{t\}is the predictive latent state, andpϕ\(Zt\+1∣Zt,ξt\)p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)specifies its transition dynamics\. The history\-dependent context encoder
qθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)plays the*state\-inference*role\. It coincides with the Bayesian filtering distribution of the corresponding HMM only under the filtering\-consistency conditions developed below\.
The remaining HMM component is the state\-to\-observation direction\. There are three ways in which this direction can be supplied\.
#### Explicit decoder\.
A probabilistic decoder
pψ\(Xt∣Zt\)p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)is the literal analogue of an HMM emission distribution\. Such a decoder may be trained jointly with the latent model or added after representation learning, depending on whether observation\-space likelihood or forecasting is part of the training objective\.
#### Inverse target encoder\.
If the target encoder is bijective on the modeled data domain, its inverse supplies a deterministic state\-to\-observation map:
ZtT=fθ¯\(Xt\),Xt=fθ¯−1\(ZtT\)\.Z\_\{t\}^\{\\mathrm\{T\}\}=f\_\{\\bar\{\\theta\}\}\(X\_\{t\}\),\\qquad X\_\{t\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}^\{\\mathrm\{T\}\}\)\.The associated emission can be represented as a Dirac kernel concentrated atfθ¯−1\(Zt\)f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}\)\. A deterministic inverse alone, however, does not provide a non\-degenerate observation density\. When normalized likelihood evaluation is required, the inverse must form part of a tractable probabilistic density model, for example through an appropriate change\-of\-variables construction or explicit observation\-noise model\. Standard JEPA target encoders are typically compressive and therefore not exactly invertible, so inverse target encoding is an alternative architectural realization rather than an assumption of the general theory\.
#### Implicit emission\.
When neither an explicit decoder nor an invertible target encoder is available, a local stochastic encoder can still induce a state\-to\-observation conditional\. We develop this construction next\.
In all three cases, the stochastic encoder itself should not be identified with the emission distribution: it maps observations to latent states, whereas an HMM emission maps latent states to observations\.
### 4\.3When no decoder is present: an implicit emission
Suppose that a local stochastic encoder is available,
qθ\(z∣x\),q\_\{\\theta\}\(z\\mid x\),and letpdata\(x\)p\_\{\\mathrm\{data\}\}\(x\)denote the observation marginal\. Define the induced latent marginal
qθ\(z\)=∫pdata\(x\)qθ\(z∣x\)𝑑x\.q\_\{\\theta\}\(z\)=\\int p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\\,dx\.Wheneverqθ\(z\)\>0q\_\{\\theta\}\(z\)\>0, define
pθimp\(x∣z\)=pdata\(x\)qθ\(z∣x\)qθ\(z\)\.p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)=\\frac\{p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\}\{q\_\{\\theta\}\(z\)\}\.\(22\)
###### Proposition 2\(Implicit emission completion\)\.
The conditional distribution in[Eq\.22](https://arxiv.org/html/2608.13621#S4.E22)is normalized and satisfies
qθ\(z∣x\)=pθimp\(x∣z\)qθ\(z\)pdata\(x\)\.q\_\{\\theta\}\(z\\mid x\)=\\frac\{p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)q\_\{\\theta\}\(z\)\}\{p\_\{\\mathrm\{data\}\}\(x\)\}\.Thus, every local stochastic encoder together with the data marginal defines a one\-time latent\-variable model for which the encoder is the exact posterior\.
A proof of normalization and Bayes consistency is provided in[SectionB\.2](https://arxiv.org/html/2608.13621#A2.SS2)\. The construction is a direct application of Bayes’ rule and establishes an exact*static*observation–state correspondence\. It does not by itself establish a sequence\-level HMM\. For that stronger claim, the induced state distributions must evolve consistently with the latent transition, and the history\-dependent context encoder must agree with the Bayesian filtering distribution induced by the transition and emission models\.
### 4\.4Four levels of correspondence
The statement that probabilistic temporal JEPA is “secretly an HMM” is not all\-or\-nothing\. The correspondence becomes progressively stronger as additional state\-space semantics are imposed\. We distinguish four levels\.
#### 1\. Computational correspondence\.
At the weakest level, the two architectures expose the same computational roles:
observation\-to\-state inference⟶state transition⟶state\-to\-observation prediction\.\\text\{observation\-to\-state inference\}\\;\\longrightarrow\\;\\text\{state transition\}\\;\\longrightarrow\\;\\text\{state\-to\-observation prediction\}\.For PIB\-VJEPA these roles are implemented by the context encoder, latent predictor, and one of the observation\-map constructions above\. This correspondence does not by itself imply a common joint distribution or training objective\.
#### 2\. Emission\-complete latent\-state representation\.
The correspondence becomes probabilistically more explicit once a valid state\-to\-observation conditional is available\. This conditional may be parameterized directly by a decoder, supplied deterministically by an invertible target encoder, or induced implicitly through[Eq\.22](https://arxiv.org/html/2608.13621#S4.E22)\. Together with a valid latent transition, these components provide the ingredients of a latent Markov model\. They do not yet guarantee, however, that the learned history encoder is the Bayesian filter of that model or that its state marginals are dynamically consistent\.
#### 3\. Sequence\-level HMM equivalence\.
A stronger statement holds when the transition, emission, latent marginals, and history\-dependent encoder are mutually consistent\. Under these conditions, the latent\-state model admits an HMM factorization at the sequence level rather than merely sharing its components\.
###### Theorem 1\(Sufficient conditions for an exact HMM representation\)\.
Consider a probabilistic temporal JEPA with latent stateZtZ\_\{t\}and observation processXtX\_\{t\}\. The full JEPA system is consistent with an exact HMM representation, meaning that its latent transition and emission define an HMM sequence model and its history\-dependent encoder coincides with the corresponding filtering distribution, provided that:
1. 1\.the latent dynamics satisfy the first\-order Markov property;
2. 2\.a valid state\-to\-observation conditional is specified by an explicit decoder, an invertible target encoder, or the implicit construction in[Eq\.22](https://arxiv.org/html/2608.13621#S4.E22);
3. 3\.the latent\-state marginals are consistent with the transition kernel; and
4. 4\.the history\-dependent encoder coincides with the filtering posterior induced by the corresponding transition and emission models\.
For the implicit\-emission construction, the local encoder must additionally satisfy the required observation\-locality and Bayes\-consistency conditions\. Under these assumptions, the resulting joint sequence distribution admits the HMM factorization in[Eq\.21](https://arxiv.org/html/2608.13621#S4.E21), and the history\-dependent encoder is the corresponding filtering distribution\. These conditions are sufficient rather than necessary and are not guaranteed by the standard JEPA training objective\.
A complete statement of the marginal\-consistency and filtering conditions, together with the proof, is provided in[AppendixE](https://arxiv.org/html/2608.13621#A5)\.
#### 4\. Model\-and\-objective equivalence\.
Given sequence\-level HMM equivalence, the strongest correspondence additionally concerns how that probabilistic model is trained\. Standard probabilistic JEPA training predicts target representations and regularizes their information content; it does not generally maximize the observation\-sequence likelihood
logp\(X1:T\)\\log p\(X\_\{1:T\}\)or optimize a conventional HMM sequence\-evidence objective\. Full model\-and\-objective equivalence requires the resulting HMM\-compatible sequence model to be trained directly by its observation\-sequence likelihood, or by the corresponding sequence\-evidence objective when exact marginalization is unavailable\. Hybrid objectives occupy an intermediate regime because they retain the original JEPA latent\-prediction objective alongside HMM\-style sequence and filtering supervision rather than replacing it\. We develop these alternatives in[AppendixD](https://arxiv.org/html/2608.13621#A4)and examine their behavior experimentally in[Section7\.4](https://arxiv.org/html/2608.13621#S7.SS4)\.
The hierarchy therefore separates four distinct claims: sharing HMM\-like computational roles, possessing an emission\-complete latent\-state representation, defining the same sequence\-level probabilistic model, and additionally training that model with an HMM\-style sequence objective\.
## 5Information Bottleneck Learning as Markovization
At the information\-theoretic level, predictive information bottleneck learning seeks a representation that compresses the past while retaining predictive information about the future\([15](https://arxiv.org/html/2608.13621#bib.bib2);[5](https://arxiv.org/html/2608.13621#bib.bib3);[1](https://arxiv.org/html/2608.13621#bib.bib4)\)\. A schematic latent\-space objective is
minqθ\(Zt∣X≤t\)I\(X≤t,Zt\)−λI\(Zt,Zt\+1\),\\min\_\{q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\}\\mathrm\{I\}\(X\_\{\\leq t\};Z\_\{t\}\)\-\\lambda\\mathrm\{I\}\(Z\_\{t\};Z\_\{t\+1\}\),\(23\)where the first term penalizes information retained from the observation history and the second rewards predictive dependence between the current and future latent states\. The trade\-off parameterλ\\lambdacontrols the relative emphasis on compression and prediction\. The practical PIB\-VJEPA objective in[Eq\.20](https://arxiv.org/html/2608.13621#S4.E20)implements this \(variational\) principle through latent prediction together with variational bottleneck regularization\([8](https://arxiv.org/html/2608.13621#bib.bib10)\)\.
Importantly, minimizing[Eq\.23](https://arxiv.org/html/2608.13621#S5.E23)does not by itself guarantee that the learned representation is predictively sufficient\. A stronger ideal target is thatZtZ\_\{t\}retain all information in the observation history that is relevant to the future\. In conditional\-independence form,
X\>t⟂X≤t\|Zt\.X\_\{\>t\}\\perp X\_\{\\leq t\}\\mid Z\_\{t\}\.\(24\)This is the predictive\-state condition: onceZtZ\_\{t\}is known, the remaining observation history provides no additional information about the future\. It is closely related to predictive\-state representations\([11](https://arxiv.org/html/2608.13621#bib.bib11)\)\.
For one\-step latent prediction, a corresponding Markov\-sufficiency condition is
I\(Zt\+1;X<t∣Zt\)=0\.\\mathrm\{I\}\(Z\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=0\.\(25\)Thus, after conditioning on the current predictive state, older observation history contains no additional information about the next latent state\. When transition\-relevant side informationξt\\xi\_\{t\}is present, the analogous diagnostic additionally conditions onξt\\xi\_\{t\}\.
###### Proposition 3\(Predictive sufficiency implies one\-step Markov sufficiency\)\.
Assume that
X\>t⟂X≤t\|Zt,X\_\{\>t\}\\perp X\_\{\\leq t\}\\mid Z\_\{t\},and that the future target state is generated from the next observation as
Zt\+1=gθ¯\(Xt\+1,Ut\+1\),Z\_\{t\+1\}=g\_\{\\bar\{\\theta\}\}\(X\_\{t\+1\},U\_\{t\+1\}\),where the target\-encoder randomness satisfies
Ut\+1⟂X≤t\|\(Xt\+1,Zt\)\.U\_\{t\+1\}\\perp X\_\{\\leq t\}\\mid\(X\_\{t\+1\},Z\_\{t\}\)\.Then
I\(Zt\+1;X<t∣Zt\)=0\.\\mathrm\{I\}\(Z\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=0\.Hence, predictive sufficiency implies that the learned state screens off older observation history from the next latent state, giving a one\-step Markov\-sufficient representation at the chosen prediction scale\.
The proof follows from the conditional data\-processing inequality and is provided in[SectionB\.3](https://arxiv.org/html/2608.13621#A2.SS3)\.
## 6Residual Predictability as a Diagnostic of Markov Sufficiency
The Markov\-sufficiency condition in[Eq\.25](https://arxiv.org/html/2608.13621#S5.E25)is difficult to verify directly in a learned, high\-dimensional representation\. A more operational approach is to ask whether prediction errors retain systematic dependence on information preceding the current state\. If older observations or latent states improve prediction after the current representation and transition\-relevant side information have been accounted for, then the current state–predictor pair has not captured all transition\-relevant information\.
Residual diagnostics test consequences of Markov sufficiency rather than the full conditional\-independence condition itself\. In particular, the diagnostics below focus primarily on conditional\-mean predictability\. Detecting residual predictability from older history therefore provides evidence against Markov sufficiency, whereas failing to detect it does not prove the full Markov property\.
### 6\.1Categorical probability innovations
For categorical MCJEPA, let
qt=qθ\(Zt∣X≤t\)q\_\{t\}=q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)denote the current\-state distribution, and let
q¯t\+1=qθ¯\(Zt\+1∣Xt\+1\)\\bar\{q\}\_\{t\+1\}=q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+1\}\\mid X\_\{t\+1\}\)denote the target\-encoder distribution at the next time step\. WriteAt=AA\_\{t\}=Afor the time\-homogeneous model andAt=Aϕ\(ξt\)A\_\{t\}=A\_\{\\phi\}\(\\xi\_\{t\}\)for an input\-conditioned transition\. The predicted next\-state distribution is
q^t\+1=qtAt\.\\widehat\{q\}\_\{t\+1\}=q\_\{t\}A\_\{t\}\.We define the categorical probability innovation as
Rt\+1=sg\(q¯t\+1\)−q^t\+1=sg\(q¯t\+1\)−qtAt\.R\_\{t\+1\}=\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\-\\widehat\{q\}\_\{t\+1\}=\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\-q\_\{t\}A\_\{t\}\.\(26\)This vector measures the discrepancy between the target\-encoder distribution and the transition\-based prediction for each latent category\. The stop\-gradient operator ensures that, when auxiliary diagnostic models are fitted to these residuals, their gradients are not propagated into the target encoder\.
To distinguish transition fitting from state sufficiency, define the current predictor information
𝒢t=σ\(qt,ξt\)\\mathcal\{G\}\_\{t\}=\\sigma\(q\_\{t\},\\xi\_\{t\}\)and the full observed\-history filtration
ℱt=σ\(X≤t,ξ≤t\)\.\\mathcal\{F\}\_\{t\}=\\sigma\(X\_\{\\leq t\},\\xi\_\{\\leq t\}\)\.For fixed model parameters,𝒢t⊆ℱt\\mathcal\{G\}\_\{t\}\\subseteq\\mathcal\{F\}\_\{t\}, sinceqtq\_\{t\}is computed from the observation history\. The transition predictor is conditionally mean\-correct with respect to its own inputs when
qtAt=𝔼\[sg\(q¯t\+1\)∣𝒢t\]\.q\_\{t\}A\_\{t\}=\\mathbb\{E\}\\left\[\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\\mid\\mathcal\{G\}\_\{t\}\\right\]\.This property is naturally associated with the forward\-KL objective used by MCJEPA\.888LetY=sg\(q¯t\+1\)∈ΔK−1Y=\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\\in\\Delta^\{K\-1\}and condition on𝒢t\\mathcal\{G\}\_\{t\}\. For any predicted distributionp∈ΔK−1p\\in\\Delta^\{K\-1\},𝔼\[KL\(Y∥p\)∣𝒢t\]=𝔼\[∑jYjlogYj\|𝒢t\]−∑j𝔼\[Yj∣𝒢t\]logpj\.\\mathbb\{E\}\[\\mathrm\{KL\}\(Y\\\|p\)\\mid\\mathcal\{G\}\_\{t\}\]=\\mathbb\{E\}\\\!\\left\[\\sum\_\{j\}Y\_\{j\}\\log Y\_\{j\}\\middle\|\\mathcal\{G\}\_\{t\}\\right\]\-\\sum\_\{j\}\\mathbb\{E\}\[Y\_\{j\}\\mid\\mathcal\{G\}\_\{t\}\]\\log p\_\{j\}\.The first term is independent ofpp, so minimizing the conditional expected KL is equivalent to minimizing the cross\-entropy withmt=𝔼\[Y∣𝒢t\]m\_\{t\}=\\mathbb\{E\}\[Y\\mid\\mathcal\{G\}\_\{t\}\]\. The population optimum is thereforep⋆=mtp^\{\\star\}=m\_\{t\}\. Hence, when the transition family can represent the optimum of the MCJEPA objective,qtAt=𝔼\[sg\(q¯t\+1\)∣𝒢t\]q\_\{t\}A\_\{t\}=\\mathbb\{E\}\[\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\\mid\\mathcal\{G\}\_\{t\}\]\. A restricted or imperfectly optimized transition family need not satisfy this equality exactly\.Under this condition,
𝔼\[Rt\+1∣𝒢t\]=0\.\\mathbb\{E\}\[R\_\{t\+1\}\\mid\\mathcal\{G\}\_\{t\}\]=0\.Thus, after conditioning on the inputs already available to the transition predictor, the residual has no remaining predictable conditional mean\.
Markov sufficiency requires a stronger invariance: older history should not alter the conditional prediction once the current predictive state and side information are known\. At the level of the target\-encoder distribution, the corresponding conditional\-mean implication is
𝔼\[sg\(q¯t\+1\)∣ℱt\]=𝔼\[sg\(q¯t\+1\)∣𝒢t\]\.\\mathbb\{E\}\\left\[\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\\mid\\mathcal\{F\}\_\{t\}\\right\]=\\mathbb\{E\}\\left\[\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+1\}\)\\mid\\mathcal\{G\}\_\{t\}\\right\]\.Combining this condition with a conditionally mean\-correct transition gives
𝔼\[Rt\+1∣ℱt\]=0\.\\mathbb\{E\}\[R\_\{t\+1\}\\mid\\mathcal\{F\}\_\{t\}\]=0\.\(27\)Assuming integrability,[Eq\.27](https://arxiv.org/html/2608.13621#S6.E27)gives the martingale\-difference property of\{Rt\+1\}\\\{R\_\{t\+1\}\\\}with respect to\{ℱt\}\\\{\\mathcal\{F\}\_\{t\}\\\}: once the current state and the available history are known, the residual has no systematic predictable component\.
A consequence of[Eq\.27](https://arxiv.org/html/2608.13621#S6.E27)is thatRt\+1R\_\{t\+1\}is uncorrelated with any square\-integrable function measurable with respect toℱt\\mathcal\{F\}\_\{t\}\. A simple linear diagnostic is therefore
𝒟lin=∑k=1Kr‖Cov\(Rt\+1,qt−k\)‖F2,\\mathcal\{D\}\_\{\\mathrm\{lin\}\}=\\sum\_\{k=1\}^\{K\_\{\\mathrm\{r\}\}\}\\left\\\|\\operatorname\{Cov\}\\left\(R\_\{t\+1\},q\_\{t\-k\}\\right\)\\right\\\|\_\{F\}^\{2\},whereKrK\_\{\\mathrm\{r\}\}is the maximum lag examined and∥⋅∥F\\\|\\cdot\\\|\_\{F\}denotes the Frobenius norm\. A large value indicates that some components of the residual remain linearly associated with earlier latent states\. A value near zero rules out only this particular form of linear dependence and does not establish Markov sufficiency\.
### 6\.2Testing incremental predictability from older history
The covariance diagnostic detects only linear dependence\. A stronger test asks whether an auxiliary model can predict the residual from older history beyond what can already be predicted from the current state and side information\. Consider a restricted residual predictor
R^t\+1\(0\)=gω0\(qt,ξt\)\\widehat\{R\}\_\{t\+1\}^\{\(0\)\}=g\_\{\\omega\_\{0\}\}\(q\_\{t\},\\xi\_\{t\}\)and a history\-augmented predictor
R^t\+1\(1\)=gω1\(qt,ξt,qt−1,…,qt−Kr\)\.\\widehat\{R\}\_\{t\+1\}^\{\(1\)\}=g\_\{\\omega\_\{1\}\}\\left\(q\_\{t\},\\xi\_\{t\},q\_\{t\-1\},\\ldots,q\_\{t\-K\_\{\\mathrm\{r\}\}\}\\right\)\.Their held\-out prediction errors can be compared through
Δhist=𝔼^\[‖Rt\+1−R^t\+1\(0\)‖22\]−𝔼^\[‖Rt\+1−R^t\+1\(1\)‖22\]\.\\displaystyle\\Delta\_\{\\mathrm\{hist\}\}=\{\}\\widehat\{\\mathbb\{E\}\}\\left\[\\left\\\|R\_\{t\+1\}\-\\widehat\{R\}\_\{t\+1\}^\{\(0\)\}\\right\\\|\_\{2\}^\{2\}\\right\]\-\\widehat\{\\mathbb\{E\}\}\\left\[\\left\\\|R\_\{t\+1\}\-\\widehat\{R\}\_\{t\+1\}^\{\(1\)\}\\right\\\|\_\{2\}^\{2\}\\right\]\.\(28\)A reliably positive value ofΔhist\\Delta\_\{\\mathrm\{hist\}\}means that older latent history improves prediction beyond\(qt,ξt\)\(q\_\{t\},\\xi\_\{t\}\)\. This provides evidence against the conditional\-mean sufficiency of the current state\-\-predictor pair999The comparison should be evaluated on held\-out data or through cross\-fitting\. The restricted and augmented auxiliary predictors should also have comparable capacity and regularization; otherwise an apparent history gain may reflect unequal model flexibility rather than genuinely additional predictive information\.\.
The same principle can be implemented by comparing restricted and history\-augmented predictors of the target itself rather than predictors of the residual\. If a common baseline prediction is used, the two formulations are equivalent because predictingRt\+1=Yt\+1−Y^t\+1R\_\{t\+1\}=Y\_\{t\+1\}\-\\widehat\{Y\}\_\{t\+1\}is equivalent to correcting the baseline predictionY^t\+1\\widehat\{Y\}\_\{t\+1\}\. Experiment 3 uses the direct\-prediction version of this diagnostic, comparing prediction fromZtZ\_\{t\}with prediction from\(Zt,Zt−1\)\(Z\_\{t\},Z\_\{t\-1\}\)\.
### 6\.3Distinguishing state insufficiency from predictor misspecification
Residual predictability can arise because of either the representation or the transition model\. First, the current representation may fail to summarize all past information relevant to predicting the future\. Second, the transition model may be too restricted or insufficiently optimized even when the representation itself is sufficient\.
The distinction between𝒢t\\mathcal\{G\}\_\{t\}andℱt\\mathcal\{F\}\_\{t\}helps separate these effects\. Predictability ofRt\+1R\_\{t\+1\}from\(qt,ξt\)\(q\_\{t\},\\xi\_\{t\}\)alone indicates that the fitted transition has not captured the conditional mean available from its own inputs\. The restricted auxiliary modelgω0g\_\{\\omega\_\{0\}\}can absorb part of this current\-input misspecification\. Additional held\-out improvement after introducing\(qt−1,…,qt−Kr\)\(q\_\{t\-1\},\\ldots,q\_\{t\-K\_\{\\mathrm\{r\}\}\}\)then asks a more specific question: does older latent history contain predictive information not recoverable from the current state and side information?
Attributing a positiveΔhist\\Delta\_\{\\mathrm\{hist\}\}specifically to representation insufficiency nevertheless requires care\. The current\-input predictor and residual correction must be sufficiently expressive and well fitted, the restricted and augmented diagnostic models should be compared under matched capacity and regularization, and evaluation should be performed out of sample\. Under these conditions, incremental predictability from older history is evidence that the current representation has omitted transition\-relevant information rather than merely that the original transition parameterization was imperfect\.
### 6\.4Continuous\-state diagnostics
For a continuous probabilistic predictor, suppose that
pϕ\(Zt\+1∣Zt,ξt\)=𝒩\(μt,Σt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)=\\mathcal\{N\}\(\\mu\_\{t\},\\Sigma\_\{t\}\),and let
Zt\+1T∼qθ¯\(Zt\+1∣Xt\+1\)Z\_\{t\+1\}^\{\\mathrm\{T\}\}\\sim q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+1\}\\mid X\_\{t\+1\}\)denote a target\-encoder latent sample\. WhenΣt\\Sigma\_\{t\}is positive definite, a standardized innovation may be defined as
Rt\+1std=Σt−1/2\(sg\(Zt\+1T\)−μt\)\.R\_\{t\+1\}^\{\\mathrm\{std\}\}=\\Sigma\_\{t\}^\{\-1/2\}\\left\(\\operatorname\{sg\}\(Z\_\{t\+1\}^\{\\mathrm\{T\}\}\)\-\\mu\_\{t\}\\right\)\.For singular or nearly singular covariance matrices, a regularized or pseudoinverse square root should be used instead\.
Under a correctly specified conditional Gaussian model, the standardized innovation has conditional mean zero and conditional covariance equal to the identity\. Markov sufficiency further implies that older history should not systematically predict this innovation once the current state and side information are given\. One may therefore examine lagged residual dependence, residual covariance, squared\-residual dependence, or auxiliary history\-prediction gains\. For non\-Gaussian predictors, score\-based diagnostics, probability\-integral\-transform diagnostics in suitable scalar settings, or appropriate multivariate calibration diagnostics can be used instead\.
These diagnostics test different consequences of correct conditional prediction\. Zero autocorrelation, for example, rules out linear temporal dependence but does not exclude nonlinear or higher\-order dependence\. Likewise, calibrated marginal probability\-integral\-transform values do not establish the conditional independence required for Markov sufficiency\.
Residual analysis should therefore be interpreted asymmetrically\. Predictability from older history provides evidence against Markov sufficiency of the current state–predictor pair\. Failure to detect such predictability means only that the chosen diagnostics do not reject sufficiency; it does not prove that the learned representation is Markov sufficient\.
## 7Experiments
Our experiments are deliberately small and diagnostic\. Rather than targeting state\-of\-the\-art forecasting performance, we test the structural claims developed in the preceding sections in settings where the latent states, transition laws, filtering distributions, and predictive sufficiency structure are known\. This allows us to separate the effect of the proposed Markov structure from model capacity and uncontrolled properties of real\-world data\.
[Table4](https://arxiv.org/html/2608.13621#S7.T4)summarizes the 4 experiments\. They follow the conceptual progression of the paper: Experiment 1 asks whether MCJEPA recovers a coherent finite\-state transition; Experiment 2 isolates the distinction between local observation evidence and filtering; Experiment 3 tests predictive compression, Markovization, and residual sufficiency; and Experiment 4 studies the strongest correspondence by training the same latent\-state architecture with an HMM sequence objective\. Full data\-generation parameters, architectures, optimization settings, and hyperparameters are deferred to Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
Table 4:Overview of the experimental questions\. All experiments use controlled synthetic processes for which the relevant latent structure is known\.#### Common protocol\.
All reported learned\-model results use five random seeds, with means and standard deviations reported across seeds\. Categorical latent\-state labels are identifiable only up to permutation\. ARI and NMI are themselves permutation invariant, whereas transition matrices and predicted categorical distributions are aligned to ground\-truth state order using a Hungarian assignment computed from the training\-set hard state assignments\. Multi\-step evaluation uses horizons
ℋ=\{1,2,4,8\}\.\\mathcal\{H\}=\\\{1,2,4,8\\\}\.
We use several common metrics to distinguish state recovery, transition recovery, predictive accuracy, and observation\-level probabilistic fit\. Full definitions are collected in Appendix[F\.1](https://arxiv.org/html/2608.13621#A6.SS1)\.
*Adjusted Rand index*\(ARI; Eq\.[47](https://arxiv.org/html/2608.13621#A6.E47)\) and*normalized mutual information*\(NMI; Eq\.[48](https://arxiv.org/html/2608.13621#A6.E48)\) measure agreement between inferred and ground\-truth state partitions\. Larger values indicate better recovery, with11corresponding to exact partition agreement\. For reference,
NMI\(S,S^\)=2I\(S,S^\)H\(S\)\+H\(S^\)\.\\operatorname\{NMI\}\(S,\\widehat\{S\}\)=\\frac\{2I\(S;\\widehat\{S\}\)\}\{H\(S\)\+H\(\\widehat\{S\}\)\}\.\(Eq\.[48](https://arxiv.org/html/2608.13621#A6.E48)\)ARI additionally corrects pairwise partition agreement for agreement expected by chance; its full expression is given in Eq\.[47](https://arxiv.org/html/2608.13621#A6.E47)\.
Transition recovery is measured by the normalized permutation\-aligned Frobenius error
ℰA=‖A^aligned−A⋆‖F‖A⋆‖F=‖P⊤A^P−A⋆‖F‖A⋆‖F\.\\mathcal\{E\}\_\{A\}=\\frac\{\\left\\\|\\widehat\{A\}\_\{\\mathrm\{aligned\}\}\-A^\{\\star\}\\right\\\|\_\{F\}\}\{\\\|A^\{\\star\}\\\|\_\{F\}\}=\\frac\{\\left\\\|P^\{\\top\}\\widehat\{A\}P\-A^\{\\star\}\\right\\\|\_\{F\}\}\{\\\|A^\{\\star\}\\\|\_\{F\}\}\.\(Eq\.[49](https://arxiv.org/html/2608.13621#A6.E49)\)Here,A^\\widehat\{A\}is the learned transition matrix,A⋆A^\{\\star\}is the ground\-truth transition matrix, andPPis the learned\-to\-ground\-truth permutation matrix obtained from training\-set state alignment\. LowerℰA\\mathcal\{E\}\_\{A\}indicates more faithful recovery of the latent dynamics\.
Predictive quality is measured using negative log\-likelihood\. When the ground\-truth future state is available, the horizon\-hhtrue\-state NLL is
ℒhstate=−𝔼t\[logq^t\+h\(St\+h\)\],\\mathcal\{L\}\_\{h\}^\{\\mathrm\{state\}\}=\-\\mathbb\{E\}\_\{t\}\\left\[\\log\\widehat\{q\}\_\{t\+h\}\(S\_\{t\+h\}\)\\right\],\(Eq\.[50](https://arxiv.org/html/2608.13621#A6.E50)\)whereSt\+hS\_\{t\+h\}is the ground\-truth latent state andq^t\+h\\widehat\{q\}\_\{t\+h\}is the predicted categorical distribution after state alignment\. Thus,q^t\+h\(St\+h\)\\widehat\{q\}\_\{t\+h\}\(S\_\{t\+h\}\)is the probability assigned to the realized future state, and lower NLL indicates better probabilistic prediction\.
When an explicit transition–emission model is available, observation\-space fit is measured by sequence NLL per time step,
ℒseq=−1T𝔼\[logpϕ,ψ\(X1:T\)\]\.\\mathcal\{L\}\_\{\\mathrm\{seq\}\}=\-\\frac\{1\}\{T\}\\mathbb\{E\}\\left\[\\log p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)\\right\]\.\(Eq\.[51](https://arxiv.org/html/2608.13621#A6.E51)\)Here,pϕ,ψ\(X1:T\)p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)is the marginal observation\-sequence density obtained after marginalizing the latent\-state trajectory under the transition and emission models\. This differs from true\-state NLL:ℒhstate\\mathcal\{L\}\_\{h\}^\{\\mathrm\{state\}\}evaluates prediction of the known synthetic latent state, whereasℒseq\\mathcal\{L\}\_\{\\mathrm\{seq\}\}evaluates the probability density assigned to the observed sequence under the complete probabilistic model\.
Experiment 1 additionally evaluates multi\-horizon structural consistency through the path\-disagreement metric
𝒟path\(h1,h2\)=𝔼t\[‖qtAh1\+h2−\(qtAh1\)Ah2‖1\]\.\\mathcal\{D\}\_\{\\mathrm\{path\}\}\(h\_\{1\},h\_\{2\}\)=\\mathbb\{E\}\_\{t\}\\left\[\\left\\\|q\_\{t\}A\_\{h\_\{1\}\+h\_\{2\}\}\-\(q\_\{t\}A\_\{h\_\{1\}\}\)A\_\{h\_\{2\}\}\\right\\\|\_\{1\}\\right\]\.\(Eq\.[52](https://arxiv.org/html/2608.13621#A6.E52)\)For MCJEPA with a single shared transition matrixAh=AhA\_\{h\}=A^\{h\}, this quantity is identically zero by construction\.
Experiment 2 additionally reports state accuracy,
Acc=1N∑t𝟏\{argmaxkqt\(k\)=St\},\\operatorname\{Acc\}=\\frac\{1\}\{N\}\\sum\_\{t\}\\mathbf\{1\}\\left\\\{\\arg\\max\_\{k\}q\_\{t\}\(k\)=S\_\{t\}\\right\\\},\(Eq\.[54](https://arxiv.org/html/2608.13621#A6.E54)\)the multiclass Brier score,
Brier=1N∑t∑k\(qt\(k\)−𝟏\{St=k\}\)2,\\operatorname\{Brier\}=\\frac\{1\}\{N\}\\sum\_\{t\}\\sum\_\{k\}\\left\(q\_\{t\}\(k\)\-\\mathbf\{1\}\\\{S\_\{t\}=k\\\}\\right\)^\{2\},\(Eq\.[55](https://arxiv.org/html/2608.13621#A6.E55)\)and mean posterior entropy,
H¯=1N∑tH\(qt\)\.\\overline\{H\}=\\frac\{1\}\{N\}\\sum\_\{t\}H\(q\_\{t\}\)\.\(Eq\.[56](https://arxiv.org/html/2608.13621#A6.E56)\)Accuracy evaluates hard state recovery, while NLL and Brier score retain information about probabilistic confidence\. Posterior entropy is descriptive rather than a stand\-alone performance criterion\.
Finally, Experiment 4 measures agreement between an exact model\-based filter and the amortized context encoder through
𝒟filter=𝔼t\[KL\(qtexact∥qtenc\)\]\.\\mathcal\{D\}\_\{\\mathrm\{filter\}\}=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(q\_\{t\}^\{\\mathrm\{exact\}\}\\,\\middle\\\|\\,q\_\{t\}^\{\\mathrm\{enc\}\}\\right\)\\right\]\.\(Eq\.[53](https://arxiv.org/html/2608.13621#A6.E53)\)Filtering KL is training aligned for the HMM\+filter and hybrid regimes because both explicitly optimize filtering distillation, so we interpret it primarily as a diagnostic of whether the context encoder has acquired the intended filtering role\.
Other experiment\-specific quantities, including effective state count, assignment entropy for collapse analysis, mutual\-information quantities in the predictive bottleneck, and residual history gainΔhist\\Delta\_\{\\mathrm\{hist\}\}, are defined where they are first introduced\. The systems and networks are intentionally small because the aim is structural diagnosis rather than scaling\. Detailed data\-generation procedures, architectures, optimization settings, metric definitions, and supplementary results are provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
### 7\.1Experiment 1: finite\-HMM recovery and Markov composition
We first test the basic MCJEPA construction on a finite HMM with four latent states and continuous observations\. We consider a separated\-emission regime, in which observations are highly informative about the state, and an ambiguous\-emission regime, in which state inference becomes more difficult\. MCJEPA uses a categorical history encoder and one shared row\-stochastic transition matrixAA, yielding
q^t\+h=qtAh\.\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\}\.We compare it with a horizon\-specific categorical JEPA that learns an independentAhA\_\{h\}for each prediction horizon, and with a correctly specified Gaussian HMM\.
#### Recovering states and transitions\.
[Table5](https://arxiv.org/html/2608.13621#S7.T5)reports ARI, aligned transition errorℰA\\mathcal\{E\}\_\{A\}, and true\-state prediction NLL at multiple horizons\. In the separated regime, all three methods recover the latent process well\. MCJEPA attains
ARI=0\.9796±0\.0055\\mathrm\{ARI\}=0\.9796\\pm 0\.0055and transition error
ℰA=0\.0191±0\.0019,\\mathcal\{E\}\_\{A\}=0\.0191\\pm 0\.0019,close to the correctly specified HMM\.
The ambiguous regime exposes the structural trade\-off more clearly\. The HMM remains strongest because it explicitly models the correct emission family and performs probabilistic filtering\. The horizon\-specific predictor obtains somewhat better state recovery than MCJEPA, but its aligned one\-step transition error is
0\.2410±0\.0029,0\.2410\\pm 0\.0029,compared with MCJEPA’s
0\.0901±0\.0091\.0\.0901\\pm 0\.0091\.Thus, independently fitting each horizon provides additional predictive flexibility but yields a substantially less faithful underlying transition law\.
Table 5:Finite\-HMM recovery\. Values are mean±\\pmstandard deviation over 5 seeds\.ℰA\\mathcal\{E\}\_\{A\}is the permutation\-aligned transition\-matrix error;ℒ1\\mathcal\{L\}\_\{1\}andℒ8\\mathcal\{L\}\_\{8\}are true\-state prediction NLL at horizons11and88\.Importantly, these results do not imply that the shared\-matrix constraint universally minimizes predictive NLL\. Under ambiguity, the independently parameterizedAhA\_\{h\}model is slightly better at several horizons\. The benefit of MCJEPA is instead structural: all horizons are generated by one transition mechanism and must therefore compose consistently\.
#### Exact path composition\.
The left panel of[Fig\.3](https://arxiv.org/html/2608.13621#S7.F3)illustrates the consequence of using a shared Markov transition\. To quantify disagreement between a direct prediction and a composed prediction with the same total horizon, we define
𝒟path\(h1,h2\)=𝔼t\[‖qtAh1\+h2−\(qtAh1\)Ah2‖1\],\\mathcal\{D\}\_\{\\mathrm\{path\}\}\(h\_\{1\},h\_\{2\}\)=\\mathbb\{E\}\_\{t\}\\left\[\\left\\\|q\_\{t\}A\_\{h\_\{1\}\+h\_\{2\}\}\-\(q\_\{t\}A\_\{h\_\{1\}\}\)A\_\{h\_\{2\}\}\\right\\\|\_\{1\}\\right\],\(29\)where the expectation is taken over valid evaluation time points\. Thus,𝒟path=0\\mathcal\{D\}\_\{\\mathrm\{path\}\}=0means that the direct and composed predicted state distributions agree exactly\.
The labels1\+11\+1,2\+22\+2, and4\+44\+4denote three decompositions of the same total prediction horizon\. Specifically,1\+11\+1compares a direct two\-step prediction with two successive one\-step predictions,
qtA2versus\(qtA1\)A1,q\_\{t\}A\_\{2\}\\quad\\text\{versus\}\\quad\(q\_\{t\}A\_\{1\}\)A\_\{1\},2\+22\+2compares a direct four\-step prediction with two successive two\-step predictions,
qtA4versus\(qtA2\)A2,q\_\{t\}A\_\{4\}\\quad\\text\{versus\}\\quad\(q\_\{t\}A\_\{2\}\)A\_\{2\},and4\+44\+4compares a direct eight\-step prediction with two successive four\-step predictions,
qtA8versus\(qtA4\)A4\.q\_\{t\}A\_\{8\}\\quad\\text\{versus\}\\quad\(q\_\{t\}A\_\{4\}\)A\_\{4\}\.
For MCJEPA,Ah=AhA\_\{h\}=A^\{h\}, so
qtAh1\+h2=\(qtAh1\)Ah2q\_\{t\}A^\{h\_\{1\}\+h\_\{2\}\}=\(q\_\{t\}A^\{h\_\{1\}\}\)A^\{h\_\{2\}\}exactly, and hence𝒟path=0\\mathcal\{D\}\_\{\\mathrm\{path\}\}=0by construction\. Independently learned horizon\-specific matricesAhA\_\{h\}, however, are not constrained to satisfy these composition identities\. In the ambiguous regime, their path disagreement is
0\.2138±0\.0041,0\.1321±0\.0053,0\.0645±0\.00720\.2138\\pm 0\.0041,\\qquad 0\.1321\\pm 0\.0053,\\qquad 0\.0645\\pm 0\.0072for the1\+11\+1,2\+22\+2, and4\+44\+4decompositions, respectively\. The corresponding separated\-regime disagreement is smaller but remains nonzero, confirming that independently trained horizon predictors need not define one coherent Markov chain\.
#### Preventing discrete\-state collapse\.
We next ablate the occupancy and entropy components of the state\-use regularizerℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}in[Eq\.10](https://arxiv.org/html/2608.13621#S2.E10)\.[Table6](https://arxiv.org/html/2608.13621#S7.T6)shows that the two terms address complementary failure modes\. Using both gives
ARI=0\.9748±0\.0134\\mathrm\{ARI\}=0\.9748\\pm 0\.0134and approximately four effective hard states\. Occupancy regularization alone maintains broad state use but leaves assignments highly uncertain, with mean assignment entropy
1\.1037±0\.1216\.1\.1037\\pm 0\.1216\.Entropy regularization alone instead makes assignments confident but usually collapses them onto a single state: four of the five runs use one effective state, and the average effective hard\-state count is only
1\.1703±0\.3807\.1\.1703\\pm 0\.3807\.Using neither term produces less severe but unstable state use and substantially weaker recovery\.
Table 6:Collapse ablation in Experiment 1\. Values are mean±\\pmstandard deviation over 5 seeds\.KeffhardK\_\{\\mathrm\{eff\}\}^\{\\mathrm\{hard\}\}is the effective number of states induced by hard assignments\. Assignment entropy measures per\-example confidence and should be interpreted jointly with state usage: very low entropy can also arise from single\-state collapse\. The occupancy and entropy terms are complementary: the former encourages global state use, while the latter encourages confident per\-example assignments\.The right panel of[Fig\.3](https://arxiv.org/html/2608.13621#S7.F3)visualizes the ARI column of[Table6](https://arxiv.org/html/2608.13621#S7.T6)\. Together, the state\-recovery and collapse diagnostics support our intended interpretation: occupancy prevents global state under\-use, whereas the entropy term prevents diffuse per\-sample assignments; both are needed to obtain confident and diverse state assignments without collapse\.


Figure 3:Structural diagnostics for Experiment 1\.*Left:*direct\-versus\-composed prediction disagreement𝒟path\\mathcal\{D\}\_\{\\mathrm\{path\}\}from[Eq\.29](https://arxiv.org/html/2608.13621#S7.E29)in the ambiguous\-emission regime\. The labels1\+11\+1,2\+22\+2, and4\+44\+4compare direct predictions at total horizons22,44, and88with predictions obtained by composing two successive predictions of horizons11,22, and44, respectively\. MCJEPA is exactly path\-consistent because all horizons are generated by powers of one shared transition matrix; independently learnedAhA\_\{h\}are not\.*Right:*visualization of the ARI column of[Table6](https://arxiv.org/html/2608.13621#S7.T6)under the four state\-regularization ablations\. Occupancy and entropy regularization are complementary, and using both gives the strongest state recovery\. Error bars denote mean±\\pmone standard deviation over 5 seeds\.The complete multi\-horizon curves, the separated\-regime path\-consistency result, and the state\-usage and assignment\-confidence ablations are provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
### 7\.2Experiment 2: filtering resolves emission ambiguity
Experiment 2 isolates the distinction between local observation evidence and filtering\. We use a persistent two\-state HMM whose emission distributions are made progressively more overlapping101010The two states have Gaussian emissions centered at−μ\-\\muand\+μ\+\\muwith common standard deviationσ\\sigma\. We control emission ambiguity through the separation ratioμ/σ∈\{2\.0,1\.25,0\.75,0\.45\}\\mu/\\sigma\\in\\\{2\.0,1\.25,0\.75,0\.45\\\}; decreasingμ/σ\\mu/\\sigmaincreases the overlap between the two emission distributions and therefore makes the current observation less informative about the latent state\.\. Because the generating model is known, we can compute both
local evidence:p\(St∣Xt\)\\text\{local evidence: \}p\(S\_\{t\}\\mid X\_\{t\}\)and
filtering:p\(St∣X≤t\)\\text\{filtering: \}p\(S\_\{t\}\\mid X\_\{\\leq t\}\)exactly\. In MCJEPA terms, these two quantities are the oracle counterparts of a local encoderqθ\(Zt∣Xt\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{t\}\)and a history\-dependent encoderqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\), respectively\. We use the exact posteriors here rather than learned encoders so that the experiment isolates the informational value of observation history without representation\-learning or optimization confounds; it is therefore not a comparison between an HMM and MCJEPA\.
[Figure4](https://arxiv.org/html/2608.13621#S7.F4)shows that the oracle filtering distribution becomes increasingly more informative than the oracle local\-evidence distribution as individual observations become ambiguous\. At the most overlapping setting, local evidence reaches state accuracy
0\.6730±0\.0037,0\.6730\\pm 0\.0037,whereas filtering reaches
0\.8433±0\.0080\.0\.8433\\pm 0\.0080\.The corresponding state NLL decreases from
0\.6022±0\.00300\.6022\\pm 0\.0030to
0\.3683±0\.0088\.0\.3683\\pm 0\.0088\.The advantage diminishes as the emissions become locally separable, as expected\.
The right panel of[Fig\.4](https://arxiv.org/html/2608.13621#S7.F4)illustrates the mechanism around a true state transition\. Local evidence fluctuates strongly with individual observations\. Filtering instead combines the current observation with the propagated state belief, remaining stable through many locally ambiguous measurements and changing when the accumulated evidence supports a transition\. This directly supports the probabilistic distinction made earlier in the paper:p\(St∣Xt\)p\(S\_\{t\}\\mid X\_\{t\}\)is local observation evidence, whereasp\(St∣X≤t\)p\(S\_\{t\}\\mid X\_\{\\leq t\}\)is the HMM filtering belief\.


Figure 4:Filtering under emission ambiguity\.*Left:*comparison of the exact local\-evidence and filtering distributions associated with the local and history\-dependent encoder roles in MCJEPA\. Filtering increasingly outperforms local evidence as emission overlap grows\. Error bars denote mean±\\pmone standard deviation over 5 seeds\.*Right:*a representative ambiguous sequence around a true state transition\. The local posterior reacts strongly to individual noisy observations, whereas the filtered belief integrates temporal evidence through the transition model\. These are oracle posteriors computed from the known generating process, rather than separately trained HMM and MCJEPA models\.The corresponding NLL curve is provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
### 7\.3Experiment 3: predictive compression and Markovization
Experiment 3 asks a simple question: can predictive compression discard unnecessary history while retaining exactly the information needed to predict the future?
We construct a binary second\-order process satisfying
p\(Xt\+1∣X≤t\)=p\(Xt\+1∣Xt−1,Xt\)\.p\(X\_\{t\+1\}\\mid X\_\{\\leq t\}\)=p\(X\_\{t\+1\}\\mid X\_\{t\-1\},X\_\{t\}\)\.\(30\)Thus, although the entire observation history is available, only the two most recent observations are needed to predictXt\+1X\_\{t\+1\}\.
We deliberately give the encoder a longer three\-step history,
Ht=\(Xt−2,Xt−1,Xt\),H\_\{t\}=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\),\(31\)which has eight possible values\. The known minimal predictive state is
Zt⋆=\(Xt−1,Xt\),Z\_\{t\}^\{\\star\}=\(X\_\{t\-1\},X\_\{t\}\),\(32\)which has only four possible values\. The older bitXt−2X\_\{t\-2\}is therefore redundant onceZt⋆Z\_\{t\}^\{\\star\}is known\.
This construction gives us a controlled ground truth for what predictive compression should do:
\(Xt−2,Xt−1,Xt\)⏟too much history⟶\(Xt−1,Xt\)⏟just enough⟶Xt⏟too little\.\\underbrace\{\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\)\}\_\{\\text\{too much history\}\}\\;\\longrightarrow\\;\\underbrace\{\(X\_\{t\-1\},X\_\{t\}\)\}\_\{\\text\{just enough\}\}\\;\\longrightarrow\\;\\underbrace\{X\_\{t\}\}\_\{\\text\{too little\}\}\.In MCJEPA terms, these are three controlled choices of latent state supplied to the predictor: an overcomplete stateZt=HtZ\_\{t\}=H\_\{t\}, the minimal sufficient stateZt=Zt⋆Z\_\{t\}=Z\_\{t\}^\{\\star\}, and an insufficient stateZt=XtZ\_\{t\}=X\_\{t\}\.
The middle representation also explains the term*Markovization*\. Although the observation process is second\-order inXtX\_\{t\}, defining
Zt⋆=\(Xt−1,Xt\)Z\_\{t\}^\{\\star\}=\(X\_\{t\-1\},X\_\{t\}\)turns it into a first\-order state process: the information needed for the next transition is contained in the current stateZt⋆Z\_\{t\}^\{\\star\}, without requiring older history\. Predictive compression should therefore removeXt−2X\_\{t\-2\}, but should not removeXt−1X\_\{t\-1\}\.
#### Does compression recover the correct predictive state?
We first compare the three controlled representations exactly\. Because the process is finite, their information quantities and optimal one\-step prediction losses can be computed without representation\-learning or optimization error\.
[Table7](https://arxiv.org/html/2608.13621#S7.T7)gives the key result\. The full three\-bit history retains
I\(Ht,Zt\)=1\.7356,\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)=1\.7356,whereas the four\-state predictive pair retains only
I\(Ht,Zt\)=1\.3378\.\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)=1\.3378\.Despite this compression, the two representations contain exactly the same information about the next observation,
I\(Zt,Xt\+1\)=0\.2807,\\mathrm\{I\}\(Z\_\{t\};X\_\{t\+1\}\)=0\.2807,and achieve the same prediction NLL,
Hence, removingXt−2X\_\{t\-2\}reduces the amount of past information stored in the state without sacrificing prediction\.
Compressing further toZt=XtZ\_\{t\}=X\_\{t\}, however, removes information that is genuinely needed\. Predictive information falls from0\.28070\.2807to0\.01920\.0192, and prediction NLL increases from0\.39780\.3978to0\.65930\.6593\. The controlled construction therefore has a known sufficiency–minimality boundary: eight states are predictively sufficient but redundant, four states are sufficient and minimal for this process, and two states are insufficient\.
Table 7:Exact representation controls for Experiment 3\. The minimal predictive pair removes redundant history while preserving all one\-step predictive information\. Compressing further to the current observation alone loses information required for prediction\.
#### Is the four\-state solution truly optimal, or just a favorable example?
The comparison above considers only three hand\-specified representations\. We therefore use the small history space to perform an exhaustive check over*every deterministic compression*of the eight possible histories represented byHtH\_\{t\}\.
A deterministic encoder
groups histories that are assigned to the same latent state\. We enumerate all such groupings and evaluate each one using
ℒPIBexp=ℒpred\+βI\(Ht,Zt\),\\mathcal\{L\}\_\{\\mathrm\{PIB\}\}^\{\\mathrm\{exp\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\beta\\,\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\),\(33\)where the first term rewards accurate prediction and the second penalizes retaining unnecessary information about the history\.
Because there are only eight possible histories, all41404140deterministic partitions can be enumerated exactly\. This provides a global deterministic reference rather than relying on a few hand\-designed candidates\. For every tested positive compression weight
0<β≤0\.014,0<\\beta\\leq 0\.014,the globally optimal deterministic representation is exactly the known four\-state predictive state
Zt⋆=\(Xt−1,Xt\)\.Z\_\{t\}^\{\\star\}=\(X\_\{t\-1\},X\_\{t\}\)\.Thus, when compression is strong enough to penalize redundant history but not so strong that predictive information is sacrificed, the predictive\-bottleneck objective selects the known minimal sufficient Markov state\.
Atβ=0\\beta=0, several predictively equivalent deterministic representations attain the same minimum prediction loss; the four\-state state is therefore not identified by prediction alone\. Onceβ\>0\\beta\>0, however, redundant stored history is penalized\. Atβ=0\.015\\beta=0\.015, the deterministic optimum changes to a two\-state representation\. Its retained predictive information decreases from0\.28070\.2807to0\.27110\.2711, indicating that compression has begun to remove information useful for prediction\. With still stronger compression, the optimum eventually collapses to a single state\. The resulting progression is therefore
redundant representation⟶minimal predictive state⟶over\-compressed state\.\\text\{redundant representation\}\\;\\longrightarrow\\;\\text\{minimal predictive state\}\\;\\longrightarrow\\;\\text\{over\-compressed state\}\.
The left panel of[Fig\.5](https://arxiv.org/html/2608.13621#S7.F5)visualizes the compression–prediction trade\-off in two complementary ways\. Each light\-blue point corresponds to one of the41404140deterministic partitions of the eight possible histories, positioned according to the amount of history information it retains,I\(Ht,Zt\)\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)on the horizontal axis, and the amount of predictive information it preserves,I\(Zt,Xt\+1\)\\mathrm\{I\}\(Z\_\{t\};X\_\{t\+1\}\)on the vertical axis\. The blue curve connects the nondominated deterministic solutions and therefore gives the exact deterministic reference frontier\.
The orange curve is obtained differently\. We initialize a stochastic encoder at the overcomplete eight\-history representation and follow a warm\-started continuation path as the compression weightβ\\betais increased\. Each orange marker shows the representation obtained at one value ofβ\\beta\. Atβ=0\\beta=0, the overcomplete initialization is retained with no compression pressure\. For subsequent values, increasingβ\\betamakes representations with smallerI\(Ht,Zt\)\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)increasingly preferable\. The learned solution is therefore encouraged to move leftward in the information plane\. Ideally, this removes redundant history while remaining near the top of the plot, where predictive information is preserved\. Ifβ\\betabecomes too large, however, compression also removes information needed for prediction and the trajectory moves downward\.
The labeledβ\\betavalues do not represent different data\-generating processes; they are different settings of the same predictive\-compression objective and trace how the learned representation changes as compression pressure increases\. The learned trajectory need not coincide with the exact blue frontier because the encoder is stochastic and is optimized by gradient descent, whereas the blue frontier is obtained by exhaustive enumeration over deterministic partitions\. We therefore use the deterministic frontier as a global reference and the orange continuation path as a practical illustration of how predictive compression behaves during learning\. Full enumeration details, the completeβ\\betasweep, continuation optimization settings, and the corresponding state\-count trajectories are provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
#### Did compression remove too much information?
The previous results identify which representations achieve a favorable trade\-off between compression and prediction\. We next ask a complementary question: can we detect when compression has gone too far and removed information that is still useful for predicting the future?
We instantiate the residual\-history diagnostic from[Section6](https://arxiv.org/html/2608.13621#S6)using one additional step of representation history\. The intuition is simple\. IfZtZ\_\{t\}already contains all information needed to predictXt\+1X\_\{t\+1\}, then additionally conditioning on the previous representation stateZt−1Z\_\{t\-1\}should not improve held\-out prediction\. Conversely, ifZt−1Z\_\{t\-1\}still contains transition\-relevant information that is absent fromZtZ\_\{t\}, then the current representation is predictively insufficient\.
We test the same three controlled representations:
Ztover=\(Xt−2,Xt−1,Xt\),Ztsuff=\(Xt−1,Xt\),Ztunder=Xt\.Z\_\{t\}^\{\\mathrm\{over\}\}=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\),\\qquad Z\_\{t\}^\{\\mathrm\{suff\}\}=\(X\_\{t\-1\},X\_\{t\}\),\\qquad Z\_\{t\}^\{\\mathrm\{under\}\}=X\_\{t\}\.
For each representation, we compare two predictors ofXt\+1X\_\{t\+1\}\. The*restricted*predictor uses only the current representation,
p^restricted=p^\(Xt\+1=1∣Zt\),\\widehat\{p\}\_\{\\mathrm\{restricted\}\}=\\widehat\{p\}\(X\_\{t\+1\}=1\\mid Z\_\{t\}\),whereas the*history\-augmented*predictor additionally receives the previous representation state,
p^history\-augmented=p^\(Xt\+1=1∣Zt,Zt−1\)\.\\widehat\{p\}\_\{\\mathrm\{history\\text\{\-\}augmented\}\}=\\widehat\{p\}\(X\_\{t\+1\}=1\\mid Z\_\{t\},Z\_\{t\-1\}\)\.We define the residual history gain as
Δhist=MSErestricted−MSEhistory\-augmented\.\\Delta\_\{\\mathrm\{hist\}\}=\\mathrm\{MSE\}\_\{\\mathrm\{restricted\}\}\-\\mathrm\{MSE\}\_\{\\mathrm\{history\\text\{\-\}augmented\}\}\.\(34\)Thus,
Δhist\>0\\Delta\_\{\\mathrm\{hist\}\}\>0means that the previous representation state contains predictive information not already captured byZtZ\_\{t\}\. By contrast,
Δhist≈0\\Delta\_\{\\mathrm\{hist\}\}\\approx 0means that adding one further step of representation history provides essentially no additional predictive benefit\. The lookup predictors, chronological train–test split, and fitting procedure are detailed in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
The three controlled representations make this diagnostic especially transparent\. For the insufficient representation
Ztunder=Xt,Z\_\{t\}^\{\\mathrm\{under\}\}=X\_\{t\},the previous representation is simply
Zt−1under=Xt−1\.Z\_\{t\-1\}^\{\\mathrm\{under\}\}=X\_\{t\-1\}\.Thus, the history\-augmented predictor restores exactly the variable omitted fromZtZ\_\{t\}that is required by the second\-order transition law\. As expected, this produces a substantial held\-out prediction gain,
Δhist=0\.1139±0\.0026\.\\Delta\_\{\\mathrm\{hist\}\}=0\.1139\\pm 0\.0026\.This is the intended positive control:Zt=XtZ\_\{t\}=X\_\{t\}is insufficient because the omittedXt−1X\_\{t\-1\}remains informative aboutXt\+1X\_\{t\+1\}\.
For the minimal sufficient representation,
Ztsuff=\(Xt−1,Xt\),Z\_\{t\}^\{\\mathrm\{suff\}\}=\(X\_\{t\-1\},X\_\{t\}\),we have
Zt−1suff=\(Xt−2,Xt−1\)\.Z\_\{t\-1\}^\{\\mathrm\{suff\}\}=\(X\_\{t\-2\},X\_\{t\-1\}\)\.SinceXt−1X\_\{t\-1\}is already contained inZtZ\_\{t\}, the only genuinely additional observation supplied byZt−1Z\_\{t\-1\}isXt−2X\_\{t\-2\}, which is redundant for predictingXt\+1X\_\{t\+1\}by construction\. Correspondingly,
Δhist≈−6\.8×10−6\.\\Delta\_\{\\mathrm\{hist\}\}\\approx\-6\.8\\times 10^\{\-6\}\.
For the overcomplete representation,
Ztover=\(Xt−2,Xt−1,Xt\),Z\_\{t\}^\{\\mathrm\{over\}\}=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\),the previous representation
Zt−1over=\(Xt−3,Xt−2,Xt−1\)Z\_\{t\-1\}^\{\\mathrm\{over\}\}=\(X\_\{t\-3\},X\_\{t\-2\},X\_\{t\-1\}\)adds only still older information beyond what is already available inZtZ\_\{t\}\. We obtain
Δhist≈−3\.7×10−5\.\\Delta\_\{\\mathrm\{hist\}\}\\approx\-3\.7\\times 10^\{\-5\}\.Both near\-zero values are negligible at the scale of the experiment\. The tiny negative values are attributable to finite\-sample fitting variation rather than a meaningful advantage of the restricted predictor\. Once the current representation already contains all transition\-relevant information, addingZt−1Z\_\{t\-1\}does not improve held\-out prediction\.
Viewed together, the three controlled cases reveal a clear sufficiency–minimality boundary\. Compressing from the overcomplete eight\-state representation to the four\-state minimal representation reducesI\(Ht,Zt\)\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)from1\.73561\.7356to1\.33781\.3378while leavingΔhist\\Delta\_\{\\mathrm\{hist\}\}effectively zero, indicating that redundant history has been removed without sacrificing predictive sufficiency\. Compressing further to the two\-state representation reducesI\(Ht,Zt\)\\mathrm\{I\}\(H\_\{t\};Z\_\{t\}\)to0\.67850\.6785, butΔhist\\Delta\_\{\\mathrm\{hist\}\}rises sharply to0\.1139±0\.00260\.1139\\pm 0\.0026: the previous representationZt−1=Xt−1Z\_\{t\-1\}=X\_\{t\-1\}now contains substantial transition\-relevant information missing fromZt=XtZ\_\{t\}=X\_\{t\}\. Thus, in this controlled example, the four\-state representation lies at the natural elbow between retaining redundant history and compressing away information required for prediction\.
This result highlights the distinction between*sufficiency*and*minimality*\. The residual\-history diagnostic tests sufficiency: it correctly identifiesZt=XtZ\_\{t\}=X\_\{t\}as missing predictive information, while both the four\-state and eight\-state representations pass because their current state already contains all information required for one\-step prediction\. The diagnostic cannot, however, determine that the eight\-state representation stores redundant history\. The predictive bottleneck supplies this complementary notion of minimality by preferring the smaller four\-state representation among predictively sufficient alternatives\.


Figure 5:Predictive compression and Markovization\.*Left:*exhaustive deterministic compression provides a global reference for the prediction–compression trade\-off, while the learned stochastic encoder traces a warm\-started continuation path as the compression weightβ\\betais increased\. Moderate compression can remove redundant history while preserving predictive information, whereas excessive compression eventually sacrifices information required for prediction\.*Right:*residual history gainΔhist\\Delta\_\{\\mathrm\{hist\}\}diagnoses predictive insufficiency by testing whether adding the previous representation stateZt−1Z\_\{t\-1\}improves held\-out prediction beyond usingZtZ\_\{t\}alone\. This augmentation substantially helps the insufficient representationZt=XtZ\_\{t\}=X\_\{t\}, for whichZt−1=Xt−1Z\_\{t\-1\}=X\_\{t\-1\}restores omitted transition\-relevant information, but provides essentially no additional predictive benefit for either the minimal sufficient or the overcomplete representation\. Error bars denote mean±\\pmone standard deviation over 5 seeds where applicable\.Together, the two diagnostics play complementary roles\. Predictive compression asks how much of the past can be discarded while preserving future prediction, thereby favoring a compact Markov state\. Residual predictability asks whether compression has discarded too much: a positiveΔhist\\Delta\_\{\\mathrm\{hist\}\}indicates that the previous representation state contains transition\-relevant information not already captured byZtZ\_\{t\}\. In this controlled process, the four\-state representation is the known minimal sufficient target: it preserves all one\-step predictive information while storing less history than the overcomplete eight\-state representation\. Detailed enumeration, optimization settings, complete compression sweeps, residual\-predictor specifications, and supplementary plots are provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
### 7\.4Experiment 4: HMM\-style training of PIB\-VJEPA
Experiment 4 examines the final and strongest level of correspondence considered in this paper:*model\-and\-objective equivalence*\. Even when JEPA and HMM formulations share an emission\-complete latent\-state representation, or satisfy the conditions for sequence\-level HMM equivalence, they need not be trained by the same objective\. Standard JEPA training optimizes prediction in latent space, whereas HMM training additionally optimizes the probability of the observed sequence through an explicit transition–emission model\.
To isolate this distinction, we use the same four\-state latent family and separated\-emission data\-generating process as Experiment 1, but train all three Experiment 4 regimes independently\. Because the latent state is categorical and the observation model is Gaussian, observation\-sequence likelihood and filtering posteriors can be evaluated exactly using the HMM forward recursion\. We compare three regimes—JEPA\-only, HMM\-style, and their hybrid—while keeping the latent\-state family and amortized context\-encoder architecture fixed\.
#### Training regimes\.
The first regime is the*JEPA latent objective*\. It uses the same MCJEPA construction as Experiment 1:
ℒJEPA=ℒMC\+ℒstate,\\mathcal\{L\}\_\{\\mathrm\{JEPA\}\}=\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{state\}\},\(35\)whereℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}is defined in[Eq\.8](https://arxiv.org/html/2608.13621#S2.E8)andℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}in[Eq\.10](https://arxiv.org/html/2608.13621#S2.E10)\. The online context encoder produces
qθ\(Zt∣X≤t\),q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\),while the EMA target encoder produces the local future target
qθ¯\(Zt\+h∣Xt\+h\)\.q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\)\.A single row\-stochastic transition matrix generates all horizons throughAhA^\{h\}\. This regime therefore represents the latent\-prediction viewpoint: no observation\-sequence likelihood or filtering target influences representation learning\. For evaluation of sequence NLL, a Gaussian observation model is fitted only after JEPA training and consequently does not influence the learned representation\.
The second regime is*HMM sequence \+ filter distillation*\. Here, the latent transition model
pϕ\(Zt\+1∣Zt\)p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\}\)and Gaussian emission model
pψ\(Xt∣Zt\)p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)are trained through exact observation\-sequence negative log\-likelihood,
ℒHMM=−𝔼\[1Tlogpϕ,ψ\(X1:T\)\],\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}=\-\\mathbb\{E\}\\left\[\\frac\{1\}\{T\}\\log p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)\\right\],\(36\)computed by the HMM forward algorithm\.
Sequence likelihood trains the generative transition–emission model, but it does not by itself require the amortized context encoder to represent the corresponding HMM filtering belief\. We therefore additionally distill the exact filtering posterior into the context encoder through
ℒfilter=𝔼t\[KL\(sg\(q~t\)∥qθ\(Zt∣X≤t\)\)\]\.\\mathcal\{L\}\_\{\\mathrm\{filter\}\}=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\(\\widetilde\{q\}\_\{t\}\)\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\right\)\\right\]\.\(37\)where
q~t=pϕ,ψ\(Zt∣X≤t\)\\widetilde\{q\}\_\{t\}=p\_\{\\phi,\\psi\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)denotes the exact filtering posterior under the current HMM transition and emission models\.111111Bothq~t\\widetilde\{q\}\_\{t\}andqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)represent a belief over the current latent state after observations through timetthave been assimilated\. The exact HMM filter first propagates the previous filtering belief through the transition kernel,πt−\(zt\)=∑zt−1pϕ\(zt∣zt−1\)q~t−1\(zt−1\),\\pi\_\{t\}^\{\-\}\(z\_\{t\}\)=\\sum\_\{z\_\{t\-1\}\}p\_\{\\phi\}\(z\_\{t\}\\mid z\_\{t\-1\}\)\\widetilde\{q\}\_\{t\-1\}\(z\_\{t\-1\}\),and then incorporates the current observation through the emission likelihood,q~t\(zt\)=pψ\(Xt∣zt\)πt−\(zt\)∑zt′pψ\(Xt∣zt′\)πt−\(zt′\)\.\\widetilde\{q\}\_\{t\}\(z\_\{t\}\)=\\frac\{p\_\{\\psi\}\(X\_\{t\}\\mid z\_\{t\}\)\\,\\pi\_\{t\}^\{\-\}\(z\_\{t\}\)\}\{\\sum\_\{z\_\{t\}^\{\\prime\}\}p\_\{\\psi\}\(X\_\{t\}\\mid z\_\{t\}^\{\\prime\}\)\\,\\pi\_\{t\}^\{\-\}\(z\_\{t\}^\{\\prime\}\)\}\.Thus,ϕ\\phidetermines how probability mass is propagated between latent states, whereasψ\\psidetermines how the current observation updates that predictive prior\. For a discrete HMM with transition matrixAA,q~t\(j\)=pψ\(Xt∣Zt=j\)∑iq~t−1\(i\)Aij∑j′pψ\(Xt∣Zt=j′\)∑iq~t−1\(i\)Aij′\.\\widetilde\{q\}\_\{t\}\(j\)=\\frac\{p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}=j\)\\sum\_\{i\}\\widetilde\{q\}\_\{t\-1\}\(i\)A\_\{ij\}\}\{\\sum\_\{j^\{\\prime\}\}p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}=j^\{\\prime\}\)\\sum\_\{i\}\\widetilde\{q\}\_\{t\-1\}\(i\)A\_\{ij^\{\\prime\}\}\}\.Henceℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}compares like with like: it distills the exact current\-state filtering belief into the amortized context encoder rather than comparing the context encoder with the pre\-observation predictive prior\.
The complete HMM\-style regime uses
ℒHMM\+filter=λseqℒHMM\+λfilterℒfilter\+ℒstate\.\\mathcal\{L\}\_\{\\mathrm\{HMM\+filter\}\}=\\lambda\_\{\\mathrm\{seq\}\}\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+\\lambda\_\{\\mathrm\{filter\}\}\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\+\\mathcal\{L\}\_\{\\mathrm\{state\}\}\.\(38\)Importantly, this regime contains*no JEPA latent\-prediction lossℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}*\.
The third regime is the*hybrid HMM \+ latent*objective:
ℒhybrid=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{hybrid\}\}=\{\}λseqℒHMM⏟sequence evidence\+λlatentℒMC⏟JEPA latent prediction\\displaystyle\\underbrace\{\\lambda\_\{\\mathrm\{seq\}\}\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\}\_\{\\text\{sequence evidence\}\}\+\\underbrace\{\\lambda\_\{\\mathrm\{latent\}\}\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\}\_\{\\text\{JEPA latent prediction\}\}\(39\)\+λfilterℒfilter⏟filtering\-posterior alignment\+ℒstate⏟state\-use regularization\.\\displaystyle\+\\underbrace\{\\lambda\_\{\\mathrm\{filter\}\}\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\}\_\{\\text\{filtering\-posterior alignment\}\}\+\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{state\}\}\}\_\{\\text\{state\-use regularization\}\}\.Importantly,ℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}is implemented in exactly the same way as in the JEPA\-only regime: the future latent target is produced by the EMA target encoder,
qθ¯\(Zt\+h∣Xt\+h\),q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\),and the context prediction is generated by the shared transition,
qθ\(Zt∣X≤t\)Ah\.q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)A^\{h\}\.The exact HMM filtering posteriorq~t\\widetilde\{q\}\_\{t\}is used only inℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}and does not replace the JEPA target inℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\. Moreover, the transition matrix used by the HMM sequence model is the same transition matrix used by the latent\-prediction objective, so both training signals act on the same latent dynamics\.
The state\-use regularizer and its coefficients are shared across all three encoder\-training regimes\. The remaining objective weights, initialization scheme, and optimization settings are given in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
#### Motivation and comparison design\.
The three regimes form a controlled objective\-level comparison\. The*JEPA latent objective*uses latent predictive alignment but no observation\-sequence likelihood or filtering target\. The*HMM sequence \+ filter distillation*regime does the converse: it uses observation\-sequence likelihood and filtering\-posterior supervision but no JEPA latent\-prediction loss\. The*hybrid*regime adds both HMM\-style signals to the same MCJEPA latent\-prediction objective used by the JEPA\-only baseline\.
This comparison addresses three related questions\. First, does adding HMM\-style sequence and filtering supervision improve MCJEPA relative to latent\-only training? Second, can the hybrid retain the genuine JEPA latent\-prediction objective while approaching the probabilistic\-model recovery achieved by HMM\-style training? Third, isℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}necessary for training this latent\-state architecture at all, or can the same architecture instead be trained through HMM\-style sequence and filtering supervision?
Table 8:Objective\-level comparison in Experiment 4\. The three objective columns indicate which training signals are active: HMM sequence likelihoodℒHMM\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}, JEPA/Markov latent predictionℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}, and filtering\-posterior distillationℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\. The state\-use regularizerℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}is shared across all three regimes\. Values are mean±\\pmstandard deviation over five seeds\. Boldface marks the numerically best mean and does not imply statistical significance\. Filtering KL is training aligned for the HMM\+filter and hybrid regimes because both explicitly optimizeℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}; for JEPA\-only it is evaluated post hoc\.
#### Results\.
[Table8](https://arxiv.org/html/2608.13621#S7.T8)reveals a clear objective\-level distinction\. The JEPA\-only model successfully learns a meaningful predictive latent state, reaching
ARI=0\.9802±0\.0053,\\mathrm\{ARI\}=0\.9802\\pm 0\.0053,but its observation\-sequence NLL is
1\.2753±0\.0069,1\.2753\\pm 0\.0069,and its transition\-recovery error is
ℰA=0\.0191±0\.0020\.\\mathcal\{E\}\_\{A\}=0\.0191\\pm 0\.0020\.This is consistent with what the objective directly supervises:ℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}trains predictive agreement in latent space, but does not directly maximize observation\-sequence likelihood or jointly train an emission model with the representation\.
Adding HMM\-style supervision produces a substantial improvement\. The hybrid retains the same EMA\-target MCJEPA loss but additionally optimizes sequence likelihood and filtering alignment\. Its ARI rises to
0\.9905±0\.0008,0\.9905\\pm 0\.0008,its sequence NLL decreases to
1\.2660±0\.0073,1\.2660\\pm 0\.0073,and its transition error falls to
ℰA=0\.0103±0\.0017\.\\mathcal\{E\}\_\{A\}=0\.0103\\pm 0\.0017\.Relative to JEPA\-only training, this corresponds to an approximately46%46\\%reduction in transition\-matrix error\. Moreover, the improvement is seed\-consistent: for each of the five random seeds, the hybrid improves over JEPA\-only on ARI, NMI, sequence NLL, transition recovery, filtering KL, and true\-state prediction NLL at every evaluated horizonh∈\{1,2,4,8\}h\\in\\\{1,2,4,8\\\}\. Thus, the gain from HMM\-style supervision is not driven by a single favorable run\.
The hybrid also nearly closes the observation\-sequence likelihood gap to HMM\-style training\. The HMM\+filter regime achieves sequence NLL
1\.265885±0\.007220,1\.265885\\pm 0\.007220,whereas the hybrid obtains
1\.266019±0\.007271\.1\.266019\\pm 0\.007271\.Their difference is only
1\.34×10−41\.34\\times 10^\{\-4\}NLL per time step, compared with a JEPA\-to\-HMM gap of approximately
9\.41×10−3\.9\.41\\times 10^\{\-3\}\.Equivalently, the hybrid closes approximately98\.6%98\.6\\%of the JEPA\-only sequence\-NLL gap to HMM\-style training while retaining the genuine JEPA latent\-prediction objective\.
Transition recovery shows the same qualitative result\. The hybrid has the numerically smallest mean error,
0\.0103±0\.0017,0\.0103\\pm 0\.0017,compared with
0\.0105±0\.00140\.0105\\pm 0\.0014for HMM\+filter and
0\.0191±0\.00200\.0191\\pm 0\.0020for JEPA\-only\. The difference between hybrid and HMM\-style training is small relative to the across\-seed variability, so we interpret the two as achieving comparable transition recovery rather than claiming that the hybrid is superior to the correctly specified HMM objective\. The important contrast is that both recover the transition substantially more faithfully than latent\-only JEPA training\.
The HMM\+filter regime provides the complementary result\. Despite containing noℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}, it achieves the strongest state recovery,
ARI=0\.9945±0\.0008,\\mathrm\{ARI\}=0\.9945\\pm 0\.0008,and
NMI=0\.9884±0\.0014,\\mathrm\{NMI\}=0\.9884\\pm 0\.0014,together with the best sequence NLL and transition recovery comparable to the hybrid\. Thus, the JEPA latent\-prediction objective is not required to train this categorical latent\-state architecture successfully: the same architecture can instead be trained using HMM\-style sequence and filtering supervision, together with the common state\-use regularizer\.
The multi\-horizon prediction results reinforce this conclusion\. At every evaluated horizonh∈\{1,2,4,8\}h\\in\\\{1,2,4,8\\\}, both HMM\-style and hybrid training achieve lower true\-state prediction NLL than JEPA\-only training\. The differences are largest at shorter horizons and diminish at longer horizons as the transition dynamics mix\. Complete multi\-horizon results are provided in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
#### Filtering as a role diagnostic\.
Filtering KL requires a different interpretation from sequence NLL and transition recovery\. The HMM\+filter regime reaches
KLfilter=0\.000060±0\.000007,\\mathrm\{KL\}\_\{\\mathrm\{filter\}\}=0\.000060\\pm 0\.000007,which is expected because its amortized encoder is explicitly trained to reproduce the exact HMM filtering posterior\. The hybrid obtains
0\.004355±0\.001015,0\.004355\\pm 0\.001015,while JEPA\-only gives
0\.025460±0\.006423\.0\.025460\\pm 0\.006423\.For the HMM\+filter and hybrid regimes this quantity is training aligned and should therefore be interpreted as a diagnostic that the intended filtering role has been learned, rather than as an independent generalization metric\. For JEPA\-only, by contrast, the filtering distribution is constructed only after fitting the post\-hoc observation model, so its filtering KL measures how closely latent\-only representation learning happens to agree with the filter induced by that fitted probabilistic model\.
The ordering is nevertheless informative about the roles induced by the different objectives\. HMM\+filter explicitly learns an amortized filter; the hybrid remains substantially aligned with that filtering interpretation while simultaneously satisfying the EMA\-target JEPA objective; and JEPA\-only has no requirement that its history encoder coincide with a Bayesian filtering belief\. The corresponding diagnostic is reported separately in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
[Figure6](https://arxiv.org/html/2608.13621#S7.F6)focuses on the two metrics that most directly expose the objective\-level distinction\. The left panel reports observation\-sequence NLL relative to HMM\-style training\. The hybrid lies almost on the HMM reference, whereas JEPA\-only retains a clear positive gap\. The right panel reports permutation\-aligned transition recovery error: HMM\-style and hybrid training form a closely matched pair, while JEPA\-only exhibits substantially larger error\.


Figure 6:Objective\-level comparison in Experiment 4\.*Left:*observation\-sequence NLL per time step relative to the HMM sequence \+ filter\-distillation regime\. Hybrid training nearly closes the entire JEPA\-to\-HMM sequence\-likelihood gap while retaining the genuine MCJEPA latent\-prediction objective\.*Right:*permutation\-aligned transition\-matrix recovery error\. HMM\-style and hybrid training achieve closely matched transition recovery, and both substantially outperform latent\-only JEPA training\. Error bars denote mean±\\pmone standard deviation over five seeds\.Taken together, these results support the distinction developed earlier between an HMM\-compatible latent\-state representation and full model\-and\-objective equivalence\. The same HMM\-compatible latent\-state architecture can support JEPA\-style latent prediction, HMM\-style probabilistic sequence learning, or a combination of the two, but sharing the underlying probabilistic model class does not imply that different training objectives recover the same fitted model\. Explicit sequence likelihood supplies transition–emission supervision that latent prediction alone does not provide, while filtering distillation connects the resulting HMM posterior back to the amortized PIB\-VJEPA context encoder\.
Experiment 4 supports two particularly important practical conclusions\. First, incorporating HMM\-style sequence and filtering supervision into MCJEPA improves recovery of the underlying probabilistic latent\-state model while retaining the original JEPA latent\-prediction objective\. The hybrid improves over JEPA\-only training on every reported metric for every seed, nearly matches HMM\-style sequence likelihood, and recovers the transition dynamics at essentially the same level as HMM\-style training\. Second, the same latent\-state architecture can be trained successfully withoutℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}: the HMM sequence \+ filter\-distillation regime attains the strongest state recovery and sequence likelihood despite omitting the JEPA latent\-prediction objective altogether\. Thus, the distinction between MCJEPA and an HMM is not determined by architecture alone; it also depends fundamentally on the objective used to train that architecture\.
More generally, the three regimes expose a continuum of training objectives on the same latent\-state family:
ℒMC⏟JEPA\-style training⟷ℒMC\+ℒHMM\+ℒfilter⏟hybrid training⟷ℒHMM\+ℒfilter⏟HMM\-style training,\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\}\_\{\\text\{JEPA\-style training\}\}\\quad\\longleftrightarrow\\quad\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\}\_\{\\text\{hybrid training\}\}\\quad\\longleftrightarrow\\quad\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\}\_\{\\text\{HMM\-style training\}\},withℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}shared as a state\-use regularizer\. The hybrid demonstrates that HMM\-style probabilistic sequence learning can be incorporated without abandoning JEPA\-style latent prediction, while the HMM\-style end of the spectrum shows that the same architecture can also be trained without latent\-prediction supervision altogether\. This objective continuum makes precise the paper’s broader claim: probabilistic temporal JEPA and HMMs can share an underlying latent Markov architecture while differing in how strongly HMM\-equivalent probabilistic semantics are enforced by training\.
Further optimization details, complete multi\-horizon results, and the training\-aligned filtering diagnostic are reported in Appendix[F](https://arxiv.org/html/2608.13621#A6)\.
### 7\.5Summary of Experiments
The four experiments test complementary and progressively stronger aspects of the proposed HMM interpretation of probabilistic temporal JEPA\. Experiment 1 establishes the finite\-state Markov structure: MCJEPA learns an explicit shared transition matrix whose powers generate all prediction horizons, thereby guaranteeing exact direct\-versus\-composed consistency\. The correctly specified Gaussian HMM remains strongest when emissions are ambiguous, while independently trained horizon\-specific predictors can gain some predictive flexibility at the cost of a less faithful and non\-compositional transition law\. The collapse ablations further show that occupancy and entropy regularization play complementary roles in stable discrete\-state learning\.
Experiment 2 isolates the inference role of the context encoder\. Using exact oracle posteriors, it shows that history\-based filteringp\(St∣X≤t\)p\(S\_\{t\}\\mid X\_\{\\leq t\}\)becomes increasingly more informative than local evidencep\(St∣Xt\)p\(S\_\{t\}\\mid X\_\{t\}\)as emissions overlap\. This supports the interpretation of a history\-dependent context encoderqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)as an amortized filtering belief rather than an emission model\.
Experiment 3 addresses how such a Markov state can be constructed\. In a controlled second\-order binary process, predictive compression removes the redundant history variableXt−2X\_\{t\-2\}and selects the known minimal sufficient stateZt⋆=\(Xt−1,Xt\)Z\_\{t\}^\{\\star\}=\(X\_\{t\-1\},X\_\{t\}\)without loss of predictive information, whereas further compression becomes insufficient\. Exhaustive enumeration of all deterministic history partitions provides a global reference for the prediction–compression trade\-off, while the residual\-history diagnostic supplies the complementary sufficiency test: adding the previous representation stateZt−1Z\_\{t\-1\}provides essentially no predictive benefit onceZtZ\_\{t\}is sufficient, but yields a large held\-out gain whenZt=XtZ\_\{t\}=X\_\{t\}has discarded the transition\-relevant variableXt−1X\_\{t\-1\}\. Together, these results separate*minimality*from*sufficiency*and show how predictive compression can Markovize an observation process by constructing a compact predictive state\.
Finally, Experiment 4 makes the objective\-level distinction explicit through a controlled comparison on the same latent\-state family\. The architecture can be trained with the original JEPA\-style latent\-prediction objective, HMM\-style sequence likelihood and filtering distillation, or a hybrid containing all three signals\. Adding HMM\-style supervision to the genuine MCJEPA objective improves state recovery, observation\-sequence likelihood, transition recovery, filtering agreement, and multi\-horizon prediction relative to latent\-only training\. The hybrid nearly closes the entire JEPA\-to\-HMM sequence\-likelihood gap and recovers the transition dynamics at essentially the same level as HMM\-style training, while retaining the original EMA\-target JEPA latent\-prediction objective\. Conversely, the HMM\-style regime achieves the strongest state recovery and sequence likelihood despite using noℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}, showing that the same latent\-state architecture can also be trained successfully without JEPA latent\-prediction supervision\.
Taken together, the experiments support a progressively stronger view of probabilistic temporal JEPA: it can instantiate a coherent latent Markov architecture; its context encoder can acquire the role of a filtering distribution; predictive compression can construct a compact sufficient Markov state; and the same architecture can be trained along a continuum from JEPA\-style latent prediction, through hybrid JEPA–HMM learning, to HMM\-style probabilistic sequence learning withoutℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\. The degree of HMM equivalence therefore depends not only on architectural structure, but also on the probabilistic components and, critically, the objective used to train them\.
## 8Discussion
#### What is “secretly an HMM”?
The central claim is structural and probabilistic, but not unconditional\. Full, time\-indexed PIB\-VJEPA exposes the same three computational roles as an HMM: inference of a latent\-state belief from observations, propagation of that state through a Markov transition, and a state\-to\-observation map\. The correspondence becomes progressively stronger across the four levels developed in this paper: computational correspondence, emission\-complete latent\-state representation, sequence\-level HMM equivalence, and model\-and\-objective equivalence\. In particular, the stochastic encoder is*not*an emission model; it plays the recognition or filtering role\. The emission direction may instead be supplied by a decoder, by the inverse of an invertible target encoder, or implicitly through the Bayes\-consistent conditional induced by a local stochastic encoder\. Reaching sequence\-level HMM equivalence further requires Markov, marginal\-consistency, and filtering\-consistency conditions, as formalized in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1); reaching model\-and\-objective equivalence additionally requires HMM\-compatible sequence\-level probabilistic training\.
#### Beyond MCJEPA: when is a general JEPA predictor Markov?
The HMM correspondence is not specific to MCJEPA, nor does a neural\-network predictor cease to be Markov merely because it is nonlinear or highly expressive\. Markovianity is a conditional\-independence property rather than a restriction on the functional form of the predictor\. A general temporal JEPA may use an arbitrary neural transition
pϕ\(Zt\+1∣Zt,ξt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\),implemented, for example, by an MLP, Transformer, mixture model, or another conditional density estimator\. It remainsfirst\-orderMarkov with respect toZtZ\_\{t\}whenever
pϕ\(Zt\+1∣Z≤t,ξ≤t\)=pϕ\(Zt\+1∣Zt,ξt\)\.p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{\\leq t\},\\xi\_\{\\leq t\}\)=p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\.The transition can therefore be arbitrarily nonlinear; for example,
pϕ\(Zt\+1∣Zt,ξt\)=𝒩\(Zt\+1,μϕ\(Zt,ξt\),Σϕ\(Zt,ξt\)\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)=\\mathcal\{N\}\\\!\\left\(Z\_\{t\+1\};\\mu\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\),\\Sigma\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\\right\),while a deterministic temporal JEPA is recovered through the Dirac kernel
pϕ\(Zt\+1∣Zt,ξt\)=δ\(Zt\+1−Pϕ\(Zt,ξt\)\)\.p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)=\\delta\\\!\\left\(Z\_\{t\+1\}\-P\_\{\\phi\}\(Z\_\{t\},\\xi\_\{t\}\)\\right\)\.MCJEPA is therefore an explicit finite\-state instantiation of a broader latent\-Markov interpretation of temporal JEPA: it replaces the general transition kernel by
pϕ\(Zt\+1=j∣Zt=i\)=Aij,p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i\)=A\_\{ij\},making the Markov property and multi\-step Chapman–Kolmogorov composition especially transparent\. Conceptually,
MCJEPA⊂Markov neural\-predictor JEPA⊂general temporal JEPA\.\\text\{MCJEPA\}\\subset\\text\{Markov neural\-predictor JEPA\}\\subset\\text\{general temporal JEPA\}\.Thus, MCJEPA demonstrates the correspondence in its simplest explicit form; it does not create or solely represent the correspondence\.
If the predictor genuinely depends on information beyondZtZ\_\{t\}, however, the latent process need not be first\-order Markov inZtZ\_\{t\}alone\. For example, a recurrent predictor may use \(Eq\.[19](https://arxiv.org/html/2608.13621#S3.E19)\)
pϕ\(Zt\+1∣Zt,Mt,ξt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},M\_\{t\},\\xi\_\{t\}\),whereMtM\_\{t\}summarizes additional history \(i\.e\. memory\), or a higher\-order predictor may depend directly on\(Zt,…,Zt−k\+1\)\(Z\_\{t\},\\ldots,Z\_\{t\-k\+1\}\)\. In such cases, exact HMM correspondence withZtZ\_\{t\}as the hidden state does not follow\. However, a first\-order representation can often be recovered byaugmenting the state\. For recurrent dynamics, one may define
Z~t=\(Zt,Mt\),\\widetilde\{Z\}\_\{t\}=\(Z\_\{t\},M\_\{t\}\),while for akkth\-order predictor one may use
Z~t=\(Zt,Zt−1,…,Zt−k\+1\)\.\\widetilde\{Z\}\_\{t\}=\(Z\_\{t\},Z\_\{t\-1\},\\ldots,Z\_\{t\-k\+1\}\)\.If the augmented state contains all transition\-relevant information from the past, then
p\(Z~t\+1∣Z~≤t,ξ≤t\)=p\(Z~t\+1∣Z~t,ξt\),p\(\\widetilde\{Z\}\_\{t\+1\}\\mid\\widetilde\{Z\}\_\{\\leq t\},\\xi\_\{\\leq t\}\)=p\(\\widetilde\{Z\}\_\{t\+1\}\\mid\\widetilde\{Z\}\_\{t\},\\xi\_\{t\}\),and the first\-order latent\-state interpretation is restored\. The correspondence developed in this paper therefore applies beyond MCJEPA to general probabilistic temporal JEPA whenever the chosen latent state, possibly after augmentation, admits such a first\-order Markov transition together with the additional emission and consistency conditions required for sequence\-level equivalence\.
#### Predictive representation learning as Markov\-state construction\.
The HMM perspective changes the interpretation of the JEPA representation itself\. Rather than viewingZtZ\_\{t\}only as a feature vector useful for predicting another feature vector, we can ask whether it constitutes a*predictive state*: does it retain the information from the past that is needed for future prediction while discarding redundant history? This distinction also separates two notions that can otherwise be conflated\. A predictor may be architecturally first\-order,
pϕ\(Zt\+1∣Zt,ξt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\),withoutZtZ\_\{t\}being a sufficient Markov representation of the underlying process\. In particular,
first\-order predictor⇏sufficient Markov state\.\\text\{first\-order predictor\}\\;\\not\\Rightarrow\\;\\text\{sufficient Markov state\}\.If older history remains predictive after conditioning onZtZ\_\{t\}, for example if
p\(Zt\+1∣Zt,Zt−1\)≠p\(Zt\+1∣Zt\),p\(Z\_\{t\+1\}\\mid Z\_\{t\},Z\_\{t\-1\}\)\\neq p\(Z\_\{t\+1\}\\mid Z\_\{t\}\),then the architecture is imposing a first\-order transition on a representation that has not fully Markovized the process\.
This gives predictive information\-bottleneck learning a state\-space interpretation\. Ideally, the representation should satisfy predictive sufficiency,
X\>t⟂X≤t\|Zt,X\_\{\>t\}\\perp X\_\{\\leq t\}\\mid Z\_\{t\},while retaining as little redundant information about the past as necessary\. Compression therefore promotes minimality, whereas predictive sufficiency prevents over\-compression\. Together they can transform a non\-Markov observation process into an approximately Markov latent process rather than merely forcing a Markov predictor onto an insufficient representation\. Experiment 3 illustrates this distinction explicitly: compression removes redundant history to recover the known minimal sufficient state, while the residual\-history diagnostic tests whether transition\-relevant information remains outside the current representation\. Thus, compression and residual predictability provide complementary tools for*learning and testing*a Markov representation rather than assuming one a priori\.
#### Architecture and objective are separate design choices\.
The HMM correspondence also clarifies a distinction that is easy to obscure: sharing an encode–transition–emit architecture does not imply sharing a training objective\. Standard JEPA training may optimize only latent predictive alignment and never maximize observation\-sequence likelihood\. Experiment 4 makes this distinction operational\. Adding HMM\-style sequence and filtering supervision to the genuine MCJEPA objective substantially improves recovery of the probabilistic latent\-state model, with the hybrid approaching HMM\-level sequence likelihood and transition recovery while retaining EMA\-target latent prediction\. Conversely, the same latent\-state architecture can be trained successfully using HMM sequence likelihood and filtering distillation withoutℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}at all\. The resulting continuum
JEPA latent prediction⟷hybrid JEPA–HMM training⟷HMM\-style sequence learning\\text\{JEPA latent prediction\}\\;\\longleftrightarrow\\;\\text\{hybrid JEPA\-\-HMM training\}\\;\\longleftrightarrow\\;\\text\{HMM\-style sequence learning\}shows that the boundary between probabilistic temporal JEPA and classical state\-space modeling is determined not only by model components, but also by which probabilistic semantics the objective enforces\.
#### Why observation reconstruction can remain optional\.
An HMM explicitly modelsp\(Xt∣St\)p\(X\_\{t\}\\mid S\_\{t\}\)because its likelihood is defined in observation space\. JEPA may instead deliberately concentrate learning on the information required for future prediction, avoiding the cost of reconstructing high\-entropy observation details that are irrelevant to the predictive task\. A decoder can be introduced when observation forecasting or sequence likelihood is required, but it need not participate in the core representation\-learning objective\. An invertible target encoder provides another realization of the observation map, although invertibility limits the encoder’s ability to discard nuisance information or reduce dimensionality and can therefore conflict with predictive compression\. The implicit\-emission construction establishes a probabilistic completion even without either explicit map, but that induced conditional need not be tractable enough for practical generation or likelihood evaluation\.
#### Scope and limitations\.
The first\-order Markov property should therefore be understood as a*representation\-design target*, not as a generic property of neural embeddings\. A neural predictor that consumes onlyZtZ\_\{t\}is architecturally first\-order, but this alone does not establish thatZtZ\_\{t\}contains all transition\-relevant information from the past\. If residual history remains predictive, the representation is insufficient at the chosen temporal scale; the state can instead be enlarged, augmented with recurrent memory, modeled with higher\-order dynamics, or predicted using an unrestricted history\-dependent model\. Likewise, finite categorical states and transition matrices improve structural interpretability but do not automatically produce semantically meaningful state labels\. Such semantics must be established through observation statistics, transition behavior, interventions, or downstream tasks\. Finally, our experiments are intentionally controlled and synthetic: they isolate composition, filtering, Markov\-state construction, and objective\-level behavior under known dynamics\. Extending these diagnostics to high\-dimensional video, control, and real\-world partially observed systems is therefore an important empirical next step\.
## 9Conclusion
We developed a state\-space interpretation of probabilistic temporal JEPA and made it concrete through*Markov\-Chain JEPA*\(MCJEPA\)\. MCJEPA replaces the latent predictor by a learned row\-stochastic transition matrix, so that multi\-step prediction is generated byAhA^\{h\}and direct and composed predictions satisfy exact Chapman\-\-Kolmogorov consistency\. Neural conditioned matrices, continuous\-state Markov kernels, and continuous\-time transitions extend this construction beyond finite homogeneous chains, while deterministic temporal JEPA appears as a degenerate transition kernel121212As shown in[Section3\.4](https://arxiv.org/html/2608.13621#S3.SS4), a deterministic predictor is a Dirac Markov kernel, and a deterministic latent representation can likewise be viewed as a point\-mass state distribution\. Thus, the latent Markov perspective is not restricted to probabilistic JEPA: probabilistic formulations expose the state\-space structure explicitly, while classical JEPA occupies its deterministic boundary\.\.
The broader contribution is to make precise when this latent Markov view becomes an HMM interpretation\. Observation\-level data correspond to HMM observations; the stochastic context encoder plays the filtering role; the probabilistic predictor defines latent transition dynamics; and a decoder, inverse target encoder, or induced implicit conditional supplies the emission direction\. We distinguish computational correspondence, emission\-complete latent\-state representation, sequence\-level HMM equivalence, and model\-and\-objective equivalence, and give sufficient conditions in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1)under which the resulting model admits an exact sequence\-level HMM representation\. Because classical deterministic JEPA is recovered through point\-mass latent distributions and Dirac transitions, the same computational and latent\-Markov correspondence extends to the classical setting in the corresponding degenerate sense, although stronger HMM equivalence still requires the emission and consistency conditions identified above\.
This perspective also yields a representation\-learning principle: predictive information bottleneck learning can be understood as seeking a compact predictive state that approximately*Markovizes*the observed process at the chosen prediction scale\. Compression promotes*minimality*by removing redundant history, while residual predictability tests*sufficiency*by detecting transition\-relevant information that remains outside the current state\. Finally, the objective\-level experiments show that the same latent\-state architecture supports a continuum from JEPA latent prediction, through hybrid JEPA–HMM learning, to HMM\-style sequence and filtering training\. Probabilistic temporal JEPA is therefore not simply an HMM under a different name; rather, it exposes an HMM\-compatible latent state\-space structure whose probabilistic semantics become progressively stronger as emission completeness, transition and marginal consistency, filtering consistency, and HMM\-style sequence training are imposed\.
## References
- Alemiet al\.\(2017\)A\. A\. Alemi, I\. Fischer, J\. V\. Dillon, and K\. MurphyDeep variational information bottleneck\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HyxQzBceg)Cited by:[§5](https://arxiv.org/html/2608.13621#S5.p1.1)\.
- Assranet al\.\(2023\)M\. Assran, Q\. Duval, I\. Misra, P\. Bojanowski, P\. Vincent, M\. Rabbat, Y\. LeCun, and N\. BallasSelf\-supervised learning from images with a joint\-embedding predictive architecture\.External Links:2301\.08243,[Link](https://arxiv.org/abs/2301.08243)Cited by:[§1](https://arxiv.org/html/2608.13621#S1.p1.1)\.
- Bardeset al\.\(2024\)A\. Bardes, Q\. Garrido, J\. Ponce, X\. Chen, M\. Rabbat, Y\. LeCun, M\. Assran, and N\. BallasRevisiting feature prediction for learning visual representations from video\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=QaCCuDfBk2)Cited by:[§1](https://arxiv.org/html/2608.13621#S1.p1.1)\.
- Bardeset al\.\(2023\)A\. Bardes, J\. Ponce, and Y\. LeCunMC\-jepa: a joint\-embedding predictive architecture for self\-supervised learning of motion and content features\.External Links:2307\.12698,[Link](https://arxiv.org/abs/2307.12698)Cited by:[footnote 3](https://arxiv.org/html/2608.13621#footnote3)\.
- Bialeket al\.\(2001\)W\. Bialek, I\. Nemenman, and N\. TishbyPredictability, complexity, and learning\.Neural Comput\.13\(11\),pp\. 2409–2463\.External Links:ISSN 0899\-7667,[Link](https://doi.org/10.1162/089976601753195969),[Document](https://dx.doi.org/10.1162/089976601753195969)Cited by:[§5](https://arxiv.org/html/2608.13621#S5.p1.1)\.
- Boyd and Vandenberghe \(2018\)S\. Boyd and L\. VandenbergheIntroduction to applied linear algebra: vectors, matrices, and least squares\.Cambridge University Press,Cambridge, UK\.External Links:ISBN 9781108424936,[Link](https://web.stanford.edu/%CB%9Cboyd/vmls/)Cited by:[footnote 4](https://arxiv.org/html/2608.13621#footnote4)\.
- Huang \(2026a\)Y\. HuangGaussian joint embeddings for self\-supervised representation learning\.External Links:2603\.26799,[Link](https://arxiv.org/abs/2603.26799)Cited by:[§3\.2](https://arxiv.org/html/2608.13621#S3.SS2.p4.1)\.
- Huang \(2026b\)Y\. HuangOn the information bottleneck of vjepa\.Note:[https://hal\.science/hal\-05622405](https://hal.science/hal-05622405)HAL preprint, hal\-05622405Cited by:[§1](https://arxiv.org/html/2608.13621#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.13621#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.13621#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2608.13621#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.13621#S4.SS1.p2.2),[§5](https://arxiv.org/html/2608.13621#S5.p1.2)\.
- Huang \(2026c\)Y\. HuangVJEPA: variational joint embedding predictive architectures as probabilistic world models\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=omJqMJt2fC)Cited by:[§1](https://arxiv.org/html/2608.13621#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.13621#S1.p1.1),[§1](https://arxiv.org/html/2608.13621#S1.p2.3),[§2\.3](https://arxiv.org/html/2608.13621#S2.SS3.p1.1),[footnote 6](https://arxiv.org/html/2608.13621#footnote6)\.
- LeCun \(2022\)Y\. LeCunA path towards autonomous machine intelligence version 0\.9\.2, 2022\-06\-27\.Open Review62\(1\),pp\. 1–62\.Cited by:[§1](https://arxiv.org/html/2608.13621#S1.p1.1)\.
- Littman and Sutton \(2001\)M\. Littman and R\. S\. SuttonPredictive representations of state\.InAdvances in Neural Information Processing Systems,T\. Dietterich, S\. Becker, and Z\. Ghahramani \(Eds\.\),Vol\.14,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2001/file/1e4d36177d71bbb3558e43af9577d70e-Paper.pdf)Cited by:[§5](https://arxiv.org/html/2608.13621#S5.p2.2)\.
- OpenAI \(2026\)OpenAIGPT\-5\.6: Frontier intelligence that scales with your ambition\.Note:[https://openai\.com/index/gpt\-5\-6/](https://openai.com/index/gpt-5-6/)OpenAI release, July 9, 2026Cited by:[Disclaimer](https://arxiv.org/html/2608.13621#Ax1.p1.1)\.
- Rabiner \(1989\)L\.R\. RabinerA tutorial on hidden markov models and selected applications in speech recognition\.Proceedings of the IEEE77\(2\),pp\. 257–286\.External Links:[Document](https://dx.doi.org/10.1109/5.18626)Cited by:[§4\.2](https://arxiv.org/html/2608.13621#S4.SS2.p1.1)\.
- Strang \(2016\)G\. StrangIntroduction to linear algebra\.5th edition,Wellesley–Cambridge Press,Wellesley, MA\.External Links:ISBN 9780980232776Cited by:[footnote 4](https://arxiv.org/html/2608.13621#footnote4)\.
- Tishbyet al\.\(1999\)N\. Tishby, F\. C\. Pereira, and W\. BialekThe information bottleneck method\.InProceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing,Cited by:[§5](https://arxiv.org/html/2608.13621#S5.p1.1)\.
## Appendix ATraining Objectives and Minimal Algorithm
For the basic time\-homogeneous MCJEPA model with one shared transition matrixAA, the training objective is
ℒ=ℒMC\+ℒstate,\\mathcal\{L\}=\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{state\}\},whereℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}is the multi\-horizon latent\-prediction objective in[Eq\.8](https://arxiv.org/html/2608.13621#S2.E8)andℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}is the discrete\-state regularizer in[Eq\.10](https://arxiv.org/html/2608.13621#S2.E10)\. Because allhh\-step predictions are generated by powers of the same matrixAA, exact path consistency follows automatically from[Proposition1](https://arxiv.org/html/2608.13621#Thmproposition1); no additional Chapman–Kolmogorov penalty is required\.
A softer alternative may instead parameterize separate horizon\-dependent transition matricesAhA\_\{h\}\. In that case, path consistency is no longer guaranteed and may be encouraged through
ℒCK=∑h1,h2‖Ah1\+h2−Ah1Ah2‖F2,\\mathcal\{L\}\_\{\\mathrm\{CK\}\}=\\sum\_\{h\_\{1\},h\_\{2\}\}\\left\\\|A\_\{h\_\{1\}\+h\_\{2\}\}\-A\_\{h\_\{1\}\}A\_\{h\_\{2\}\}\\right\\\|\_\{F\}^\{2\},with an additional weightλCK≥0\\lambda\_\{\\mathrm\{CK\}\}\\geq 0\. This penalty belongs only to the horizon\-dependent variant and is unnecessary for the shared\-AAMCJEPA used in the main experiments\.
For the basic shared\-AAmodel, a minimal training step is:
1. 1\.sample a time indextt, prediction horizonh∈ℋh\\in\\mathcal\{H\}, observation historyX≤tX\_\{\\leq t\}, future targetXt\+hX\_\{t\+h\}, and any required side information;
2. 2\.compute the current\-state distribution qt=qθ\(Zt∣X≤t\);q\_\{t\}=q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\);
3. 3\.compute the EMA target distribution q¯t\+h=qθ¯\(Zt\+h∣Xt\+h\);\\bar\{q\}\_\{t\+h\}=q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\);
4. 4\.propagate the current state through the shared transition matrix, q^t\+h=qtAh;\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\};
5. 5\.evaluate the latent\-prediction loss in[Eq\.8](https://arxiv.org/html/2608.13621#S2.E8)usingsg\(q¯t\+h\)\\operatorname\{sg\}\(\\bar\{q\}\_\{t\+h\}\)as the target, and add the state\-use regularizerℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\};
6. 6\.update the online encoder parametersθ\\thetaand the learnable transition parameters, then update the target encoder by exponential moving average, θ¯←τθ¯\+\(1−τ\)θ,τ∈\[0,1\)\.\\bar\{\\theta\}\\leftarrow\\tau\\bar\{\\theta\}\+\(1\-\\tau\)\\theta,\\qquad\\tau\\in\[0,1\)\.
For the conditioned discrete\-state model in[Eq\.12](https://arxiv.org/html/2608.13621#S3.E12), Step 4 is replaced by the ordered transition composition
q^t\+h=qt∏j=0h−1Aϕ\(ξt\+j\),\\widehat\{q\}\_\{t\+h\}=q\_\{t\}\\prod\_\{j=0\}^\{h\-1\}A\_\{\\phi\}\(\\xi\_\{t\+j\}\),as defined in[Eq\.13](https://arxiv.org/html/2608.13621#S3.E13)\. The remainder of the training procedure is unchanged\.
The residual quantities introduced in[Section6](https://arxiv.org/html/2608.13621#S6)are used as held\-out diagnostics of state and transition sufficiency rather than as part of the default MCJEPA training objective\. In particular, residual\-history gain is evaluated after fitting the representation and transition model so that residual predictability can diagnose information omitted from the current state without directly training the representation to satisfy the diagnostic\.
## Appendix BProofs
### B\.1Proof of Proposition[1](https://arxiv.org/html/2608.13621#Thmproposition1)
For nonnegative integersh1h\_\{1\}andh2h\_\{2\}, the definition of matrix powers together with associativity of matrix multiplication gives
Ah1\+h2=Ah1Ah2\.A^\{h\_\{1\}\+h\_\{2\}\}=A^\{h\_\{1\}\}A^\{h\_\{2\}\}\.Left\-multiplying by the current state distributionqtq\_\{t\}yields
qtAh1\+h2=\(qtAh1\)Ah2,q\_\{t\}A^\{h\_\{1\}\+h\_\{2\}\}=\\left\(q\_\{t\}A^\{h\_\{1\}\}\\right\)A^\{h\_\{2\}\},which is[Eq\.9](https://arxiv.org/html/2608.13621#S2.E9)\.
More generally, let a total horizonhhbe partitioned into nonnegative integers
h=h1\+⋯\+hm\.h=h\_\{1\}\+\\cdots\+h\_\{m\}\.Repeated application of the same identity gives
Ah=Ah1⋯Ahm,A^\{h\}=A^\{h\_\{1\}\}\\cdots A^\{h\_\{m\}\},and therefore
qtAh=\(⋯\(\(qtAh1\)Ah2\)⋯\)Ahm\.q\_\{t\}A^\{h\}=\\bigl\(\\cdots\(\(q\_\{t\}A^\{h\_\{1\}\}\)A^\{h\_\{2\}\}\)\\cdots\\bigr\)A^\{h\_\{m\}\}\.Hence every decomposition of the same total horizon produces the same predictive distribution, proving the proposition\.
### B\.2Proof of Proposition[2](https://arxiv.org/html/2608.13621#Thmproposition2)
Recall from[Eq\.22](https://arxiv.org/html/2608.13621#S4.E22)that, wheneverqθ\(z\)\>0q\_\{\\theta\}\(z\)\>0,
pθimp\(x∣z\)=pdata\(x\)qθ\(z∣x\)qθ\(z\),p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)=\\frac\{p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\}\{q\_\{\\theta\}\(z\)\},where
qθ\(z\)=∫pdata\(x\)qθ\(z∣x\)𝑑x\.q\_\{\\theta\}\(z\)=\\int p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\\,dx\.
For anyzzwithqθ\(z\)\>0q\_\{\\theta\}\(z\)\>0,
∫pθimp\(x∣z\)𝑑x\\displaystyle\\int p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)\\,dx=1qθ\(z\)∫pdata\(x\)qθ\(z∣x\)𝑑x\\displaystyle=\\frac\{1\}\{q\_\{\\theta\}\(z\)\}\\int p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\\,dx=qθ\(z\)qθ\(z\)=1\.\\displaystyle=\\frac\{q\_\{\\theta\}\(z\)\}\{q\_\{\\theta\}\(z\)\}=1\.Thus,pθimp\(x∣z\)p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)is a normalized conditional distribution\. For discrete observations, the corresponding integrals are replaced by sums\.
Moreover, for anyxxin the support ofpdatap\_\{\\mathrm\{data\}\}and anyzzwithqθ\(z\)\>0q\_\{\\theta\}\(z\)\>0, direct substitution gives
pθimp\(x∣z\)qθ\(z\)pdata\(x\)\\displaystyle\\frac\{p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)q\_\{\\theta\}\(z\)\}\{p\_\{\\mathrm\{data\}\}\(x\)\}=pdata\(x\)qθ\(z∣x\)qθ\(z\)qθ\(z\)pdata\(x\)\\displaystyle=\\frac\{p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\}\{q\_\{\\theta\}\(z\)\}\\frac\{q\_\{\\theta\}\(z\)\}\{p\_\{\\mathrm\{data\}\}\(x\)\}=qθ\(z∣x\)\.\\displaystyle=q\_\{\\theta\}\(z\\mid x\)\.Henceqθ\(z∣x\)q\_\{\\theta\}\(z\\mid x\)is exactly the posterior associated with the priorqθ\(z\)q\_\{\\theta\}\(z\)and the implicit emissionpθimp\(x∣z\)p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)\.
Equivalently, the resulting one\-time joint distribution satisfies
pθimp\(x∣z\)qθ\(z\)=pdata\(x\)qθ\(z∣x\)\.p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)q\_\{\\theta\}\(z\)=p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\.Its observation marginal ispdata\(x\)p\_\{\\mathrm\{data\}\}\(x\), so the construction defines a valid static latent\-variable model\. As emphasized in the main text, this one\-time Bayes completion does not by itself establish a sequence\-level HMM; that stronger result additionally requires transition, marginal, and filtering consistency\.
### B\.3Proof of Proposition[3](https://arxiv.org/html/2608.13621#Thmproposition3)
By assumption, the future target state is generated as
Zt\+1=gθ¯\(Xt\+1,Ut\+1\),Z\_\{t\+1\}=g\_\{\\bar\{\\theta\}\}\(X\_\{t\+1\},U\_\{t\+1\}\),soZt\+1Z\_\{t\+1\}is a stochastic post\-processing of\(Xt\+1,Ut\+1\)\(X\_\{t\+1\},U\_\{t\+1\}\)\. The conditional data\-processing inequality therefore gives
I\(Zt\+1;X<t∣Zt\)≤I\(Xt\+1,Ut\+1;X<t∣Zt\)\.\\mathrm\{I\}\(Z\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)\\leq\\mathrm\{I\}\(X\_\{t\+1\},U\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)\.
By the chain rule for conditional mutual information,
I\(Xt\+1,Ut\+1;X<t∣Zt\)=\\displaystyle\\mathrm\{I\}\(X\_\{t\+1\},U\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=\{\}I\(Xt\+1;X<t∣Zt\)\\displaystyle\\mathrm\{I\}\(X\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)\+I\(Ut\+1;X<t∣Xt\+1,Zt\)\.\\displaystyle\+\\mathrm\{I\}\(U\_\{t\+1\};X\_\{<t\}\\mid X\_\{t\+1\},Z\_\{t\}\)\.
Predictive sufficiency,
X\>t⟂X≤t\|Zt,X\_\{\>t\}\\perp X\_\{\\leq t\}\\mid Z\_\{t\},implies
I\(Xt\+1;X<t∣Zt\)=0,\\mathrm\{I\}\(X\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=0,becauseXt\+1X\_\{t\+1\}is contained inX\>tX\_\{\>t\}andX<tX\_\{<t\}is contained inX≤tX\_\{\\leq t\}\.
Likewise, the assumed conditional independence of the target\-encoder randomness,
Ut\+1⟂X≤t\|\(Xt\+1,Zt\),U\_\{t\+1\}\\perp X\_\{\\leq t\}\\mid\(X\_\{t\+1\},Z\_\{t\}\),implies
I\(Ut\+1;X<t∣Xt\+1,Zt\)=0\.\\mathrm\{I\}\(U\_\{t\+1\};X\_\{<t\}\\mid X\_\{t\+1\},Z\_\{t\}\)=0\.Consequently,
I\(Xt\+1,Ut\+1;X<t∣Zt\)=0\.\\mathrm\{I\}\(X\_\{t\+1\},U\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=0\.
Combining this equality with the conditional data\-processing inequality and the nonnegativity of conditional mutual information yields
I\(Zt\+1;X<t∣Zt\)=0,\\mathrm\{I\}\(Z\_\{t\+1\};X\_\{<t\}\\mid Z\_\{t\}\)=0,which proves the proposition\. Thus, under predictive sufficiency and the stated target\-encoder independence condition, the current representationZtZ\_\{t\}screens off older observation history from the next latent stateZt\+1Z\_\{t\+1\}\. For a deterministic target encoder, the auxiliary randomnessUt\+1U\_\{t\+1\}can be omitted\.
## Appendix CThree Realizations of the Emission Direction
The main text describes three alternative ways to complete the state\-to\-observation direction of the latent\-state model\. These constructions should not be interpreted as three progressively stronger notions of equivalence\. Rather, each can supply the emission component required for an emission\-complete latent\-state representation\. Exact sequence\-level HMM equivalence additionally requires the transition, marginal\-consistency, and filtering\-consistency conditions developed in[AppendixE](https://arxiv.org/html/2608.13621#A5)\.
### C\.1Explicit decoder
The most direct realization introduces a probabilistic decoder
pψ\(Xt∣Zt\),p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\),which has the same state\-to\-observation direction as an HMM emission model\.
When both the latent state and observation space are finite, this conditional may be represented by an emission matrixBB, for example
Bjk=pψ\(Xt=k∣Zt=j\)\.B\_\{jk\}=p\_\{\\psi\}\(X\_\{t\}=k\\mid Z\_\{t\}=j\)\.For a categorical latent state with continuous observations, each latent state instead indexes an observation density\. More generally, for images, signals, or other high\-dimensional observations,pψ\(Xt∣Zt\)p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)may be parameterized by a Gaussian, discretized logistic, autoregressive, diffusion\-based, or other suitable conditional observation model\.
If the decoder is fitted only after JEPA representation learning while the latent model is held fixed, it acts as a post\-hoc observation model or probe\. If it participates jointly in training, it becomes part of the generative latent\-state model and can contribute directly to observation\-sequence likelihood\.
### C\.2Invertible target encoder
A second realization is available when the target encoder
fθ¯:𝒳→𝒮f\_\{\\bar\{\\theta\}\}:\\mathcal\{X\}\\rightarrow\\mathcal\{S\}is bijective on the modeled data domain\. Its target representation satisfies
ZtT=fθ¯\(Xt\),Xt=fθ¯−1\(ZtT\)\.Z\_\{t\}^\{\\mathrm\{T\}\}=f\_\{\\bar\{\\theta\}\}\(X\_\{t\}\),\\qquad X\_\{t\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}^\{\\mathrm\{T\}\}\)\.Identifying the latent state with this target representation gives the deterministic state\-to\-observation kernel
p\(Xt∣Zt\)=δ\(Xt−fθ¯−1\(Zt\)\)\.p\(X\_\{t\}\\mid Z\_\{t\}\)=\\delta\\\!\\left\(X\_\{t\}\-f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}\)\\right\)\.Thus, invertibility supplies the required state\-to\-observation direction without introducing a separate decoder\.
A deterministic inverse should nevertheless be distinguished from a non\-degenerate probabilistic emission\. Iffθ¯f\_\{\\bar\{\\theta\}\}is a tractable invertible density model, the change\-of\-variables formula can be used to evaluate the observation density induced by a latent density\. The conditional mapXt=fθ¯−1\(Zt\)X\_\{t\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}\)itself remains deterministic, however\. A non\-degenerate conditional emission can instead be obtained by augmenting the inverse map with an observation\-noise model, for example
Xt=fθ¯−1\(Zt\)\+εt,X\_\{t\}=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(Z\_\{t\}\)\+\\varepsilon\_\{t\},with a specified noise distribution forεt\\varepsilon\_\{t\}\.
Exact invertibility imposes strong architectural constraints\. In particular, it prevents unrestricted dimensionality reduction and may require the representation to preserve observation details that a predictive information bottleneck would otherwise discard\. Standard compressed JEPA encoders therefore need not admit this construction\.
### C\.3Implicit emission
When no explicit decoder is parameterized and the target encoder is not invertible, a local stochastic encoder can still induce a state\-to\-observation conditional\. As defined in[Eq\.22](https://arxiv.org/html/2608.13621#S4.E22),
pθimp\(x∣z\)=pdata\(x\)qθ\(z∣x\)qθ\(z\),p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)=\\frac\{p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\}\{q\_\{\\theta\}\(z\)\},forqθ\(z\)\>0q\_\{\\theta\}\(z\)\>0\. As shown in the preceding proof, this conditional is normalized and, together withqθ\(z\)q\_\{\\theta\}\(z\), reproduces the one\-time joint distribution
pθimp\(x∣z\)qθ\(z\)=pdata\(x\)qθ\(z∣x\)\.p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)q\_\{\\theta\}\(z\)=p\_\{\\mathrm\{data\}\}\(x\)q\_\{\\theta\}\(z\\mid x\)\.
This construction provides an exact static probabilistic completion of the observation–state relationship, but it does not automatically provide a practical generative model\. In particular,pθimp\(x∣z\)p\_\{\\theta\}^\{\\mathrm\{imp\}\}\(x\\mid z\)depends explicitly on the generally unknown data marginalpdata\(x\)p\_\{\\mathrm\{data\}\}\(x\), so direct sampling and likelihood evaluation may be intractable\.
Moreover, the construction uses a*local*encoderqθ\(z∣x\)q\_\{\\theta\}\(z\\mid x\)\. An arbitrary history\-dependent encoderqθ\(Zt∣X≤t\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)cannot simply be reinterpreted as an emission model\. To obtain an exact sequence\-level HMM from the implicit construction, the resulting one\-time conditionals must additionally be consistent with the latent transition and with the Bayesian filtering recursion, as formalized in[AppendixE](https://arxiv.org/html/2608.13621#A5)\.
## Appendix DHMM\-Style Training Objectives for PIB\-VJEPA
The HMM correspondence suggests an alternative to purely latent\-space JEPA training\. Once a valid observation model is available, the latent transition and emission can be trained from observation\-sequence likelihood, while the history\-dependent context encoder can be aligned with the corresponding Bayesian filtering distribution\. This appendix summarizes the probabilistic objectives underlying the model\-and\-objective correspondence developed in the main text\.
### D\.1Sequence likelihood
Suppose that the latent\-state model is equipped with an initial\-state distributionp0\(Z1\)p\_\{0\}\(Z\_\{1\}\), transition model
pϕ\(Zt\+1∣Zt,ξt\),p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\),and explicit emission model
pψ\(Xt∣Zt\)\.p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)\.The resulting conditional sequence model is
pϕ,ψ\(X1:T,Z1:T∣ξ1:T−1\)\\displaystyle p\_\{\\phi,\\psi\}\\left\(X\_\{1:T\},Z\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\\right\)=p0\(Z1\)∏t=1Tpψ\(Xt∣Zt\)∏t=1T−1pϕ\(Zt\+1∣Zt,ξt\)\.\\displaystyle=p\_\{0\}\(Z\_\{1\}\)\\prod\_\{t=1\}^\{T\}p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)\\prod\_\{t=1\}^\{T\-1\}p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\.Marginalizing the latent trajectory gives the observation\-sequence evidence
pϕ,ψ\(X1:T∣ξ1:T−1\)=∫pϕ,ψ\(X1:T,Z1:T∣ξ1:T−1\)dZ1:T\.p\_\{\\phi,\\psi\}\\left\(X\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\\right\)=\\int p\_\{\\phi,\\psi\}\\left\(X\_\{1:T\},Z\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\\right\)\\,dZ\_\{1:T\}\.Training from this evidence directly constrains the transition–emission model in observation space, in contrast to the standard JEPA objective, which is imposed primarily in latent space\.
### D\.2Exact likelihood and filtering for categorical MCJEPA
For categorical statesZt∈\{1,…,K\}Z\_\{t\}\\in\\\{1,\\ldots,K\\\}, the sequence likelihood and filtering posterior can be evaluated exactly by the HMM forward recursion\. Let
πj=p0\(Z1=j\),bψ\(Xt∣j\)=pψ\(Xt∣Zt=j\),\\pi\_\{j\}=p\_\{0\}\(Z\_\{1\}=j\),\\qquad b\_\{\\psi\}\(X\_\{t\}\\mid j\)=p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}=j\),and
\(At\)ij=pϕ\(Zt\+1=j∣Zt=i,ξt\)\.\(A\_\{t\}\)\_\{ij\}=p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i,\\xi\_\{t\}\)\.For the time\-homogeneous MCJEPA model,At=AA\_\{t\}=A\.
Define the forward message
αt\(j\)=pϕ,ψ\(X1:t,Zt=j∣ξ1:t−1\)\.\\alpha\_\{t\}\(j\)=p\_\{\\phi,\\psi\}\\left\(X\_\{1:t\},Z\_\{t\}=j\\mid\\xi\_\{1:t\-1\}\\right\)\.It satisfies
α1\(j\)=πjbψ\(X1∣j\)\\alpha\_\{1\}\(j\)=\\pi\_\{j\}b\_\{\\psi\}\(X\_\{1\}\\mid j\)and
αt\+1\(j\)=bψ\(Xt\+1∣j\)∑i=1Kαt\(i\)\(At\)ij\.\\alpha\_\{t\+1\}\(j\)=b\_\{\\psi\}\(X\_\{t\+1\}\\mid j\)\\sum\_\{i=1\}^\{K\}\\alpha\_\{t\}\(i\)\(A\_\{t\}\)\_\{ij\}\.The sequence evidence is therefore
pϕ,ψ\(X1:T∣ξ1:T−1\)=∑j=1KαT\(j\)\.p\_\{\\phi,\\psi\}\\left\(X\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\\right\)=\\sum\_\{j=1\}^\{K\}\\alpha\_\{T\}\(j\)\.In practice, the recursion is evaluated in log space or with normalized forward messages for numerical stability\.
Normalizing the forward messages also gives the exact filtering posterior
q~t=pϕ,ψ\(Zt∣X≤t,ξ<t\)\.\\widetilde\{q\}\_\{t\}=p\_\{\\phi,\\psi\}\\left\(Z\_\{t\}\\mid X\_\{\\leq t\},\\xi\_\{<t\}\\right\)\.The history\-dependent PIB\-VJEPA context encoder can then be trained as an amortized filter using
ℒfilter=𝔼t\[KL\(sg\(q~t\)∥qθ\(Zt∣X≤t\)\)\]\.\\mathcal\{L\}\_\{\\mathrm\{filter\}\}=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\(\\widetilde\{q\}\_\{t\}\)\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\right\)\\right\]\.Thus, the encoder learns to approximate in one forward pass the current\-state posterior that the HMM computes recursively\.
This distinction is important: sequence likelihood trains the transition and emission model, whereas filtering distillation trains the context encoder to reproduce the corresponding filtering belief\. Sequence likelihood alone does not require an independently parameterized history encoder to equal that filter\.
### D\.3Continuous\-state extension
For continuous or nonlinear latent states, exact marginalization ofZ1:TZ\_\{1:T\}is generally unavailable\. Introducing an approximate sequence posterior
qη\(Z1:T∣X1:T,ξ1:T−1\)q\_\{\\eta\}\\left\(Z\_\{1:T\}\\mid X\_\{1:T\},\\xi\_\{1:T\-1\}\\right\)gives the standard variational lower bound
logpϕ,ψ\(X1:T∣ξ1:T−1\)≥𝔼qη\[\\displaystyle\\log p\_\{\\phi,\\psi\}\\left\(X\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\\right\)\\geq\\mathbb\{E\}\_\{q\_\{\\eta\}\}\\Bigg\[∑t=1Tlogpψ\(Xt∣Zt\)\+logp0\(Z1\)\\displaystyle\\sum\_\{t=1\}^\{T\}\\log p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)\+\\log p\_\{0\}\(Z\_\{1\}\)\+∑t=1T−1logpϕ\(Zt\+1∣Zt,ξt\)−logqη\(Z1:T∣X1:T,ξ1:T−1\)\]\.\\displaystyle\+\\sum\_\{t=1\}^\{T\-1\}\\log p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\},\\xi\_\{t\}\)\-\\log q\_\{\\eta\}\\left\(Z\_\{1:T\}\\mid X\_\{1:T\},\\xi\_\{1:T\-1\}\\right\)\\Bigg\]\.The approximate posterior may be causal when online filtering is required or smoothing when full\-sequence information is available during training\. We include this extension to show how the same model\-and\-objective interpretation extends beyond the finite categorical setting; the experiments in this paper use exact finite\-state inference\.
### D\.4Relation to JEPA and hybrid training
HMM\-style sequence learning and JEPA latent prediction are distinct objectives even when they operate on the same latent\-state architecture\. In the categorical setting studied in Experiment 4, the main text compares
ℒMC⏟JEPA latent predictionandℒHMM\+ℒfilter⏟HMM\-style training,\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\}\_\{\\text\{JEPA latent prediction\}\}\\qquad\\text\{and\}\\qquad\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\}\_\{\\text\{HMM\-style training\}\},withℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}used as a shared state\-use regularizer\. The hybrid regime combines all three signals,
ℒHMM\+ℒMC\+ℒfilter,\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{filter\}\},with the corresponding weights given in[Sections7\.4](https://arxiv.org/html/2608.13621#S7.SS4)and[F](https://arxiv.org/html/2608.13621#A6)\.
Importantly, the exact HMM filtering posterior is used only as the target ofℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\. It does not replace the EMA future target inℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}: the JEPA component retains the same target\-encoder construction as the JEPA\-only regime\. The transition matrix is shared between the HMM sequence objective and the MCJEPA latent\-prediction objective, so the two training signals constrain the same latent dynamics from observation\-space and representation\-space perspectives, respectively\.
Consequently, adding an emission model alone does not make PIB\-VJEPA training identical to HMM training\. Model\-and\-objective equivalence additionally requires observation\-sequence evidence and the corresponding sequence\-inference semantics to participate in the learning objective\.
## Appendix EExact HMM Representation Conditions
This appendix makes precise the sufficient conditions in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1)\. We state the construction for a discrete latent state for clarity; the same argument extends to general Markov kernels by replacing sums with integrals\. Side informationξt\\xi\_\{t\}is treated as observed, so all sequence distributions below are conditional onξ1:T−1\\xi\_\{1:T\-1\}\.
Let the latent transition be
At\(i,j\)=pϕ\(Zt\+1=j∣Zt=i,ξt\),A\_\{t\}\(i,j\)=p\_\{\\phi\}\(Z\_\{t\+1\}=j\\mid Z\_\{t\}=i,\\xi\_\{t\}\),and let
denote a valid state\-to\-observation conditional\. This emission may be supplied by an explicit decoder, an invertible target encoder interpreted as a deterministic kernel, or the implicit construction described below\. Given an initial distributionρ1\\rho\_\{1\}, these components define
p\(z1:T,x1:T∣ξ1:T−1\)=ρ1\(z1\)∏t=1Tbt\(xt∣zt\)∏t=1T−1At\(zt,zt\+1\),\\displaystyle p\(z\_\{1:T\},x\_\{1:T\}\\mid\\xi\_\{1:T\-1\}\)\\qquad=\\rho\_\{1\}\(z\_\{1\}\)\\prod\_\{t=1\}^\{T\}b\_\{t\}\(x\_\{t\}\\mid z\_\{t\}\)\\prod\_\{t=1\}^\{T\-1\}A\_\{t\}\(z\_\{t\},z\_\{t\+1\}\),\(40\)which is the conditional HMM factorization\.
### E\.1Marginal and filtering consistency
Let
ρt\(j\)=p\(Zt=j∣ξ1:t−1\)\\rho\_\{t\}\(j\)=p\(Z\_\{t\}=j\\mid\\xi\_\{1:t\-1\}\)denote the latent marginal before observingXtX\_\{t\}, conditional on the side\-information history\. Dynamic consistency requires
ρt\+1\(j\)=∑iρt\(i\)At\(i,j\)\.\\rho\_\{t\+1\}\(j\)=\\sum\_\{i\}\\rho\_\{t\}\(i\)A\_\{t\}\(i,j\)\.\(41\)Thus, the one\-time latent marginals must be generated by the same transition kernel used by the temporal model\.
For a realized observation history, let
qt\(j\)=p\(Zt=j∣X≤t,ξ<t\)q\_\{t\}\(j\)=p\(Z\_\{t\}=j\\mid X\_\{\\leq t\},\\xi\_\{<t\}\)denote the filtering distribution\. Its transition\-based predictive prior is
πt\(j\)=∑iqt−1\(i\)At−1\(i,j\),\\pi\_\{t\}\(j\)=\\sum\_\{i\}q\_\{t\-1\}\(i\)A\_\{t\-1\}\(i,j\),withπ1=ρ1\\pi\_\{1\}=\\rho\_\{1\}\. Bayes’ rule then gives the filtering recursion
qt\(j\)=bt\(Xt∣j\)πt\(j\)∑kbt\(Xt∣k\)πt\(k\)\.q\_\{t\}\(j\)=\\frac\{b\_\{t\}\(X\_\{t\}\\mid j\)\\pi\_\{t\}\(j\)\}\{\\sum\_\{k\}b\_\{t\}\(X\_\{t\}\\mid k\)\\pi\_\{t\}\(k\)\}\.\(42\)
The history\-dependent PIB\-VJEPA encoder is filtering\-consistent when
qθ\(Zt∣X≤t,ξ<t\)=qtq\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\},\\xi\_\{<t\}\)=q\_\{t\}\(43\)for the transition and emission model under consideration\. An arbitrary history encoder need not satisfy this equality\.
### E\.2Explicit and invertible emissions
If a decoder directly specifies
bt\(x∣z\)=pψ\(x∣z\),b\_\{t\}\(x\\mid z\)=p\_\{\\psi\}\(x\\mid z\),then[Eq\.40](https://arxiv.org/html/2608.13621#A5.E40)follows immediately from the initial distribution, first\-order transition, and emission model\. If the target encoder is invertible, the deterministic kernel induced by
x=fθ¯−1\(z\)x=f\_\{\\bar\{\\theta\}\}^\{\-1\}\(z\)plays the same role\. In either case, if the latent marginals satisfy[Eq\.41](https://arxiv.org/html/2608.13621#A5.E41)and the history encoder satisfies[Eq\.43](https://arxiv.org/html/2608.13621#A5.E43), the resulting temporal JEPA admits the sequence\-level HMM interpretation stated in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1)\.
### E\.3Implicit\-emission case
The less direct case begins with a local evidence encoder
eθ,t\(z∣x\)e\_\{\\theta,t\}\(z\\mid x\)that depends only on the current observation\. Letpt\(x\)p\_\{t\}\(x\)be the one\-time observation marginal and define its induced latent marginal
ρt\(z\)=∫pt\(x\)eθ,t\(z∣x\)𝑑x\.\\rho\_\{t\}\(z\)=\\int p\_\{t\}\(x\)e\_\{\\theta,t\}\(z\\mid x\)\\,dx\.\(44\)Forρt\(z\)\>0\\rho\_\{t\}\(z\)\>0, define
btimp\(x∣z\)=pt\(x\)eθ,t\(z∣x\)ρt\(z\)\.b\_\{t\}^\{\\mathrm\{imp\}\}\(x\\mid z\)=\\frac\{p\_\{t\}\(x\)e\_\{\\theta,t\}\(z\\mid x\)\}\{\\rho\_\{t\}\(z\)\}\.By the argument in the proof of implicit emission completion, this is a normalized state\-to\-observation conditional and satisfies
eθ,t\(z∣x\)=btimp\(x∣z\)ρt\(z\)pt\(x\)\.e\_\{\\theta,t\}\(z\\mid x\)=\\frac\{b\_\{t\}^\{\\mathrm\{imp\}\}\(x\\mid z\)\\rho\_\{t\}\(z\)\}\{p\_\{t\}\(x\)\}\.
The locality assumption is important:eθ,t\(z∣xt\)e\_\{\\theta,t\}\(z\\mid x\_\{t\}\)supplies the observation\-dependent evidence factor, whereas the history\-dependent context encoder represents the filtering belief\. The two should not be identified\.
Substituting the implicit emission into the filtering recursion gives
qt\(z\)\\displaystyle q\_\{t\}\(z\)∝πt\(z\)btimp\(xt∣z\)\\displaystyle\\propto\\pi\_\{t\}\(z\)b\_\{t\}^\{\\mathrm\{imp\}\}\(x\_\{t\}\\mid z\)=πt\(z\)pt\(xt\)eθ,t\(z∣xt\)ρt\(z\)\\displaystyle=\\pi\_\{t\}\(z\)\\frac\{p\_\{t\}\(x\_\{t\}\)e\_\{\\theta,t\}\(z\\mid x\_\{t\}\)\}\{\\rho\_\{t\}\(z\)\}∝πt\(z\)eθ,t\(z∣xt\)ρt\(z\)\.\\displaystyle\\propto\\pi\_\{t\}\(z\)\\frac\{e\_\{\\theta,t\}\(z\\mid x\_\{t\}\)\}\{\\rho\_\{t\}\(z\)\}\.Hence the history\-dependent filtering distribution may equivalently be written as
qt\(z\)∝πt\(z\)eθ,t\(z∣xt\)ρt\(z\)\.q\_\{t\}\(z\)\\propto\\pi\_\{t\}\(z\)\\frac\{e\_\{\\theta,t\}\(z\\mid x\_\{t\}\)\}\{\\rho\_\{t\}\(z\)\}\.\(45\)
For an exact sequence\-level interpretation, the induced marginals in[Eq\.44](https://arxiv.org/html/2608.13621#A5.E44)must additionally be dynamically consistent with the transition:
ρt\+1\(z′\)=∑zρt\(z\)At\(z,z′\)\.\\rho\_\{t\+1\}\(z^\{\\prime\}\)=\\sum\_\{z\}\\rho\_\{t\}\(z\)A\_\{t\}\(z,z^\{\\prime\}\)\.\(46\)Finally, the PIB\-VJEPA history encoder must coincide with the filtering distribution generated by[Eq\.45](https://arxiv.org/html/2608.13621#A5.E45)\.
### E\.4Completion of the proof for Theorem\.[1](https://arxiv.org/html/2608.13621#Thmtheorem1)
We can now verify the four sufficient conditions in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1)\. First,AtA\_\{t\}defines first\-order latent Markov dynamics\. Second, one of the three constructions above supplies a valid state\-to\-observation conditional\. Third,[Eq\.41](https://arxiv.org/html/2608.13621#A5.E41), or[Eq\.46](https://arxiv.org/html/2608.13621#A5.E46)in the implicit case, ensures that the latent marginals evolve under the same transition kernel\. Fourth,[Eq\.43](https://arxiv.org/html/2608.13621#A5.E43)identifies the history\-dependent context encoder with the Bayesian filtering posterior of that transition–emission model\.
Therefore the joint sequence distribution is precisely[Eq\.40](https://arxiv.org/html/2608.13621#A5.E40), and the context encoder represents its filtering distribution\. This proves the sufficient\-condition statement in[Theorem1](https://arxiv.org/html/2608.13621#Thmtheorem1)\.
These conditions are stronger than architectural correspondence alone\. In particular, a valid transition and emission specify an HMM\-compatible generative model, but an arbitrary JEPA history encoder need not equal its Bayesian filter, and independently induced one\-time latent marginals need not evolve according to the learned transition\. The additional consistency conditions are what promote an emission\-complete latent\-state representation to exact sequence\-level HMM equivalence\.
## Appendix FExperimental Details
This appendix provides the data\-generation procedures, model architectures, optimization settings, evaluation metrics, and supplementary results for the experiments in[Section7](https://arxiv.org/html/2608.13621#S7)\. The experiments are deliberately small and synthetic because their purpose is to isolate the structural claims of the paper under known latent dynamics rather than to benchmark large\-scale forecasting performance\.
Unless otherwise stated, experiments involving sampled data or learned models use five random seeds,
\{0,1,2,3,4\},\\\{0,1,2,3,4\\\},and report mean±\\pmone standard deviation across seeds\. Exact finite calculations in Experiment 3, such as deterministic\-partition enumeration, are deterministic and are therefore reported without seed variability\. We do not perform formal hypothesis tests; the experiments are intended as controlled structural diagnostics, and across\-seed variability is reported to expose sampling and optimization variability\.
For the finite\-HMM experiments, the paper configuration uses400400training sequences and160160test sequences of length8080, with prediction horizons
ℋ=\{1,2,4,8\}\.\\mathcal\{H\}=\\\{1,2,4,8\\\}\.No separate validation split is used because hyperparameter selection is not the purpose of these controlled diagnostics\. Within a seed, competing methods are evaluated on the same generated data whenever a paired comparison is intended\. Experiment 4 independently regenerates the separated\-emission data used in Experiment 1 with the same data\-generating process and seeds, but none of the fitted Experiment 1 models is reused\.
### F\.1Evaluation metrics
We collect here the evaluation metrics used across the experiments\. This also separates*representation recovery*,*transition recovery*,*predictive performance*, and*probabilistic\-model fit*, which measure different aspects of the proposed correspondence\.
#### State recovery: ARI and NMI\.
When ground\-truth latent states are available, we convert each learned categorical distributionqtq\_\{t\}to a hard state assignment
S^t=argmaxkqt\(k\)\.\\widehat\{S\}\_\{t\}=\\arg\\max\_\{k\}q\_\{t\}\(k\)\.We report the adjusted Rand index \(ARI\) and normalized mutual information \(NMI\) between the learned assignmentsS^t\\widehat\{S\}\_\{t\}and ground\-truth statesStS\_\{t\}\.
For a contingency table with entriesnijn\_\{ij\}, row sumsaia\_\{i\}, column sumsbjb\_\{j\}, and total sample sizeNN, ARI is
ARI=∑ij\(nij2\)−\(∑i\(ai2\)\)\(∑j\(bj2\)\)\(N2\)12\[∑i\(ai2\)\+∑j\(bj2\)\]−\(∑i\(ai2\)\)\(∑j\(bj2\)\)\(N2\)\.\\operatorname\{ARI\}=\\frac\{\\displaystyle\\sum\_\{ij\}\\binom\{n\_\{ij\}\}\{2\}\-\\frac\{\\left\(\\sum\_\{i\}\\binom\{a\_\{i\}\}\{2\}\\right\)\\left\(\\sum\_\{j\}\\binom\{b\_\{j\}\}\{2\}\\right\)\}\{\\binom\{N\}\{2\}\}\}\{\\displaystyle\\frac\{1\}\{2\}\\left\[\\sum\_\{i\}\\binom\{a\_\{i\}\}\{2\}\+\\sum\_\{j\}\\binom\{b\_\{j\}\}\{2\}\\right\]\-\\frac\{\\left\(\\sum\_\{i\}\\binom\{a\_\{i\}\}\{2\}\\right\)\\left\(\\sum\_\{j\}\\binom\{b\_\{j\}\}\{2\}\\right\)\}\{\\binom\{N\}\{2\}\}\}\.\(47\)ARI corrects the ordinary Rand index for agreement expected by chance\. A value of11denotes identical partitions, while values near00correspond to chance\-level agreement under the adjustment\.
NMI is computed using the arithmetic normalization,
NMI\(S,S^\)=2I\(S,S^\)H\(S\)\+H\(S^\)\.\\operatorname\{NMI\}\(S,\\widehat\{S\}\)=\\frac\{2I\(S;\\widehat\{S\}\)\}\{H\(S\)\+H\(\\widehat\{S\}\)\}\.\(48\)NMI lies in\[0,1\]\[0,1\], with larger values indicating greater shared information between the learned and ground\-truth state partitions\. Both ARI and NMI are invariant to permutation of categorical state labels\.
#### Permutation alignment\.
Although ARI and NMI do not require label alignment, transition matrices and predicted categorical probabilities do\. We therefore compute a Hungarian assignment on the*training\-set*hard state assignments\. LetMMdenote the resulting permutation matrix mapping learned\-state order to ground\-truth\-state order\. A learned transition matrixA^\\widehat\{A\}is aligned as
A^aligned=M⊤A^M\.\\widehat\{A\}\_\{\\mathrm\{aligned\}\}=M^\{\\top\}\\widehat\{A\}M\.The mapping is fitted only on training assignments and then held fixed for test evaluation\.
#### Transition recovery\.
When the true transition matrixA⋆A^\{\\star\}is known, we measure normalized Frobenius error,
ℰA=‖A^aligned−A⋆‖F‖A⋆‖F\.\\mathcal\{E\}\_\{A\}=\\frac\{\\left\\\|\\widehat\{A\}\_\{\\mathrm\{aligned\}\}\-A^\{\\star\}\\right\\\|\_\{F\}\}\{\\left\\\|A^\{\\star\}\\right\\\|\_\{F\}\}\.\(49\)Lower values indicate more faithful recovery of the underlying Markov transition law\. This metric evaluates the learned dynamics themselves rather than only their downstream predictive consequences\.
#### True\-state prediction NLL\.
For prediction horizonhh, letq^t\+h\\widehat\{q\}\_\{t\+h\}denote the predicted categorical distribution after alignment to ground\-truth state order\. We report
ℒhstate=−𝔼t\[logq^t\+h\(St\+h\)\]\.\\mathcal\{L\}\_\{h\}^\{\\mathrm\{state\}\}=\-\\mathbb\{E\}\_\{t\}\\left\[\\log\\widehat\{q\}\_\{t\+h\}\(S\_\{t\+h\}\)\\right\]\.\(50\)Thus the metric measures the probability assigned to the actual future latent state\. Lower values are better\. For the shared\-transition models,
q^t\+h=qtAh\.\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\}\.
#### Observation\-sequence NLL\.
For models equipped with a transition–emission likelihood, observation\-space fit is measured by negative log\-likelihood per time step,
ℒseq=−1T𝔼\[logpϕ,ψ\(X1:T\)\],\\mathcal\{L\}\_\{\\mathrm\{seq\}\}=\-\\frac\{1\}\{T\}\\mathbb\{E\}\\left\[\\log p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)\\right\],\(51\)wherepϕ,ψ\(X1:T\)p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)denotes the marginal observation\-sequence density induced by the latent transition and emission models after marginalizing the latent\-state sequence\. In the finite\-state setting used in our experiments,
pϕ,ψ\(X1:T\)=∑Z1:Tpϕ\(Z1\)\[∏t=1Tpψ\(Xt∣Zt\)\]\[∏t=1T−1pϕ\(Zt\+1∣Zt\)\]\.p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)=\\sum\_\{Z\_\{1:T\}\}p\_\{\\phi\}\(Z\_\{1\}\)\\left\[\\prod\_\{t=1\}^\{T\}p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}\)\\right\]\\left\[\\prod\_\{t=1\}^\{T\-1\}p\_\{\\phi\}\(Z\_\{t\+1\}\\mid Z\_\{t\}\)\\right\]\.Thus, the likelihood integrates out the unobserved latent trajectory rather than conditioning on the ground\-truth latent states\. In practice, this marginalization is evaluated exactly and efficiently by the HMM forward algorithm rather than by explicitly enumerating all possible latent\-state sequences\.
Unlike state\-prediction NLL, this metric evaluates the probability density assigned to the observed sequence rather than the probability assigned to the known synthetic latent state\. In Experiment 4, the same mathematical quantity plays different roles across training regimes\. For the HMM\+filter and hybrid regimes, it is optimized during training asℒHMM\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}and subsequently evaluated on held\-out sequences\. For the JEPA\-only regime, no observation\-sequence likelihood is optimized during representation learning;ℒseq\\mathcal\{L\}\_\{\\mathrm\{seq\}\}is computed only after fitting the post\-hoc emission model\. It therefore serves as a common evaluation metric across the three regimes rather than a common training objective\.
#### Path disagreement\.
To measure whether direct and composed multi\-step predictions agree, we use
𝒟path\(h1,h2\)=𝔼t\[‖qtAh1\+h2−\(qtAh1\)Ah2‖1\]\.\\mathcal\{D\}\_\{\\mathrm\{path\}\}\(h\_\{1\},h\_\{2\}\)=\\mathbb\{E\}\_\{t\}\\left\[\\left\\\|q\_\{t\}A\_\{h\_\{1\}\+h\_\{2\}\}\-\(q\_\{t\}A\_\{h\_\{1\}\}\)A\_\{h\_\{2\}\}\\right\\\|\_\{1\}\\right\]\.\(52\)For MCJEPA with one shared transition matrix,
and therefore
𝒟path\(h1,h2\)=0\\mathcal\{D\}\_\{\\mathrm\{path\}\}\(h\_\{1\},h\_\{2\}\)=0algebraically, up to numerical precision\. For independently learned horizon\-specific matrices no such guarantee exists\.
#### Filtering KL\.
When both an exact model\-based filtering distribution and an amortized context\-encoder distribution are available, we report
𝒟filter=𝔼t\[KL\(qtexact∥qtenc\)\]\.\\mathcal\{D\}\_\{\\mathrm\{filter\}\}=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(q\_\{t\}^\{\\mathrm\{exact\}\}\\,\\middle\\\|\\,q\_\{t\}^\{\\mathrm\{enc\}\}\\right\)\\right\]\.\(53\)Lower values mean that the amortized encoder more closely reproduces the corresponding filtering belief\. In Experiment 4 this quantity is training aligned for the HMM\+filter and hybrid regimes because both explicitly optimize filtering distillation; it is therefore interpreted as a role diagnostic rather than an independent generalization metric\.
#### Accuracy, Brier score, and entropy\.
Experiment 2 additionally reports state accuracy,
Acc=1N∑t𝟏\{argmaxkqt\(k\)=St\},\\operatorname\{Acc\}=\\frac\{1\}\{N\}\\sum\_\{t\}\\mathbf\{1\}\\left\\\{\\arg\\max\_\{k\}q\_\{t\}\(k\)=S\_\{t\}\\right\\\},\(54\)the multiclass Brier score,
Brier=1N∑t∑k\(qt\(k\)−𝟏\{St=k\}\)2,\\operatorname\{Brier\}=\\frac\{1\}\{N\}\\sum\_\{t\}\\sum\_\{k\}\\left\(q\_\{t\}\(k\)\-\\mathbf\{1\}\\\{S\_\{t\}=k\\\}\\right\)^\{2\},\(55\)and mean posterior entropy,
H¯=1N∑tH\(qt\)\.\\overline\{H\}=\\frac\{1\}\{N\}\\sum\_\{t\}H\(q\_\{t\}\)\.\(56\)Accuracy measures hard classification correctness, whereas NLL and Brier score retain information about probabilistic confidence\. Posterior entropy is descriptive and should not be interpreted as a performance metric by itself\.
### F\.2Experiment 1: finite\-HMM recovery and Markov composition
#### Data generation\.
We generate observations from a four\-state stationary Gaussian HMM\. The ground\-truth transition matrix is
A⋆=\[0\.850\.100\.050\.000\.050\.850\.100\.000\.000\.050\.850\.100\.100\.000\.050\.85\]\.A^\{\\star\}=\\begin\{bmatrix\}0\.85&0\.10&0\.05&0\.00\\\\ 0\.05&0\.85&0\.10&0\.00\\\\ 0\.00&0\.05&0\.85&0\.10\\\\ 0\.10&0\.00&0\.05&0\.85\\end\{bmatrix\}\.The initial state is sampled from the stationary distribution ofA⋆A^\{\\star\}\. Conditional on stateSt=kS\_\{t\}=k, the two\-dimensional observation is generated as
Xt\|St=k∼𝒩\(μk,σ2I2\),X\_\{t\}\\mid S\_\{t\}=k\\sim\\mathcal\{N\}\(\\mu\_\{k\},\\sigma^\{2\}I\_\{2\}\),with state means
μ1=\(−1,−1\),μ2=\(−1,1\),μ3=\(1,1\),μ4=\(1,−1\)\.\\mu\_\{1\}=\(\-1,\-1\),\\qquad\\mu\_\{2\}=\(\-1,1\),\\qquad\\mu\_\{3\}=\(1,1\),\\qquad\\mu\_\{4\}=\(1,\-1\)\.We consider two emission regimes:
σ=0\.35\(separated\),σ=0\.80\(ambiguous\)\.\\sigma=0\.35\\quad\\text\{\(separated\)\},\\qquad\\sigma=0\.80\\quad\\text\{\(ambiguous\)\}\.For each seed and regime we independently generate400400training sequences and160160test sequences, each of length8080\.
#### MCJEPA architecture\.
The online context encoder is a one\-layer GRU with hidden dimension4848, followed by a linear projection toK=4K=4logits and a softmax:
qθ\(Zt∣X≤t\)=softmax\(Wht\+b\)\.q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)=\\operatorname\{softmax\}\\left\(Wh\_\{t\}\+b\\right\)\.The target encoder has the same architecture but processes eachXtX\_\{t\}as an independent length\-one sequence, producing a local target distribution\. Its parameters are initialized from the online encoder and subsequently updated by exponential moving average\.
For the shared\-transition MCJEPA model, a trainable4×44\\times 4logit matrix is row\-normalized by softmax,
A=softmaxrow\(LA\),A=\\operatorname\{softmax\}\_\{\\mathrm\{row\}\}\(L\_\{A\}\),and horizon\-hhprediction uses
q^t\+h=qtAh\.\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\}\.The transition logits are initialized with a mild diagonal bias,
The horizon\-specific baseline uses the same online and target encoders but replaces the shared matrix with independent row\-stochastic matrices
A1,A2,A4,A8\.A\_\{1\},\\quad A\_\{2\},\\quad A\_\{4\},\\quad A\_\{8\}\.Each matrix is initialized with the same diagonal logit bias, but no constraint requires
#### Warm start and optimization\.
To make the small synthetic recovery experiment stable and reproducible, both categorical JEPA variants receive an unsupervised K\-means warm start\. K\-means withK=4K=4and1010initializations is fitted to individual training observations; ground\-truth states are never used\. The online encoder is then trained for100100warm\-start updates with Adam at learning rate
to predict the K\-means assignments, after which the target encoder is copied from the online encoder\.
The main MCJEPA optimization runs for180180epochs with mini\-batches of6464sequences and Adam learning rate
The prediction loss averages the target\-to\-prediction KL divergence overℋ=\{1,2,4,8\}\\mathcal\{H\}=\\\{1,2,4,8\\\}:
ℒpred=1\|ℋ\|∑h∈ℋ𝔼\[KL\(q¯t\+h∥qtAh\)\],\\mathcal\{L\}\_\{\\mathrm\{pred\}\}=\\frac\{1\}\{\|\\mathcal\{H\}\|\}\\sum\_\{h\\in\\mathcal\{H\}\}\\mathbb\{E\}\\left\[\\mathrm\{KL\}\\left\(\\bar\{q\}\_\{t\+h\}\\,\\middle\\\|\\,q\_\{t\}A\_\{h\}\\right\)\\right\],whereAh=AhA\_\{h\}=A^\{h\}for MCJEPA andAhA\_\{h\}is independently learned for the horizon\-specific baseline\.
The state\-use terms are
ℒocc\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{occ\}\}=KL\(q¯∥Unif\(K\)\),\\displaystyle=\\mathrm\{KL\}\\left\(\\bar\{q\}\\middle\\\|\\operatorname\{Unif\}\(K\)\\right\),ℒent\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ent\}\}=𝔼t\[H\(qt\)\],\\displaystyle=\\mathbb\{E\}\_\{t\}\\left\[H\(q\_\{t\}\)\\right\],with
λocc=0\.30,λent=0\.02\.\\lambda\_\{\\mathrm\{occ\}\}=0\.30,\\qquad\\lambda\_\{\\mathrm\{ent\}\}=0\.02\.Becauseℒent\\mathcal\{L\}\_\{\\mathrm\{ent\}\}is minimized, it encourages confident per\-example assignments\. The target encoder uses EMA coefficient
Gradients are clipped to norm55\.
#### Gaussian\-HMM baseline\.
The HMM baseline uses the correctly specified four\-state family with a learned initial\-state distribution, row\-stochastic transition matrix, and state\-conditional diagonal Gaussian emissions\. The emission means are initialized from K\-means cluster centers, and the diagonal variances are initialized from within\-cluster variances with a small additive floor\.
The HMM is trained directly through the observation\-sequence NLL defined above, computed exactly by the forward algorithm\. We use Adam with learning rate
for350350optimization steps and clip gradients to norm1010\.
#### Permutation alignment and evaluation\.
The Hungarian alignment and common metrics follow[SectionF\.1](https://arxiv.org/html/2608.13621#A6.SS1)\. ARI and NMI evaluate recovery of the hidden\-state partition, transition error evaluates recovery ofA⋆A^\{\\star\}, and true\-state NLL evaluates future\-state prediction at
h∈\{1,2,4,8\}\.h\\in\\\{1,2,4,8\\\}\.
For path consistency we evaluate
\(h1,h2\)∈\{\(1,1\),\(2,2\),\(4,4\)\}\.\(h\_\{1\},h\_\{2\}\)\\in\\\{\(1,1\),\(2,2\),\(4,4\)\\\}\.For the shared\-transition model,
qtAh1\+h2=\(qtAh1\)Ah2q\_\{t\}A^\{h\_\{1\}\+h\_\{2\}\}=\(q\_\{t\}A^\{h\_\{1\}\}\)A^\{h\_\{2\}\}exactly, so path disagreement is zero by construction\. The horizon\-specific baseline has no corresponding constraint\.
#### State\-usage metrics\.
In addition to the common evaluation metrics, we monitor both soft and hard effective state counts\. If
q¯=1N∑nqn\\bar\{q\}=\\frac\{1\}\{N\}\\sum\_\{n\}q\_\{n\}is the average soft assignment distribution, then
Keffsoft=exp\(H\(q¯\)\)\.K\_\{\\mathrm\{eff\}\}^\{\\mathrm\{soft\}\}=\\exp\\left\(H\(\\bar\{q\}\)\\right\)\.For hard assignments, letp^k\\widehat\{p\}\_\{k\}be the empirical frequency of statekk\. We define
Keffhard=exp\(−∑k:p^k\>0p^klogp^k\)\.K\_\{\\mathrm\{eff\}\}^\{\\mathrm\{hard\}\}=\\exp\\left\(\-\\sum\_\{k:\\widehat\{p\}\_\{k\}\>0\}\\widehat\{p\}\_\{k\}\\log\\widehat\{p\}\_\{k\}\\right\)\.The reported assignment entropy is
H¯assign=𝔼t\[H\(qt\)\]\.\\overline\{H\}\_\{\\mathrm\{assign\}\}=\\mathbb\{E\}\_\{t\}\[H\(q\_\{t\}\)\]\.
Effective state count measures diversity of state usage, whereas assignment entropy measures confidence of individual assignments\. Low assignment entropy is not desirable by itself: an encoder that confidently maps every observation to one state also has low entropy\. State\-use diversity and assignment confidence must therefore be interpreted jointly\.
#### Collapse ablations\.
The collapse diagnostic is run separately from the warm\-started recovery experiment\. We generate a new separated\-emission dataset withσ=0\.35\\sigma=0\.35and train the shared\-AAmodel from random initialization, deliberately omitting the K\-means warm start\. The four settings are
\(λocc,λent\)∈\{\(0,0\),\(0\.30,0\),\(0,0\.02\),\(0\.30,0\.02\)\}\.\(\\lambda\_\{\\mathrm\{occ\}\},\\lambda\_\{\\mathrm\{ent\}\}\)\\in\\left\\\{\(0,0\),\\,\(0\.30,0\),\\,\(0,0\.02\),\\,\(0\.30,0\.02\)\\right\\\}\.Each model is trained for180180epochs with batch size6464; for this diagnostic the EMA coefficient isτ=0\.99\\tau=0\.99\. This deliberately creates a more collapse\-prone optimization problem and isolates the complementary roles of the two penalties: occupancy regularization discourages global state under\-use, whereas entropy regularization encourages confident per\-example assignments\.
#### Supplementary results\.
[Figure7](https://arxiv.org/html/2608.13621#A6.F7)reports the complete multi\-horizon prediction curves omitted from the main text\. In the separated regime, all three models remain close across horizons\. Under ambiguous emissions, the correctly specified HMM remains strongest, while the horizon\-specific and shared\-transition JEPA models exhibit similar predictive NLL despite their substantially different structural consistency\.


Figure 7:Complete multi\-horizon state\-prediction results for Experiment 1\.*Left:*separated emissions\.*Right:*ambiguous emissions\. Lower true\-state NLL is better\. Error bars denote mean±\\pmone standard deviation over five seeds\.The separated\-emission path\-consistency result is shown in[Fig\.8](https://arxiv.org/html/2608.13621#A6.F8)\. As in the ambiguous regime, the shared\-AAmodel is exactly compositionally consistent, whereas independently trained horizon\-specific matrices exhibit nonzero direct\-versus\-composed disagreement\.
Figure 8:Direct\-versus\-composed prediction disagreement in the separated\-emission regime\. MCJEPA has zero path disagreement by construction because all horizons are powers of one shared transition matrix\.[Figure9](https://arxiv.org/html/2608.13621#A6.F9)supplements the main\-text ARI ablation with state\-usage and assignment\-confidence diagnostics\.


Figure 9:Additional collapse diagnostics for Experiment 1\.*Left:*effective number of hard states\.*Right:*mean per\-example assignment entropy\. Occupancy regularization primarily promotes broad global state usage, whereas entropy regularization promotes confident assignments; neither diagnostic should be interpreted in isolation\.
### F\.3Experiment 2: filtering under emission ambiguity
#### Data generation\.
Experiment 2 uses a persistent two\-state HMM with transition matrix
A=\[0\.970\.030\.030\.97\]\.A=\\begin\{bmatrix\}0\.97&0\.03\\\\ 0\.03&0\.97\\end\{bmatrix\}\.Its stationary distribution is uniform,
π=\(0\.5,0\.5\)\.\\pi=\(0\.5,0\.5\)\.The scalar observation model is
Xt\|St=\{𝒩\(−μ,σ2\),St=0,𝒩\(\+μ,σ2\),St=1,X\_\{t\}\\mid S\_\{t\}=\\begin\{cases\}\\mathcal\{N\}\(\-\\mu,\\sigma^\{2\}\),&S\_\{t\}=0,\\\\ \\mathcal\{N\}\(\+\\mu,\\sigma^\{2\}\),&S\_\{t\}=1,\\end\{cases\}with
Emission ambiguity is controlled by
μσ∈\{2\.0,1\.25,0\.75,0\.45\}\.\\frac\{\\mu\}\{\\sigma\}\\in\\\{2\.0,1\.25,0\.75,0\.45\\\}\.For every seed and separation value we generate160160sequences of length8080, corresponding to12,80012\{,\}800state–observation pairs per setting\. Because the comparison uses exact oracle posteriors, there is no learned train/test model split in this experiment; independent random seeds provide repeated sampled datasets\.
#### Exact local evidence\.
The local posterior uses only the current observation and the stationary state prior:
qtlocal\(k\)=p\(St=k∣Xt\)=πkp\(Xt∣St=k\)∑jπjp\(Xt∣St=j\)\.q\_\{t\}^\{\\mathrm\{local\}\}\(k\)=p\(S\_\{t\}=k\\mid X\_\{t\}\)=\\frac\{\\pi\_\{k\}p\(X\_\{t\}\\mid S\_\{t\}=k\)\}\{\\sum\_\{j\}\\pi\_\{j\}p\(X\_\{t\}\\mid S\_\{t\}=j\)\}\.This is the oracle counterpart of a local encoderqθ\(Zt∣Xt\)q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{t\}\)\.
#### Exact filtering\.
The filtering posterior incorporates both the propagated previous belief and the current emission evidence\. At the first step,
q0filter\(k\)∝πkp\(X0∣S0=k\)\.q\_\{0\}^\{\\mathrm\{filter\}\}\(k\)\\propto\\pi\_\{k\}p\(X\_\{0\}\\mid S\_\{0\}=k\)\.Thereafter,
qtfilter∝\(qt−1filterA\)⊙p\(Xt∣St\),q\_\{t\}^\{\\mathrm\{filter\}\}\\propto\\left\(q\_\{t\-1\}^\{\\mathrm\{filter\}\}A\\right\)\\odot p\(X\_\{t\}\\mid S\_\{t\}\),followed by normalization across the two states\. This quantity is exactly
p\(St∣X≤t\)\.p\(S\_\{t\}\\mid X\_\{\\leq t\}\)\.
No learned HMM and MCJEPA models are being compared in Experiment 2\. Both curves are oracle calculations under the same known generating process\. This design isolates the informational value of temporal history from representation\-learning and optimization effects\.
#### Evaluation\.
We use the accuracy, state NLL, Brier score, and posterior entropy defined in[SectionF\.1](https://arxiv.org/html/2608.13621#A6.SS1)\. Accuracy gives the most immediately interpretable state\-recovery comparison, while NLL measures whether the posterior assigns high probability to the realized state\. Brier score provides a complementary proper probabilistic score, and entropy records posterior confidence\.
#### Representative sequence selection\.
The representative trajectory in[Fig\.4](https://arxiv.org/html/2608.13621#S7.F4)is selected only for visualization; all quantitative results use all generated sequences\. We use the most ambiguous setting,
μ/σ=0\.45,\\mu/\\sigma=0\.45,from the first seed and search4545\-step windows centered on genuine latent\-state transitions\. Windows containing one to three true state switches receive a small preference, and among candidate windows we favor those in which filtering gives a larger realized\-state NLL improvement over local evidence\. This produces a transition\-rich example that visibly illustrates the mechanism quantified by the aggregate experiment rather than selecting the first sequence arbitrarily\.
#### Supplementary result\.
[Figure10](https://arxiv.org/html/2608.13621#A6.F10)gives the complete state\-NLL comparison across emission separations\. The filtering advantage grows asμ/σ\\mu/\\sigmadecreases, matching the accuracy trend reported in the main text\.
Figure 10:State NLL for exact local evidence and exact filtering under increasing emission ambiguity\. Smallerμ/σ\\mu/\\sigmacorresponds to stronger overlap between the two Gaussian emissions\. Error bars denote mean±\\pmone standard deviation over five independently generated datasets\.
### F\.4Experiment 3: predictive compression and Markovization
#### Second\-order binary process\.
Let
The data\-generating process is
p\(Xt\+1=1∣Xt−1,Xt\)=\{0\.10,\(Xt−1,Xt\)=\(0,0\),0\.90,\(Xt−1,Xt\)=\(0,1\),0\.80,\(Xt−1,Xt\)=\(1,0\),0\.20,\(Xt−1,Xt\)=\(1,1\)\.p\(X\_\{t\+1\}=1\\mid X\_\{t\-1\},X\_\{t\}\)=\\begin\{cases\}0\.10,&\(X\_\{t\-1\},X\_\{t\}\)=\(0,0\),\\\\ 0\.90,&\(X\_\{t\-1\},X\_\{t\}\)=\(0,1\),\\\\ 0\.80,&\(X\_\{t\-1\},X\_\{t\}\)=\(1,0\),\\\\ 0\.20,&\(X\_\{t\-1\},X\_\{t\}\)=\(1,1\)\.\\end\{cases\}The corresponding first\-order transition matrix on pair states
\(00,01,10,11\)\(00,01,10,11\)is
Apair=\[0\.90\.100000\.10\.90\.20\.800000\.80\.2\]\.A\_\{\\mathrm\{pair\}\}=\\begin\{bmatrix\}0\.9&0\.1&0&0\\\\ 0&0&0\.1&0\.9\\\\ 0\.2&0\.8&0&0\\\\ 0&0&0\.8&0\.2\\end\{bmatrix\}\.Its stationary distribution is
πpair=\(1641,841,841,941\)≈\(0\.3902,0\.1951,0\.1951,0\.2195\)\.\\pi\_\{\\mathrm\{pair\}\}=\\left\(\\frac\{16\}\{41\},\\frac\{8\}\{41\},\\frac\{8\}\{41\},\\frac\{9\}\{41\}\\right\)\\approx\(0\.3902,0\.1951,0\.1951,0\.2195\)\.The exact joint distribution of
Ht=\(Xt−2,Xt−1,Xt\)H\_\{t\}=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\)andYt=Xt\+1Y\_\{t\}=X\_\{t\+1\}is constructed analytically from this stationary pair chain\. Consequently, the exact representation controls and deterministic frontier do not require Monte Carlo estimation\.
#### Exact representation controls\.
A deterministic representation is a mapping
For such a mapping,
I\(H,Z\)=H\(Z\),I\(H;Z\)=H\(Z\),becauseH\(Z∣H\)=0H\(Z\\mid H\)=0\. For each representation we construct the exact joint distributionp\(z,y\)p\(z,y\)and compute
I\(Z,Y\)=∑z,yp\(z,y\)logp\(z,y\)p\(z\)p\(y\)\.I\(Z;Y\)=\\sum\_\{z,y\}p\(z,y\)\\log\\frac\{p\(z,y\)\}\{p\(z\)p\(y\)\}\.All logarithms are natural, so information quantities and NLLs are measured in nats\.
The optimal one\-step probabilistic predictor for a fixed representation is the exact conditional distributionp\(Y∣Z\)p\(Y\\mid Z\)\. Its prediction NLL is therefore
ℒpred=−∑z,yp\(z,y\)logp\(y∣z\)=H\(Y∣Z\)\.\\mathcal\{L\}\_\{\\mathrm\{pred\}\}=\-\\sum\_\{z,y\}p\(z,y\)\\log p\(y\\mid z\)=H\(Y\\mid Z\)\.
The three control mappings are
Zover\\displaystyle Z^\{\\mathrm\{over\}\}=\(Xt−2,Xt−1,Xt\),\\displaystyle=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\),Zsuff\\displaystyle Z^\{\\mathrm\{suff\}\}=\(Xt−1,Xt\),\\displaystyle=\(X\_\{t\-1\},X\_\{t\}\),Zunder\\displaystyle Z^\{\\mathrm\{under\}\}=Xt\.\\displaystyle=X\_\{t\}\.The first has eight possible values, the second four, and the third two\. Their exact information and prediction quantities are reproduced in[Table9](https://arxiv.org/html/2608.13621#A6.T9)for completeness\.
Table 9:Exact structural controls for Experiment 3\. The four\-state predictive pair removes redundant history while preserving all one\-step predictive information\.
#### Enumeration of deterministic partitions\.
BecauseHtH\_\{t\}has only eight possible values, every deterministic compression can be enumerated\. A deterministic encoder identifies histories that share the same output label and therefore corresponds to a set partition of the eight histories\. The number of such partitions is the eighth Bell number,
The Bell numberBnB\_\{n\}counts the number of partitions of annn\-element set into nonempty unlabeled subsets\. Here, each of the eight possible three\-bit histories is an element and each subset collects histories assigned to the same latent state\. The implementation uses restricted\-growth strings to enumerate each partition exactly once, thereby eliminating duplicates caused solely by relabeling latent states\.
For every partition we compute
I\(H,Z\),I\(Z,Y\),H\(Y∣Z\)I\(H;Z\),\\qquad I\(Z;Y\),\\qquad H\(Y\\mid Z\)exactly\. A representation belongs to the deterministic Pareto frontier if there is no other deterministic representation with no largerI\(H,Z\)I\(H;Z\)and strictly largerI\(Z,Y\)I\(Z;Y\)\. A numerical tolerance of10−1210^\{\-12\}is used when constructing this frontier\. Of the41404140deterministic partitions,3636are nondominated\.
#### Exact compression sweep\.
For each compression coefficientβ\\beta, every deterministic partition is scored using
ℒPIBexp=H\(Y∣Z\)\+βI\(H,Z\)\.\\mathcal\{L\}\_\{\\mathrm\{PIB\}\}^\{\\mathrm\{exp\}\}=H\(Y\\mid Z\)\+\\beta I\(H;Z\)\.The tested values are
β∈\{0,0\.0005,0\.001,0\.003,0\.005,0\.01,0\.012,0\.014,0\.015,0\.02,0\.03,0\.10,0\.30,1\.0\}\.\\beta\\in\\\{0,\\,0\.0005,\\,0\.001,\\,0\.003,\\,0\.005,\\,0\.01,\\,0\.012,\\,0\.014,\\,0\.015,\\,0\.02,\\,0\.03,\\,0\.10,\\,0\.30,\\,1\.0\\\}\.
At score ties, using tolerance10−1010^\{\-10\}, we first select the candidate with lowerI\(H,Z\)I\(H;Z\), then the candidate with fewer occupied states, and finally largerI\(Z,Y\)I\(Z;Y\)\. This convention matters atβ=0\\beta=0: several deterministic representations achieve the same minimum prediction loss, so the four\-state minimal sufficient representation is the*reported tie\-broken representative*, not a uniquely preferred solution of prediction alone\.
The exact optimum evolves as shown in[Table10](https://arxiv.org/html/2608.13621#A6.T10)\. For every tested positive value throughβ=0\.014\\beta=0\.014, the optimum is the known four\-state predictive pair\. Atβ=0\.015\\beta=0\.015, the optimum switches to a two\-state compression and sacrifices a small amount of predictive information\. Atβ=1\\beta=1, complete compression to a single state becomes optimal\.
Table 10:Exact deterministic compression regimes in Experiment 3\. Atβ=0\\beta=0, the four\-state entry is the lower\-information representative selected among predictively tied optima\.
#### Learned stochastic continuation\.
The orange continuation curve in[Fig\.5](https://arxiv.org/html/2608.13621#S7.F5)is generated separately from the exhaustive deterministic search\. Its purpose is to show how gradient optimization behaves as compression pressure is gradually increased\.
Because the underlying problem is finite, the learned encoder is represented directly as a categorical table
qθ\(z∣h\),h∈\{0,…,7\},z∈\{0,…,7\},q\_\{\\theta\}\(z\\mid h\),\\qquad h\\in\\\{0,\\ldots,7\\\},\\qquad z\\in\\\{0,\\ldots,7\\\},rather than by a neural sequence encoder\. This removes architectural capacity as a confound\. The predictor is a second categorical table,
pω\(Y∣Z\)\.p\_\{\\omega\}\(Y\\mid Z\)\.
The encoder is initialized near the overcomplete identity mappingZ=HZ=H\. Specifically, the diagonal encoder logits are initialized to\+8\+8, the off\-diagonal logits to−8\-8, and independent noise of scale10−310^\{\-3\}is added to break exact symmetry\. The predictor is initialized from the exact conditional distributionp\(Y∣H\)p\(Y\\mid H\)associated with the identity representation\.
For each testedβ\\beta, the model minimizes
ℒlearned=ℒpred\+βI\(H,Z\)\\mathcal\{L\}\_\{\\mathrm\{learned\}\}=\\mathcal\{L\}\_\{\\mathrm\{pred\}\}\+\\beta I\(H;Z\)using Adam with learning rate
Each nonzero compression stage receives22002200gradient steps\. The solution at one value ofβ\\betainitializes the next, so the orange curve is a*continuation path*rather than a collection of independently initialized models\. Atβ=0\\beta=0, the deliberately overcomplete eight\-state initialization is retained without an optimization stage\. Five independent perturbation seeds are used, and the plotted curve reports their mean with standard deviations\.
This distinguishes the twoβ=0\\beta=0constructions in the experiment\. The exact deterministic sweep reports the lower\-information four\-state solution after tie\-breaking among equally predictive partitions, whereas the learned continuation deliberately begins from the overcomplete eight\-state solution so that the compression trajectory can be observed\.
The learned encoder is stochastic, whereas the exact blue frontier contains only deterministic mappings\. The learned continuation is therefore interpreted*relative to*the deterministic global reference, not as an optimization method expected to lie exactly on that frontier\.
#### Residual\-predictability diagnostic\.
The residual diagnostic uses five independently generated sequences of length120,000120\{,\}000, initialized from the stationary pair\-state distribution\. Each sequence is divided chronologically into70%70\\%training and30%30\\%held\-out evaluation data\.
The restricted predictor estimates
p^0=p^\(Xt\+1=1∣Zt\),\\widehat\{p\}\_\{0\}=\\widehat\{p\}\(X\_\{t\+1\}=1\\mid Z\_\{t\}\),whereas the history\-augmented predictor receives one additional step of representation history,
p^1=p^\(Xt\+1=1∣Zt,Zt−1\)\.\\widehat\{p\}\_\{1\}=\\widehat\{p\}\(X\_\{t\+1\}=1\\mid Z\_\{t\},Z\_\{t\-1\}\)\.Both are discrete lookup estimators fitted on the training portion with Laplace smoothing parameter
For the three controlled representations, the augmented inputs specialize to
Ztunder=Xt\\displaystyle Z\_\{t\}^\{\\mathrm\{under\}\}=X\_\{t\}:\\displaystyle:\(Zt,Zt−1\)\\displaystyle\(Z\_\{t\},Z\_\{t\-1\}\)=\(Xt,Xt−1\),\\displaystyle=\(X\_\{t\},X\_\{t\-1\}\),Ztsuff=\(Xt−1,Xt\)\\displaystyle Z\_\{t\}^\{\\mathrm\{suff\}\}=\(X\_\{t\-1\},X\_\{t\}\):\\displaystyle:\(Zt,Zt−1\)\\displaystyle\(Z\_\{t\},Z\_\{t\-1\}\)=\(\(Xt−1,Xt\),\(Xt−2,Xt−1\)\),\\displaystyle=\\bigl\(\(X\_\{t\-1\},X\_\{t\}\),\(X\_\{t\-2\},X\_\{t\-1\}\)\\bigr\),Ztover=\(Xt−2,Xt−1,Xt\)\\displaystyle Z\_\{t\}^\{\\mathrm\{over\}\}=\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\):\\displaystyle:\(Zt,Zt−1\)\\displaystyle\(Z\_\{t\},Z\_\{t\-1\}\)=\(\(Xt−2,Xt−1,Xt\),\(Xt−3,Xt−2,Xt−1\)\)\.\\displaystyle=\\bigl\(\(X\_\{t\-2\},X\_\{t\-1\},X\_\{t\}\),\(X\_\{t\-3\},X\_\{t\-2\},X\_\{t\-1\}\)\\bigr\)\.
This makes the positive and negative controls transparent\. ForZt=XtZ\_\{t\}=X\_\{t\}, the augmentationZt−1=Xt−1Z\_\{t\-1\}=X\_\{t\-1\}restores exactly the transition\-relevant variable omitted from the current state\. ForZt=\(Xt−1,Xt\)Z\_\{t\}=\(X\_\{t\-1\},X\_\{t\}\), the only genuinely new observation supplied byZt−1Z\_\{t\-1\}isXt−2X\_\{t\-2\}, which is redundant under the data\-generating process\. The overcomplete representation already contains still more history, so adding its previous state should likewise provide no one\-step predictive benefit\.
On the held\-out portion we compute
MSErestricted\\displaystyle\\mathrm\{MSE\}\_\{\\mathrm\{restricted\}\}=𝔼\[\(Xt\+1−p^0\)2\],\\displaystyle=\\mathbb\{E\}\\left\[\(X\_\{t\+1\}\-\\widehat\{p\}\_\{0\}\)^\{2\}\\right\],MSEhistory\-augmented\\displaystyle\\mathrm\{MSE\}\_\{\\mathrm\{history\\text\{\-\}augmented\}\}=𝔼\[\(Xt\+1−p^1\)2\],\\displaystyle=\\mathbb\{E\}\\left\[\(X\_\{t\+1\}\-\\widehat\{p\}\_\{1\}\)^\{2\}\\right\],and define
Δhist=MSErestricted−MSEhistory\-augmented\.\\Delta\_\{\\mathrm\{hist\}\}=\\mathrm\{MSE\}\_\{\\mathrm\{restricted\}\}\-\\mathrm\{MSE\}\_\{\\mathrm\{history\\text\{\-\}augmented\}\}\.Thus, positiveΔhist\\Delta\_\{\\mathrm\{hist\}\}means thatZt−1Z\_\{t\-1\}contains predictive information absent fromZtZ\_\{t\}\. Values near zero indicate that one further step of representation history does not improve held\-out prediction\.
The measured gains are
0\.1139±0\.00260\.1139\\pm 0\.0026for the insufficient representation,
−6\.8×10−6\-6\.8\\times 10^\{\-6\}for the minimal sufficient representation, and
−3\.7×10−5\-3\.7\\times 10^\{\-5\}for the overcomplete representation\. The tiny negative values are finite\-sample fitting variation and are effectively zero at the scale of the experiment\.
Operationally, this diagnostic measures the held\-out gain from addingZt−1Z\_\{t\-1\}rather than fitting a separate neural regressor to residuals\. It realizes the same sufficiency principle used in the main text: if omitted history still carries transition\-relevant information, augmenting the predictor with an earlier representation state should reduce held\-out prediction error\.
#### Supplementary results\.
[Figure11](https://arxiv.org/html/2608.13621#A6.F11)shows the exact deterministic frontier without the learned continuation overlay\.
Figure 11:Exact deterministic compression frontier for Experiment 3\. All41404140deterministic partitions of the eight histories are evaluated exactly\. The nondominated frontier gives the best achievable deterministic trade\-offs between retaining history information and preserving information aboutXt\+1X\_\{t\+1\}\.[Figure12](https://arxiv.org/html/2608.13621#A6.F12)shows the number of occupied hard states along the compression sweep for both the exact deterministic optimum and the learned continuation\. The learned warm\-started path can differ from the global deterministic optimum because of stochastic parameterization and optimization path dependence\.
Figure 12:Number of occupied hard states as the compression weightβ\\betaincreases\. The dashed reference gives the exact deterministic optimum, while the learned stochastic continuation follows the warm\-started gradient trajectory\.
### F\.5Experiment 4: HMM\-style training of PIB\-VJEPA
#### Independent rerun with a shared latent\-state family\.
Experiment 4 is implemented and run independently from Experiments 1–3\. It regenerates the same separated\-emission data used in Experiment 1 using the same ground\-truth transition matrix, Gaussian state means, emission standard deviation
and data seeds\. For top\-level seeds∈\{0,…,4\}s\\in\\\{0,\\ldots,4\\\}, the training generator uses
and the test generator uses
Consequently, Experiment 4 sees the same paired data realizations as the separated condition of Experiment 1, but none of the fitted Experiment 1 models is reused\. All three Experiment 4 regimes are trained afresh\.
Each seed contains400400training sequences and160160test sequences of length8080\. All regimes useK=4K=4categorical latent states and the same basic one\-layer GRU context\-encoder family with hidden dimension4848\.
#### Common initialization and state\-use regularization\.
All categorical encoder regimes use an unsupervised K\-means warm start withK=4K=4and1010K\-means initializations; ground\-truth states are never used\. The context encoder is pretrained for100100updates at learning rate
to reproduce the K\-means assignments\.
The Gaussian HMM components in the HMM\+filter and hybrid regimes are initialized from the same K\-means partition\. State\-conditional means are initialized from cluster centers, diagonal variances from within\-cluster variances plus0\.050\.05, the initial\-state logits are initialized uniformly, and the transition logits receive a diagonal bias
LA=1\.5I4\.L\_\{A\}=1\.5I\_\{4\}\.
The same state\-use regularization is applied to all three amortized context encoders:
ℒstate=0\.30ℒocc\+0\.02ℒent\.\\mathcal\{L\}\_\{\\mathrm\{state\}\}=0\.30\\,\\mathcal\{L\}\_\{\\mathrm\{occ\}\}\+0\.02\\,\\mathcal\{L\}\_\{\\mathrm\{ent\}\}\.Sharing these coefficients removes state\-regularization strength as a confound in the objective\-level comparison\.
#### JEPA latent regime\.
The JEPA\-only regime is trained from scratch using the same MCJEPA construction and optimization settings as the shared\-AAmodel in Experiment 1\. Its objective is
ℒJEPA=ℒMC\+ℒstate,\\mathcal\{L\}\_\{\\mathrm\{JEPA\}\}=\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+\\mathcal\{L\}\_\{\\mathrm\{state\}\},where
ℒMC=1\|ℋ\|∑h∈ℋ𝔼t\[KL\(sg\(qθ¯\(Zt\+h∣Xt\+h\)\)∥qθ\(Zt∣X≤t\)Ah\)\]\.\\mathcal\{L\}\_\{\\mathrm\{MC\}\}=\\frac\{1\}\{\|\\mathcal\{H\}\|\}\\sum\_\{h\\in\\mathcal\{H\}\}\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\\bigl\(q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\)\\bigr\)\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)A^\{h\}\\right\)\\right\]\.
The context encoder is trained for180180epochs with mini\-batches of6464sequences, Adam learning rate
EMA coefficient
and gradient clipping at norm55\.
No observation model participates in this training\. To evaluate observation\-sequence NLL afterward, we fit a diagonal Gaussian emission distribution to each learned latent state using the soft training assignments:
μ^k=∑n,tqnt\(k\)Xnt∑n,tqnt\(k\)\.\\widehat\{\\mu\}\_\{k\}=\\frac\{\\sum\_\{n,t\}q\_\{nt\}\(k\)X\_\{nt\}\}\{\\sum\_\{n,t\}q\_\{nt\}\(k\)\}\.The diagonal variance is the corresponding soft\-assignment\-weighted second moment aroundμ^k\\widehat\{\\mu\}\_\{k\}, with a minimum variance of0\.030\.03\. The initial\-state distribution is estimated from the mean encoder distribution at the first time step\. These post\-hoc parameters do not backpropagate into either the JEPA encoder or transition matrix\.
The fitted observation model and learned transition are then treated as a fixed HMM solely for evaluation of test sequence NLL and the post\-hoc filtering diagnostic\.
#### HMM sequence \+ filter\-distillation regime\.
The second regime jointly maintains a Gaussian HMM and an amortized GRU context encoder but contains no JEPA latent\-prediction loss\. The HMM contributes
ℒHMM=−1T𝔼\[logpϕ,ψ\(X1:T\)\]\.\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}=\-\\frac\{1\}\{T\}\\mathbb\{E\}\\left\[\\log p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)\\right\]\.At each update, its current exact filtering posterior is computed and detached,
q~t=sg\(pϕ,ψ\(Zt∣X≤t\)\),\\widetilde\{q\}\_\{t\}=\\operatorname\{sg\}\\left\(p\_\{\\phi,\\psi\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\right\),and the context encoder is trained through
ℒfilter=𝔼t\[KL\(q~t∥qθ\(Zt∣X≤t\)\)\]\.\\mathcal\{L\}\_\{\\mathrm\{filter\}\}=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\widetilde\{q\}\_\{t\}\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\right\)\\right\]\.The implemented objective is
ℒHMM\+filter=1\.0ℒHMM\+0\.5ℒfilter\+0\.30ℒocc\+0\.02ℒent\.\\mathcal\{L\}\_\{\\mathrm\{HMM\+filter\}\}=1\.0\\,\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+0\.5\\,\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\+0\.30\\,\\mathcal\{L\}\_\{\\mathrm\{occ\}\}\+0\.02\\,\\mathcal\{L\}\_\{\\mathrm\{ent\}\}\.
The HMM and encoder are jointly optimized for600600full\-data updates using Adam with learning rate
Gradients are clipped to norm1010\. Because the filtering target is detached,ℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}updates the amortized encoder but not the HMM parameters\. Likewise,ℒstate\\mathcal\{L\}\_\{\\mathrm\{state\}\}acts only on the encoder\. Thus, the generative transition and emission parameters are learned through sequence likelihood, while the context encoder learns to amortize the corresponding Bayesian filtering operation\.
#### Hybrid HMM \+ latent regime\.
The hybrid uses the same Gaussian HMM and amortized context\-encoder families as the preceding regime but adds the genuine MCJEPA latent\-prediction objective\. A separate target encoder is initialized from the context encoder after K\-means pretraining and subsequently updated only by EMA\.
Its three principal losses are
ℒHMM\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}=−1T𝔼\[logpϕ,ψ\(X1:T\)\],\\displaystyle=\-\\frac\{1\}\{T\}\\mathbb\{E\}\\left\[\\log p\_\{\\phi,\\psi\}\(X\_\{1:T\}\)\\right\],ℒfilter\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{filter\}\}=𝔼t\[KL\(sg\(q~t\)∥qθ\(Zt∣X≤t\)\)\],\\displaystyle=\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\(\\widetilde\{q\}\_\{t\}\)\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)\\right\)\\right\],ℒMC\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{MC\}\}=1\|ℋ\|∑h∈ℋ𝔼t\[KL\(sg\(qθ¯\(Zt\+h∣Xt\+h\)\)∥qθ\(Zt∣X≤t\)Ah\)\]\.\\displaystyle=\\frac\{1\}\{\|\\mathcal\{H\}\|\}\\sum\_\{h\\in\\mathcal\{H\}\}\\mathbb\{E\}\_\{t\}\\left\[\\mathrm\{KL\}\\left\(\\operatorname\{sg\}\\bigl\(q\_\{\\bar\{\\theta\}\}\(Z\_\{t\+h\}\\mid X\_\{t\+h\}\)\\bigr\)\\,\\middle\\\|\\,q\_\{\\theta\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)A^\{h\}\\right\)\\right\]\.
The critical implementation distinction is that the HMM filtering posterior
q~t=pϕ,ψ\(Zt∣X≤t\)\\widetilde\{q\}\_\{t\}=p\_\{\\phi,\\psi\}\(Z\_\{t\}\\mid X\_\{\\leq t\}\)is used*only*byℒfilter\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\. It is not substituted for the future JEPA target inℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\. Instead, the latter uses the same EMA local\-target construction as the JEPA\-only baseline\. The hybrid comparison is therefore objective\-faithful: itsℒMC\\mathcal\{L\}\_\{\\mathrm\{MC\}\}remains the original JEPA latent\-prediction signal\.
The transition matrixAAis shared between the HMM and JEPA objectives\. Hence the same latent dynamics are trained simultaneously by observation\-sequence evidence and latent predictive alignment\. The complete implemented objective is
ℒhybrid=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{hybrid\}\}=\{\}1\.0ℒHMM\+1\.0ℒMC\+0\.5ℒfilter\\displaystyle 1\.0\\,\\mathcal\{L\}\_\{\\mathrm\{HMM\}\}\+1\.0\\,\\mathcal\{L\}\_\{\\mathrm\{MC\}\}\+0\.5\\,\\mathcal\{L\}\_\{\\mathrm\{filter\}\}\+0\.30ℒocc\+0\.02ℒent\.\\displaystyle\+0\.30\\,\\mathcal\{L\}\_\{\\mathrm\{occ\}\}\+0\.02\\,\\mathcal\{L\}\_\{\\mathrm\{ent\}\}\.The model is trained for600600full\-data updates with Adam learning rate
EMA coefficient
and gradient clipping at norm1010\.
#### Which parameters are trained by each objective?
For clarity,[Table11](https://arxiv.org/html/2608.13621#A6.T11)summarizes the effective parameter flow\. Both the exact filtering target and EMA JEPA target are stop\-gradient quantities\.
Table 11:Effective parameter updates in the revised Experiment 4 implementation\. The target encoder receives no gradient and is updated only by EMA\.This separation clarifies the interpretation of the comparison\. In HMM\+filter training, the generative parameters are learned from sequence evidence and the context encoder amortizes the resulting filter\. In hybrid training, the transition additionally receives the JEPA latent\-prediction signal, while the emission and initial\-state parameters remain trained through sequence evidence\.
#### Forward algorithm and numerical stabilization\.
All HMM sequence likelihoods are evaluated exactly in log space\. Let
ℓt\(k\)=logpψ\(Xt∣Zt=k\)\.\\ell\_\{t\}\(k\)=\\log p\_\{\\psi\}\(X\_\{t\}\\mid Z\_\{t\}=k\)\.The forward recursion is initialized as
α1\(k\)=logπk\+ℓ1\(k\)\\alpha\_\{1\}\(k\)=\\log\\pi\_\{k\}\+\\ell\_\{1\}\(k\)and updated by
αt\(j\)=ℓt\(j\)\+logsumexpi\[αt−1\(i\)\+logAij\]\.\\alpha\_\{t\}\(j\)=\\ell\_\{t\}\(j\)\+\\operatorname\{logsumexp\}\_\{i\}\\left\[\\alpha\_\{t\-1\}\(i\)\+\\log A\_\{ij\}\\right\]\.The sequence log\-likelihood is
logp\(X1:T\)=logsumexpkαT\(k\)\.\\log p\(X\_\{1:T\}\)=\\operatorname\{logsumexp\}\_\{k\}\\alpha\_\{T\}\(k\)\.
Learned HMM log variances are clamped to
before evaluating Gaussian emissions\. A numerical floor
ε=10−8\\varepsilon=10^\{\-8\}is used when taking logarithms or normalizing probabilities\.
#### Evaluation protocol\.
The common metrics and Hungarian alignment follow[SectionF\.1](https://arxiv.org/html/2608.13621#A6.SS1)\. For every regime, the alignment is obtained from training\-set hard assignments and held fixed during test evaluation\. We report test ARI and NMI, transition errorℰA\\mathcal\{E\}\_\{A\}, observation\-sequence NLL per time step, filtering KL, and true\-state prediction NLL at
h∈\{1,2,4,8\}\.h\\in\\\{1,2,4,8\\\}\.
The three groups of metrics have distinct interpretations\. ARI and NMI evaluate representation recovery; transition error and true\-state NLL evaluate learned latent dynamics; sequence NLL evaluates the complete transition–emission model\. Filtering KL is treated separately because it is an explicitly optimized quantity for two of the three regimes\.
#### Paired sequence\-evidence comparison\.
All three regimes within a seed are evaluated on exactly the same test realization\. The sequence\-evidence figure therefore uses a paired difference\. For regimerrand seedss, we compute
ΔNLLr,s=NLLr,s−NLLHMM\+filter,s\\Delta\\mathrm\{NLL\}\_\{r,s\}=\\mathrm\{NLL\}\_\{r,s\}\-\\mathrm\{NLL\}\_\{\\mathrm\{HMM\+filter\},s\}before averaging across seeds\. This removes variability caused by different sampled test sequences and makes the objective\-induced difference easier to see\.
Table 12:Paired observation\-sequence NLL difference relative to HMM\+filter training\. Differences are computed within each seed before aggregation\. Lower is better\.Thus, the hybrid retains only a very small sequence\-evidence gap relative to HMM\-style training, whereas the latent\-only JEPA model remains clearly separated\. We use this paired comparison descriptively rather than as a formal hypothesis test\.
#### Multi\-horizon prediction\.
For each regime we propagate the current context distribution through powers of its learned transition matrix,
q^t\+h=qtAh,\\widehat\{q\}\_\{t\+h\}=q\_\{t\}A^\{h\},align the result to ground\-truth state order, and evaluate true\-state NLL\. The complete results are shown in[Table13](https://arxiv.org/html/2608.13621#A6.T13)\.
Table 13:True\-state multi\-horizon prediction NLL in Experiment 4\. Values are mean±\\pmstandard deviation over five seeds\. Lower is better\.Both HMM\-style and hybrid training improve over JEPA\-only at every evaluated horizon\. The differences are largest at shorter horizons and narrow byh=8h=8, where repeated application of the transition matrix increasingly mixes the predictive state distribution\.
#### Filtering agreement\.
Filtering KL is defined in[SectionF\.1](https://arxiv.org/html/2608.13621#A6.SS1)\. The exact reference is the filtering distribution implied by the probabilistic model associated with each regime\. For HMM\+filter and hybrid training, this is the exact filter of their jointly trained Gaussian HMM\. For JEPA\-only, it is the filter obtained after fitting the post\-hoc Gaussian observation model to the learned JEPA states\.
The resulting values are
0\.000060±0\.0000070\.000060\\pm 0\.000007for HMM\+filter,
0\.004355±0\.0010150\.004355\\pm 0\.001015for the hybrid, and
0\.025460±0\.0064230\.025460\\pm 0\.006423for JEPA\-only\.
These quantities do not compare every model with one common external filtering oracle\. Moreover, filtering KL is explicitly optimized for HMM\+filter and hybrid training, so for those regimes it is a*training\-aligned role diagnostic*\. For JEPA\-only it is instead a post\-hoc diagnostic of how closely the learned context representation happens to agree with the filtering distribution induced by its fitted observation model\.
#### Seed\-wise consistency\.
The aggregate improvement of hybrid training over JEPA\-only is not produced by one favorable seed\. For each of the five paired runs, hybrid training improves over JEPA\-only in ARI, NMI, observation\-sequence NLL, transition error, filtering KL, and true\-state prediction NLL at every evaluated horizonh∈\{1,2,4,8\}h\\in\\\{1,2,4,8\\\}\. We report this pattern descriptively and do not infer formal statistical significance from five seeds\.
#### Supplementary filtering diagnostic\.
[Figure13](https://arxiv.org/html/2608.13621#A6.F13)reports filtering agreement separately from the two generative\-model metrics emphasized in the main text\.
Figure 13:Filtering agreement across the three Experiment 4 training regimes\. The quantity isKL\(qtexact∥qtenc\)\\mathrm\{KL\}\(q\_\{t\}^\{\\mathrm\{exact\}\}\\\|q\_\{t\}^\{\\mathrm\{enc\}\}\)averaged over test time points\. HMM\+filter and hybrid training explicitly optimize filtering alignment, so their values should be interpreted as training\-aligned role diagnostics\. JEPA\-only is evaluated post hoc using the filtering distribution induced by its fitted observation model\.
### F\.6Reproducibility and computation
The Python code which implements these 4 experiments can be found at this Github repo:https://github\.com/YongchaoHuang/HMM\-JEPA\.
Five fixed seeds
\{0,1,2,3,4\}\.\\\{0,1,2,3,4\\\}\.were used\. The implementation uses single\-precision PyTorch tensors,
torch\.float32,\\texttt\{torch\.float32\},and automatically selects CUDA when available, otherwise falling back to CPU\.
#### Experiments 1–3\.
The principal paper\-mode settings are
Ntrain=400,Ntest=160,T=80,dhidden=48,B=64\.N\_\{\\mathrm\{train\}\}=400,\\qquad N\_\{\\mathrm\{test\}\}=160,\\qquad T=80,\\qquad d\_\{\\mathrm\{hidden\}\}=48,\\qquad B=64\.The main optimization budgets are
180MCJEPA epochs,100MCJEPA warm\-start steps,350Gaussian\-HMM steps,180\\ \\text\{MCJEPA epochs\},\\qquad 100\\ \\text\{MCJEPA warm\-start steps\},\\qquad 350\\ \\text\{Gaussian\-HMM steps\},and
2200updates per nonzero Experiment 3 compression stage\.2200\\ \\text\{updates per nonzero Experiment~3 compression stage\}\.Experiment 3 uses five independent perturbation seeds for the learned information\-bottleneck continuation and five independently generated long sequences for the residual diagnostic\.
#### Experiment 4\.
The standalone Experiment 4 script uses
Ntrain=400,Ntest=160,T=80,dhidden=48\.N\_\{\\mathrm\{train\}\}=400,\\qquad N\_\{\\mathrm\{test\}\}=160,\\qquad T=80,\\qquad d\_\{\\mathrm\{hidden\}\}=48\.
The JEPA\-only baseline uses180180epochs with batch size6464, learning rate3×10−33\\times 10^\{\-3\},100100K\-means warm\-start updates, and EMA coefficient0\.9950\.995\. The HMM\+filter and hybrid regimes each use600600full\-data joint updates with Adam learning rate
Their common objective coefficients are
λseq=1\.0,λlatent=1\.0,λfilter=0\.5,\\lambda\_\{\\mathrm\{seq\}\}=1\.0,\\qquad\\lambda\_\{\\mathrm\{latent\}\}=1\.0,\\qquad\\lambda\_\{\\mathrm\{filter\}\}=0\.5,whereλlatent\\lambda\_\{\\mathrm\{latent\}\}applies only to the hybrid, together with
λocc=0\.30,λent=0\.02\.\\lambda\_\{\\mathrm\{occ\}\}=0\.30,\\qquad\\lambda\_\{\\mathrm\{ent\}\}=0\.02\.The hybrid EMA coefficient is
The JEPA\-only model is initialized with top\-level seedss, while the independently trained HMM\+filter and hybrid regimes use deterministic seed offsets associated withssso that each run remains reproducible while avoiding accidental reuse of identical parameter initialization streams\.
Because the JEPA\-only objective is naturally optimized with sequence mini\-batches whereas the differentiable HMM sequence objective is evaluated on the full training collection in the HMM\+filter and hybrid implementations, Experiment 4 should be interpreted as an*objective\-behavior diagnostic*, not as a compute\-matched optimization\-efficiency benchmark\. The latent\-state family, data realization, context\-encoder family, state\-use regularization, and evaluation protocol are controlled across regimes\.
#### Randomness and numerical reproducibility\.
Randomness is seeded for Python’srandommodule, NumPy, PyTorch, and all available CUDA devices\. The data generators use fixed seed offsets so that training data, test data, collapse diagnostics, Experiment 2 filtering datasets, Experiment 3 continuation runs, residual\-diagnostic sequences, and Experiment 4 model initializations can be reproduced independently from the top\-level seed\.
We do not enforce PyTorch deterministic\-algorithm mode, so exact bitwise reproducibility across different CUDA libraries or hardware is not guaranteed\. The scripts do not record the specific accelerator model or wall\-clock runtime, and package versions are not pinned in the experimental source; we therefore do not report hardware\-specific timing claims\. The implementation depends on NumPy, Pandas, PyTorch, scikit\-learn, SciPy, and Matplotlib\.
## Disclaimer
This work was developed with assistance from ChatGPT\([12](https://arxiv.org/html/2608.13621#bib.bib15)\)in idea development, technical formulation, writing, experimental design, and coding\. The central idea, i\.e\. the correspondence between probabilistic temporal JEPA and hidden Markov models, was originally and independently proposed by the author, while ChatGPT contributed to its subsequent development\. The presentation of this work, e\.g\. appearance of experimental results, is therefore different from previous work\. The author estimates the overall contributions split as approximately 60%:40% between the author and ChatGPT\. At the time of writing, the author does not expect an AI system to independently discover this research direction and refine it without substantial and careful human input, guidance, examination, correction and refinement\. The work therefore reflects a hybrid mode of human–AI research collaboration, in which the human researcher provides the originating insight, direction, judgement, and verification, while the AI assists with elaboration and execution\. All mathematical statements, technical claims, experimental procedures, results, and contents in main texts were manually reviewed and verified by the author, who takes full responsibility for the final work\. Nevertheless, errors or inaccuracies may remain, and readers are encouraged to interpret the claims and results with appropriate caution\.Similar Articles
The Annotated JEPA
A step-by-step annotated implementation and explanation of Joint Embedding Predictive Architectures (JEPA) for self-supervised learning, covering I-JEPA, V-JEPA, and LeJEPA.
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.
One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
This paper introduces Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture that lets every robot in a swarm predict the same future collective state from local observations and limited messages. The method shows label-efficient improvements in prediction error and inter-robot agreement, plus planning-relevant value estimation.
The 90-year-old idea behind JEPA models: Canonical Correlation Analysis
This blog post explains the connection between JEPA (Joint Embedding Predictive Architecture) models and Canonical Correlation Analysis (CCA), a statistical method from 1936, arguing that CCA is the conceptual precursor to JEPA and that the idea of maximizing correlation in embedding space dates back to Hotelling.
I built Micro-JEPA: A lightweight JEPA (Joint Embedding Predictive Architecture) in Python
Micro-JEPA is a lightweight Python implementation of the Joint Embedding Predictive Architecture (JEPA), enabling an agent to learn environment representations, predict future states in latent space, and plan actions to avoid obstacles.