Tag
Contrastive World Models propose a new approach for learning latent dynamics without pixel reconstruction, using a contrastive objective to improve robustness and efficiency in visually complex environments for model-based reinforcement learning.
Flow-JEPA introduces a conditional flow matching approach to JEPA world models, improving robustness and accuracy in predicting future latent states under noisy conditions.
XP-JEPA introduces cross-predictive physical grounding to improve latent dynamics for better forecastable control in world models without requiring physical inputs at test time.
This paper proposes a minimal 'advantage-style' action channel for latent world models that cancels action-independent distractor variation by subtracting the mean effect over actions, improving controllability without auxiliary losses or reconstruction.
This paper proposes a neural ODE-based regularization method that enforces latent embeddings in reinforcement learning agents to follow consistent ODE flows, aligning representation learning with environment dynamics and yielding performance gains on Atari and gridworld benchmarks.
Introduces Latent Dynamics Reasoning (LDR), a video world model that integrates kinematic dynamics in a structured latent space, enabling extrapolation of learned dynamics far beyond training distributions while using far fewer parameters and running much faster than video diffusion baselines.
Introduces Quantum-Structured World Models (QSWMs), a quantum-inspired framework for predictive world modeling with structured latent states, and evaluates them on elementary cellular automata against classical baselines.
SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.
This paper introduces Latent Lie-Poisson Neural Networks (LLPNNs), a structure-preserving framework for learning Lie-Poisson dynamics directly from observable data, using geometric methods and Magnus-based Lie-group updates. It demonstrates strong accuracy and robustness on rigid body, underwater vehicle, and optimal control examples.
This paper empirically investigates how multi-horizon latent consistency affects transition geometry in world models, using an expansion proxy on Moving-MNIST, Pendulum, CartPole, and KTH Actions. It finds that soft consistency can push passive video dynamics toward contraction but not action-conditioned domains.
This paper presents a goal-agnostic control framework for partial differential equations using a joint-embedding predictive architecture (JEPA) with a lightweight 2D ViT encoder and action-conditioned latent dynamics, showing that using a learned physical observable probe outperforms raw latent distance for control tasks.
This paper models Chain-of-Thought reasoning as a switching dynamical system, showing that reasoning fine-tuning globally reorganizes latent policy states, leading to improved multi-step reasoning. The proposed framework combines time-aware contrastive learning with discrete regime discovery, and experiments demonstrate that fine-tuned models exhibit richer latent-policy organization with functional specialization.
LIDAR-AD proposes a decoder-free latent-interaction world model for autonomous driving that uses redundancy-reduced latent alignment and residual action updates to improve risk-aware state abstraction and long-horizon dynamics prediction, outperforming baseline world models in simulated and real-world scenarios.
FLARE is a forced latent autoencoder that discovers compact response coordinates and sparse input-dependent latent dynamics from high-dimensional observations of forced physical systems, enabling long-horizon forecasting under unseen inputs.
Valdi combines end-to-end online training with latent diffusion dynamics for fast, uncertain dynamics prediction in model predictive control for reinforcement learning, showing preliminary results on the CarRacing environment.
This paper studies when conservation laws can be certified in learned latent world models, proposing bounded horizons that guarantee how long rollouts stay on physical invariant level sets using measurable model defects.
This paper introduces attention-free latent memory and dynamic re-encoding to improve long-horizon predictions in Koopman autoencoders, reducing error accumulation on benchmark dynamical systems.
LaWAM enables efficient robot control by predicting compact latent visual subgoals instead of expensive video generation, achieving state-of-the-art success rates with up to 24x lower latency than pixel-space world action models.
This paper argues that large language models struggle with causal reasoning and long-horizon planning due to a mismatch between sequence prediction and reasoning over latent environment dynamics, and introduces the Latent Dynamics Inference perspective along with the Flux environment to study these limitations.
GPLD introduces a gradient-penalized latent dynamics regularizer for DreamerV3 to enforce local smoothness in transition learning, improving sample efficiency on continuous control tasks, especially complex locomotion.