Tag
A paper introduces TC-LeWM, applying regularization losses like SIGReg and VISReg to temporal latent residuals to enable multi-task learning in LeWM, significantly improving success rates on the LIBERO robot arm benchmark.
This paper introduces a physics-aware autoencoder-based latent-space framework for reduced-order forward modeling and variational parameter estimation in parametric dynamical systems, demonstrated on computational fluid dynamics benchmarks. The method enables differentiable surrogate-based inverse modeling and shows improved calibration robustness under realistic noisy or partial observations.
Introduces MaSRead, a content-addressed reading mechanism for replicated latent stores where agents share KV cache fragments, enabling later queries to reliably retrieve cached reasoning via opaque keyed tag sets and hard attention masks.
This paper introduces the Latent Critic, a lightweight LoRA adapter that detects hallucinated agent actions in real time by restructuring the transformer's residual stream into localized natural-language feedback, achieving 0.966 AUROC and enabling self-correction.
This paper proposes Dreamer-SAC, a model-based reinforcement learning framework that integrates a recurrent state-space world model with soft actor-critic in latent space for sample-efficient autonomous driving. It outperforms DreamerV3, SAC, and PPO baselines while requiring fewer real environment interactions.
This paper investigates how molecular generative models internally organize molecular identity in their latent spaces, revealing piecewise-constant regions and coarse-to-fine boundaries across three architectures.
This arXiv paper investigates whether simple linear transformations can translate representations across nine heterogeneous text embedding models, finding that shared structure and transferability depend jointly on architecture, training objective, pooling, and data distribution, challenging the notion of universal latent compatibility.
This paper introduces HiLP, a hierarchical representation training method that adds multi-scale self-predictive learning to transformer pretraining, aiming to reduce compounding error and improve long-horizon reasoning and speculative decoding efficiency.
This paper shows that matching a marginal Gaussian prior in factorized generative models does not prevent conditional style leakage, where style latents carry class information. Multiple remedies are explored, but the authors conclude that marginal statistics alone cannot certify class-invariance.
This paper introduces ODEWorld, a continuous-time latent world model using Physical-Time Flow (PT-Flow) that learns a latent velocity field parameterized by an ordinary differential equation, enabling arbitrary temporal resolution, backward prediction, and improved planning for video generation and robotic control.
LatentFlow is a visual analytics system for analyzing latent spaces in molecular graph neural networks, helping scientists understand how clusters of molecular embeddings evolve across layers and model states.
The article explores the idea that human language is a designed latent space that constrains reasoning into pattern matching and symbol manipulation, and argues that LLMs cannot create new languages to solve problems beyond existing language.
Induction Labs introduces imagination models, a new foundation model architecture that learns from internet-scale video. Their first model, Photon-1, learns to use a computer by watching 18 years of screen recordings without action labels, achieving better results at 30× lower pretraining cost than Gemini 3.1 Flash.
MUX proposes a method for lossless continuous reasoning by distilling discrete reasoning steps into multiplexed latent tokens that encode a superposition of subwords, achieving higher bandwidth and enabling parallel exploration in language model reasoning tasks.
Discusses a Latent Space podcast episode where Anjney Midha explains why AI labs with unlimited GPUs still fail, drawing on his experience at amppublic and a16z.
Kevin Kelly believes that AI's greatest value lies in exploring possibilities within latent space, rather than merely answering questions. The latent space becomes a new creative medium, emphasizing the importance of judgment and cross-domain connections.
This paper proposes modeling the CLIP latent space using Mixtures of von Mises–Fisher distributions on the unit hypersphere, capturing its directional and multimodal structure better than Gaussian assumptions. The model improves long-tailed and out-of-distribution detection and provides a semantic decomposition of CLIP embeddings.
SinAE introduces a single-architecture flow-matching autoencoder using vanilla Transformers that achieves near-lossless reconstruction across molecules, crystals, and proteins, enabling cross-domain training and strong generative performance on standard benchmarks.
This paper introduces VideoRAE, a representation autoencoder that leverages frozen video foundation models to create compact, reconstruction-capable, and generation-friendly video latents. It achieves state-of-the-art results on UCF-101 with faster convergence than competing autoencoders.
Kevin Kelly argues that the latent space within AI models represents a compressed form of all human knowledge, and that this latent space will become a new medium for creativity, enabling novel artistic and scientific exploration.