Tag
BLARM is a feed-forward method for animating 3D meshes from monocular video by blending latent rigid motion primitives, enabling temporally coherent animations without explicit rigs.
This paper proposes Intention Distillation (INDI) to distill behavior intent into the action decoder of Vision-Language-Action models, improving performance on benchmarks like SimplerEnv-Bridge and real-world tasks.
The paper introduces 'Kilobyte models', an extreme compression paradigm where neural network weights are regenerated from a seed and a quantized latent vector, reducing storage to kilobytes by leveraging Mapping Networks and seeded block-wise bases.
AURORA-LM introduces a continuous-latent diffusion language model that separates decodable text representation from distribution modeling, achieving strong performance on OpenWebText and XSum while scaling to 1B parameters.
MUGEN introduces a unified motion-language framework that avoids discrete codebooks and iterative decoding, using a single adaptive-length autoencoder with continuous latent slots and one-shot generation to achieve efficient, high-quality text-to-motion and motion-to-text performance across HumanML3D and SnapMoGen benchmarks.
LATO.2 is a factorized flow matching framework for 3D mesh generation that decomposes the process into a vertex flow followed by a connectivity flow conditioned on the vertices, achieving state-of-the-art geometric fidelity and connectivity quality.
Identifies two coupled causes of posterior collapse in VAEs and introduces λ-VAE, a modification to the reparameterization step that equalizes variance across latent dimensions, reducing collapse and improving information capacity.
AdaJEPA introduces an adaptive latent world model that continuously updates during test-time via closed-loop model predictive control, significantly improving planning success under distribution shift.
This paper demonstrates that diffusion language models (DLMs) internally represent denoising progress as a latent timestep signal in their residual streams, which can be probed and steered to modulate model confidence and entropy.
BAAI released the Orca paper describing a multimodal latent world model that learns a unified world representation first, then decodes into text, images, or actions using frozen backbones and tiny decoders, with weights coming soon.
This paper introduces the concept of the audit gap between behavioral safety and representation-level robustness in LLMs, proposing an intervention-based evaluation framework and the Latent Vulnerability Score (LVS) to measure hidden vulnerabilities.
Introduces Thoughts-as-Planning, a framework that models chain-of-thought optimization as sequential decision-making using latent world models and reinforcement learning, outperforming existing methods in efficiency and generalization.
Latte introduces a framework that represents personalization as forecasting a peer-anchored relative preference state using latent trajectories, injecting a soft token into a frozen LLM to achieve personalized generation. It outperforms existing personalization methods on Amazon Reviews 2023 and MemoryCD datasets.
ATLAS presents a visual reasoning framework that combines agentic operations and latent representations using functional tokens, enabling efficient training via next-token prediction and reinforcement learning while avoiding intermediate image generation.
This paper establishes nonparametric identifiability guarantees for extracting task-relevant representations from generalist models, proving that task structure is identifiable across time steps and latent representations are identifiable within each step under sparsity regularization.