Tag
The paper introduces AC-MTM, a contrastive inverse dynamics method to prevent encoder collapse in JEPA world models, achieving improved performance on multi-object tasks without Gaussian constraints.
This paper introduces decision-metric alignment diagnostics and action-conditioned objectives to improve latent world models for model-predictive control. The proposed DA-LeWM method enhances convergence and success rates in planning tasks.
This paper proposes 'Mirror Learning', a framework for imitation learning from third-person observation that uses a fine-tuned video diffusion model for perspective transformation and an inverse dynamics model to synthesize pseudo first-person expert data, showing that this mirror data alone can train effective policies and improve behavior cloning.
RynnWorld-4D is a generative world model that co-produces future RGB, depth, and optical flow from a single RGB-D image and language instruction using a unified diffusion process, enabling efficient robotic manipulation through inverse dynamics policy learning. It achieves state-of-the-art on real-world bimanual manipulation tasks.
The paper introduces ACID, a method that enhances decision-time planning in world models by incorporating cycle action consistency via an inverse dynamics model, ensuring predicted transitions are realizable and improving planning efficiency across multiple tasks with reduced computation.
Task-Agnostic Pretraining (TAP) decomposes VLA training into self-supervised motor skill learning from unlabeled interaction data, then lightweight language grounding, achieving strong performance with minimal expert demonstrations. It matches or outperforms models trained on millions of expert trajectories while being robust to real-world perturbations.
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
This paper proposes a method to bridge the simulation-to-real-world gap in robotics by learning a deep inverse dynamics model that maps desired next states (from simulation) to appropriate real-world actions. The approach is evaluated against baselines like output error control and Gaussian dynamics adaptation.