μ_0: A Scalable 3D Interaction-Trace World Model
Summary
μ_0 is a scalable world model that predicts smooth 3D trajectories for interaction points, enabling embodiment-agnostic robot learning without action labels by using a TraceExtract system for supervision.
View Cached Full Text
Cached at: 06/15/26, 09:05 AM
Paper page - μ_0: A Scalable 3D Interaction-Trace World Model
Source: https://huggingface.co/papers/2606.13769
Abstract
A scalable world model called μ₀ uses 3D traces to predict smooth trajectories for key interaction points, enabling embodiment-agnostic robot learning without action labels.
World modelsthat capture how actions induce physical change enable scalable robot learning without reliance on embodiment-specific action labels. Pixel-space video models provide broad visual priors but expend model capacity on dense appearance reconstruction, while direct action models require embodiment-specific labels that hinder scalability. We present μ_0, a scalable world model based on3D traces. Rather than predicting dense pixels or directly modeling actions, μ_0 forecasts smooth 3D trajectories for salient interaction points such as objects, tools, hands, and contact regions, yielding a compact, embodiment-agnostic motion interface. To enable training from diverse video sources, ourTraceExtractsystem automatically extracts 3D supervision by selecting keypoints, constructing globally aligned traces, and associating motion segments with hierarchical language captions. ThisTraceExtractsupervision pretrains μ_0 by combining a pretrainedvision-language backbonewith amodular trace expert, which represents each query viaB-spline control pointsand predicts future traces. Experiments show that μ_0 outperforms baselines in both 2D and 3D trace prediction, including trace prediction models and tokenized VLM methods. Because μ_0 is frozen and reusable, it can be paired with action experts for downstream robot embodiments. Despiteaction-free pretraining, the resultingtrace-conditioned policiesachieve performance competitive withVLA modelspretrained with action supervision, such as π_0. These results establish3D tracesas a scalable and transferable representation for cross-embodiment manipulation.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2606\.13769
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.13769 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.13769 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.13769 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation
τ_0-WM is a unified video-action world model for robotic manipulation that integrates policy learning, video prediction, and action evaluation using a shared video diffusion backbone. It shows superior performance on challenging long-horizon and fine-grained tasks.
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
Masked Visual Actions for Unified World Modeling
Introduces Masked Visual Actions, a pixel-space control interface that expresses actions as partially revealed trajectories, enabling a single model to act as forward dynamics model, recover robot behavior, and support model-based planning and inverse modeling with only 15 hours of training data.
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from text, images, and videos through specialized modules for panorama generation, trajectory planning, and scene composition, achieving state-of-the-art performance among open-source approaches.
PhiZero: A World Model Built Around Physical Language
PhiZero is a physical world model that learns a compact discrete representation called 'physical language' from videos and uses it to reason about world state transitions before rendering future videos, improving physical coherence in generation and understanding tasks.