OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
Summary
OASIS is a simulation-data-driven framework for humanoid loco-manipulation that uses 3D generative models and hierarchical visuomotor policies. It achieves better zero-shot performance than real-robot training by leveraging domain randomization in simulation.
View Cached Full Text
Cached at: 06/09/26, 08:43 AM
Paper page - OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
Source: https://huggingface.co/papers/2606.08548
Abstract
A simulation-data-driven framework for humanoid loco-manipulation that uses 3D generative models to create realistic assets and hierarchical visuomotor policies trained on simulated data achieves better zero-shot performance than real-robot training.
Recent progress in robot manipulation has been largely driven by learning from large-scale demonstrations. For humanoid robot loco-manipulation tasks, however, existing data sources force an unsatisfying tradeoff between trajectory quality and scalability. Real-worldteleoperationprovides the highest-quality trajectories but requires dedicated physical space and time-consuming scene resets. Simulation offers an alternative way out of this dilemma: it can produce clean, embodiment-aligned data at scale without any physical hardware. In this paper, we propose OASIS, asimulation-data-driven frameworkforhumanoid loco-manipulation. OASIS automatically reconstructs realistic object assets from real-world images using a3D generative model. Based on these assets, trajectories are first collected throughteleoperationin simulation, and then augmented under diversedomain randomizations in a post-processing stage. With the resulting simulation data, we further design ahierarchical visuomotor policyforhumanoid loco-manipulation. Extensive experiments on the real humanoid robot show that, under zero-shot deployment, the policy trained on our simulation data achieves higher success rates on most tasks than that trained on real-robotteleoperationdata, owing largely to the broad lighting and environmental variations covered by our simulation rendering, which real-robot data fails to capture. The project page is available at https://oasis-humanoid.github.io/.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2606\.08548
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.08548 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.08548 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.08548 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
PRISM is a real-to-sim-to-real framework that amplifies a few real human-object interaction videos into hundreds of diverse counterfactual videos via V2V generation, then reconstructs physically plausible motions to train a generalizable humanoid loco-manipulation policy deployed on a real robot without real-world fine-tuning.
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
GRAIL generates diverse humanoid manipulation and locomotion data using 3D assets and video foundation models, enabling effective sim-to-real transfer for humanoid robot control with high real-world success rates.
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
This paper introduces LUCID, a hierarchical model-based reinforcement learning framework for long-horizon humanoid loco-manipulation. It learns reusable latent skills and a macro-dynamics world model, enabling high-level planning via imagined rollouts and improving success rates in simulated multi-object rearrangement tasks.
@omarrayyann: Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-w…
FetchMan is a vision-based humanoid policy trained entirely in simulation, capable of zero-shot transfer to diverse real-world scenes and objects for loco-manipulation tasks.
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Open-AoE is an open, community-oriented egocentric manipulation dataset and toolchain that spans from smartphone capture to model training, providing approximately 2,000 hours of manipulation video with annotations and downstream tools for embodied learning.