Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
Summary
PRISM is a real-to-sim-to-real framework that amplifies a few real human-object interaction videos into hundreds of diverse counterfactual videos via V2V generation, then reconstructs physically plausible motions to train a generalizable humanoid loco-manipulation policy deployed on a real robot without real-world fine-tuning.
View Cached Full Text
Cached at: 10/01/26, 04:22 AM
Paper page - Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
Source: https://huggingface.co/papers/2609.38172
Abstract
Teachinghumanoidsloco-manipulationskills,suchascarryingdiverseobjects,viavisualimitationisapromisingpathtowardgeneralistrobots.However,collectingdiverse,high-qualityinteractionvideos,suchasclipsthatclearlyshowaperson’sfullbodyandunoccludedinteractionswithobjects,posesapracticalbarriertoscalingthisapproach.WeproposePRISM,areal-to-sim-to-realframeworkthatovercomesthislimitationbyamplifyingahandfulofrealvideosintoalarge,diversetrainingset.PRISMfirstgenerateshundredsofdiverse“counterfactual“human-objectinteractionvideosviavideo-to-video(V2V)generationfromafewexemplarrealvideos.Ourcontact-anchoredreal-to-simpipelinethenreconstructsbothhumanandobjectmotions,retargetingthisimperfectvideodataintophysicallyplausibletrajectories.Theintra-classvariabilityacrossthesecounterfactualvideosletsustrainasinglepolicythatgeneralizestounseenobjectswithineachcategory.Wedemonstratethefullpipelinebydeployingthispolicyonarealrobotwithoutanyreal-worldfine-tuning.Usingonlyonboarddepthobservations,ourhumanoidpicksup,carries,anddropsobjects,includingboxes,barrels,bins,andballs,acrossnovelinstances,sizes,andinitialconfigurations.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.38172
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38172 in a model README.md to link it from this page.
Datasets citing this paper1
#### Amazon-FAR/far-prism-data Updatedabout 3 hours ago • 5 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38172 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
GRAIL generates diverse humanoid manipulation and locomotion data using 3D assets and video foundation models, enabling effective sim-to-real transfer for humanoid robot control with high real-world success rates.
OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
OASIS is a simulation-data-driven framework for humanoid loco-manipulation that uses 3D generative models and hierarchical visuomotor policies. It achieves better zero-shot performance than real-robot training by leveraging domain randomization in simulation.
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
This paper introduces LUCID, a hierarchical model-based reinforcement learning framework for long-horizon humanoid loco-manipulation. It learns reusable latent skills and a macro-dynamics world model, enabling high-level planning via imagined rollouts and improving success rates in simulated multi-object rearrangement tasks.
ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
ReImagine introduces an image-first approach to controllable high-quality human video generation, combining SMPL-X motion guidance with video diffusion models to decouple appearance from temporal consistency.
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
CoInteract introduces an end-to-end Diffusion Transformer framework that jointly models RGB appearance and HOI geometry to generate physically-plausible human-object interaction videos with stable hands/faces and zero inference overhead.