Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Hugging Face Daily Papers Papers

Summary

PRISM is a real-to-sim-to-real framework that amplifies a few real human-object interaction videos into hundreds of diverse counterfactual videos via V2V generation, then reconstructs physically plausible motions to train a generalizable humanoid loco-manipulation policy deployed on a real robot without real-world fine-tuning.

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:22 AM

Paper page - Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Source: https://huggingface.co/papers/2609.38172

Abstract

Teachinghumanoidsloco-manipulationskills,suchascarryingdiverseobjects,viavisualimitationisapromisingpathtowardgeneralistrobots.However,collectingdiverse,high-qualityinteractionvideos,suchasclipsthatclearlyshowaperson’sfullbodyandunoccludedinteractionswithobjects,posesapracticalbarriertoscalingthisapproach.WeproposePRISM,areal-to-sim-to-realframeworkthatovercomesthislimitationbyamplifyingahandfulofrealvideosintoalarge,diversetrainingset.PRISMfirstgenerateshundredsofdiverse“counterfactual“human-objectinteractionvideosviavideo-to-video(V2V)generationfromafewexemplarrealvideos.Ourcontact-anchoredreal-to-simpipelinethenreconstructsbothhumanandobjectmotions,retargetingthisimperfectvideodataintophysicallyplausibletrajectories.Theintra-classvariabilityacrossthesecounterfactualvideosletsustrainasinglepolicythatgeneralizestounseenobjectswithineachcategory.Wedemonstratethefullpipelinebydeployingthispolicyonarealrobotwithoutanyreal-worldfine-tuning.Usingonlyonboarddepthobservations,ourhumanoidpicksup,carries,anddropsobjects,includingboxes,barrels,bins,andballs,acrossnovelinstances,sizes,andinitialconfigurations.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.38172

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38172 in a model README.md to link it from this page.

Datasets citing this paper1

#### Amazon-FAR/far-prism-data Updatedabout 3 hours ago • 5 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38172 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles