VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction

Hugging Face Daily Papers Papers

Summary

VideoPhysEdit is a training-free pipeline for physical counterfactual video editing that uses rigid-body physical scene reconstruction to simulate edits and generate accurate downstream motions and interactions.

Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored. We formulate this problem as physical counterfactual video editing (PCVE), which aims to generate a counterfactual video depicting the resulting motion and interactions given a source video, a physical edit, and its execution frame. PCVE is challenging because it requires understanding scene physics and inferring the downstream motion and interactions induced by a physical intervention, while paired factual and counterfactual data and dedicated evaluation metrics are lacking. We introduce VideoPhysEdit, a new training-free pipeline for PCVE in rigid-body scenes. It makes physical reasoning explicit through a novel physical scene reconstruction method that recovers a scene reproducing the observed motion and interactions under simulation, enabling the pipeline to apply physical edits as interventions and use the resulting trajectories to guide counterfactual video generation. We further construct PCVE-RigidBench, a synthetic benchmark with paired source and counterfactual target videos and physical ground truth, and introduce the Physical Edit Score. VideoPhysEdit achieves substantially higher physical edit accuracy than open-source methods and commercial models while maintaining competitive visual fidelity. Its Physical Edit Score is 0.376, the only positive score among the compared methods. Qualitative comparisons on real videos further show that VideoPhysEdit applies to real-world scenes and better depicts the downstream motion and interactions induced by the edits than the compared methods. Code: https://github.com/Hammour-steak/VideoPhysEdit
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:22 AM

Paper page - VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction

Source: https://huggingface.co/papers/2609.35134

Abstract

Videoeditinghasadvancedsubstantiallyinrecentyears,withmethodsincreasinglyaccountingforthevisualconsequencesofedits,suchaschangestoshadowsandocclusions.However,thephysicalconsequencesofedits,includingchangestosubsequentmotionandinteractions,remainlessexplored.Weformulatethisproblemasphysicalcounterfactualvideoediting(PCVE),whichaimstogenerateacounterfactualvideodepictingtheresultingmotionandinteractionsgivenasourcevideo,aphysicaledit,anditsexecutionframe.PCVEischallengingbecauseitrequiresunderstandingscenephysicsandinferringthedownstreammotionandinteractionsinducedbyaphysicalintervention,whilepairedfactualandcounterfactualdataanddedicatedevaluationmetricsarelacking.WeintroduceVideoPhysEdit,anewtraining-freepipelineforPCVEinrigid-bodyscenes.Itmakesphysicalreasoningexplicitthroughanovelphysicalscenereconstructionmethodthatrecoversascenereproducingtheobservedmotionandinteractionsundersimulation,enablingthepipelinetoapplyphysicaleditsasinterventionsandusetheresultingtrajectoriestoguidecounterfactualvideogeneration.WefurtherconstructPCVE-RigidBench,asyntheticbenchmarkwithpairedsourceandcounterfactualtargetvideosandphysicalgroundtruth,andintroducethePhysicalEditScore.VideoPhysEditachievessubstantiallyhigherphysicaleditaccuracythanopen-sourcemethodsandcommercialmodelswhilemaintainingcompetitivevisualfidelity.ItsPhysicalEditScoreis0.376,theonlypositivescoreamongthecomparedmethods.QualitativecomparisonsonrealvideosfurthershowthatVideoPhysEditappliestoreal-worldscenesandbetterdepictsthedownstreammotionandinteractionsinducedbytheeditsthanthecomparedmethods.Code:https://github.com/Hammour-steak/VideoPhysEdit

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.35134

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35134 in a model README.md to link it from this page.

Datasets citing this paper1

#### ccmoony/PCVE-RigidBench Updatedabout 3 hours ago • 430 • 5

Spaces citing this paper1

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video

Hugging Face Daily Papers

EgoPhys introduces a framework to construct deformable physical digital twins from egocentric RGB video using generalizable priors and a compact codebook, enabling zero-shot generalization to unseen objects without per-spring optimization. The system is demonstrated on a real robot, showing that egocentric human play video can serve as internal world representation for deformable-object planning.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Hugging Face Daily Papers

PhysisForcing is a training framework that enhances embodied video generation for robotic manipulation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based architecture, achieving notable improvements on benchmarks.