VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction
Summary
VideoPhysEdit is a training-free pipeline for physical counterfactual video editing that uses rigid-body physical scene reconstruction to simulate edits and generate accurate downstream motions and interactions.
View Cached Full Text
Cached at: 09/30/26, 04:22 AM
Paper page - VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction
Source: https://huggingface.co/papers/2609.35134
Abstract
Videoeditinghasadvancedsubstantiallyinrecentyears,withmethodsincreasinglyaccountingforthevisualconsequencesofedits,suchaschangestoshadowsandocclusions.However,thephysicalconsequencesofedits,includingchangestosubsequentmotionandinteractions,remainlessexplored.Weformulatethisproblemasphysicalcounterfactualvideoediting(PCVE),whichaimstogenerateacounterfactualvideodepictingtheresultingmotionandinteractionsgivenasourcevideo,aphysicaledit,anditsexecutionframe.PCVEischallengingbecauseitrequiresunderstandingscenephysicsandinferringthedownstreammotionandinteractionsinducedbyaphysicalintervention,whilepairedfactualandcounterfactualdataanddedicatedevaluationmetricsarelacking.WeintroduceVideoPhysEdit,anewtraining-freepipelineforPCVEinrigid-bodyscenes.Itmakesphysicalreasoningexplicitthroughanovelphysicalscenereconstructionmethodthatrecoversascenereproducingtheobservedmotionandinteractionsundersimulation,enablingthepipelinetoapplyphysicaleditsasinterventionsandusetheresultingtrajectoriestoguidecounterfactualvideogeneration.WefurtherconstructPCVE-RigidBench,asyntheticbenchmarkwithpairedsourceandcounterfactualtargetvideosandphysicalgroundtruth,andintroducethePhysicalEditScore.VideoPhysEditachievessubstantiallyhigherphysicaleditaccuracythanopen-sourcemethodsandcommercialmodelswhilemaintainingcompetitivevisualfidelity.ItsPhysicalEditScoreis0.376,theonlypositivescoreamongthecomparedmethods.QualitativecomparisonsonrealvideosfurthershowthatVideoPhysEditappliestoreal-worldscenesandbetterdepictsthedownstreammotionandinteractionsinducedbytheeditsthanthecomparedmethods.Code:https://github.com/Hammour-steak/VideoPhysEdit
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.35134
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35134 in a model README.md to link it from this page.
Datasets citing this paper1
#### ccmoony/PCVE-RigidBench Updatedabout 3 hours ago • 430 • 5
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
EgoPhys introduces a framework to construct deformable physical digital twins from egocentric RGB video using generalizable priors and a compact codebook, enabling zero-shot generalization to unseen objects without per-spring optimization. The system is demonstrated on a real robot, showing that egocentric human play video can serve as internal world representation for deformable-object planning.
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
PhyMotion proposes a physics-grounded reward system that evaluates kinematic plausibility, contact consistency, and dynamic feasibility of human motion in generated videos, achieving stronger correlation with human judgment and improving motion realism in RL-based post-training.
@ZimingLiu11: Physics is editable in world models, but only up to a critical depth. The model "makes up its mind" about physics at so…
The paper introduces causal writability in video models, showing that correct physical motion remains available but becomes uneditable after a critical depth, which impacts how training corrects errors.
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
PhysisForcing is a training framework that enhances embodied video generation for robotic manipulation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based architecture, achieving notable improvements on benchmarks.
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
EditVid is a unified training-free video editing framework that supports instruction-guided and subject-guided edits using sparse causal memory, token injection, and soft latent blending, achieving high fidelity and outperforming baseline methods in benchmarks.