DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
Summary
DeformSmith is a framework for generating interactive, physically credible deformable assets for robot manipulation from text or images, using physics-guided hierarchical generation to improve quality and plausibility.
View Cached Full Text
Cached at: 09/21/26, 03:18 AM
Paper page - DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
Source: https://huggingface.co/papers/2609.18620 Published on Sep 17
·
Submitted byhttps://huggingface.co/Canlee
Can Lion Sep 21
Abstract
Creatingdeformableassetsforrobotmanipulationrequiresjointlyspecifyingtheirgeometry,appearance,andphysicalproperties.Thisisespeciallychallengingfordeformableobjects,sincetextandimagesprovidelimitedevidenceabouthowtheydeformandrespondtocontact,yettheseresponsesdirectlyaffecttheirsuitabilityforinteraction.Automatedgenerationthereforeneedstoresolvecoupledphysicalrequirementsanduseinteractionevidencetoguideconstructionandrefinement.WepresentDeformSmith,aframeworkthatenablesautomatedgenerationofinteractive,physicallycredibledeformableassetsfromtextorasingleimage.Throughhierarchicalagenticconstructionandasharedphysics-groundedharness,itprogressivelybuilds,tests,andrefinesgeometry,physicalmodels,materialbehavior,androbotinteractionuntiltheresultingassetisreadyforsimulationandmanipulation.Robotinteractionclosesthegenerationloopthroughmanipulationfeedbackandreplayableinteractiondata.ResultsshowthatDeformSmithgeneratesassetswithbettervisualqualityandphysicalplausibilitythanstate-of-the-artbaselines,includingPhysGen3D,PhysGM,andPhysX-Omni,whilesupportingthesynthesisofdataforroboticmanipulationofdeformableobjects.Projectpage:https://can-lee.github.io/deformsmith-web/
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2609\.18620
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.18620 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.18620 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.18620 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
PhysisForcing is a training framework that enhances embodied video generation for robotic manipulation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based architecture, achieving notable improvements on benchmarks.
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
DeVI introduces a framework that turns text-conditioned synthetic videos into physically plausible dexterous robot control via a hybrid 3D-2D tracking reward, enabling zero-shot generalization to unseen objects.
PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World
PhysForge is a two-stage framework that generates interactive 3D assets with grounded physics and kinematic parameters, addressing the bottleneck of static geometry in virtual worlds.
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
EgoPhys introduces a framework to construct deformable physical digital twins from egocentric RGB video using generalizable priors and a compact codebook, enabling zero-shot generalization to unseen objects without per-spring optimization. The system is demonstrated on a real robot, showing that egocentric human play video can serve as internal world representation for deformable-object planning.
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Deform360 is a large-scale visuotactile dataset with 198 objects and over 215 hours of observations for studying deformable object dynamics, enabling comparison between 2D video and 3D particle world models for robotic manipulation.