DeltaWAM: Delta World Action Models for Bimanual Manipulation
Summary
DeltaWAM introduces delta-based world-action models for bimanual manipulation, enhancing efficiency and performance by predicting visual changes and actions. It demonstrates improved success rates and reduced computational overhead.
View Cached Full Text
Cached at: 09/25/26, 07:44 AM
Paper page - DeltaWAM: Delta World Action Models for Bimanual Manipulation
Source: https://huggingface.co/papers/2609.28811
Abstract
World-actionmodels(WAMs)transfervisualandmotionpriorsfrompretrainedvideogeneratorstorobotcontrolbyjointlymodelingvisualdynamicsandactions.ExistingWAMs,however,predictdensefutureframesduringtraining,repeatedlymodelinglargelyunchangedcontentandcouplingaction-conditioneddynamicstonuisanceappearancevariations.Atinference,processingeachcompleteobservationwiththeheavyvideoexpertbottlenecksfew-stepactiongeneration.Accordingly,weproposeDeltaWAM,whichjointlypredictsvisualdeltasandactionsusingdense-anchor,sparse-delta,andactionstreams,withthreearchitecturesthatdifferinrepresentationandcomputationsharing.WefurtherdevelopStreamingDeltaMemory(SDM),whichupdatescachedanchorcontextwithcompactobserveddeltas,reducingheavyvideo-expertprocessing.OnRoboTwin,DeltaWAMwithSDMimprovesaveragesuccessoverFast-WAMfrom81.3%to85.4%inthecleansettingandfrom75.8%to83.9%undervisualrandomization.ThethreearchitecturesreducetrainingFLOPsby17.78-23.77%,whileSDMreducesone-stepinferencelatencyandFLOPsby36.57%and31.55%,respectively;real-worldevaluationsfurthershowthehighestoverallsuccessrateandnormalizedprogressamongtheevaluatedpolicies.Code:https://github.com/AIGeeksGroup/DeltaWAM.Website:https://aigeeksgroup.github.io/DeltaWAM.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.28811
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.28811 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.28811 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.28811 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
ST-WAM proposes a semantic-temporal world action model that uses DINOv3 features as a shared semantic representation to improve robot manipulation robustness under visual distribution shifts, achieving 98.7% on LIBERO and 92.8% on RoboTwin 2.0, with significant gains over Fast-WAM in zero-shot settings.
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
The paper introduces DECOWAM, a decoupled whole-body world-action model for legged mobile manipulation that improves video and action prediction performance over existing models like FastWAM through dedicated conditional interfaces and a new dataset.
AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
AHA-WAM is an asynchronous world-action model that uses dual Diffusion Transformers to decouple world prediction from action execution, achieving efficient long-horizon planning and real-time control. It achieves state-of-the-art performance on robotic manipulation tasks with up to 92.8% success on RoboTwin and 78.3% on real-world tasks, while reaching 24.17 Hz closed-loop control.
Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
Light-WAM is a lightweight world action model for efficient robot manipulation that uses a compact video backbone and downsampled latent space for future-video supervision, achieving high performance with low inference latency.
BadWAM: When World-Action Models Dream Right but Act Wrong
BadWAM introduces a framework for adversarial attacks on World-Action Models (WAMs), breaking the alignment between imagination and action via small visual perturbations. The attacks significantly reduce task success rates, exposing a vulnerability in this class of models.