DeltaWAM: Delta World Action Models for Bimanual Manipulation

Hugging Face Daily Papers Papers

Summary

DeltaWAM introduces delta-based world-action models for bimanual manipulation, enhancing efficiency and performance by predicting visual changes and actions. It demonstrates improved success rates and reduced computational overhead.

World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing. We further develop Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing. On RoboTwin, DeltaWAM with SDM improves average success over Fast-WAM from 81.3% to 85.4% in the clean setting and from 75.8% to 83.9% under visual randomization. The three architectures reduce training FLOPs by 17.78-23.77%, while SDM reduces one-step inference latency and FLOPs by 36.57% and 31.55%, respectively; real-world evaluations further show the highest overall success rate and normalized progress among the evaluated policies. Code: https://github.com/AIGeeksGroup/DeltaWAM. Website: https://aigeeksgroup.github.io/DeltaWAM.
Original Article
View Cached Full Text

Cached at: 09/25/26, 07:44 AM

Paper page - DeltaWAM: Delta World Action Models for Bimanual Manipulation

Source: https://huggingface.co/papers/2609.28811

Abstract

World-actionmodels(WAMs)transfervisualandmotionpriorsfrompretrainedvideogeneratorstorobotcontrolbyjointlymodelingvisualdynamicsandactions.ExistingWAMs,however,predictdensefutureframesduringtraining,repeatedlymodelinglargelyunchangedcontentandcouplingaction-conditioneddynamicstonuisanceappearancevariations.Atinference,processingeachcompleteobservationwiththeheavyvideoexpertbottlenecksfew-stepactiongeneration.Accordingly,weproposeDeltaWAM,whichjointlypredictsvisualdeltasandactionsusingdense-anchor,sparse-delta,andactionstreams,withthreearchitecturesthatdifferinrepresentationandcomputationsharing.WefurtherdevelopStreamingDeltaMemory(SDM),whichupdatescachedanchorcontextwithcompactobserveddeltas,reducingheavyvideo-expertprocessing.OnRoboTwin,DeltaWAMwithSDMimprovesaveragesuccessoverFast-WAMfrom81.3%to85.4%inthecleansettingandfrom75.8%to83.9%undervisualrandomization.ThethreearchitecturesreducetrainingFLOPsby17.78-23.77%,whileSDMreducesone-stepinferencelatencyandFLOPsby36.57%and31.55%,respectively;real-worldevaluationsfurthershowthehighestoverallsuccessrateandnormalizedprogressamongtheevaluatedpolicies.Code:https://github.com/AIGeeksGroup/DeltaWAM.Website:https://aigeeksgroup.github.io/DeltaWAM.

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.28811

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.28811 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.28811 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.28811 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

Hugging Face Daily Papers

AHA-WAM is an asynchronous world-action model that uses dual Diffusion Transformers to decouple world prediction from action execution, achieving efficient long-horizon planning and real-time control. It achieves state-of-the-art performance on robotic manipulation tasks with up to 92.8% success on RoboTwin and 78.3% on real-world tasks, while reaching 24.17 Hz closed-loop control.

BadWAM: When World-Action Models Dream Right but Act Wrong

Hugging Face Daily Papers

BadWAM introduces a framework for adversarial attacks on World-Action Models (WAMs), breaking the alignment between imagination and action via small visual perturbations. The attacks significantly reduce task success rates, exposing a vulnerability in this class of models.