Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
Summary
Introduces Discrete-WAM, a unified discrete latent vision-action world policy that enables compositional causal reasoning and counterfactual reasoning in autonomous driving through aligned discrete tokens and a shared discrete diffusion framework.
View Cached Full Text
Cached at: 06/05/26, 06:07 AM
Paper page - Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
Source: https://huggingface.co/papers/2606.05645 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
Discrete-WAM introduces a unified discrete latent vision-action world policy that enables compositional causal reasoning and counterfactual reasoning in autonomous driving through aligned discrete tokens and a shared discrete diffusion framework.
Autonomous drivingrequires reasoning about how ego actions shape the evolution of the surrounding world. However, most end-to-end methods rely on direct state-to-action mappings, capturing correlations without explicitly modeling action-conditioned dynamics. Conversely, continuous-latentworld modelsoften lack compositional structure forcausal reasoningacross counterfactual futures. We introduce Discrete-WAM, a unified latent vision-action world policy that represents future visual states and ego actions as aligneddiscrete tokens, enabling compositionalcausal reasoningacross alternative futures. Built upon this unified discrete alignment, Discrete-WAM establishes a shareddiscrete diffusion frameworkwith unified generative tasks, jointly formulating world modeling, world-action policy, and hierarchical decision-enabled policy, supportingcompositional generalizationacross diverse driving scenarios. Experiments on large-scale autonomous-driving benchmarks show that Discrete-WAM achieves competitive performance while supporting controllable generation andcounterfactual reasoning, offering a principled path toward more reliable decision-making.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2606\.05645
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.05645 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.05645 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.05645 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
Light-WAM is a lightweight world action model for efficient robot manipulation that uses a compact video backbone and downsampled latent space for future-video supervision, achieving high performance with low inference latency.
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
SimWAM is a simple yet effective World Action Model for end-to-end autonomous driving that uses video generation purely as a training signal, achieving state-of-the-art 91.5 PDMS on NAVSIM while reducing inference latency.
DWM: Separating World Effects from Actions in Latent World Models
Introduces DWM, a framework that decomposes latent world model transitions into action-driven and action-invariant (world effect) components, improving planning success on benchmarks with persistent world effects.
RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
RepWAM introduces a world action modeling approach using representation visual-action tokenizers, aiming to learn unified visual and action representations for planning and control.
The DAWN of World-Action Interactive Models
This paper introduces DAWN, a latent generative baseline for World-Action Interactive Models (WAIMs) that jointly models scene evolution and action generation through recursive refinement, achieving strong long-horizon planning in autonomous driving scenarios.