IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
Summary
IMPACT is a scalable framework for training interaction-aware world models by using cross-attention as an internal interaction map to reweight denoising supervision, improving performance without external representations or inference-time changes.
View Cached Full Text
Cached at: 09/02/26, 05:57 AM
# Attention Is the Interaction Map forScalable Interaction-Aware World Model Training
Source: [https://arxiv.org/html/2609.00161](https://arxiv.org/html/2609.00161)
\\aaai@corrmultitrueRongze Tang1,2, Jianjie Fang3, Zhaolu Wang3, Ziyou Wang3, Xvyuan Liu3, Haisheng Su4, Xin Zhang4, Wei Wu4, Chen Gao2,3\\corresponding, Yong Li3, Zhibo Chen1,2\\corresponding
###### Abstract
World models have made remarkable progress in action\-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions\. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics\. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability\. We instead revisit the training objective and identify asupervision\-allocation mismatchunder the globally averaged mean squared error \(MSE\) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic\-object regions critical to interaction generation disproportionately under\-supervised\. Motivated by this observation, we introduceIMPACT, a scalableInteraction\-awareModel training framework withPrior\-guidedAttentionCalibration andTargeting\. IMPACT uses cross\-attention associated with manipulated\-object tokens as an internal spatiotemporal prior for action\-conditioned changes\. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference\-time modifications\. Extensive experiments on robot\-arm and human\-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE\-trained baselines, improving interaction fidelity, physical plausibility, and visual quality\.
1University of Science and Technology of China2Zhongguancun Academy
3Tsinghua University4Manifold AI
chgao96@gmail\.com, chenzhibo@ustc\.edu\.cnattr /Border \[0 0 0\] user /Subtype /Link /A << /S /URI /URI \(https://embodiedcity\.github\.io/IMPACT/\) \>\>Project Pageattr /Border \[0 0 0\] user /Subtype /Link /A << /S /URI /URI \(https://github\.com/EmbodiedCity/IMPACT\.code\) \>\>Code
Figure 1:Overview of IMPACT\. \(a\) Existing approaches rely on costly and unscalable external dense representations\. \(b\) Standard world\-model training with uniformly averaged MSE suffers from a supervision\-allocation mismatch, under\-supervising sparse interaction regions\. \(c\) IMPACT calibrates object\-conditioned cross\-attention with prediction errors into an interaction map for efficient and scalable interaction\-targeted training\.\\pdfdestname impact\.fig:overview xyz## Introduction
\\pdfdest
name impact\.sec:introduction xyz
World models have made remarkable progress in predicting future observations\([Zhang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib23);[Kim et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib33);[Fang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib34)\)and supporting embodied simulation\([Wang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib26);[Gao et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib27)\)\. Building on large\-scale pretrained video generation backbones\([Yang et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib5);[Kong et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib6);[Wan Team et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib7);[NVIDIA et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib8)\), they can simulate how an environment evolves under a commanded action, generating action responses in interactive environments\([AlayaWorld Team et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib21);[Hu et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib22);[Fang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib38)\)and robotic manipulation tasks\([Zhang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib23);[Bi et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib24);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib25)\)\. Their imagined futures provide scalable experience for long\-horizon planning\([Wang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib26);[Gao et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib27)\), policy learning\([Hu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib28);[Zhen et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib29);[Su et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib35)\), and policy evaluation\([Shang et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib30);[Zhu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib32)\)\. In this role, world models should go beyond visual coherence to simulate physically plausible interactions, yet still suffer from object deformation, discontinuous motion, weak action coupling and inconsistent contact\([Zhang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib25)\)\.
As shown in attr /Border \[0 0 0\] goto name impact\.fig:overviewFigure[1](https://arxiv.org/html/2609.00161#S0.F1)a, existing approaches typically address these failures by constraining generation with external representations of interaction dynamics\. Motion\-based methods use optical flow or point trajectories to describe object dynamics\([Gao et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib27);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib25)\)and geometry\-based methods introduce depth, surface normals, reconstructed scenes, or articulated hand meshes\([Zhen et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib29);[Kim et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib33)\)\. However, obtaining such spatiotemporally dense representations through auxiliary estimators or manual annotations is costly, and the resulting supervision is bounded by the accuracy of these external signals, limiting both training scalability and the achievable interaction quality\.
We instead revisit the standard training objective of video world models and identify a*supervision\-allocation mismatch*, as illustrated in attr /Border \[0 0 0\] goto name impact\.fig:overviewFigure[1](https://arxiv.org/html/2609.00161#S0.F1)b\. Inherited from general video generation, this globally averaged mean squared error \(MSE\) denoising objective uniformly weights all spatiotemporal positions\([Po et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib31);[Zhu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib32);[Kim et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib33)\)\. Under uniform weighting, each region contributes to optimization according to its spatial extent rather than its functional importance\. Prevalent static content thus dominates the training signal, leaving sparse dynamic\-object regions that carry action\-conditioned changes disproportionately under\-supervised\. As a result, models may reduce the global denoising loss and generate visually coherent videos while leaving interaction regions under\-optimized\.
Our key insight is that world models built on large\-scale pretrained video generation backbones already contain a spatiotemporal prior for interaction regions\. For manipulation instructions, cross\-attention aligns language tokens with spatiotemporal video representations, enabling the attention maps associated with manipulated\-object tokens to serve as a spatiotemporal prior for regions likely to undergo action\-conditioned changes\. The refined prior provides a natural basis for reweighting denoising supervision, allowing sparse interaction regions to contribute to gradient optimization according to their functional importance rather than their spatial extent, thereby improving interaction generation\. Because this prior is obtained from the model’s standard forward pass, it requires no external dense representations, scales readily with training data, and leaves inference unchanged\.
Based on this insight, we introduceIMPACT, a scalableInteraction\-awareModel training framework withPrior\-guidedAttentionCalibration andTargeting\. As shown in attr /Border \[0 0 0\] goto name impact\.fig:overviewFigure[1](https://arxiv.org/html/2609.00161#S0.F1)c, the core idea is to treat object\-conditioned cross\-attention as an interaction prior, turn it into an interaction map, and use this map to reweight denoising supervision toward interaction regions, thereby mitigating the supervision\-allocation mismatch and improving interaction generation\. IMPACT realizes this within a single training step through two complementary components\. In the forward pass,Attention Distribution Sampling\(ADS\) aggregates the cross\-attention of the manipulated\-object tokens into a proposal distribution, samples multiple candidate regions, and weights each by its detached local prediction error, calibrating the attention prior with the model’s current prediction difficulty to form a precise interaction map\. In the backward pass,Interaction\-Weighted Supervision\(IWS\) uses this map to strengthen denoising supervision on interaction regions while preserving the global objective, and routes gradients so that the interaction\-weighted objective updates the non\-cross\-attention DiT parameters while the attention\-producing parameters follow the global objective, preventing the prior from collapsing onto its own signal\.
We evaluate IMPACT in two settings: robotic\-arm manipulation on WorldArena\([Shang et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib30)\)and human\-hand manipulation on EgoDex\([Hoque et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib10)\)\. Across both settings, IMPACT surpasses standard uniformly\-supervised training \(MSE\) under the same backbone as well as various baselines, delivering higher interaction fidelity and physical plausibility across diverse embodied scenarios and action conditions\.
In summary:
- •We identify a supervision\-allocation mismatch in world model training and introduce IMPACT, a scalable training framework that leverages the object\-conditioned cross\-attention prior to reweight denoising supervision toward interaction regions, without requiring any external dense representations\.
- •We realize IMPACT through two complementary designs: ADS evaluates object\-conditioned attention proposals using detached local prediction errors and converts these errors into weights to form an interaction map, while IWS uses this map to target denoising supervision, preserve the global objective, and decouple region estimation from regional optimization\.
- •We demonstrate across robotic\-arm and human\-hand manipulation with different control signals and DiT backbones that IMPACT consistently outperforms uniform\-MSE training, improving both the physical consistency and visual quality of generated videos\.
## Related Work
\\pdfdest
name impact\.sec:related xyz
#### Interactive video generation and world models
Building on large\-scale pretrained video generation backbones, including CogVideoX, HunyuanVideo, Wan, and Cosmos\([Yang et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib5);[Kong et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib6);[Wan Team et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib7);[NVIDIA et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib8)\), recent world models inherit rich representations of visual appearance and temporal dynamics, enabling high\-fidelity prediction and coherent motion modeling, but these capabilities alone are insufficient to make them useful embodied simulators\. To bridge this gap, recent work introduces diverse control conditions that make world models interactive, including natural\-language instructions\([Xiang et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib43);[Zhang et al\. 2026a](https://arxiv.org/html/2609.00161#bib.bib23);[Zhao et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib36);[Zhao et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib37)\), game and camera controls\([Bruce et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib1);[He et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib44)\), robot action trajectories\([NVIDIA et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib8);[Zhu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib32)\), and articulated hand or body controls\([Wang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib45);[Gao et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib42)\)\. However, these methods primarily target overall visual quality and controllable generation, while providing limited constraints on interaction dynamics, resulting in physically implausible behavior within interaction regions\.
#### External representations priors
To address this limitation, recent methods introduce spatiotemporally dense representations such as optical flow, depth maps, or reconstructed 3D structure to explicitly constrain the generation process\. Motion\-based approaches exploit optical flow, point trajectories, or latent temporal discrepancies to emphasize dynamic regions\([Fang et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib16);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib25);[Wu et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib13)\)\. Geometry\-based methods introduce depth, cross\-view 3D structure, or projected robot kinematics as auxiliary targets or structured conditions\([Tian et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib14);[Yang et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib15);[Liu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib17)\)\. However, constructing these representations often requires external motion, depth, segmentation, and video\-understanding models, sometimes together with manually verified annotations\([Fang et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib16);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00161#bib.bib25);[Yan et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib18);[Luo et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib19)\)\. These preprocessing costs grow with the size and duration of the training corpus, thereby limiting training scalability\.
## Method
\\pdfdest
name impact\.sec:method xyz
attr /Border \[0 0 0\] goto name impact\.fig:frameworkFigure[2](https://arxiv.org/html/2609.00161#Sx3.F2)presents an overview ofIMPACT\. Given a training sample, IMPACT identifies the manipulated\-object tokens and uses their cross\-attention as a spatial prior\. Attention Distribution Sampling\(ADS\) samples candidate regions from this prior and weights them by detached local prediction errors to construct an interaction map, which Interaction\-Weighted Supervision\(IWS\) uses to target denoising supervision toward interaction\-relevant tokens\. Gradient routing optimizes the cross\-attention parameters with the original global objective and the remaining DiT parameters with the interaction\-weighted objective, while all additional operations are training\-only and incur no inference\-time overhead\.
Figure 2:Framework of IMPACT\. Object\-token grounding \(left\): a frozen Qwen2\.5\-0\.5B extracts the manipulated\-object phrase from the instruction and aligns it to the corresponding object tokens\. Training pipeline \(right\): within the DiT blocks, the cross\-attention of these object tokens forms a proposal distribution \(cross\-attention map\), from which𝐌1,…,𝐌8\\mathbf\{M\}\_\{1\},\\dots,\\mathbf\{M\}\_\{8\}candidate regions are sampled and calibrated by their detached local prediction errors into an interaction map\. The map targets denoising supervision toward interaction regions, and gradient routing optimizes the cross\-attention parameters with the global objectiveℒMSE\\mathcal\{L\}\_\{\\mathrm\{MSE\}\}and the remaining DiT parameters with the interaction\-weighted objectiveℒIWS\\mathcal\{L\}\_\{\\mathrm\{IWS\}\}\.\\pdfdestname impact\.fig:framework xyz### Preliminaries
\\pdfdest
name impact\.sec:preliminaries xyz
We build on a latent video diffusion transformer trained with flow matching\([Lipman et al\. 2023](https://arxiv.org/html/2609.00161#bib.bib4)\)\. Given a video𝐱0\\mathbf\{x\}\_\{0\}, a frozen variational autoencoder maps it into a latent representation𝐳0∈ℝC×T×H×W\\mathbf\{z\}\_\{0\}\\in\\mathbb\{R\}^\{C\\times T\\times H\\times W\}, whereTT,HH, andWWdenote the temporal and spatial dimensions\. The conditioning information is denoted by𝐜=\(𝐲,𝐱ref,𝐚\)\\mathbf\{c\}=\(\\mathbf\{y\},\\mathbf\{x\}^\{\\mathrm\{ref\}\},\\mathbf\{a\}\), where𝐲\\mathbf\{y\}is the language instruction,𝐱ref\\mathbf\{x\}^\{\\mathrm\{ref\}\}is the reference observation, and𝐚\\mathbf\{a\}represents the corresponding control signal, such as hand poses or robot trajectories\.
For a sampled noise levelσt∈\[0,1\]\\sigma\_\{t\}\\in\[0,1\]and Gaussian noiseϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\), the clean latent is interpolated with noise as
𝐳t=\(1−σt\)𝐳0\+σtϵ\.\\pdfdestnameimpact\.eq:flowinterpolationxyz\\mathbf\{z\}\_\{t\}=\(1\-\\sigma\_\{t\}\)\\mathbf\{z\}\_\{0\}\+\\sigma\_\{t\}\\boldsymbol\{\\epsilon\}\.\\pdfdest name\{impact\.eq:flow\_\{i\}nterpolation\}xyz\(1\)The diffusion transformervθv\_\{\\theta\}takes\(𝐳t,t,𝐜\)\(\\mathbf\{z\}\_\{t\},t,\\mathbf\{c\}\)as input and predicts the velocity targetϵ−𝐳0\\boldsymbol\{\\epsilon\}\-\\mathbf\{z\}\_\{0\}\. Let
Ω=\{1,…,T\}×\{1,…,H\}×\{1,…,W\}\\Omega=\\\{1,\\ldots,T\\\}\\times\\\{1,\\ldots,H\\\}\\times\\\{1,\\ldots,W\\\}\(2\)denote the set of spatiotemporal latent positions\. The prediction error at positionp∈Ωp\\in\\Omegais
ℓp=1C‖vθ\(𝐳t,t,𝐜\)p−\(ϵ−𝐳0\)p‖22\.\\pdfdestnameimpact\.eq:localerrorxyz\\ell\_\{p\}=\\frac\{1\}\{C\}\\left\\\|v\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t,\\mathbf\{c\}\)\_\{p\}\-\(\\boldsymbol\{\\epsilon\}\-\\mathbf\{z\}\_\{0\}\)\_\{p\}\\right\\\|\_\{2\}^\{2\}\.\\pdfdest name\{impact\.eq:local\_\{e\}rror\}xyz\(3\)Standard flow\-matching training uniformly averages this error over all spatiotemporal positions:
ℒglobal\(θ\)=𝔼𝐳0,ϵ,t\[1\|Ω\|∑p∈Ωℓp\]\.\\pdfdestnameimpact\.eq:globalobjectivexyz\\mathcal\{L\}\_\{\\mathrm\{global\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\},t\}\\left\[\\frac\{1\}\{\|\\Omega\|\}\\sum\_\{p\\in\\Omega\}\\ell\_\{p\}\\right\]\.\\pdfdest name\{impact\.eq:global\_\{o\}bjective\}xyz\(4\)Consequently, the contribution of a region to optimization is largely determined by its spatial extent\. Small interaction regions therefore receive no additional emphasis despite containing rapid motion, contact transitions, and action\-conditioned state changes\. IMPACT addresses this supervision\-allocation mismatch without changing the underlying flow\-matching formulation\.
### Object\-Token Grounding
\\pdfdest
name impact\.sec:object\_grounding xyz
The interaction region depends on the object being manipulated\. We therefore begin by grounding the manipulated object in the language instruction\. For an instruction𝐲\\mathbf\{y\}, a frozen Qwen2\.5\-0\.5B model\([Qwen et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib20)\)extracts the noun phrase that denotes the manipulated object\. For example, given the instruction “Right gripper stacks blue bowl on white tabletop,” the extracted phrase is “blue bowl\.”
Let
𝐞=\(𝐞1,…,𝐞N\)\\mathbf\{e\}=\(\\mathbf\{e\}\_\{1\},\\ldots,\\mathbf\{e\}\_\{N\}\)\(5\)denote the sequence of text embeddings produced by the tokenizer and text encoder\. We align the extracted object phrase with this token sequence and denote its token positions by
𝒪⊆\{1,…,N\}\.\\pdfdestnameimpact\.eq:objecttokensetxyz\\mathcal\{O\}\\subseteq\\\{1,\\ldots,N\\\}\.\\pdfdest name\{impact\.eq:object\_\{t\}oken\_\{s\}et\}xyz\(6\)When an object phrase is divided into multiple subword tokens, all matched positions are included in𝒪\\mathcal\{O\}\. This object\-token set provides the semantic anchor used by ADS to extract an object\-conditioned spatial distribution\. The grounding stage operates only on the instruction and therefore requires neither visual annotations nor external spatial estimators\.
### Forward: Attention Distribution Sampling
\\pdfdest
name impact\.sec:ads xyz
ADS converts the semantic object grounding into an interaction map through three operations: object\-conditioned attention aggregation, candidate\-region sampling, and prediction\-error\-based weighting\.
#### Object\-conditioned attention distribution\.
Within each DiT cross\-attention layer, visual queries attend to the text\-token sequence\. For blockll, attention headhh, visual positionpp, and text\-token positionjj, the cross\-attention probability is
Ap,j\(l,h\)=softmaxj\(𝐪p\(l,h\)𝐤j\(l,h\)⊤d\),\\pdfdestnameimpact\.eq:crossattentionxyzA\_\{p,j\}^\{\(l,h\)\}=\\operatorname\{softmax\}\_\{j\}\\left\(\\frac\{\\mathbf\{q\}\_\{p\}^\{\(l,h\)\}\\mathbf\{k\}\_\{j\}^\{\(l,h\)\\top\}\}\{\\sqrt\{d\}\}\\right\),\\pdfdest name\{impact\.eq:cross\_\{a\}ttention\}xyz\(7\)whereddis the key dimension\. We sum the probability mass assigned to the object\-token set𝒪\\mathcal\{O\}and average it over the selected blocksℬ\\mathcal\{B\}and attention headsℋ\\mathcal\{H\}:
Ap=1\|ℬ\|\|ℋ\|∑l∈ℬ∑h∈ℋ∑j∈𝒪Ap,j\(l,h\)\.\\pdfdestnameimpact\.eq:objectattentionxyzA\_\{p\}=\\frac\{1\}\{\|\\mathcal\{B\}\|\\,\|\\mathcal\{H\}\|\}\\sum\_\{l\\in\\mathcal\{B\}\}\\sum\_\{h\\in\\mathcal\{H\}\}\\sum\_\{j\\in\\mathcal\{O\}\}A\_\{p,j\}^\{\(l,h\)\}\.\\pdfdest name\{impact\.eq:object\_\{a\}ttention\}xyz\(8\)The resulting map𝐀∈\[0,1\]T×H×W\\mathbf\{A\}\\in\[0,1\]^\{T\\times H\\times W\}forms an object\-conditioned proposal distribution over the video latent\. Because it is obtained from the standard conditional forward pass, constructing𝐀\\mathbf\{A\}requires no additional visual model or spatial annotation\.
#### Candidate\-region sampling\.
Instead of using a single deterministic region, ADS samplesKKspatially coherent candidates around the attention distribution\. We first detach and temper the attention map:
𝐀~=clip\(sg\[𝐀\]κ,ε,1−ε\),\\pdfdestnameimpact\.eq:attentiontemperingxyz\\widetilde\{\\mathbf\{A\}\}=\\operatorname\{clip\}\\left\(\\operatorname\{sg\}\[\\mathbf\{A\}\]^\{\\,\\kappa\},\\varepsilon,1\-\\varepsilon\\right\),\\pdfdest name\{impact\.eq:attention\_\{t\}empering\}xyz\(9\)wheresg\[⋅\]\\operatorname\{sg\}\[\\cdot\]denotes stop\-gradient andκ\\kappacontrols the concentration of the proposal distribution\. We then transform it into logits:
𝐙=log𝐀~1−𝐀~\.\\pdfdestnameimpact\.eq:attentionlogitsxyz\\mathbf\{Z\}=\\log\\frac\{\\widetilde\{\\mathbf\{A\}\}\}\{1\-\\widetilde\{\\mathbf\{A\}\}\}\.\\pdfdest name\{impact\.eq:attention\_\{l\}ogits\}xyz\(10\)
For each candidatek∈\{1,…,K\}k\\in\\\{1,\\ldots,K\\\}, we sample Gaussian noise on a coarse spatiotemporal grid and trilinearly interpolate it to the resolution of𝐀\\mathbf\{A\}, producing a smooth perturbation field𝜼k\\boldsymbol\{\\eta\}\_\{k\}\. A soft candidate is generated as
𝐒k=sigmoid\(𝐙\+σ𝜼kτ\),\\pdfdestnameimpact\.eq:softcandidatexyz\\mathbf\{S\}\_\{k\}=\\operatorname\{sigmoid\}\\left\(\\frac\{\\mathbf\{Z\}\+\\sigma\\boldsymbol\{\\eta\}\_\{k\}\}\{\\tau\}\\right\),\\pdfdest name\{impact\.eq:soft\_\{c\}andidate\}xyz\(11\)whereσ\\sigmacontrols the perturbation magnitude andτ\\tauis the sampling temperature\. Sampling noise on a coarse grid produces coherent spatial variations rather than independent token\-wise perturbations\.
To prevent candidate scores from being dominated by differences in region size, every candidate is converted into a fixed\-area binary region\. Specifically, we retain the largestρ\|Ω\|\\rho\|\\Omega\|values of𝐒k\\mathbf\{S\}\_\{k\}:
Mk,p=𝕀\[Sk,p≥TopKThreshold\(𝐒k,ρ\|Ω\|\)\],\\pdfdestnameimpact\.eq:fixedareacandidatexyzM\_\{k,p\}=\\mathbb\{I\}\\left\[S\_\{k,p\}\\geq\\operatorname\{TopKThreshold\}\(\\mathbf\{S\}\_\{k\},\\rho\|\\Omega\|\)\\right\],\\pdfdest name\{impact\.eq:fixed\_\{a\}rea\_\{c\}andidate\}xyz\(12\)yielding𝐌k∈\{0,1\}T×H×W\\mathbf\{M\}\_\{k\}\\in\\\{0,1\\\}^\{T\\times H\\times W\}\. All candidates consequently cover the same number of latent positions and remain directly comparable\.
#### Prediction\-error\-based weighting\.
ADS evaluates every candidate using the mean local prediction error within that region:
dk=∑p∈ΩMk,psg\[ℓp\]∑p∈ΩMk,p\+ε\.\\pdfdestnameimpact\.eq:candidatescorexyzd\_\{k\}=\\frac\{\\sum\_\{p\\in\\Omega\}M\_\{k,p\}\\operatorname\{sg\}\[\\ell\_\{p\}\]\}\{\\sum\_\{p\\in\\Omega\}M\_\{k,p\}\+\\varepsilon\}\.\\pdfdest name\{impact\.eq:candidate\_\{s\}core\}xyz\(13\)A higher value ofdkd\_\{k\}indicates that the candidate covers content that is currently more difficult for the model to predict\. The candidate errors are converted into normalized weights:
αk=exp\(βdk\)∑j=1Kexp\(βdj\),\\pdfdestnameimpact\.eq:candidateweightxyz\\alpha\_\{k\}=\\frac\{\\exp\(\\beta d\_\{k\}\)\}\{\\sum\_\{j=1\}^\{K\}\\exp\(\\beta d\_\{j\}\)\},\\pdfdest name\{impact\.eq:candidate\_\{w\}eight\}xyz\(14\)whereβ\\betacontrols the concentration of the weighting distribution\. The candidate regions are then weighted to form the interaction map
𝐆=clip\(∑k=1Kαk𝐌k,0,1\)\.\\pdfdestnameimpact\.eq:interactionmapxyz\\mathbf\{G\}=\\operatorname\{clip\}\\left\(\\sum\_\{k=1\}^\{K\}\\alpha\_\{k\}\\mathbf\{M\}\_\{k\},0,1\\right\)\.\\pdfdest name\{impact\.eq:interaction\_\{m\}ap\}xyz\(15\)Through this process, the object\-conditioned attention determines where candidate regions are sampled, while the current local prediction error determines their relative contribution to𝐆\\mathbf\{G\}\. The entire ADS pipeline is computed without gradient tracking, preventing the model from directly modifying the proposal distribution or candidate weights to reduce the regional objective\.
In our implementation, we useK=8K=8, an attention tempering exponentκ=0\.65\\kappa=0\.65, perturbation scaleσ=0\.75\\sigma=0\.75, sampling temperatureτ=1\\tau=1, and a coarse noise grid of size5×8×125\\times 8\\times 12\. Each candidate retainsρ=0\.10\\rho=0\.10of the spatiotemporal latent positions, and candidate errors are converted into weights usingβ=0\.7\\beta=0\.7\.
### Backward: Interaction\-Weighted Supervision
\\pdfdest
name impact\.sec:iws xyz
IWS uses the interaction map𝐆\\mathbf\{G\}to increase the contribution of interaction\-relevant positions during denoising optimization\. For each positionp∈Ωp\\in\\Omega, we define
wp=1\+\(γ−1\)sg\[Gp\],\\pdfdestnameimpact\.eq:interactionweightxyzw\_\{p\}=1\+\(\\gamma\-1\)\\operatorname\{sg\}\[G\_\{p\}\],\\pdfdest name\{impact\.eq:interaction\_\{w\}eight\}xyz\(16\)whereγ≥1\\gamma\\geq 1controls the maximum regional emphasis\. BecauseGp∈\[0,1\]G\_\{p\}\\in\[0,1\], the resulting weight satisfieswp∈\[1,γ\]w\_\{p\}\\in\[1,\\gamma\]\. Positions outside the estimated interaction region retain their original unit weight, while positions with high interaction\-map values receive stronger supervision\.
The interaction\-weighted denoising objective is
ℒIWS\(θ\)=𝔼𝐳0,ϵ,t\[∑p∈Ωwpℓp∑p∈Ωwp\]\.\\pdfdestnameimpact\.eq:iwsobjectivexyz\\mathcal\{L\}\_\{\\mathrm\{IWS\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\mathbf\{z\}\_\{0\},\\boldsymbol\{\\epsilon\},t\}\\left\[\\frac\{\\sum\_\{p\\in\\Omega\}w\_\{p\}\\ell\_\{p\}\}\{\\sum\_\{p\\in\\Omega\}w\_\{p\}\}\\right\]\.\\pdfdest name\{impact\.eq:iws\_\{o\}bjective\}xyz\(17\)The normalization by the total weight prevents the loss scale from growing with the size or magnitude of the emphasized region\. Unlike a masked regional loss, attr /Border \[0 0 0\] goto name impact\.eq:iws\_objectiveEquation[17](https://arxiv.org/html/2609.00161#Sx3.E17)retains supervision over the complete spatiotemporal field and only changes its spatial allocation\. We useγ=5\\gamma=5in the main experiments\.
#### Gradient\-decoupled optimization\.
The interaction map is derived from cross\-attention, creating a dependency between region estimation and the objective guided by that region\. Although𝐆\\mathbf\{G\}is detached, allowingℒIWS\\mathcal\{L\}\_\{\\mathrm\{IWS\}\}to update the attention\-producing parameters could still cause the cross\-attention representation itself to adapt to the regional objective in subsequent training steps\. We therefore separate the DiT parameters into the cross\-attention parametersθca\\theta\_\{\\mathrm\{ca\}\}and all remaining parametersθrest\\theta\_\{\\mathrm\{rest\}\}\. Their gradients are routed according to
∇θcaℒIMPACT=∇θcaℒglobal,\\pdfdestnameimpact\.eq:crossattentionroutingxyz\\nabla\_\{\\theta\_\{\\mathrm\{ca\}\}\}\\mathcal\{L\}\_\{IMPACT\}=\\nabla\_\{\\theta\_\{\\mathrm\{ca\}\}\}\\mathcal\{L\}\_\{\\mathrm\{global\}\},\\pdfdest name\{impact\.eq:cross\_\{a\}ttention\_\{r\}outing\}xyz\(18\)and
∇θrestℒIMPACT=∇θrestℒIWS\.\\pdfdestnameimpact\.eq:remainingparameterroutingxyz\\nabla\_\{\\theta\_\{\\mathrm\{rest\}\}\}\\mathcal\{L\}\_\{IMPACT\}=\\nabla\_\{\\theta\_\{\\mathrm\{rest\}\}\}\\mathcal\{L\}\_\{\\mathrm\{IWS\}\}\.\\pdfdest name\{impact\.eq:remaining\_\{p\}arameter\_\{r\}outing\}xyz\(19\)In practice, this routing is implemented through two backward passes with parameter\-group gradient hooks\. The global backward pass retains gradients only forθca\\theta\_\{\\mathrm\{ca\}\}, while the IWS backward pass retains gradients only forθrest\\theta\_\{\\mathrm\{rest\}\}\. Consequently, the cross\-attention used to estimate interaction regions remains governed by the original uniformly supervised objective, whereas the remaining DiT parameters learn from the spatially reallocated supervision\. This separation prevents the regional objective from directly optimizing its own localization signal and decouples interaction\-region estimation from interaction\-focused optimization\.
#### Training and inference cost\.
IMPACT reuses the cross\-attention probabilities and token\-wise prediction errors already produced during standard world\-model training\. Its additional computation consists primarily of samplingKKcandidate masks, evaluating their masked mean errors, and constructing the interaction\-weighted objective\. No component of ADS or IWS is used at inference time, so the trained world model preserves the original architecture and inference procedure\.
## Experiments
\\pdfdest
name impact\.sec:experiments xyz
Table 1:Generation quality on WorldArena\. Models are grouped into general, embodied, and representation\-guided models, with Wan 2\.2\-based and Cosmos\-based variants listed separately\. We report the overall EWMScore and six aggregate dimensions, each on a\[0,100\]\[0,100\]scale where higher is better \(↑\\uparrow\)\. Boldface and underlining denote the best and second\-best results in each column\.\\pdfdestname impact\.tab:main xyzTable 2:Generation quality on EgoDex\. Boldface and underlining denote the best and second\-best results in each column\.\\pdfdestname impact\.tab:egodex xyz### Setup
\\pdfdest
name impact\.sec:setup xyz
#### Implementation details\.
We evaluate IMPACT in two manipulation settings to demonstrate the generality of our method across interaction types, applying it to the Wan2\.2 and Cosmos\-Predict 2\.5 backbone in both\. For robot\-arm manipulation, the model is trained on∼\\sim350K 17\-frame videos collected fromRoboTwin\([Chen et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib9)\), conditioned on 14\-DoF dual\-arm action trajectories, injected through an additional action encoder\. For human\-hand manipulation, we train on∼\\sim256K clips drawn fromEgoDex\([Hoque et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib10)\)across 118 tasks, using 81\-frame videos conditioned on hand\-pose videos temporally aligned with the RGB frames, injected through the shared VAE encoder without any additional encoder\. Both settings are trained at 720p and optimization is identical across the two settings: we use bf16 mixed precision with FSDP over 8 GPUs, a per\-device batch size of 1 with 4\-step gradient accumulation \(effective global batch size1×8×4=321\\times 8\\times 4=32\), a constant learning rate of2×10−52\\times 10^\{\-5\}after 100 warm\-up steps, and we train for one epoch\. IMPACT hyperparameters are fixed across both settings:K=8K=8,σ=0\.75\\sigma=0\.75, coarse grid5×8×125\\times 8\\times 12,ρ=0\.10\\rho=0\.10,κ=0\.65\\kappa=0\.65,β=0\.7\\beta=0\.7,τ=1\\tau=1,γ=5\\gamma=5\.
#### Benchmarks and metrics\.
For robot\-arm manipulation, we evaluate onWorldArena\([Shang et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib30)\), which scores dual\-arm manipulation along six dimensions: Visual Quality, Motion Quality, Content Consistency, Physics Adherence, 3D Accuracy, and Controllability, spanning 16 normalized metrics, and condenses overall generation quality into a singleEWMScore\(the mean of the 16 metrics\)\. We report EWMScore together with the six aggregate dimensions in attr /Border \[0 0 0\] goto name impact\.tab:mainTable[1](https://arxiv.org/html/2609.00161#Sx4.T1), and provide all 16 metrics in the technical appendix\. For human\-hand manipulation, we evaluate on theEgoDex\([Hoque et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib10)\)test set along two axes \(attr /Border \[0 0 0\] goto name impact\.tab:egodexTable[2](https://arxiv.org/html/2609.00161#Sx4.T2)\):*visual*metrics: FVD\([Unterthiner et al\. 2018](https://arxiv.org/html/2609.00161#bib.bib3)\)and FID\([Heusel et al\. 2017](https://arxiv.org/html/2609.00161#bib.bib2)\), covering temporal coherence and per\-frame appearance quality; and*hand\-interaction*metrics: CLIP\-Hand for the local appearance and semantics of the hand and nearby manipulated object, and Hand IoU for the coarse 2D position and scale of the generated hand\([Sun et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib11)\)\.
#### Baselines\.
For robot\-arm manipulation, all models follow the WorldArena evaluation settings\. Existing baselines comprise general world models: CogVideoX\([Yang et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib5)\), Wan 2\.6\([Wan Team et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib7)\), and Veo 3\.1; embodied world models: GigaWorld\-0, Genie Envisioner, Vidar, IRASim\([Zhu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib32)\), and CtrlWorld; and representation\-guided models: TesserAct\([Zhen et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib29)\), RoboMaster, and WoW\. We additionally report two backbone\-specific groups to evaluate IMPACT: Wan 2\.2\-based models include Wan 2\.2\([Wan Team et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib7)\), MSE\-trained Wan 2\.2\-AC, and Wan 2\.2\-AC with IMPACT; Cosmos\-based models include Cosmos\-Predict 2\.5 \(text\)\([NVIDIA et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib8)\), Cosmos\-Predict 2\.5 \(action\), and its IMPACT variant\. For human\-hand manipulation, baselines comprise general video world models: HunyuanVideo\-1\.5\([Wu et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib39)\)and Cosmos\-Predict 2\.5\([NVIDIA et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib8)\)and pose\-controlled models: MimicMotion\([Zhang et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib40)\), MagicPose\([Chang et al\. 2024](https://arxiv.org/html/2609.00161#bib.bib41)\), VACE\([Jiang et al\. 2025](https://arxiv.org/html/2609.00161#bib.bib12)\), and LOME\([Gao et al\. 2026](https://arxiv.org/html/2609.00161#bib.bib42)\)\.

Figure 3:Qualitative comparison on both interaction settings\. Left: robot\-arm manipulation on WorldArena; right: first\-person human\-hand manipulation on EgoDex\. For each setting we show generated frames for a representative instruction against various baselines\. In both settings, IMPACT follows the instruction more faithfully and renders the interaction\. Red boxes highlight artifacts and inconsistencies in the interaction regions of the baseline generations\.\\pdfdestname impact\.fig:qualitative xyz
Figure 4:Visualization of ADS\. On one robot\-arm training forward pass, the light\-to\-dark color bar encodes increasing candidate errors and thus the relative contribution of each candidate to the aggregate\.\\pdfdestname impact\.fig:ads\_calibration xyz
### Quantitative Analysis
\\pdfdest
name impact\.sec:quantitative\_analysis xyz
#### Robot\-arm manipulation\.
attr /Border \[0 0 0\] goto name impact\.tab:mainTable[1](https://arxiv.org/html/2609.00161#Sx4.T1)shows that IMPACT improves both backbone families\. On Cosmos\-Predict 2\.5 \(action\), it raises EWMScore from 55\.91 to 62\.53 \(\+6\.62 points, 11\.8%\), achieving the best overall result, the best Controllability \(71\.19\), and the second\-best Motion Quality \(68\.22\)\. Interaction Quality increases from 0\.5500 to 0\.6360 and Action Following from 0\.0133 to 0\.6260, although the remaining aggregate dimensions decline\. On Wan 2\.2\-AC, IMPACT raises EWMScore from 58\.65 to 62\.46 \(\+3\.81 points, 6\.5%\) and improves Visual Quality, Motion Quality, Physics Adherence, 3D Accuracy, and Controllability, attaining the best Physics Adherence \(55\.87\) and 3D Accuracy \(92\.56\) and the second\-best Visual Quality \(60\.60\)\. The best IMPACT result exceeds Wan 2\.6, CtrlWorld, and WoW by 0\.67, 2\.83, and 7\.65 points, respectively\.
#### Human\-hand manipulation\.
For human\-hand manipulation, attr /Border \[0 0 0\] goto name impact\.tab:egodexTable[2](https://arxiv.org/html/2609.00161#Sx4.T2)reports that IMPACT attains the best results on visual metrics across both general video world models and pose\-controlled models, cutting FVD from 366\.12 to 110\.94 and FID from 44\.71 to 5\.79 over the action\-conditioned Wan 2\.2\-AC on the same backbone\. It also leads on both hand\-interaction metrics, improving CLIP\-Hand from 0\.921 to 0\.952 and Hand IoU from 0\.693 to 0\.772 over Wan 2\.2\-AC on the same backbone\. These gains indicate stronger local interaction fidelity and hand localization than both the MSE\-trained counterpart and the pose\-controlled baselines\. Together with the robot\-arm results, these improvements demonstrate that IMPACT delivers consistent gains across different DiT backbones and control signals\.
### Qualitative analysis
\\pdfdest
name impact\.sec:Qualitative analysis xyz
#### Generation comparison\.
attr /Border \[0 0 0\] goto name impact\.fig:qualitativeFigure[4](https://arxiv.org/html/2609.00161#Sx4.F4)compares generations in robot\-arm and human\-hand manipulation\. In both cases, the baselines follow the instruction loosely and tend to blur or distort the contact region, whether the gripper–object contact for the robot arm or the hand–object contact for the human hand, and some drift away from the target\. In contrast, IMPACT produces the specified interaction with a sharper contact region and more coherent object dynamics, while keeping the surrounding scene stable\. This clearly demonstrates the effectiveness of IMPACT for interaction\-region generation\.
#### ADS calibration\.
attr /Border \[0 0 0\] goto name impact\.fig:ads\_calibrationFigure[4](https://arxiv.org/html/2609.00161#Sx4.F4)illustrates how ADS calibrates the attention prior in a representative training example\. The raw object\-conditioned cross\-attention map𝐀\\mathbf\{A\}is diffuse, spreading across the robot arm, workspace, and background\. ADS samplesK=8K\{=\}8candidate regions𝐌1,…,𝐌8\\mathbf\{M\}\_\{1\},\\dots,\\mathbf\{M\}\_\{8\}from this prior and evaluates them using detached local prediction errors \(attr /Border \[0 0 0\] goto name impact\.eq:candidate\_scoreEq\.[13](https://arxiv.org/html/2609.00161#Sx3.E13)\)\. Candidates covering the contact region receive higher weights than those dominated by the static background\. Their weighted aggregation \(attr /Border \[0 0 0\] goto name impact\.eq:interaction\_mapEq\.[15](https://arxiv.org/html/2609.00161#Sx3.E15)\) produces an interaction map𝐆\\mathbf\{G\}that is more concentrated around the interacting arm, gripper, and manipulated object\. This example illustrates how ADS refines a coarse attention prior into a more targeted supervision map for IWS\.
### Ablation Studies
\\pdfdest
name impact\.sec:ablation xyz
#### Component ablation: IWS and ADS\.
attr /Border \[0 0 0\] goto name impact\.tab:ads\_samplingTable[3](https://arxiv.org/html/2609.00161#Sx4.T3)shows the effect of separating the components on WorldArena and reveals their complementarity\. Starting from the AC backbone \(58\.65 EWMScore\), IWS provides the more direct gain \(\+2\.89 to 61\.54\), since it primarily addresses the supervision\-allocation mismatch\. On top of this, ADS further calibrates the cross\-attention prior and delivers an additional improvement \(\+0\.92 to 62\.46\) over weighting the raw, coarse prior directly\. Integrated within a single training step, the two components jointly strengthen interaction\-region generation, improving EWMScore by 3\.81 points overall\.
Table 3:Per\-component ablation on WorldArena\.\\pdfdestname impact\.tab:ads\_sampling xyz
## Conclusion
\\pdfdest
name impact\.sec:conclusion xyz
In this work, we introduced IMPACT, a scalable framework that addresses the supervision\-allocation mismatch by converting object\-conditioned cross\-attention into targeted denoising supervision\. ADS calibrates attention proposals with detached local prediction errors, while IWS reweights training with the resulting interaction map, requiring neither external spatial signals nor inference\-time changes\. Experiments on robot\-arm and human\-hand manipulation show consistent gains over uniform MSE training and strong baselines, demonstrating the effectiveness and scalability of IMPACT for interaction\-aware world model training\.
## References
- AlayaWorld Team, K\. Zhang, C\. Li, Y\. Zhan, Y\. Ge, Y\. Yin, J\. Tan, K\. He, L\. Fan, R\. Liu, X\. Xu, X\. Chu, Z\. Li, Z\. Lin, Z\. Wang, Z\. Meng, and Z\. GaoAlayaWorld: long\-horizon and playable video world generation\.External Links:2607\.06291,[Link](https://arxiv.org/abs/2607.06291)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Biet al\.\(2026\)H\. Bi, H\. Tan, S\. Xie, Z\. Wang, S\. Huang, H\. Liu, R\. Zhao, Y\. Feng, C\. Xiang, Y\. Rong, H\. Zhao, H\. Liu, Z\. Su, L\. Ma, H\. Su, and J\. ZhuMotus: a unified latent action world model\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 35101–35113\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2026/html/Bi_Motus_A_Unified_Latent_Action_World_Model_CVPR_2026_paper.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Bruceet al\.\(2024\)J\. Bruce, M\. Dennis, A\. Edwards, J\. Parker\-Holder, Y\. Shi, E\. Hughes, M\. Lai, A\. Mavalankar, R\. Steigerwald, C\. Apps, Y\. Aytar, S\. Bechtle, F\. Behbahani, S\. C\. Y\. Chan, N\. Heess, L\. Gonzalez, S\. Osindero, S\. Ozair, S\. Reed, J\. Zhang, K\. Zolna, J\. Clune, N\. d\. Freitas, T\. Rocktaschel, and D\. HafnerGenie: generative interactive environments\.arXiv preprint arXiv:2402\.15391\.Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Changet al\.\(2024\)D\. Chang, Y\. Shi, Q\. Gao, H\. Xu, J\. Fu, G\. Song, Q\. Yan, Y\. Zhu, X\. Yang, and M\. SoleymaniMagicPose: realistic human poses and facial expressions retargeting with identity\-aware diffusion\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 6263–6285\.External Links:[Link](https://proceedings.mlr.press/v235/chang24d.html)Cited by:[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Chenet al\.\(2025\)T\. Chen, Z\. Chen, B\. Chen, Z\. Cai, Y\. Liu, Z\. Li, Q\. Liang, X\. Lin, Y\. Ge, Z\. Gu, W\. Deng, Y\. Guo, T\. Nian, X\. Xie, Q\. Chen, K\. Su, T\. Xu, G\. Liu, M\. Hu, H\. Gao, K\. Wang, Z\. Liang, Y\. Qin, X\. Yang, P\. Luo, and Y\. MuRoboTwin 2\.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation\.arXiv preprint arXiv:2506\.18088\.Cited by:[Implementation details\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px1.p1.1)\.
- Fanget al\.\(2026a\)J\. Fang, Y\. Lei, Q\. Wan, Z\. Wang, Y\. Huang, Y\. Xu, B\. Zhao, W\. Zhang, C\. Gao, X\. Chen, and Y\. LiiWorld\-Bench: a benchmark for interactive world models with a unified action generation framework\.External Links:2605\.03941,[Link](https://arxiv.org/abs/2605.03941)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Fanget al\.\(2026b\)J\. Fang, Y\. Xu, Z\. Wang, C\. Gao, Y\. Huang, Z\. Wang, R\. Tang, M\. Jia, B\. Zhao, W\. Zhang, X\. Zhang, H\. Su, Y\. Shang, W\. Wu, X\. Chen, and Y\. LiWorldscape\-MoE: a unified mixture\-of\-experts world model for scalable heterogeneous action control\.External Links:2607\.03964,[Link](https://arxiv.org/abs/2607.03964)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Fanget al\.\(2025\)Y\. Fang, K\. Ranasinghe, L\. Xue, H\. Zhou, J\. Tan, R\. Xu, S\. Heinecke, C\. Xiong, S\. Savarese, D\. Szafir, M\. Ding, M\. S\. Ryoo, and J\. C\. NieblesRobotic VLA benefits from joint learning with motion image diffusion\.arXiv preprint arXiv:2512\.18007\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Gaoet al\.\(2025\)C\. Gao, H\. Zhang, Z\. Xu, Z\. Cai, and L\. ShaoFLIP: flow\-centric generative planning as general\-purpose manipulation world model\.InInternational Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/37bf751554eb05465ed5dc60dced3f38-Abstract-Conference.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p3.1)\.
- Gaoet al\.\(2026\)Q\. Gao, J\. Yang, Q\. Xu, L\. Chen, and Y\. WangLOME: learning human\-object manipulation with action\-conditioned egocentric world model\.External Links:2603\.27449,[Link](https://arxiv.org/abs/2603.27449)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Heet al\.\(2025\)X\. He, C\. Peng, Z\. Liu, B\. Wang, Y\. Zhang, Q\. Cui, F\. Kang, B\. Jiang, M\. An, Y\. Ren, B\. Xu, H\. Guo, K\. Gong, C\. Wu, W\. Li, X\. Song, Y\. Liu, E\. Li, and Y\. ZhouMatrix\-Game 2\.0: an open\-source, real\-time, and streaming interactive world model\.External Links:2508\.13009,[Link](https://arxiv.org/abs/2508.13009)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Heuselet al\.\(2017\)M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. HochreiterGANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InAdvances in Neural Information Processing Systems,Cited by:[Benchmarks and metrics\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Hoqueet al\.\(2025\)R\. Hoque, P\. Huang, D\. J\. Yoon, M\. Sivapurapu, and J\. ZhangEgoDex: learning dexterous manipulation from large\-scale egocentric video\.External Links:2505\.11709,[Link](https://arxiv.org/abs/2505.11709)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p7.1),[Implementation details\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px1.p1.1),[Benchmarks and metrics\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Huet al\.\(2026\)A\. Hu, V\. Volhejn, A\. R\. Rahary, C\. Mulder, A\. Makkar, A\. Royer, M\. Orsini, A\. Liao, A\. Jelley, E\. Alonso, F\. Laurent, F\. Norén, J\. Swingos, J\. Hünermann, K\. Rollins, L\. Hosseini, M\. Le Cauchois, M\. Peter, P\. de Witte, T\. Brown, V\. Micheli, M\. Böhle, G\. de Marmiesse, V\. Sharmanska, L\. Specia, M\. Black, and P\. PérezMultiplayer interactive world models with representation autoencoders\.External Links:2607\.05352,[Link](https://arxiv.org/abs/2607.05352)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Huet al\.\(2025\)Y\. Hu, Y\. Guo, P\. Wang, X\. Chen, Y\. Wang, J\. Zhang, K\. Sreenath, C\. Lu, and J\. ChenVideo prediction policy: a generalist robot policy with predictive visual representations\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 24328–24346\.External Links:[Link](https://proceedings.mlr.press/v267/hu25g.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Jianget al\.\(2025\)Z\. Jiang, Z\. Han, C\. Mao, J\. Zhang, Y\. Pan, and Y\. LiuVACE: all\-in\-one video creation and editing\.arXiv preprint arXiv:2503\.07598\.Cited by:[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Kimet al\.\(2026\)B\. Kim, T\. Kim, J\. Lee, and H\. JooDexterous world models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 29663–29673\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2026/html/Kim_Dexterous_World_Models_CVPR_2026_paper.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p3.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p4.1)\.
- Konget al\.\(2024\)W\. Kong, Q\. Tian, Z\. Zhang, R\. Min, Z\. Dai, J\. Zhou, J\. Xiong, X\. Li, B\. Wu, J\. Zhang,et al\.HunyuanVideo: a systematic framework for large video generative models\.arXiv preprint arXiv:2412\.03603\.Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Lipmanet al\.\(2023\)Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. LeFlow matching for generative modeling\.InInternational Conference on Learning Representations,Cited by:[Preliminaries](https://arxiv.org/html/2609.00161#Sx3.SSx1.p2.1)\.
- Liuet al\.\(2025\)Z\. Liu, S\. Li, E\. Cousineau, S\. Feng, B\. Burchfiel, and S\. SongGeometry\-aware 4d video generation for robot manipulation\.arXiv preprint arXiv:2507\.01099\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Luoet al\.\(2026\)X\. Luo, X\. Xin, T\. Feng, X\. Guo, M\. Jin, and J\. MaCoInteract: physically\-consistent human\-object interaction video synthesis via spatially\-structured co\-generation\.arXiv preprint arXiv:2604\.19636\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- NVIDIAet al\.\(2025\)NVIDIA, N\. Agarwal, A\. Ali, M\. Bala, Y\. Balaji, E\. Barker, T\. Cai, P\. Chattopadhyay, Y\. Chen, Y\. Cui,et al\.Cosmos world foundation model platform for physical ai\.arXiv preprint arXiv:2501\.03575\.Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Poet al\.\(2025\)R\. Po, Y\. Nitzan, R\. Zhang, B\. Chen, T\. Dao, E\. Shechtman, G\. Wetzstein, and X\. HuangLong\-context state\-space video world models\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 8733–8744\.External Links:[Link](https://openaccess.thecvf.com/content/ICCV2025/html/Po_Long-Context_State-Space_Video_World_Models_ICCV_2025_paper.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p4.1)\.
- Qwenet al\.\(2025\)Qwen, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[Object\-Token Grounding](https://arxiv.org/html/2609.00161#Sx3.SSx2.p2.1)\.
- Shanget al\.\(2026\)Y\. Shang, Z\. Li, Y\. Ma, W\. Su, X\. Jin, Z\. Wang, L\. Jin, X\. Zhang, Y\. Tang, H\. Su, C\. Gao, W\. Wu, X\. Liu, D\. Shah, Z\. Zhang, Z\. Chen, J\. Zhu, Y\. Tian, T\. Chua, W\. Zhu, and Y\. LiWorldArena: a unified benchmark for evaluating perception and functional utility of embodied world models\.External Links:2602\.08971,[Link](https://arxiv.org/abs/2602.08971)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p7.1),[Benchmarks and metrics\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Suet al\.\(2026\)H\. Su, Z\. Liu, X\. Jin, H\. Dou, C\. Hu, B\. Li, Z\. Liu, R\. Xu, J\. Fang, X\. Zhang, Z\. Yang, X\. Yang, C\. Gao, J\. Yan, Y\. Li, and W\. WuWorldScape Policy 2\.0: empowering steerable world action modeling with reasoning\-augmented memory\.External Links:2607\.18840,[Link](https://arxiv.org/abs/2607.18840)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Sunet al\.\(2026\)Z\. Sun, Z\. Du, X\. Yang, and Z\. WuHandWorld: hand\-centric unified video action generation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 15976–15985\.Cited by:[Benchmarks and metrics\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Tianet al\.\(2026\)Y\. Tian, Y\. Jin, B\. Yu, Y\. Shi, H\. Wu, C\. H\. Liu, K\. Chen, and C\. HuangSTARRY: spatial\-temporal action\-centric world modeling for robotic manipulation\.arXiv preprint arXiv:2604\.26848\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Unterthineret al\.\(2018\)T\. Unterthiner, S\. van Steenkiste, K\. Kurach, R\. Marinier, M\. Michalski, and S\. GellyTowards accurate generative models of video: a new metric and challenges\.arXiv preprint arXiv:1812\.01717\.Cited by:[Benchmarks and metrics\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Wan Teamet al\.\(2025\)Wan Team, A\. Wang, B\. Ai, B\. Wen, C\. Mao, C\. Xie, D\. Chen, F\. Yu, H\. Zhao, J\. Yang,et al\.Wan: open and advanced large\-scale video generative models\.arXiv preprint arXiv:2503\.20314\.Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Wanget al\.\(2026a\)A\. N\. Wang, T\. Darrell, P\. Izmailov, Y\. Bai, and A\. BarLifting embodied world models for planning and control\.External Links:2604\.26182,[Link](https://arxiv.org/abs/2604.26182)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1)\.
- Wanget al\.\(2026b\)Y\. Wang, W\. Ouyang, T\. Wei, Y\. Dong, Z\. Shen, and X\. PanHand2World: autoregressive egocentric interaction generation via free\-space hand gestures\.External Links:2602\.09600,[Link](https://arxiv.org/abs/2602.09600)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2025\)B\. Wu, C\. Zou, C\. Li, D\. Huang, F\. Yang, H\. Tan, J\. Peng, J\. Wu, J\. Xiong, J\. Jiang, Linus, Patrol, P\. Zhang, P\. Chen, P\. Zhao, Q\. Tian, S\. Liu, W\. Kong, W\. Wang, X\. He, X\. Li, X\. Deng, X\. Zhe, Y\. Li, Y\. Long, Y\. Peng, Y\. Wu, Y\. Liu, Z\. Wang, Z\. Dai, B\. Peng, C\. Li, G\. Gong, G\. Xiao, J\. Tian, J\. Lin, J\. Liu, J\. Zhang, J\. Lian, K\. Pan, L\. Wang, L\. Niu, M\. Chen, M\. Chen, M\. Zheng, M\. Yang, Q\. Hu, Q\. Yang, Q\. Xiao, R\. Wu, R\. Xu, R\. Yuan, S\. Sang, S\. Huang, S\. Gong, S\. Huang, W\. Guo, X\. Yuan, X\. Chen, X\. Hu, W\. Sun, X\. Wu, X\. Ren, X\. Yuan, X\. Mi, Y\. Zhang, Y\. Sun, Y\. Lu, Y\. Li, Y\. Huang, Y\. Tang, Y\. Li, Y\. Deng, Y\. Zhou, Z\. Hu, Z\. Liu, Z\. Yang, Z\. Yang, Z\. Lu, Z\. Zhou, and Z\. ZhongHunyuanVideo 1\.5 technical report\.External Links:2511\.18870,[Link](https://arxiv.org/abs/2511.18870)Cited by:[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Wuet al\.\(2026\)M\. Wu, B\. Song, R\. Lin, C\. Zhu, X\. Feng, J\. Wu, X\. Chu, and K\. HuangLatent temporal discrepancy as motion prior: a loss\-weighting strategy for dynamic fidelity in t2v\.arXiv preprint arXiv:2601\.20504\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Xianget al\.\(2024\)J\. Xiang, G\. Liu, Y\. Gu, Q\. Gao, Y\. Ning, Y\. Zha, Z\. Feng, T\. Tao, S\. Hao, Y\. Shi, Z\. Liu, E\. P\. Xing, and Z\. HuPandora: towards general world model with natural language actions and video states\.External Links:2406\.09455,[Link](https://arxiv.org/abs/2406.09455)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Yanet al\.\(2025\)H\. Yan, H\. Yu, Z\. Zhong, W\. Yuan, X\. Gong, Z\. Luo, C\. Heyu, J\. Li, W\. Song, S\. Zhou, and H\. LiOpen\-world hand\-object interaction video generation based on structure and contact\-aware representation\.arXiv preprint arXiv:2512\.01677\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2026\)Z\. Yang, Y\. Jin, L\. Qi, C\. Huang, and K\. ChenEA\-WM: event\-aware generative world model with structured kinematic\-to\-visual action fields\.arXiv preprint arXiv:2605\.06192\.Cited by:[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2024\)Z\. Yang, J\. Teng, W\. Zheng, M\. Ding, S\. Huang, J\. Xu, Y\. Yang, W\. Hong, X\. Zhang, G\. Feng, D\. Yin, X\. Gu, Y\. Zhang, W\. Wang, Y\. Cheng, T\. Liu, B\. Xu, Y\. Dong, and J\. TangCogVideoX: text\-to\-video diffusion models with an expert transformer\.arXiv preprint arXiv:2408\.06072\.Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Zhanget al\.\(2026a\)J\. Zhang, X\. Chen, A\. Chen, C\. Lv, D\. Li, G\. Zhou, H\. Yin, H\. Yuan, H\. Li, J\. Li, J\. Zhang, J\. Zhou, K\. Gao, K\. Yan, L\. Jiang, N\. Tang, P\. Lin, Q\. Peng, S\. Yin, T\. Wu, T\. Yan, X\. Xu, Y\. Shu, Y\. Zhang, Y\. Wang, Y\. Wang, Y\. Chen, Y\. Xu, Y\. Huang, Y\. Chen, Z\. Zhang, Z\. Wang, Z\. Lei, Z\. Liang, Z\. Liu, Z\. Zhou, X\. Chen, and C\. WuQwen\-RobotWorld technical report: unifying embodied world modeling through language\-conditioned video generation\.External Links:2606\.17030,[Link](https://arxiv.org/abs/2606.17030)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2026b\)P\. Zhang, Y\. Deng, S\. Sun, J\. Ma, D\. Wang, J\. Du, Z\. Pan, Y\. Huang, H\. Liang, S\. Huang, R\. Zhang, E\. Xie, M\. Liu, and D\. ZhouPhysisForcing: physics reinforced world simulator for robotic manipulation\.External Links:2606\.28128,[Link](https://arxiv.org/abs/2606.28128)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p3.1),[External representations priors](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025\)Y\. Zhang, J\. Gu, L\. Wang, H\. Wang, J\. Cheng, Y\. Zhu, and F\. ZouMimicMotion: high\-quality human motion video generation with confidence\-aware pose guidance\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 74896–74910\.External Links:[Link](https://proceedings.mlr.press/v267/zhang25v.html)Cited by:[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Zhaoet al\.\(2025\)B\. Zhao, R\. Tang, M\. Jia, Z\. Wang, F\. Man, X\. Zhang, Y\. Shang, W\. Zhang, W\. Wu, C\. Gao, X\. Chen, and Y\. LiAirScape: an aerial generative world model with motion controllability\.External Links:2507\.08885,[Link](https://arxiv.org/abs/2507.08885)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2026\)B\. Zhao, J\. Xu, W\. Feng, X\. Zhang, Z\. Wang, H\. Wang, S\. Ji, Z\. Wang, J\. Fang, Z\. Zheng, W\. Zhang, Y\. Shang, W\. Wu, C\. Gao, X\. Chen, and Y\. LiWorldVLN: autoregressive world action model for aerial vision\-language navigation\.External Links:2605\.15964,[Link](https://arxiv.org/abs/2605.15964)Cited by:[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1)\.
- Zhenet al\.\(2025\)H\. Zhen, Q\. Sun, H\. Zhang, J\. Li, S\. Zhou, Y\. Du, and C\. GanLearning 4d embodied world models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 5337–5347\.External Links:[Link](https://openaccess.thecvf.com/content/ICCV2025/html/Zhen_Learning_4D_Embodied_World_Models_ICCV_2025_paper.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p3.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
- Zhuet al\.\(2025\)F\. Zhu, H\. Wu, S\. Guo, Y\. Liu, C\. Cheang, and T\. KongIRASim: a fine\-grained world model for robot manipulation\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 9834–9844\.External Links:[Link](https://openaccess.thecvf.com/content/ICCV2025/html/Zhu_IRASim_A_Fine-Grained_World_Model_for_Robot_Manipulation_ICCV_2025_paper.html)Cited by:[Introduction](https://arxiv.org/html/2609.00161#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.00161#Sx1.p4.1),[Interactive video generation and world models](https://arxiv.org/html/2609.00161#Sx2.SS0.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.00161#Sx4.SSx1.SSS0.Px3.p1.1)\.
## A\. Data Construction and Process Details
\\pdfdest
name impact\.sec:data\_details xyz
### A\.1 RoboTwin Data
\\pdfdest
name impact\.sec:robotwin\_data xyz
#### Data source and modalities\.
The robot\-arm training corpus comprises simulated bimanual manipulation rollouts from RoboTwin 2\.0, generated with the SAPIEN physics engine and the dual\-arm Aloha\-AgileX embodiment\. Each episode provides synchronized RGB observations, camera calibration, robot states, end\-effector poses, gripper states, language instructions, and planned action trajectories\. We use the head\-camera stream at1280×7041280\\times 704resolution and 10 fps, obtained by sampling every third frame from the 30 fps source video\. The synchronized action at each frame is
\[q1:6L,gL,q1:6R,gR\]∈ℝ14,\\pdfdestnameimpact\.eq:robotwinactionorderxyz\[\\,q^\{L\}\_\{1:6\},\\ g^\{L\},\\ q^\{R\}\_\{1:6\},\\ g^\{R\}\\,\]\\in\\mathbb\{R\}^\{14\},\\pdfdest name\{impact\.eq:robotwin\_\{a\}ction\_\{o\}rder\}xyz\(20\)whereqLq^\{L\}andqRq^\{R\}denote the six joint positions of the left and right arms, andgLg^\{L\}andgRg^\{R\}denote their gripper coordinates\.
#### Task and visual diversity\.
The training set contains both clean and domain\-randomized rollouts\. The randomized regime introduces collision\-aware distractors, samples table and background appearances from more than 12,000 textures, varies illumination, perturbs the tabletop height by up to 0\.03 m, and displaces the head camera\. Clean backgrounds and extreme illumination each occur in 2% of randomized rollouts\. The language instructions also vary across equivalent task executions\.
The corpus covers 50 tasks spanning pick\-and\-place, container placement, stacking, object ranking, tool use, switch operation, articulated\-object manipulation, bimanual handover, and coordinated dual\-arm manipulation\. The task distribution is non\-uniform and ranges from 99 to 28,117 training windows per task\. The four largest tasks are color\-based block ranking \(28,117 windows\), size\-based block ranking \(27,170\), bottle placement into a container \(16,881\), and block handover \(16,723\)\.
#### Window construction and action normalization\.
Each training example contains 17 consecutive frames with stride one and the corresponding 17 action vectors\. Boundary examples are temporally resampled or padded to preserve this fixed length\. For action dimensionjj, we normalize using the first and ninety\-ninth dataset percentiles,p01,jp\_\{01,j\}andp99,jp\_\{99,j\}:
a~j=clip\(2aj−p01,jp99,j−p01,j−1,−1,1\)\.\\pdfdestnameimpact\.eq:actionnormalizationxyz\\widetilde\{a\}\_\{j\}=\\operatorname\{clip\}\\\!\\left\(2\\frac\{a\_\{j\}\-p\_\{01,j\}\}\{p\_\{99,j\}\-p\_\{01,j\}\}\-1,\\,\-1,\\,1\\right\)\.\\pdfdest name\{impact\.eq:action\_\{n\}ormalization\}xyz\(21\)RGB frames and actions always share the same temporal indices\. Spatial bucket sampling and matched crop\-and\-resize augmentation produce the 720p training inputs\.
### A\.2 EgoDex Data
\\pdfdest
name impact\.sec:egodex\_data xyz
#### Data source and modalities\.
The human\-hand training corpus contains approximately 256K egocentric clips from EgoDex spanning 118 tasks\. Each clip is paired with camera intrinsics, per\-frame camera poses, articulated hand transforms and a language instruction\. Each RGB clip is paired with a rendered hand\-pose video constructed from the capture\-time 3D hand tracks\.
The tasks cover tabletop setup and cleanup, pick\-and\-place, stacking, assembly and disassembly, folding and wrapping, cooking, washing, device insertion and removal, writing, drawing, and tool use\.
#### Temporal clip construction\.
Each source clip is converted into an 81\-frame training clip\. Clips longer than 81 frames are represented by 81 uniformly spaced indices including both endpoints; clips shorter than 81 frames repeat the final index\. The same index sequence is applied to RGB and hand\-pose videos, preserving frame\-level correspondence between the target and condition\. Source videos are1920×10801920\\times 1080at 30 fps, and contiguous 81\-frame clips therefore span 2\.7 seconds\. Training uses matched spatial transformations at 720p\.
### A\.3 Object\-Token Grounding
\\pdfdest
name impact\.sec:grounding\_details xyz
#### Manipulated\-object extraction\.
We annotate each training example once with Qwen2\.5\-0\.5B\-Instruct\. For both RoboTwin and EgoDex, the task instruction is the sole annotation input; no video frame, pose signal, temporal metadata, or external visual annotation is used\. The model extracts compact noun phrases for the manipulated objects while preserving discriminative attributes such as color, material, shape, label, or container type\. Agents, body parts, cameras, supporting surfaces, backgrounds, and non\-interacting scene elements are excluded\.
#### Complete annotation prompt\.
RoboTwin and EgoDex use the same prompt template below, with*\{instruction\}*replaced by the task instruction for the current training clip\.
System message *You are a precise language annotation assistant for manipulation training\. Given only a task instruction, identify the concrete physical objects directly manipulated by the agent\. Return only valid JSON, without Markdown, explanations, or code fences\.**Use compact English noun phrases and preserve discriminative attributes stated in the instruction, including color, material, shape, printed label, object part, and container type\. Exclude the robot, hands, fingers, arms, people, cameras, frames, tables, workspaces, backgrounds, and generic scene regions unless the instruction explicitly identifies one of them as the manipulated object\. Do not infer objects that are not stated in the instruction\.*User message *Task instruction:* *\{instruction\}**Return the following JSON object:**\{**“interacting\_objects”: \[**\{“object”: “compact manipulated\-object phrase”\}**\]**\}**Include every manipulated object explicitly stated in the instruction\. Preserve the order in which the objects appear\. If the instruction refers to the same object more than once, return it once\. Return JSON only\.*
#### Mapping object phrases to text tokens\.
The extracted object phrase is aligned with its occurrence in the original task instruction after tokenization\. All subword tokens associated with the object phrase are included in the object\-token set𝒪\\mathcal\{O\}\. We then aggregate the cross\-attention associated with𝒪\\mathcal\{O\}across attention heads and transformer layers to obtain the object\-conditioned attention map𝐀\\mathbf\{A\}used by ADS\.
Table 4:Representative instruction\-only object\-grounding examples\.\\pdfdestname impact\.tab:grounding\_examples xyz
### A\.4 Hand\-Pose Video Construction
\\pdfdest
name impact\.sec:hand\_pose\_construction xyz
#### Pose representation\.
EgoDex provides capture\-time 3D hand and finger tracks together with per\-frame camera calibration\. We use these annotations directly and do not estimate hand pose from the compressed RGB video\. Each hand is represented by 21 three\-dimensional joints, accompanied by a presence mask\. The camera metadata provides a4×44\\times 4world\-to\-camera transformation and four intrinsic parameters for every frame\.
#### Projection and rendering\.
For framett, a homogeneous world point𝐩¯\\bar\{\\mathbf\{p\}\}is transformed by the world\-to\-camera matrix𝐄t\\mathbf\{E\}\_\{t\}and projected using the intrinsics\(fx,fy,cx,cy\)t\(f\_\{x\},f\_\{y\},c\_\{x\},c\_\{y\}\)\_\{t\}:
𝐩c=𝐄t𝐩¯,u=fxpcx/pcz\+cx,v=fypcy/pcz\+cy\.\\mathbf\{p\}\_\{c\}=\\mathbf\{E\}\_\{t\}\\bar\{\\mathbf\{p\}\},\\qquad u=f\_\{x\}p\_\{c\}^\{x\}/p\_\{c\}^\{z\}\+c\_\{x\},\\quad v=f\_\{y\}p\_\{c\}^\{y\}/p\_\{c\}^\{z\}\+c\_\{y\}\.\(22\)Points with depthpcz≤0\.01p\_\{c\}^\{z\}\\leq 0\.01are excluded\. The 21 joints follow the wrist, thumb, index, middle, ring, and little\-finger ordering\. Each finger is connected to the wrist and rendered as an anti\-aliased skeleton on a black background\. We draw green edges with width 2, blue joints with radius 3, and fingertips with radius 5\. Pose videos preserve the RGB resolution and are encoded at 30 fps using H\.264 with the YUV420p pixel format and a constant\-rate factor of 23\.
#### Temporal and spatial alignment\.
RGB and pose videos contain the same number of decoded frames and use the same 81\-frame index sequence for every training segment\. The two modalities therefore remain aligned by frame index throughout temporal sampling\. Identical crop and resize parameters are subsequently applied to both modalities\.
## B\. Implementation Details
\\pdfdest
name impact\.sec:implementation\_details xyz
### B\.1 Robot\-Arm Models
\\pdfdest
name impact\.sec:robot\_architecture xyz
#### Backbone\.
We instantiate the robot\-arm setting with two action\-conditioned diffusion transformer backbones: Wan 2\.2 TI2V 5B and Cosmos\-Predict 2\.5 \(action\)\. The Wan 2\.2 backbone contains 30 transformer blocks with hidden widthd=3072d=3072, 24 attention heads, and an FFN width of 14,336\. It receives a 17\-frame RGB sequence together with the first\-frame visual condition and the synchronized 14\-DoF action trajectory\. The second model is initialized from Cosmos\-Predict2\.5\-2B and retains its action\-conditioning interface for the same robot\-arm setting\.
#### Action encoder\.
Let𝐚∈ℝ17×14\\mathbf\{a\}\\in\\mathbb\{R\}^\{17\\times 14\}denote the normalized action trajectory\. We flatten the complete trajectory into 238 scalars and process it with two independent multilayer perceptrons of identical topology:
MLPq\(𝐚\)=Wq,2GELU\(Wq,1vec\(𝐚\)\+bq,1\)\+bq,2\\operatorname\{MLP\}\_\{q\}\(\\mathbf\{a\}\)=W\_\{q,2\}\\,\\operatorname\{GELU\}\(W\_\{q,1\}\\operatorname\{vec\}\(\\mathbf\{a\}\)\+b\_\{q,1\}\)\+b\_\{q,2\}\(23\)with hidden width4d=12,2884d=12\{,\}288\. The first encoder produces add\-dimensional vector that is added to the timestep embedding\. The second produces6d=18,4326d=18\{,\}432values, reshaped to6×d6\\times d, which modulate the six adaptive normalization components in every transformer block\. Learned binary condition embeddings distinguish action\-conditioned and action\-dropped examples\. Action information is therefore injected globally through the timestep and adaptive normalization pathways\.
#### IMPACT configuration\.
For both backbones, IMPACT constructs its attention anchor from the object\-token cross\-attention produced by the native transformer\. ADS samplesK=8K=8candidate masks by adding Gaussian perturbations withσ=0\.75\\sigma=0\.75to the logit of the detached attention anchor on a5×8×125\\times 8\\times 12coarse grid\. The perturbed maps are trilinearly upsampled to the latent resolution, transformed with sampling temperatureτ=1\\tau=1and anchor powerκ=0\.65\\kappa=0\.65, and thresholded to retain the topρ=10%\\rho=10\\%of positions\. Detached regional MSEs score the candidate masks, and centered scores are converted into aggregation weights with coefficientβ=0\.7\\beta=0\.7\. IWS uses the resulting soft interaction map to assign per\-position weights in\[1,γ\]\[1,\\gamma\], withγ=5\\gamma=5\.
### B\.2 Human\-Hand Model
\\pdfdest
name impact\.sec:human\_architecture xyz
#### Pose\-conditioned input\.
The human\-hand model uses the same 5B transformer and conditions on an 81\-frame hand\-pose video\. RGB targets, the first\-frame reference, and pose frames are encoded by the shared Wan VAE without an additional pose encoder\. The VAE produces 48\-channel latents with temporal and spatial compression factors of 4, 16, and 16, yielding 21 latent timesteps for an 81\-frame sequence\.
The first latent timestep retains the RGB reference condition\. At subsequent timesteps, the 48\-channel reference latent is replaced by the temporally aligned 48\-channel pose latent\. Four binary mask channels are concatenated to form a 52\-channel condition𝐲\\mathbf\{y\}\. Before 3D patch embedding, the model concatenates the 48\-channel noisy video latent𝐱\\mathbf\{x\}with𝐲\\mathbf\{y\}, giving 100 input channels:
𝐡0=PatchEmbed\(\[𝐱;𝐲mask\+pose/ref\]\)\.\\mathbf\{h\}\_\{0\}=\\operatorname\{PatchEmbed\}\\bigl\(\[\\mathbf\{x\};\\mathbf\{y\}\_\{\\rm mask\+pose/ref\}\]\\bigr\)\.\(24\)The action mask applies pose replacement only to conditioned samples\. RGB and pose inputs share the same temporal indices and spatial transformation\.
## C\. Evaluation Details
\\pdfdest
name impact\.sec:evaluation\_details xyz
### C\.1 Evaluation Protocols and Baseline Versions
\\pdfdest
name impact\.sec:baseline\_protocol xyz
WorldArena evaluates 500 held\-out episodes from 50 RoboTwin 2\.0 tasks\. All decoded submissions have a minimum resolution of640×480640\\times 480and a frame rate of 24 fps\. Text\-conditioned submissions contain 121 frames\. Action\-conditioned submissions follow the benchmark action sequence and match the corresponding reference trajectory length\. EgoDex models are evaluated at their release\-specific inference settings, as detailed in attr /Border \[0 0 0\] goto name impact\.tab:baseline\_specsTable[5](https://arxiv.org/html/2609.00161#Sx8.T5)\.
MethodVersionResolutionSec\. / FPS*General video generation*HunyuanVideo\-1\.525\.11\.20848×480848\\times 4805 / 24Cosmos\-Predict 2\.525\.10\.061280×7041280\\times 7045 / 16*Pose\-controlled video generation*MimicMotion24\.07\.081024×5761024\\times 5764\.8 / 15MagicPose24\.04\.03512×512512\\times 5124 / 15VACE25\.03\.11720×1080720\\times 10805 / 16LOME26\.04\.05832×480832\\times 4805 / 15*Wan 2\.2\-based models*Wan 2\.225\.07\.281280×7041280\\times 7045 / 24Wan 2\.2\-AC / \+IMPACT–1280×7041280\\times 7043\.4 / 24
Table 5:Detailed inference configurations of video generation models evaluated on EgoDex\. All models use image\-to\-video generation; “–” denotes our adapted variants\.\\pdfdestname impact\.tab:baseline\_specs xyz
Metric computation uses the ordered decoded frames from each output\. For Wan 2\.2\-AC and IMPACT, the 81 output frames and the pose condition use identical temporal indices\. Each model uses its official sampling schedule and guidance configuration\.
### C\.2 WorldArena Metric Definitions
\\pdfdest
name impact\.sec:worldarena\_metrics xyz
WorldArena normalizes each raw metric to\[0,1\]\[0,1\]using empirically selected boundaries and reports EWMScore as100100times the arithmetic mean of the 16 normalized metrics\. Thus, all displayed entries are higher\-is\-better, including metrics whose underlying raw quantity is an error\.
Table 6:Meaning and implementation of all 16 WorldArena metrics\.\\pdfdestname impact\.tab:worldarena\_metric\_definitions xyz#### Full 16\-metric results\.
\\pdfdest
name impact\.sec:full\_worldarena\_results xyz
Tables attr /Border \[0 0 0\] goto name impact\.tab:worldarena\_full\_quality[8](https://arxiv.org/html/2609.00161#Sx8.T8)and attr /Border \[0 0 0\] goto name impact\.tab:worldarena\_full\_task[8](https://arxiv.org/html/2609.00161#Sx8.T8)report all 16 normalized WorldArena metrics for the same models and ordering\. The first covers the generation\-quality dimensions \(visual quality, motion quality, content consistency\) and the overall EWMScore; the second covers the task\-oriented dimensions \(physics adherence, 3D accuracy, controllability\)\.
Table 7:Full WorldArena results \(1/2\): EWMScore and the nine generation\-quality metrics\. EWMScore is the mean of all 16 metrics on a\[0,100\]\[0,100\]scale; components are normalized to\[0,1\]\[0,1\]\. Higher is better \(↑\\uparrow\); boldface and underlining mark the best and second\-best per column\.\\pdfdestname impact\.tab:worldarena\_full\_quality xyzTable 8:Full WorldArena results \(2/2\): the task\-oriented metrics\.\\pdfdestname impact\.tab:worldarena\_full\_task xyz
### C\.3 Additional Robot\-Arm Cases
\\pdfdest
name impact\.sec:additional\_robot\_cases xyz In this section, we present more qualitative results of robot\-arm manipulation on WorldArena\.
![[Uncaptioned image]](https://arxiv.org/html/2609.00161v1/figures/robot_more_cases/case_0038.jpg)
Prompt:Pick up theprinted sneakerand place it on the blue mat\.
Figure 5:Printed\-sneaker placement comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0038 xyz
![[Uncaptioned image]](https://arxiv.org/html/2609.00161v1/figures/robot_more_cases/case_0042.jpg)
Prompt:Pick up thebrown\-and\-white bottleand place it on the blue mat\.
Figure 6:Bottle placement comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0042 xyz

Prompt:Pick up thered blockand place it on the blue target\.
Figure 7:Block placement comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0047 xyz
Prompt:Grasp thelidded potwith both grippers and lift it\.
Figure 8:Bimanual pot\-lifting comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0059 xyz
Prompt:Pick up thegreen blockand stack it on the red block\.
Figure 9:Block\-stacking comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0080 xyz
Prompt:Pick up thebrown shoeand place it on the blue mat\.
Figure 10:Shoe placement comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0108 xyz
Prompt:Pick up oneblue bowland stack it inside the other blue bowl\.
Figure 11:Bowl\-stacking comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0143 xyz
Prompt:Press theblue service bell\.
Figure 12:Service\-bell pressing comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0169 xyz
Prompt:Place thetoy hamburger and French friesin the tray\.
Figure 13:Food\-toy placement comparison\. Rows show WoW, CogVideoX, CtrlWorld, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation, instruction, and action trajectory; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:robot\_case\_0220 xyz
Prompt:Pick up theblue elephant toyfrom beside the black case\.
Figure 14:Elephant\-toy pickup comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_01 xyz
### C\.4 Additional Human\-Hand Cases
\\pdfdest
name impact\.sec:additional\_hand\_cases xyz In this section, we present more qualitative results of human\-hand manipulation on EgoDex\.
![[Uncaptioned image]](https://arxiv.org/html/2609.00161v1/case_02.png)
Prompt:Grasp thetransparent containerand lift it from the shelf\.
Figure 15:Container\-lifting comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_02 xyz
![[Uncaptioned image]](https://arxiv.org/html/2609.00161v1/case_03.png)
Prompt:Grasp thewhite cylindrical objectand lift it from the base\.
Figure 16:Cylindrical\-object lifting comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_03 xyz

Prompt:Pick up thetwo white drawersand stack them together\.
Figure 17:Drawer\-stacking comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_04 xyz
Prompt:Grasp and adjust thecolorful block structure\.
Figure 18:Block\-structure adjustment comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_05 xyz
Prompt:Pick up thered bookfrom the table\.
Figure 19:Book\-pickup comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_06 xyz
Prompt:Stack theorange bowlon the white bowl\.
Figure 20:Bowl\-stacking comparison\. Rows show HunyuanVideo\-1\.5, VACE, MimicMotion, Wan 2\.2\-AC, and IMPACT \(Ours\) under the same initial observation and task instruction; columns are uniformly sampled frames from each rollout\.\\pdfdestname impact\.fig:hand\_case\_07 xyzSimilar Articles
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
INTACT is an end-to-end unified JEPA that learns the intent-to-action mapping directly, enabling search-free world model control. It achieves 95.33% direct macro success rate across four visual-control tasks with zero test-time search and ~300x lower planning latency.
TaskSense: Focusing on What Matters in World Models
TaskSense introduces a task-centric world modeling framework that uses stochastic spatial attention conditioned on previous latent states and an auxiliary inverse-dynamics objective to focus on control-relevant regions, improving robustness to visual distractions compared to DreamerV3.
Interdomain Attention: Beyond Token-Level Key-Value Memory
Proposes Interdomain Attention, a new method that integrates state space models into attention via kernel methods, achieving efficient long-context modeling with a fixed-size state and outperforming SSMs and softmax attention in language modeling experiments up to 1.3B parameters.
ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
ActWorld proposes a chunk-autoregressive world model with hierarchical action-aware memory to support object interaction alongside navigation, addressing data and memory bottlenecks in existing interactive world models.
World in World: Explore the World with World Models
The paper presents World in World, a training-free interface that enables flexible camera and time control in frozen autoregressive video world models by using correspondence-guided queries and evidence-wise attention guidance.