WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

Hugging Face Daily Papers Papers

Summary

WorldAttention proposes an efficient attention architecture with Hybrid Sparse Attention and Hierarchical KV Cache for interactive video world models, achieving state-of-the-art performance on benchmarks like VBench-Long and InterVBench.

Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primarily rely on sliding-window mechanisms to bound computational complexity. However, this approach inherently sacrifices historical context, undermining the long-range interactive capabilities. Conversely, maintaining a full-history cache remains computationally prohibitive and memory-intensive: the quadratic complexity of attention leads to excessive computational overhead, while the linear growth of the KV cache inevitably leads to GPU memory saturation. To overcome these limitations, we propose WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. First, we introduce Hybrid Sparse Attention (HSA), which integrates linear global attention supplemented with head-adaptive sparse attention. Additionally, we design a Hierarchical KV Cache (HKV) that organizes historical KV pairs into semantically indexed pages across multi-tier memory, enabling fine-grained retrieval and controlled GPU residency. These two designs are supported by tailored kernels to effectively translate their theoretical efficiency into real-world performance. Extensive experiments on VBench-Long and InterVBench demonstrate that WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:22 AM

Paper page - WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

Source: https://huggingface.co/papers/2609.34606

Abstract

Leveragingtheparadigmofautoregressivediffusion,text-conditionedinteractivevideoworldmodelsaimtosimulatetemporallycoherentenvironmentsguidedbytextualinstructions.Whileenablinglow-latency,long-durationgenerationispivotalforembodiedAIandsimulation-basedplanning,currentframeworksprimarilyrelyonsliding-windowmechanismstoboundcomputationalcomplexity.However,thisapproachinherentlysacrificeshistoricalcontext,underminingthelong-rangeinteractivecapabilities.Conversely,maintainingafull-historycacheremainscomputationallyprohibitiveandmemory-intensive:thequadraticcomplexityofattentionleadstoexcessivecomputationaloverhead,whilethelineargrowthoftheKVcacheinevitablyleadstoGPUmemorysaturation.Toovercometheselimitations,weproposeWorldAttention,asystem-orientedattentionarchitecturethatachieveshighefficiencythroughtheco-designofspecializedattentionkernelsandhierarchicalKVcachemanagement.First,weintroduceHybridSparseAttention(HSA),whichintegrateslinearglobalattentionsupplementedwithhead-adaptivesparseattention.Additionally,wedesignaHierarchicalKVCache(HKV)thatorganizeshistoricalKVpairsintosemanticallyindexedpagesacrossmulti-tiermemory,enablingfine-grainedretrievalandcontrolledGPUresidency.Thesetwodesignsaresupportedbytailoredkernelstoeffectivelytranslatetheirtheoreticalefficiencyintoreal-worldperformance.ExtensiveexperimentsonVBench-LongandInterVBenchdemonstratethatWorldAttentionconsistentlysurpassespriorstate-of-the-artmethods,achievingsubjectconsistencyscoresof0.9472onVBench-Longand0.9668onInterVBench,respectively.

View arXiv pageView PDFProject pageGitHub7Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34606 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34606 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34606 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Addressable Memory for Video World Models

Hugging Face Daily Papers

This paper introduces WorldTrace, a training-free memory framework for long-horizon video world models that keeps compressed cache addressable, plus LoopBench, a benchmark for episodic recall after long detours. It improves temporal consistency by +15.5% and episodic recall by +19.5% on LoopBench.

ReWorld: An Interactive World Model with Long-Horizon Memory

Hugging Face Daily Papers

ReWorld presents an interactive world model that decouples short-horizon control from long-horizon memory using mixed attention windows and distribution-matching LoRA distillation, enabling real-time video generation with high action fidelity and recall.

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Hugging Face Daily Papers

FVAttn is a training-free sparse attention system that uses runtime load balancing to improve distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism for video generation, achieving up to 4.41x attention speedup and 2.11x inference speedup over FlashAttention on step-distilled Wan2.2 I2V while maintaining competitive video quality.

Memory in Video World Models (6 minute read)

TLDR AI

A best-paper research from NVIDIA and collaborators introduces WorldTrace, a training-free framework that keeps compressed memory addressable in autoregressive video world models by assigning fixed slot-rank positions, enabling coherent long rollouts and long-range recall beyond the training horizon.