WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
Summary
WorldAttention proposes an efficient attention architecture with Hybrid Sparse Attention and Hierarchical KV Cache for interactive video world models, achieving state-of-the-art performance on benchmarks like VBench-Long and InterVBench.
View Cached Full Text
Cached at: 09/30/26, 04:22 AM
Paper page - WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
Source: https://huggingface.co/papers/2609.34606
Abstract
Leveragingtheparadigmofautoregressivediffusion,text-conditionedinteractivevideoworldmodelsaimtosimulatetemporallycoherentenvironmentsguidedbytextualinstructions.Whileenablinglow-latency,long-durationgenerationispivotalforembodiedAIandsimulation-basedplanning,currentframeworksprimarilyrelyonsliding-windowmechanismstoboundcomputationalcomplexity.However,thisapproachinherentlysacrificeshistoricalcontext,underminingthelong-rangeinteractivecapabilities.Conversely,maintainingafull-historycacheremainscomputationallyprohibitiveandmemory-intensive:thequadraticcomplexityofattentionleadstoexcessivecomputationaloverhead,whilethelineargrowthoftheKVcacheinevitablyleadstoGPUmemorysaturation.Toovercometheselimitations,weproposeWorldAttention,asystem-orientedattentionarchitecturethatachieveshighefficiencythroughtheco-designofspecializedattentionkernelsandhierarchicalKVcachemanagement.First,weintroduceHybridSparseAttention(HSA),whichintegrateslinearglobalattentionsupplementedwithhead-adaptivesparseattention.Additionally,wedesignaHierarchicalKVCache(HKV)thatorganizeshistoricalKVpairsintosemanticallyindexedpagesacrossmulti-tiermemory,enablingfine-grainedretrievalandcontrolledGPUresidency.Thesetwodesignsaresupportedbytailoredkernelstoeffectivelytranslatetheirtheoreticalefficiencyintoreal-worldperformance.ExtensiveexperimentsonVBench-LongandInterVBenchdemonstratethatWorldAttentionconsistentlysurpassespriorstate-of-the-artmethods,achievingsubjectconsistencyscoresof0.9472onVBench-Longand0.9668onInterVBench,respectively.
View arXiv pageView PDFProject pageGitHub7Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34606 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34606 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34606 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Addressable Memory for Video World Models
This paper introduces WorldTrace, a training-free memory framework for long-horizon video world models that keeps compressed cache addressable, plus LoopBench, a benchmark for episodic recall after long detours. It improves temporal consistency by +15.5% and episodic recall by +19.5% on LoopBench.
ReWorld: An Interactive World Model with Long-Horizon Memory
ReWorld presents an interactive world model that decouples short-horizon control from long-horizon memory using mixed attention windows and distribution-matching LoRA distillation, enabling real-time video generation with high action fidelity and recall.
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
FVAttn is a training-free sparse attention system that uses runtime load balancing to improve distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism for video generation, achieving up to 4.41x attention speedup and 2.11x inference speedup over FlashAttention on step-distilled Wan2.2 I2V while maintaining competitive video quality.
Memory in Video World Models (6 minute read)
A best-paper research from NVIDIA and collaborators introduces WorldTrace, a training-free framework that keeps compressed memory addressable in autoregressive video world models by assigning fixed slot-rank positions, enabling coherent long rollouts and long-range recall beyond the training horizon.
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
This paper introduces MBench, a benchmark for evaluating the memory capabilities of video world models across entity, environment, and causal consistency over long temporal horizons.