FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Hugging Face Daily Papers Papers

Summary

FVAttn is a training-free sparse attention system that uses runtime load balancing to improve distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism for video generation, achieving up to 4.41x attention speedup and 2.11x inference speedup over FlashAttention on step-distilled Wan2.2 I2V while maintaining competitive video quality.

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present , a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. uses Top-p routing, a Top-k safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, reduces average load imbalance from 1.34 to 1.08 and delivers a 4.41times attention speedup over FlashAttention, while achieving a 2.02--2.11times DiT inference speedup with competitive video quality.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:08 AM

Paper page - FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Source: https://huggingface.co/papers/2607.16190

Abstract

VideoDiffusionTransformersprocesslongspatio-temporalsequences,makingself-attentionthemainbottleneckinhigh-resolutionvideogeneration.Training-freesparseattentionreducesthiscost,butadaptiveTop-proutingcreatesunevenper-headworkloadsundermulti-GPUsequenceparallelism.Theresultingworkloadheterogeneityturnssparseattentionintoarank-levelstragglerproblem.Wepresent,atraining-freesparse-attentionsystemthatimprovesthedistributedexecutionefficiencyofadaptivesparseattentionundermulti-GPUsequenceparallelism.usesTop-prouting,aTop-ksafetyfloor,andvideo-awareblockorganizationasthesparse-routingfrontend,thenrepairsthematerializedmaskatruntime.RuntimeLoadBalancingmigratesasmallnumberofheavyheadsviaP2Pcommunicationtoshortenthecurrentcriticalpath.Slack-AwareSparseAugmentationfillsresidualnon-critical-rankslackwithadditionalhigh-valueblocks,whileoverlaphidesschedulingandmigrationoverheadbehindexistingcomputation.Onstep-distilledWan2.2I2V,reducesaverageloadimbalancefrom1.34to1.08anddeliversa4.41timesattentionspeedupoverFlashAttention,whileachievinga2.02--2.11timesDiTinferencespeedupwithcompetitivevideoquality.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2607\.16190

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.16190 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.16190 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.16190 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

arXiv cs.CL

SparDA proposes a decoupled sparse attention architecture that adds a lightweight 'Forecast' projection to predict future KV cache needs, enabling lookahead prefetching from CPU to GPU and reducing selection overhead. On 8B sparse-pretrained models, it achieves up to 1.25× prefill and 1.7× decode speedup, with up to 5.3× higher decode throughput over non-offload baselines.