Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Summary
Q-ARVD is a novel quantization framework to reduce inference costs of autoregressive video diffusion models by addressing frame-wise sensitivity imbalance and weight outlier patterns.
View Cached Full Text
Cached at: 05/22/26, 06:23 AM
Paper page - Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Source: https://huggingface.co/papers/2605.21072
Abstract
Autoregressive video diffusion models face high inference costs that limit practical deployment, prompting the development of Q-ARVD, a novel quantization framework addressing frame-wise sensitivity imbalance and weight outlier patterns specific to these models.
Autoregressive video diffusion models(ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and world modeling. Despite their potential, the substantial inference cost of ARVDs remains a major obstacle to practical deployment, making modelquantizationa natural direction for improving efficiency. However,quantizationfor ARVDs remains largely unexplored. Our empirical analysis shows that directly applying existingquantizationschemes developed for standarddiffusion transformersto ARVDs leads to suboptimal performance, revealingquantizationbehaviors that differ from those observed in bidirectional diffusion models. In this paper, we identify two critical challenges in quantizing ARVDs: (C1) Highly unbalancedframe-wise quantization sensitivity.Error accumulationduring autoregressive generation can induce severely skewedquantizationsensitivity across frames, following an exponential-like decay pattern. (C2) Prominent and heterogeneous outlier patterns in weights.Weight distributionsexhibit pronouncedoutlier channels, whose patterns vary substantially across layer types and block depths. To address these issues, we propose Q-ARVD, a novel framework for accurate ARVDquantization. (S1) To tackle the highly unbalanced frame-wise sensitivity, Q-ARVD incorporates afinal-quality aware frame-weightingmechanism into thequantizationobjective. (S2) To prevent heterogeneous outliers from degrading performance, Q-ARVD introduces an outlier-awareadaptive dual-scale quantization, which automatically detects the presence and quantity ofoutlier channelsfor an arbitrary layer, and isolates them to protect normal channels. Extensive experiments demonstrate the superiority of Q-ARVD.
View arXiv pageView PDFGitHub9Add to collection
Get this paper in your agent:
hf papers read 2605\.21072
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.21072 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.21072 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.21072 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Show HN: Shoehorn – Quantize any model down to run on your machine
Shoehorn is a developer tool that quantizes language models to perfectly fit within a machine's available memory, achieving up to 99.99% memory usage efficiency for local inference.
Muse Glimmer is a memory hierarchy disguised as a 30B Transformer
Meta's Muse Glimmer is a 30B multimodal Transformer model designed for autonomous agentic tasks on consumer hardware, using a memory hierarchy and quantization to fit within 24-32 GB envelopes.
The GOAT of local LLM youtube is back
A YouTube creator returns after a hiatus to teach building a distributed training framework from first principles, focusing on advanced AI topics like DeepSeek, MoE, and MLA, with an emphasis on developing problem-solving skills and self-confidence.
p-Spin Glass Network Efficient Single-Batch Continual Learning
Introduces the p-Spin Glass Network, a novel architecture for sequence models that achieves memory efficiency, sample efficiency, and single-batch stability, enabling continual learning and edge AI applications.
Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks
This paper introduces NanoEQL, a fully white-box neural network that deciphers size determination mechanisms in nanocrystal synthesis through interpretable equations and scalars, advancing rational design and chemical reaction analysis.