FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
Summary
FadeMem introduces a distance-aware key-value memory consolidation mechanism that organizes historical video data into a temporal hierarchy, improving long-video generation under fixed cache constraints.
View Cached Full Text
Cached at: 06/10/26, 01:44 PM
Paper page - FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
Source: https://huggingface.co/papers/2606.10671 Published on Jun 9
·
Submitted byhttps://huggingface.co/Simase
YLon Jun 10
Abstract
FadeMem introduces a distance-aware key-value memory consolidation mechanism that organizes historical video data into a temporal hierarchy, improving long-video generation by preserving recent context and long-range anchors under fixed cache constraints.
Autoregressive video generatorssynthesize long videos by generating successive temporal segments, but their historicalKV cachegrows with video length. Existing bounded-cache methods reduce this cost with local windows, sink tokens, or compressed memory states, yet they usually assign fixed roles to different parts of the history. We propose FadeMem, a distance-aware KVmemory consolidationmechanism that organizes historical KV blocks into atemporal hierarchyunder a fixed cache budget. This design is motivated by frequency-dependenttemporal decay: fine details decorrelate quickly, while coarse scene structure and identity remain useful over longer horizons. During generation, new history is inserted as fine-grained entries, while older adjacent entries are progressively merged under a power-law temporal allocation schedule, yielding a dense-near, sparse-far memory within one cache. Without architectural changes, FadeMem preserves recent context for short-term dynamics and compact long-range anchors for identity and scene coherence. Experiments show improvedsubject consistency,background stability, andtemporal coherenceover existing bounded-cache strategies.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2606\.10671
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.10671 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.10671 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.10671 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
The Past Frames the Future: Memory for Autoregressive Video Generation
This paper provides a comprehensive review of memory mechanisms for autoregressive video generation, addressing the challenge of maintaining historical information across extended sequences to ensure temporal persistence.
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
VideoMLA replaces per-head KV caches in video diffusion models with a shared low-rank latent and decoupled 3D-RoPE positional keys, reducing per-token KV memory by 92.7% and improving throughput by 1.23x on a B200 while maintaining quality on VBench benchmarks.
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
This paper introduces Forcing-KV, a hybrid KV cache compression strategy for autoregressive video diffusion models that separates attention heads into static and dynamic categories, achieving up to 2.82x speedup at 1080P resolution while maintaining output quality.
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
DecMem introduces a decoupled memory architecture with Sparse Global Memory and Anchored Local Memory to achieve consistent minute-long video generation, outperforming state-of-the-art methods.
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model
Introduces Flex-Forcing, a unified training and inference framework that allows video diffusion models to operate under both bidirectional and autoregressive regimes via a flexible chunking mechanism over temporal and denoising steps, achieving better video quality, long-video stability, and faster inference.