FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
Summary
A novel inference-time method for long video generation using overlapping sliding windows with Tweedie matching and stochastic early-phase sampling to improve temporal consistency and visual quality without additional training.
View Cached Full Text
Cached at: 05/22/26, 06:35 AM
Paper page - FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
Source: https://huggingface.co/papers/2605.20910
Abstract
A novel inference-time method for long video generation using overlapping sliding windows with Tweedie matching and stochastic early-phase sampling to improve temporal consistency and visual quality.
Extending the generation horizon ofvideo diffusion modelsto long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions ofbidirectional models, which are tightly coupled to specific architectures and suffer from quality degradation over long horizons, andautoregressive models, which accumulate drift errors due toexposure biasand tend to produce repetitive motion patterns. To address these issues, we propose a novel but simple inference-time approach for longvideo generationthat is architecture-agnostic and requires no additional training. Our method generates long videos via overlappingsliding windows, where predicted clean samples from adjacent windows are blended viaTweedie matchingto enforce both manifold constraint andtemporal consistencyacross overlap regions.Stochastic early-phase samplingthen synchronizes per-window trajectories by injecting fresh noise after eachTweedie matchingcorrection in the high-noise phase, before transitioning todeterministic ODE samplingto preserve fine-grained visual fidelity. Applied to variousvideo generationmodels, our method generates videos several times longer than the native window length while outperforming both training-free and autoregressive baselines intemporal consistencyand visual quality, and further extends toaudio-video joint generationandtext-to-3DGSwithout any fine-tuning.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2605\.20910
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.20910 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.20910 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.20910 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models
CoRe introduces a co-evolving latent reward framework that continually refits the reward model on the generator's current samples while anchoring to real-video preferences, preventing latent reward hacking and quality collapse in video diffusion models like Wan2.1-T2V-1.3B.
Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives
The paper introduces Coffee, a plug-and-play framework that guides discrete diffusion models at inference time using compiled finite-state objectives, avoiding exponential enumeration of token completions while supporting hard constraints and learned soft objectives across symbolic, language, and biological benchmarks.
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Proposes SplitMoE, a split-role sparse architecture to break uniformity in mixture-of-experts for scaling video diffusion models, improving convergence speed and video generation quality.
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
LocUS is a novel method for targeted activation steering in large language models that restricts interventions to specific attention heads and a subspace of the unembedding matrix, improving performance while preserving general capabilities.
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
The paper introduces target-speaker unlearning ASR (TSU-ASR) and proposes a novel Enrollment-Conditioned Gating module for dynamic opt-out of speakers during inference in LLM-based ASR, enhancing privacy in online conferencing.