long-video

Tag

Cards List
#long-video

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

arXiv cs.AI · 6d ago Cached

SCOUT introduces a recovery-aware agentic framework with an adaptive exploration-exploitation policy for ultra-long egocentric video reasoning, trained via UPS-GRPO, a uncertainty-prioritized RL method with turn-level advantage decomposition. It achieves state-of-the-art results on ultra-long egocentric benchmarks and remains competitive on shorter long-video settings.

0 favorites 0 likes
#long-video

Self Gradient Forcing: Native Long Video Extrapolation

Papers with Code Trending · 2026-07-22 Cached

Proposes Self Gradient Forcing (SGF), a two-pass training strategy for autoregressive video diffusion models that provides missing supervision for writing useful context memory, enabling strong long-video extrapolation even from short training windows.

0 favorites 0 likes
#long-video

@svpino: This model can generate coherent 1+ hour videos across multiple scenes without skipping a beat. I read their paper so y…

X AI KOLs Following · 2026-07-13 Cached

This tweet explains LingBot-World-Infinity, an open-weight video generation model that uses a training technique to recover from errors, enabling coherent hour-long videos across multiple scenes.

0 favorites 0 likes
#long-video

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

Hugging Face Daily Papers · 2026-07-09 Cached

OPSD-V improves few-step autoregressive video diffusion models by using real long-video data as temporal context during training, providing dense trajectory-level supervision that enhances visual quality and motion dynamics without altering inference mechanisms.

0 favorites 0 likes
#long-video

@IndieDevHailey: HKU open-sources ViMax — AI one-person film crew, one-click to produce a complete long-form video! One-sentence idea → automatically generates a finished production with script, storyboard, voiceover, and consistent characters! Solves these pain points: - No more short clips of a few seconds where characters change faces - Multi-agent collaboration: screenwriter writes story, director controls shots, producer ensures consistency, generator outputs video...

X AI KOLs Timeline · 2026-07-06 Cached

HKU open-sourced ViMax, a multi-agent collaborative long video generation tool that can generate a coherent video with script, storyboard, voiceover, and consistent characters from a single sentence, solving issues like short video fragmentation and character inconsistency. The developer also introduced Taste-Skill, a frontend framework that improves the aesthetics of AI-generated interfaces.

0 favorites 0 likes
#long-video

@VraserX: Seedance 2.5 making consistent videos up to 3 minutes in one go is honestly insane. If this pace continues, we’ll be ma…

X AI KOLs Following · 2026-07-03 Cached

Seedance 2.5 enables consistent AI-generated videos up to 3 minutes in one go, marking a significant leap in AI video generation that could soon lead to TV episodes and feature films.

0 favorites 0 likes
#long-video

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

arXiv cs.AI · 2026-06-12 Cached

This paper introduces V-RAGBench, a benchmark for evaluating retrieval-augmented generation over long egocentric videos, and CARVE, a method that adaptively selects retrieval configurations per chunk to improve VideoRAG performance.

0 favorites 0 likes
#long-video

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

arXiv cs.AI · 2026-06-09 Cached

OmniMem introduces a modality-aware memory allocation and perturbation-aware selection strategy for streaming audio-visual LLMs, achieving 2-4% absolute accuracy gains over compression baselines on long-video benchmarks.

0 favorites 0 likes
#long-video

@yukangchen_: Excited to share our new blog: Scaling Video Training with Parallelism https://research.nvidia.com/labs/eai/blogs/scali…

X AI KOLs Following · 2026-06-08 Cached

This blog from NVIDIA Research discusses how sequence parallelism can scale long-video training systems for both understanding and generation, addressing the challenge of fitting very long video sequences across multiple GPUs.

0 favorites 0 likes
#long-video

jdopensource/JoyAI-Echo

Hugging Face Models Trending · 2026-06-02 Cached

JD Open Source releases JoyAI-Echo (Echo-LongVideo), a text-to-audio-video diffusion model capable of generating minute-level multi-shot videos with consistent character identity and voice, using DMD distillation for 7.5x speedup.

0 favorites 0 likes
#long-video

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

Hugging Face Daily Papers · 2026-06-01 Cached

LongLive-RAG formulates long video generation as a retrieval-augmented generation problem, using a dynamic memory of previously generated latents to reduce error accumulation and identity drift, achieving improved quality across multiple autoregressive backbones.

0 favorites 0 likes
#long-video

LVSA: Training-Free Sparse Attention for Long Video Diffusion

Hugging Face Daily Papers · 2026-05-29 Cached

LVSA introduces a training-free sparse attention mechanism for video diffusion models, reducing compute up to 3.17x while enabling generation beyond training horizons without quality loss.

0 favorites 0 likes
#long-video

Keye-VL-2.0-30B-A3B -- Introducing DSA attention into multimodality for the first time

Reddit r/LocalLLaMA · 2026-05-26

Kwai releases Keye-VL-2.0-30B-A3B, a 30B-class multimodal base model that introduces DSA attention to multimodality for the first time, targeting long-video understanding and agent capabilities.

0 favorites 0 likes
#long-video

Real-Time Long Video Generation (GitHub Repo)

TLDR AI · 2026-05-20 Cached

NVlabs releases LongLive 2.0, a parallel infrastructure for real-time long video generation using NVFP4 quantization, supporting both training and inference. It achieves 45.7 FPS and is accepted at ICLR 2026.

0 favorites 0 likes
#long-video

Long Video Generation (4 minute read)

TLDR AI · 2026-05-12 Cached

The article introduces A²RD, a novel architecture for generating consistent long videos using agentic autoregressive diffusion. It proposes a Retrieve–Synthesize–Refine–Update cycle and a new benchmark, LVBench-C, to address semantic drift in long-horizon video synthesis.

0 favorites 0 likes
← Back to home

Submit Feedback