Tag
BFL has introduced FLUX 3, a multi-modal AI model capable of generating images, videos, and audio.
seedance2 is an AI model that converts depth-video inputs into full video outputs.
This paper introduces a training-free method to improve revisit consistency in autoregressive generative rendering by using temporal and spatial correspondences from the 3D engine to maintain consistent appearance when the camera revisits locations.
SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.
GraphVid introduces a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs, outperforming prior methods with significant reductions in FID and FVD.
AlayaWorld is a full-stack, open-source video world model capable of generating 720p, 24 FPS streaming video with camera control, enabling dynamic scene creation.
Proposes Self Gradient Forcing (SGF), a two-pass training strategy for autoregressive video diffusion models that provides missing supervision for writing useful context memory, enabling strong long-video extrapolation even from short training windows.
Japan's AIdea Labs released AnimeGen, a free AI model for anime-style video generation, capable of text-to-video and image-to-video, with commercial use allowed.
ShotPlan introduces learnable planning tokens with fractional temporal rotary position embeddings for cinematic multi-shot video generation, enabling explicit shot-level planning and achieving superior inter-shot consistency.
HOMIE is a framework for human-object centric video personalization that integrates MLLM features to improve subject fidelity and interaction patterns, achieving state-of-the-art performance.
HarmoHOI is a unified diffusion framework that jointly generates synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks, achieving state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency.
FVAttn is a training-free sparse attention system that uses runtime load balancing to improve distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism for video generation, achieving up to 4.41x attention speedup and 2.11x inference speedup over FlashAttention on step-distilled Wan2.2 I2V while maintaining competitive video quality.
Apple-π is a benchmark that evaluates video generation models on their ability to reason about physical laws through a three-stage protocol: perception, formulation, deduction. It includes 400 videos covering classical mechanics tasks and reveals current models fall short of reliable law-grounded world simulation.
Kimi K3 has been released, showcased in a video created using the model.
Google Vids introduces Gemini Omni for generating and editing videos via natural language, and personal avatars that let users create digital versions of themselves from a selfie and voice recording without filming.
Kimi K3, a new AI model, can generate complete motion graphic videos from a single prompt without any tools or visual feedback, outperforming previous methods.
html-video is an open-source tool that renders HTML+CSS+data into high-quality MP4 videos, supports AI voiceover and background music, has 21 built-in templates, and is compatible with mainstream AI agents.
MeanFlowNFT introduces a forward-process reinforcement learning method for average-velocity generators, enabling efficient alignment with human preferences while preserving fast few-step sampling. Experiments show it outperforms prior RL-tuned few-step generators on most metrics and even surpasses multi-step RL-tuned diffusion models.
HDR is a unified framework integrating hierarchical latents into causal video generation for multi-step visual reasoning, achieving better reasoning consistency, lower latency, and strong data efficiency compared to baselines.
This video was entirely produced by AI agents within Raft, with only human direction, showcasing agent-driven content creation.