Tag
SCoPE is a model from TencentARC that adds camera sightlines as positional coordinates to a pretrained video diffusion transformer, enabling camera trajectory control while preserving the image-to-video prior. The release includes a self-contained checkpoint for Wan2.2-I2V-A14B inference.
Introduces Blur2Vid, a method that generates video from motion-blurred images, published at SIGGRAPH Asia 2025.
Hugging Face page for the Minimax-h3-Turbo video generation model, with instructions for using it via Diffusers and Colab/Kaggle notebooks.
MiniMax released the open-weights H3 video model, the first open model to top an AI video ranking, ranking first in video editing and second in text-to-video. The 33B parameter model handles text, images, video, and audio, with some components like 2K resolution and H3-Context-IR remaining closed.
MiniMax H3 tops the Artificial Analysis video editing leaderboard, offering instruction-based editing on existing clips with multimodal input and competitive pricing, while planning open weights for commercial use.
Comfy-Org repackaged MiniMax-H3 model files for ComfyUI, including diffusion models, text encoders, and VAEs, with workflow templates for text-to-video, image-to-video, and reference-to-video generation.
Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly learns from images, videos, and audio, enabling unified generation and understanding across modalities with early access now available.
GraphVid introduces a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs, outperforming prior methods with significant reductions in FID and FVD.
Japan's AIdea Labs released AnimeGen, a free AI model for anime-style video generation, capable of text-to-video and image-to-video, with commercial use allowed.
This paper presents CineMobile, a method for efficient on-device image-to-video generation that achieves a 40x speedup over the teacher model through distillation-guided pruning, diffusion distillation, and hybrid quantization, enabling cinematic camera motion effects on mobile devices.
Pixlie is an AI video studio that enables text and image to video conversion with real control.
Go-with-the-Track unifies motion control and reference image compositing in video generation using point-track embeddings with spatial-aware encoding and video diffusion transformers, achieving superior motion and reference control in a single model.
PhaseLock is a training-free framework that preserves motion priors from early-step inference to improve physical consistency in image-to-video diffusion models, achieving 6.2 point improvement with minimal overhead.
AAD-1 introduces asymmetric adversarial distillation with phased training to achieve one-step autoregressive video generation, outperforming prior methods on VBench.
ByteDance released Bernini, an open-source video generation and editing model on Hugging Face that rivals top closed-source models.
Grok-Imagine-Video-1.5-Preview (720p) has secured first place in the Image-to-Video Arena, improving by 52 points over its predecessor and surpassing top video models like Seedance-2.0 and HappyHorse.
NVIDIA releases Cosmos3-Super-Image2Video, a model that generates temporally coherent video sequences from an input image and text instructions, part of the Cosmos 3 omnimodal world model platform for Physical AI applications.
SANA-WM is an efficient 2.6B-parameter open-source world model for minute-scale video generation with precise camera control. It uses a hybrid linear diffusion transformer and a two-stage pipeline to produce 720p videos from images and text prompts.
Dhee is a new agentic video generation AI that creates videos from a single description, handling prompting, image generation, and assembly, with shot-by-shot editing for refinement.
Grok AI now offers image generation, X search, and image-to-video capabilities to all subscribers with a normal subscription.