Tag
MiniMax launches H3, an open multimodal generation model that handles text, images, video, and audio, generating up to 15 seconds of 2K video with native stereo sound, and plans to open-source the weights.
Comfy-Org repackaged MiniMax-H3 model files for ComfyUI, including diffusion models, text encoders, and VAEs, with workflow templates for text-to-video, image-to-video, and reference-to-video generation.
FilmBench is a new benchmark for cinematic video generation, using prompts derived from award-winning films across 20 genres and a three-level cinematic taxonomy with 35+ sub-metrics. It includes an open-source automatic evaluation agent (FilmOps) that reproduces human model rankings with high correlation, revealing significant gaps in dynamic aesthetics and multi-shot performance compared to prior web-style benchmarks.
Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly learns from images, videos, and audio, enabling unified generation and understanding across modalities with early access now available.
Lightricks releases LTX-2.5, an open-weights world model for generating synchronized video and audio from text, image, and video inputs, with features like native multishot generation and a new diffusion video decoder.
Neill Blomkamp released a 13-minute sci-fi short 'Nightborne' made entirely with ByteDance's Seedance 2.0 text-to-video AI. The Verge criticizes it as slop, noting the AI's limitations and lack of creative voice.
This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.
Japan's AIdea Labs released AnimeGen, a free AI model for anime-style video generation, capable of text-to-video and image-to-video, with commercial use allowed.
ByteDance is set to launch Seedance 2.5, an AI video model capable of generating up to 3-minute videos, with features like 30-second scenes, multimodal references, and integration into Dreamina and CapCut.
This is a detailed tutorial on how to create Chinese self-media videos using Claude Code and the Remotion framework. Users only need to provide ideas and prompts, and the AI handles setting up the environment, writing code, generating voiceovers, and rendering the final MP4 video. The tutorial covers the complete process from pipeline setup and script writing to video generation and adjustments.
Mimo is a creator platform that generates short dramas from a single sentence, and includes an operation backend.
On the Codex platform, using plugins like HyperFrames (ideation), ElevenLabs (AI dubbing), Remotion (procedural animation), and Captions (subtitles), you can build an automated pipeline that turns text into video.
Physics Question Scene Graph (PQSG) is a hierarchical question-based pipeline using VLMs to evaluate video generation models' physical plausibility with fine-grained violation detection. It introduces the FinePhyEval dataset and shows higher correlation with human judgments than prior work.
DomainShuttle introduces a method for open domain subject-driven text-to-video generation, achieving high fidelity and flexibility across in-domain and cross-domain scenarios using domain-aware modeling and dual RoPE schemes.
MaineCoon is a 22B real-time text-to-audio-video model that achieves up to 47.5 FPS on a single H100 GPU, enabling low-cost, long-duration streaming with synchronized speech and visuals for live AI characters.
TapVid is a tool that allows users to turn any idea into a motion video.
Pixlie is an AI video studio that enables text and image to video conversion with real control.
LooseControlVideo introduces a framework for intuitive 3D spatial control in text-to-video generation using sparse oriented 3D boxes as proxies, achieving superior trajectory accuracy and occlusion handling. It fine-tunes a Wan 2.2 backbone and demonstrates significant improvements over existing methods on multiple benchmarks.
Introducing the open-source project Pixelle-Video: a fully automated AI short video engine. Input a topic and it automatically generates a video with script, images, voiceover, and background music. Supports local and cloud models, modular design allows flexible replacement of each component model.
ElevenLabs introduces Avatars in ElevenCreative, a dedicated entry point for generating talking-head videos, enabling users to create AI-driven avatar videos with realistic speech and lip-sync.