Tag
CaM-Wolf is the first multimodal AI agent for social deduction games like Werewolf, integrating video perception and generation with a causal-aware reasoner trained via reinforcement learning to handle hidden roles and social reasoning. Experiments and user studies show improved gameplay and human-AI interaction quality.
Announcement of an open-source SOTA video generation model, MiniMax H3, offering commercial-grade generation, unbeatable cost efficiency, and open weights.
MiniMax launches H3, an open multimodal generation model that handles text, images, video, and audio, generating up to 15 seconds of 2K video with native stereo sound, and plans to open-source the weights.
Comfy-Org repackaged MiniMax-H3 model files for ComfyUI, including diffusion models, text encoders, and VAEs, with workflow templates for text-to-video, image-to-video, and reference-to-video generation.
This paper introduces ODEWorld, a continuous-time latent world model using Physical-Time Flow (PT-Flow) that learns a latent velocity field parameterized by an ordinary differential equation, enabling arbitrary temporal resolution, backward prediction, and improved planning for video generation and robotic control.
PhiZero is a physical world model that learns a compact discrete representation called 'physical language' from videos and uses it to reason about world state transitions before rendering future videos, improving physical coherence in generation and understanding tasks.
NVIDIA introduces Parallel Decoding Distillation (PDD) for accelerating image and video generation, enabling high-quality outputs with fewer neural function evaluations on models like LTX-2.3 and Wan2.1-14B.
LeapTalk proposes a novel framework that overcomes the latency-quality trade-off in talking head generation via single-step bridge distillation, enabling stable real-time generation at up to 200 FPS with reduced identity drift.
Introduces VideoCoCo, an agentic dual-engine framework that uses executable Blender code as a chain-of-thought intermediate representation for physically-consistent video generation, achieving state-of-the-art scores on PhyGenBench and VBench-2.0.
StatePlay proposes a state-aware game world model that jointly predicts visual content and game states to generate mechanics-consistent game rollouts, achieving 18.6% improvement in mechanics fidelity.
Mirage Avatar X by Captions creates a realistic digital twin from just 10 seconds of video, preserving identity, expressions, and mannerisms without quality degradation.
TopviewAI's Film Studio provides six controls for AI video generation, including performance direction, camera control, and 3D blocking, allowing filmmakers to direct shots rather than just prompt clips.
Parallel Decoding Distillation (PDD) is a trajectory-based distillation method that accelerates image and video generation by predicting multiple denoising steps per network evaluation, achieving state-of-the-art performance with 4-8 NFEs on models like LTX-2.3, Wan14B, and Qwen-Image.
Sol-Attn introduces a training-free method to sparsify attention for video generation inference, achieving over 2x speedup by dynamically selecting key-value blocks during online softmax with minimal quality loss.
FilmBench is a new benchmark for cinematic video generation, using prompts derived from award-winning films across 20 genres and a three-level cinematic taxonomy with 35+ sub-metrics. It includes an open-source automatic evaluation agent (FilmOps) that reproduces human model rankings with high correlation, revealing significant gaps in dynamic aesthetics and multi-shot performance compared to prior web-style benchmarks.
NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.
A GitHub repository video-shotcraft provides 106 shot recipes, 162 styles, and 161 dynamic previews organized as Agent Skills for AI video generation using Remotion and compatible with Claude Code and Codex.
LTX-Video is an open-source Python repository by Lightricks for generating and conditioning videos locally using LTX-Video models, with support for text/image inputs, multi-condition workflows, and integration with ComfyUI and Diffusers.
ComfyUI is an open-source, node-based AI creation engine that lets builders and visual professionals design complex generation workflows for image, video, audio, and 3D without coding, with partial re-execution and broad model support.
Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.