Tag
AlayaVista is a camera-controllable streaming video world model that decouples panoramic scene evolution from perspective video synthesis, supported by the large-scale MUGEN dataset for high-fidelity interactive world modeling.
The paper presents World in World, a training-free interface that enables flexible camera and time control in frozen autoregressive video world models by using correspondence-guided queries and evidence-wise attention guidance.
FlashRender is a few-step generative rendering framework that accelerates video synthesis by aligning representations and using mean-flow objectives, achieving comparable quality to multi-step methods with significantly reduced sampling cost.
Atlas by World Labs is an omni world model that generates camera-controlled HD video from text, images, video, and 3D inputs, reconstructs scenes, and simulates space-time for robotics, available in early access.
World Labs introduces Atlas, a multimodal world model that can generate, reconstruct, and simulate 3D environments with camera control and spatial consistency.
This paper presents JoyAI-Echo-1.5, a unified audio-visual generation system for long-form video and interactive worlds, using cross-shot memory and geometry-aware control to maintain coherence and persistence.
SCoPE is a model from TencentARC that adds camera sightlines as positional coordinates to a pretrained video diffusion transformer, enabling camera trajectory control while preserving the image-to-video prior. The release includes a self-contained checkpoint for Wan2.2-I2V-A14B inference.
TopviewAI's Film Studio provides six controls for AI video generation, including performance direction, camera control, and 3D blocking, allowing filmmakers to direct shots rather than just prompt clips.
Wonder is a general-purpose video world model that enables real-time, camera-controllable world exploration from an image or conditional video. It introduces camera conditioning via dense coordinate fields, a sparse attention memory mechanism, and techniques to improve distillation, allowing minute-scale video generation at 16 FPS.
AlayaWorld is a full-stack, open-source video world model capable of generating 720p, 24 FPS streaming video with camera control, enabling dynamic scene creation.
Introduces ViewSuite, a benchmark with 6DoF camera control and ~165K tasks for evaluating VLMs' ability to plan camera moves. Finds a planning gap where models can track but not compose plans, and proposes View Graph Distillation (RL-Graph-SFT) to improve success from 2.5% to 47.8%.
Holo-World presents a unified controllable video world model that generates videos from a single image with explicit control over camera, object motion, and weather. It introduces a novel dataset and techniques to preserve scene structure while transferring to target weather states.
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model that supports camera navigation, scene persistence, and promptable events across multiple domains, using novel techniques like E-PRoPE, causal forcing, and memory-conditioned scene persistence to achieve controllable long-horizon generation.
Track2View generates novel camera viewpoints from videos by conditioning a video diffusion transformer on paired 3D point tracks, achieving state-of-the-art visual quality and significant reductions in rotation and translation errors.
minWM is a full-stack open-source framework that converts bidirectional video diffusion models into real-time interactive video world models with controllable camera, low-latency rollout, and modular architecture.
Geo-Align presents a reinforcement learning framework for camera-controlled video re-rendering that improves generalization through scale-aware perceptual rewards and metric 3D estimation for camera trajectory extraction.
SANA-WM is a 2.6B-parameter open-source world model that generates high-fidelity 720p minute-scale videos with precise camera control, achieving industrial-level quality while significantly reducing computational requirements.