camera-control

Tag

Cards List
#camera-control

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

Hugging Face Daily Papers · 3d ago Cached

AlayaVista is a camera-controllable streaming video world model that decouples panoramic scene evolution from perspective video synthesis, supported by the large-scale MUGEN dataset for high-fidelity interactive world modeling.

0 favorites 0 likes
#camera-control

World in World: Explore the World with World Models

Hugging Face Daily Papers · 6d ago Cached

The paper presents World in World, a training-free interface that enables flexible camera and time control in frozen autoregressive video world models by using correspondence-guided queries and evidence-wise attention guidance.

0 favorites 0 likes
#camera-control

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

Hugging Face Daily Papers · 2026-09-03 Cached

FlashRender is a few-step generative rendering framework that accelerates video synthesis by aligning representations and using mean-flow objectives, achieving comparable quality to multi-step methods with significantly reduced sampling cost.

0 favorites 0 likes
#camera-control

Atlas by World Labs

Product Hunt · 2026-09-02 Cached

Atlas by World Labs is an omni world model that generates camera-controlled HD video from text, images, video, and 3D inputs, reconstructs scenes, and simulates space-time for robotics, available in early access.

0 favorites 0 likes
#camera-control

Atlas: A World Model for Spatial Intelligence

Hacker News Top · 2026-09-01 Cached

World Labs introduces Atlas, a multimodal world model that can generate, reconstruct, and simulate 3D environments with camera control and spatial consistency.

0 favorites 0 likes
#camera-control

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Hugging Face Daily Papers · 2026-08-24 Cached

This paper presents JoyAI-Echo-1.5, a unified audio-visual generation system for long-form video and interactive worlds, using cross-shot memory and geometry-aware control to maintain coherence and persistence.

0 favorites 0 likes
#camera-control

@_akhaliq: SCoPE Sightline-Coordinate Positional Encoding for Video Diffusion Transformers model: https://huggingface.co/TencentAR…

X AI KOLs Timeline · 2026-08-13 Cached

SCoPE is a model from TencentARC that adds camera sightlines as positional coordinates to a pretrained video diffusion transformer, enabling camera trajectory control while preserving the image-to-video prior. The release includes a self-contained checkpoint for Wan2.2-I2V-A14B inference.

0 favorites 0 likes
#camera-control

@alex_prompter: If you make AI video, you know the loop. Write a prompt, generate, reroll, and repeat until your credits run out. @Topv…

X AI KOLs Timeline · 2026-07-28 Cached

TopviewAI's Film Studio provides six controls for AI video generation, including performance direction, camera control, and 3D blocking, allowing filmmakers to direct shots rather than just prompt clips.

0 favorites 0 likes
#camera-control

Wonder: Video World Model Done Better

Hugging Face Daily Papers · 2026-07-28 Cached

Wonder is a general-purpose video world model that enables real-time, camera-controllable world exploration from an image or conditional video. It introduces camera conditioning via dense coordinate fields, a sparse attention memory mechanism, and techniques to improve distillation, allowing minute-scale video generation at 16 FPS.

0 favorites 0 likes
#camera-control

AlayaWorld is a full-stack, open-source video world model that supports 720p, 24 FPS streaming video generation with camera control

Reddit r/singularity · 2026-07-22

AlayaWorld is a full-stack, open-source video world model capable of generating 720p, 24 FPS streaming video with camera control, enabling dynamic scene creation.

0 favorites 0 likes
#camera-control

@ManlingLi_: Planning with the views: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We int…

X AI KOLs Following · 2026-06-18 Cached

Introduces ViewSuite, a benchmark with 6DoF camera control and ~165K tasks for evaluating VLMs' ability to plan camera moves. Finds a planning gap where models can track but not compose plans, and proposes View Graph Distillation (RL-Graph-SFT) to improve success from 2.5% to 47.8%.

0 favorites 0 likes
#camera-control

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Hugging Face Daily Papers · 2026-06-18 Cached

Holo-World presents a unified controllable video world model that generates videos from a single image with explicit control over camera, object motion, and weather. It introduces a novel dataset and techniques to preserve scene structure while transferring to target weather states.

0 favorites 0 likes
#camera-control

DreamX-World 1.0: A General-Purpose Interactive World Model

Hugging Face Daily Papers · 2026-06-15 Cached

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model that supports camera navigation, scene persistence, and promptable events across multiple domains, using novel techniques like E-PRoPE, causal forcing, and memory-conditioned scene persistence to achieve controllable long-horizon generation.

0 favorites 0 likes
#camera-control

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

Hugging Face Daily Papers · 2026-06-14 Cached

Track2View generates novel camera viewpoints from videos by conditioning a video diffusion transformer on paired 3D point tracks, achieving state-of-the-art visual quality and significant reductions in rotation and translation errors.

0 favorites 0 likes
#camera-control

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

Hugging Face Daily Papers · 2026-05-28 Cached

minWM is a full-stack open-source framework that converts bidirectional video diffusion models into real-time interactive video world models with controllable camera, low-latency rollout, and modular architecture.

0 favorites 0 likes
#camera-control

Geo-Align: Video Generation Alignment via Metric Geometry Reward

Hugging Face Daily Papers · 2026-05-22 Cached

Geo-Align presents a reinforcement learning framework for camera-controlled video re-rendering that improves generalization through scale-aware perceptual rewards and metric 3D estimation for camera trajectory extraction.

0 favorites 0 likes
#camera-control

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

Hugging Face Daily Papers · 2026-05-14 Cached

SANA-WM is a 2.6B-parameter open-source world model that generates high-fidelity 720p minute-scale videos with precise camera control, achieving industrial-level quality while significantly reducing computational requirements.

0 favorites 0 likes
← Back to home

Submit Feedback