Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Hugging Face Daily Papers Papers

Summary

Vidu S2 presents real-time interactive AI models for avatar and video editing, enabling high-resolution spatial video generation with dynamic reference updates and superior performance over baselines.

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.
Original Article
View Cached Full Text

Cached at: 09/15/26, 02:38 AM

Paper page - Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Source: https://huggingface.co/papers/2609.11638 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Vidu S2 introduces real-time interactive avatar and video editing models that support high-resolution spatial video generation and dynamic reference updates.

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-timevideo editingmodel. Moreover, we explore the feasibility of real-timespatial video generationfor both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation withdynamic referencesthat can be updated at any moment, and strongerinstruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, includingstyle rendering, clothing replacement,character replacement, andbackground replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2609\.11638

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.11638 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.11638 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.11638 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Vidu S1: A Real-Time Interactive Video Generation Model

Hugging Face Daily Papers

Vidu S1 is a real-time interactive video generation model that enables voice-controlled digital character animation with infinite-length output and high frame rate on consumer GPUs, achieving state-of-the-art performance.

Avatar V: Scaling Video-Reference Avatar Video Generation

Hugging Face Daily Papers

Avatar V is a production-scale framework for generating behaviorally recognizable avatar videos conditioned on full video references, introducing sparse reference attention and motion representation streams to achieve state-of-the-art identity preservation and lip synchronization.

Sora 2 is here

OpenAI Blog

OpenAI has released Sora 2, an advanced video generation model representing a significant advancement in AI-powered content creation capabilities.

Video generation models as world simulators

OpenAI Blog

OpenAI's technical report on Sora describes a video generation model that unifies diverse visual data through visual patches, enabling large-scale training of generative models capable of producing high-definition videos up to one minute long across variable durations, aspect ratios, and resolutions.