Vidu S1: A Real-Time Interactive Video Generation Model
Summary
Vidu S1 is a real-time interactive video generation model that enables voice-controlled digital character animation with infinite-length output and high frame rate on consumer GPUs, achieving state-of-the-art performance.
View Cached Full Text
Cached at: 07/10/26, 06:16 AM
Paper page - Vidu S1: A Real-Time Interactive Video Generation Model
Source: https://huggingface.co/papers/2607.03118 Published on Jul 3
#1 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware.
We introduce Vidu S1, a real-time interactive video generation model supportingvoice controlofdigital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-lengthreal-time video generationwithout blurring, drift, or visual distortion. Built withTurboDiffusionandTurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regularconsumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.
View arXiv pageView PDFProject pageGitHub43Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.03118 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.03118 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.03118 in a Space README.md to link it from this page.
Collections including this paper2
Similar Articles
Video generation models as world simulators
OpenAI's technical report on Sora describes a video generation model that unifies diverse visual data through visual patches, enabling large-scale training of generative models capable of producing high-definition videos up to one minute long across variable durations, aspect ratios, and resolutions.
GraphVid: Interactive Graph-Controllable Video Generation
GraphVid introduces a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs, outperforming prior methods with significant reductions in FID and FVD.
@svpino: This model can generate coherent 1+ hour videos across multiple scenes without skipping a beat. I read their paper so y…
This tweet explains LingBot-World-Infinity, an open-weight video generation model that uses a training technique to recover from errors, enabling coherent hour-long videos across multiple scenes.
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
UltraViT is a latency-optimized vision encoder for large vision-language models, designed for on-device deployment with a pyramidal architecture and a two-stage generative pre-training strategy, achieving state-of-the-art performance at 1.7x speed.
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.