Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
Summary
The paper proposes Super Star, a real-time framework for online co-speech gesture generation in digital humans using a causal multimodal autoregressive model with streaming speech and user feedback for continual adaptation.
View Cached Full Text
Cached at: 08/27/26, 03:16 AM
Paper page - Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
Source: https://huggingface.co/papers/2608.24909
Abstract
A real-time framework for online co-speech gesture generation uses a causal multimodal autoregressive model with streaming speech and motion history, supported by synthetic dialogue data and continual user-feedback adaptation.
Existingco-speech gesture generationmethods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate onlineco-speech gesture generationfor interactive digital humans and propose a real-time interactive framework that couples astreaming speech responsemodule with anonline gesture generationmodule. Specifically, the gesture generator is designed as acausal multimodal autoregressive modelthat predicts body motion from streaming response speech andmotion history, enabling low-latency and speech-aligned gesture synthesis without access to future speech. To support this setting, we further propose an offline data synthesis pipeline tailored to virtual companion scenarios, which leverages topic- and emotion-aware subject corpora to construct diverse human-agent dialogues and then generates co-speech gestures conditioned on the agent responses. Moreover, to bridge the gap between offline data construction and online deployment, we establish aself-evolving training loopby incorporatinguser feedbackcollected during online interaction into the data generation process, enabling continual adaptation to user preferences. Extensive experiments demonstrate that our framework achieves superior betterlatency-quality trade-off, stronger speech-motion synchronization, and higher user preference than competitive existing baselines. Project Page: https://super-star-2026.github.io/
View arXiv pageView PDFProject pageGitHub5Add to collection
Get this paper in your agent:
hf papers read 2608\.24909
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.24909 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.24909 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.24909 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
ARDY introduces a streaming generation framework for real-time, high-fidelity 3D human motion generation controlled by text and kinematic constraints, using a hybrid representation and two-stage autoregressive transformer denoiser.
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation
This paper introduces Agentic ASR, an interactive speech recognition framework that uses semantic correction and reasoning-based editing to reduce semantic errors through multi-turn refinement. It also proposes a new sentence-level semantic error rate metric and an interactive simulation system for benchmarking.
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Wan-Streamer is a unified end-to-end multimodal model for real-time audio-visual interaction using causal attention and integrated processing of visual, audio, and text modalities, achieving sub-second latency.
Are AI social apps moving from text chat to real-time video interfaces?
A discussion about the evolution of AI social apps from text chat to real-time video interfaces, highlighting Mel's multimodal interaction stack and the technical challenges of latency, lip sync, and orchestration.
Higgsfield just launched what they call the first fully automated AI agent for video - real shift or just another hype?
Higgsfield launched Supercomputer, described as the first fully automated AI agent for end-to-end video creation, capable of planning, generating, and distributing multi-minute videos from a single chat interface, though currently buggy with coherence issues in longer outputs.