Tag
The paper proposes Super Star, a real-time framework for online co-speech gesture generation in digital humans using a causal multimodal autoregressive model with streaming speech and user feedback for continual adaptation.