Tag
The paper presents Puppeteer, a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, posture-aware, and physically grounded gestures.
This paper proposes semantic motion anchors, natural-language abstractions of gesture motion for co-speech gesture retrieval and synthesis. The method discretizes 3D gestures into body-hand motion primitives and grounds them in transcripts, achieving significant improvements in text-to-gesture retrieval and user preference in generation.