Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Summary
The paper presents Puppeteer, a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, posture-aware, and physically grounded gestures.
View Cached Full Text
Cached at: 09/10/26, 06:11 AM
Paper page - Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Source: https://huggingface.co/papers/2609.00369
Abstract
Puppeteer is a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, physically grounded gestures.
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, aposture-aware,object-groundedco-speech gesture diffusion modeloperating in a causal latent space. We decompose long gestures into structured primitives and learn acausal variational autoencoderthat encodes them into temporally orderedlatent tokens, each depending only on the past. We then performconditional diffusiondirectly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicittemporal controland supports tasks such asgesture in-betweeningand gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also createdSceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enablingobject-groundedgesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enablingobject-groundedgesture synthesis.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.00369
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.00369 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.00369 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.00369 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
This paper identifies an embodiment gap in humanoid co-speech motion generation caused by human-centric pipelines, and proposes PhysDrift, an embodiment-aware framework that directly predicts executable humanoid joint trajectories from speech, improving speech-motion alignment and physical plausibility.
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
This paper proposes semantic motion anchors, natural-language abstractions of gesture motion for co-speech gesture retrieval and synthesis. The method discretizes 3D gestures into body-hand motion primitives and grounds them in transcripts, achieving significant improvements in text-to-gesture retrieval and user preference in generation.
Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans
The paper proposes Super Star, a real-time framework for online co-speech gesture generation in digital humans using a causal multimodal autoregressive model with streaming speech and user feedback for continual adaptation.
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
Introduces Pegasus, a low-resource framework that translates human demonstration videos into robot-executable data using graph-based task representation, hierarchical affordance latent space, and closed-loop physics verification, aiming to turn hardware data collection into scalable knowledge transfer.
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
CoInteract introduces an end-to-end Diffusion Transformer framework that jointly models RGB appearance and HOI geometry to generate physically-plausible human-object interaction videos with stable hands/faces and zero inference overhead.