PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Summary
This paper identifies an embodiment gap in humanoid co-speech motion generation caused by human-centric pipelines, and proposes PhysDrift, an embodiment-aware framework that directly predicts executable humanoid joint trajectories from speech, improving speech-motion alignment and physical plausibility.
View Cached Full Text
Cached at: 06/20/26, 02:34 PM
# Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Source: [https://arxiv.org/html/2606.19935](https://arxiv.org/html/2606.19935)
Zhangzhao Liang1, Xiaofen Xing11, Mingyue Yang2, Wenlve Zhou3, Xiangmin Xu13 1South China University of Technology2DexForce Technology3Foshan University 1Corresponding Author
###### Abstract
Humanoid robots require co\-speech motions that are not only expressive and speech\-aligned, but also physically executable under embodiment constraints\. Existing co\-speech generation pipelines are predominantly human\-centric: motions are first generated in human\-body representations such as SMPL\-X and subsequently retargeted to humanoid robots\. In this work, we identify a fundamental embodiment gap in this paradigm, where the mismatch between human motion manifolds and humanoid embodiment constraints disrupts embodiment consistency during motion transfer and physical execution\. Through extensive analysis, we show that although retargeting can preserve coarse motion semantics, it significantly compresses motion diversity and weakens prosody\-motion synchronization, limiting expressive humanoid behaviors\. To address this problem, we first propose IK\-EER, a prosody\-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech\-motion temporal alignment during retargeting\. Building upon the curated robot\-native motion dataset, we further introduce PhysDrift, an embodiment\-aware co\-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human\-body representations\. Unlike conventional human\-centric pipelines, PhysDrift maintains embodiment consistency throughout both training and inference while incorporating physical regularization to stabilize robot motion dynamics\. Extensive experiments and real\-world humanoid deployment demonstrate that embodiment\-aware robot\-native generation substantially improves speech\-motion alignment, physical plausibility, motion smoothness, inference efficiency, and real\-time interaction capability\. The results further reveal that robot\-native motion representations are fundamentally more suitable than human\-centric intermediates for embodied co\-speech interaction in humanoid robots\.
Figure 1:From human motion capture data to humanoid robot co\-speech motion\. The entire pipeline consists of four steps\. First, data that clearly violates physical laws or severely distorts limbs in human motions is filtered out\. Next, the proposed IK\-EER maps human motions onto humanoid robots to obtain co\-speech motions in the robot’s native motion space, while further removing data that does not match the robot’s attributes\. Then, a PhysDrift generation model is trained using the robot’s co\-speech motions\. Finally, the whole\-body controller takes the motion generated by PhysDrift as a reference to execute actions on the real robot\.## IIntroduction
Humanoid robots are expected to engage in natural face\-to\-face interaction with humans\[[4](https://arxiv.org/html/2606.19935#bib.bib43),[11](https://arxiv.org/html/2606.19935#bib.bib44),[9](https://arxiv.org/html/2606.19935#bib.bib45),[42](https://arxiv.org/html/2606.19935#bib.bib46),[22](https://arxiv.org/html/2606.19935#bib.bib47)\], where speech is tightly coupled with expressive body motion\. Co\-speech motions play a critical role in this process by conveying emphasis, rhythm, emotion, and conversational intent beyond verbal content alone\. Recent advances in generative modeling have substantially improved the realism and diversity of human co\-speech motion synthesis\[[38](https://arxiv.org/html/2606.19935#bib.bib67),[13](https://arxiv.org/html/2606.19935#bib.bib68)\]\. These approaches demonstrate impressive capability in learning speech\-motion correspondence from large\-scale human motion datasets and have shown promising results for virtual avatars and digital humans\.
Despite this progress, transferring co\-speech motion generation from virtual humans to physical humanoid robots remains fundamentally challenging\. Similar to the motion of general humanoid robots, existing pipelines\[[47](https://arxiv.org/html/2606.19935#bib.bib69),[44](https://arxiv.org/html/2606.19935#bib.bib64)\]are predominantly human\-centric: motion is first generated in human\-body representations such as SMPL\-X\[[36](https://arxiv.org/html/2606.19935#bib.bib36)\]and subsequently retargeted onto humanoid embodiments through inverse kinematics or optimization\-based motion transfer\. While effective for animation, this paradigm implicitly assumes that human motion representations are compatible with humanoid embodiment constraints\. However, humanoid robots differ substantially from humans in kinematic structure, joint limits, actuation capability, balance constraints, and feasible motion manifolds\. As a result, motion representations learned in human\-centric latent spaces are not naturally aligned with the physically executable motion space of humanoid robots\[[20](https://arxiv.org/html/2606.19935#bib.bib56)\]\.
In this work, we identify and formalize this discrepancy as anembodiment gapin humanoid co\-speech motion generation\. Unlike prior works that primarily focus on motion retargeting accuracy or kinematic feasibility, we show that the embodiment gap manifests more fundamentally as a distribution mismatch between human motion manifolds and robot\-executable motion manifolds\. Our experiments reveal that although modern retargeting methods can largely preserve coarse motion semantics and avoid severe joint violations, the retargeting process significantly compresses motion diversity and weakens prosody\-motion synchronization\. Consequently, expressive speech\-driven motions learned in human\-centric representations become progressively distorted during humanoid reconstruction and physical execution\.
A natural solution is to abandon human\-centric intermediate representations and directly model humanoid co\-speech behavior in robot joint space\. However, robot\-native generation introduces a new challenge\. Without embodiment\-aware constraints, highly expressive generative models, particularly flow\-based models\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\], can easily produce physically unstable motions with excessive jerk, contact artifacts, and joint\-limit violations, despite achieving strong distributional metrics\. Therefore, effective humanoid co\-speech motion generation requires not only expressive generative capability, but also embodiment\-consistent physical regularization\.
To address these challenges, we propose PhysDrift, an embodiment\-aware robot\-native framework for humanoid co\-speech motion generation\. Instead of relying on intermediate human\-body representations, PhysDrift directly predicts executable humanoid joint trajectories from speech\. To construct high\-quality robot\-native training data, we further introduce Inverse Kinematics\-Energy Envelope Retargeting \(IK\-EER\), a prosody\-preserving humanoid motion curation framework that jointly optimizes kinematic feasibility and speech\-motion temporal alignment during retargeting\. Building upon the curated dataset, PhysDrift incorporates embodiment\-aware regularization to stabilize generated motion dynamics while preserving speech\-motion alignment and real\-time generation capability\. Fig\.[1](https://arxiv.org/html/2606.19935#S0.F1)provides an overview of the research presented in this paper\.
Extensive experiments across motion quality, physical feasibility, speech alignment, and deployment efficiency demonstrate that embodiment\-aware robot\-native generation substantially outperforms conventional human\-centric pipelines\. In particular, our method achieves superior speech\-motion synchronization, smoother physical dynamics, significantly faster inference speed, and more stable humanoid execution while maintaining expressive motion diversity\. We further validate the proposed framework through real\-world humanoid deployment, demonstrating robust real\-time co\-speech interaction capability\.
Our contributions are summarized as follows:
- •We identify and formalize the embodiment gap in humanoid co\-speech motion generation, showing that human\-centric motion representations fundamentally disrupt embodiment consistency during retargeting and physical execution\.
- •We propose IK\-EER, a prosody\-preserving humanoid motion curation framework that jointly considers kinematic feasibility and speech\-motion temporal alignment during retargeting\.
- •We propose PhysDrift, an embodiment\-aware robot\-native co\-speech motion generation framework that directly predicts executable humanoid joint trajectories from speech without relying on intermediate human\-body representations\.
- •We demonstrate through extensive experiments and humanoid deployment that embodiment\-aware generation substantially improves physical plausibility, motion smoothness, speech alignment, inference efficiency, and real\-time interaction capability for humanoid co\-speech motion\.
## IIRelated Work
In this section, we review prior work from three perspectives closely related to humanoid co\-speech motion generation: human\-centric co\-speech generation, motion retargeting for humanoid robotics, and robot\-native humanoid motion generation\. Unlike conventional categorizations, we particularly focus on how existing methods model the relationship between motion representation and physical embodiment, which forms the core challenge addressed in this work\.
### II\-AHuman\-Centric Co\-Speech Motion Generation
Recent progress in co\-speech motion generation has been primarily driven by the computer vision and graphics communities\. Existing approaches typically formulate the task as generating human body motion conditioned on speech signals, where motions are represented in human skeletal spaces or parametric body models such as SMPL\-X\. Large\-scale conversational motion datasets have substantially accelerated this research direction\. BEAT\[[30](https://arxiv.org/html/2606.19935#bib.bib8)\]and BEAT2\[[29](https://arxiv.org/html/2606.19935#bib.bib9)\]provide multilingual conversational motion capture datasets with synchronized speech and body motion annotations\. ZeroEGGS\[[16](https://arxiv.org/html/2606.19935#bib.bib10)\]further contributes high\-fidelity expressive speaking styles for gesture synthesis\. In parallel, large\-scale Internet video datasets such as AVSpeech\[[12](https://arxiv.org/html/2606.19935#bib.bib11)\]and the TED Gesture Dataset\[[48](https://arxiv.org/html/2606.19935#bib.bib24)\]enable scalable learning of speech\-driven gestures from unconstrained audio\-visual data\. Building upon these datasets, recent generative models have achieved remarkable progress in producing realistic and semantically meaningful human gestures\. HOP\[[6](https://arxiv.org/html/2606.19935#bib.bib48)\]models multimodal interactions among speech, text, and gesture dynamics\. SemGes\[[31](https://arxiv.org/html/2606.19935#bib.bib49)\]improves semantic consistency through local\-global constraints\. DIDiffGes\[[8](https://arxiv.org/html/2606.19935#bib.bib50)\]introduces an efficient decoupled diffusion framework for gesture generation, while MotionCraft\[[3](https://arxiv.org/html/2606.19935#bib.bib51)\]employs unified diffusion transformers for multimodal whole\-body motion synthesis\. Despite their impressive visual quality, these methods are fundamentally designed for digital humans rather than physically embodied humanoid robots\. More importantly, their learned motion representations are intrinsically human\-centric, assuming motion manifolds defined by human kinematics and morphology\. Such representations do not explicitly account for embodiment\-specific constraints including robot joint structures, actuator limitations, balance constraints, or physically executable humanoid motion spaces\. Consequently, transferring these generated motions onto humanoid robots inevitably requires an additional embodiment transfer stage, introducing a mismatch between human motion representations and robot\-executable behaviors\.
### II\-BMotion Retargeting for Humanoid Robotics
Motion retargeting aims to transfer human motion onto humanoid embodiments while preserving motion semantics and physical feasibility\. In humanoid robotics, retargeting has become a practical strategy for leveraging large\-scale human motion priors to improve robot motion generation and control\. Traditional humanoid motion generation approaches\[[23](https://arxiv.org/html/2606.19935#bib.bib25),[14](https://arxiv.org/html/2606.19935#bib.bib26),[40](https://arxiv.org/html/2606.19935#bib.bib27),[24](https://arxiv.org/html/2606.19935#bib.bib28),[39](https://arxiv.org/html/2606.19935#bib.bib29)\]mainly rely on trajectory optimization or manually designed controllers\. Although these methods achieve stable motion control, they often struggle to reproduce expressive and socially meaningful human motion dynamics\. To address this limitation, recent retargeting frameworks attempt to transfer natural human movements onto humanoid embodiments\. Recent studies have significantly improved the physical plausibility of retargeted humanoid motion\. Jeong et al\.\[[19](https://arxiv.org/html/2606.19935#bib.bib31)\]proposed a standardized retargeting framework addressing self\-collision and contact consistency\. Exbody\[[7](https://arxiv.org/html/2606.19935#bib.bib33)\]introduced local joint mapping strategies for transferring human motion onto the Unitree H1 humanoid platform\. Lu et al\.\[[33](https://arxiv.org/html/2606.19935#bib.bib32)\]refined inverse\-kinematics\-based upper\-body retargeting to preserve natural motion characteristics under embodiment differences\. Mao et al\.\[[35](https://arxiv.org/html/2606.19935#bib.bib30)\]further demonstrated scalable retargeting pipelines for converting Internet\-scale human motions into executable humanoid datasets\. However, existing retargeting methods primarily optimize spatial pose reconstruction and physical feasibility while paying limited attention to embodiment consistency in speech\-driven interaction\. In particular, co\-speech motions depends not only on pose accuracy but also on subtle temporal coupling between speech prosody and motion dynamics\. Retargeting processes based on frame\-wise inverse kinematics, smoothing, or projection into feasible robot subspaces often distort expressive motion distributions and weaken prosody\-motion synchronization\. As a result, motions that remain physically executable after retargeting may still lose expressive diversity and conversational naturalness during humanoid interaction\.
### II\-CRobot\-Native Humanoid Motion Generation
More recent research has begun exploring robot\-native humanoid motion generation frameworks that operate directly within robot action spaces instead of relying entirely on human\-body intermediates\. These methods aim to incorporate embodiment constraints during generation and control more explicitly\. Several approaches explore language\-conditioned humanoid generation and interaction\. Harmon\[[21](https://arxiv.org/html/2606.19935#bib.bib37)\]combines human motion priors with vision\-language reasoning for whole\-body humanoid generation\. Xu et al\.\[[45](https://arxiv.org/html/2606.19935#bib.bib38)\]proposed a text\-driven humanoid motion generation framework on the NAO platform using angle\-space representations and reinforcement learning controllers\. Bao et al\.\[[2](https://arxiv.org/html/2606.19935#bib.bib39)\]introduced a hierarchical framework integrating intention reasoning and diffusion\-based social gesture generation for humanoid interaction\. Parallel advances have also emerged in imitation learning and diffusion\-policy\-based humanoid control\. Ze et al\.\[[50](https://arxiv.org/html/2606.19935#bib.bib40)\]combined teleoperation and diffusion policies for full\-body humanoid skill learning\. HOVER\[[17](https://arxiv.org/html/2606.19935#bib.bib41)\]proposed a neural whole\-body humanoid controller, while ManiDP\[[27](https://arxiv.org/html/2606.19935#bib.bib42)\]introduced manipulability\-aware diffusion policies for bimanual humanoid manipulation\. Although these methods operate more closely to robot embodiments, most focus on locomotion, manipulation, or general task\-oriented motion rather than speech\-driven social interaction\. Moreover, many approaches still rely on human\-derived motion priors, reference trajectories, or intermediate representations during training\. Consequently, the problem of jointly preserving speech alignment, expressive motion dynamics, and embodiment\-consistent physical execution remains insufficiently explored in humanoid co\-speech motion generation\.
### II\-DSummary
Existing co\-speech motion generation methods predominantly rely on human\-centric motion representations and post\-hoc retargeting pipelines, while existing humanoid motion generation approaches mainly focus on task\-oriented behaviors rather than expressive speech\-driven interaction\. As a result, current methods lack an explicit mechanism for maintaining embodiment consistency between speech prosody, expressive motion dynamics, and physically executable humanoid behavior\. In contrast, this work proposes an embodiment\-aware robot\-native co\-speech generation framework that jointly considers humanoid motion curation, embodiment\-consistent motion representation, and physically grounded speech\-driven generation directly within humanoid joint space\.
Figure 2:The feasible region of the end effectors and the joint rotation space between SMPL\-X and Unitree\-G1 are illustrated\. In \(a\), the end effectors are the wrist and foot, which are common to both G1 and SMPL\-X\. \(b\) shows the joint rotation space of the shoulders, wrists, and hips for SMPL\-X and Unitree\-G1\.
## IIIEmbodiment Gap Analysis and Problem Formulation
Embodiment Gap\.The primary distinction between SMPL\-X pose representation and robotic joint configuration lies in the nature of rotational freedom and physical constraints\. For a given joint, SMPL\-X uses an axis\-angle representation:
𝐫=θ𝐮^,\\mathbf\{r\}=\\theta\\mathbf\{\\hat\{u\}\},\(1\)where𝐮^∈ℝ3\\mathbf\{\\hat\{u\}\}\\in\\mathbb\{R\}^\{3\}denotes the unit rotation axis satisfying‖𝐮^‖=1\\\|\\mathbf\{\\hat\{u\}\}\\\|=1,∥⋅∥\\\|\\cdot\\\|denotes the Euclidean norm, andθ\\thetais the rotation magnitude\. This formulation allows arbitrary 3D orientations without consideration of mechanical limits or collisions, effectively sampling from the mathematical rotation groupSO\(3\)SO\(3\)\. Consequently, SMPL\-X can represent poses that are physically impossible for a real human or a robot, including extreme rotations or self\-intersecting limbs\.
In contrast, a robot’s joint configuration is defined as
qi∈\[qimin,qimax\],i=1,…,Njoints,q\_\{i\}\\in\[q\_\{i\}^\{\\min\},q\_\{i\}^\{\\max\}\],\\quad i=1,\\dots,N\_\{\\text\{joints\}\},\(2\)where eachqiq\_\{i\}corresponds to a mechanically constrained degree of freedom \(DoF\), often restricted to rotation about a fixed axis\.qiminq\_\{i\}^\{\\min\}andqimaxq\_\{i\}^\{\\max\}denote the rotation range of the joint\. These limits arise from actuator capabilities, structural design, collision avoidance, and stability requirements\. Unlike SMPL\-X, robotic joints cannot arbitrarily rotate, and multiple axis rotations require solving inverse kinematics under constraints\.
The embodiment gap is summarized as Table[I](https://arxiv.org/html/2606.19935#S3.T1)highlighting that while SMPL\-X represents idealized human\-like orientations, robotic joints encode physically realizable actuator configurations\.
TABLE I:The difference in body structure between SMPL\-X and Unitree\-G1Body\# JointsDoFRotation Rep\.Rotation RestrictSMPL\-X24 Body \+ 30 Fingers162Axis\-Angle×\\timesG11 Root \+ 29 Body29Joint Angle✓
Problem Formulation\.Humanoid co\-speech motion generation aims to synthesize expressive body motion conditioned on speech signals while maintaining physical executability on humanoid platforms\[[29](https://arxiv.org/html/2606.19935#bib.bib9),[21](https://arxiv.org/html/2606.19935#bib.bib37)\]\. Given an input speech sequenceA=\{at\}t=1T,A=\\\{a\_\{t\}\\\}\_\{t=1\}^\{T\},the goal is to generate a corresponding humanoid motion sequenceQ=\{qt\}t=1TQ=\\\{q\_\{t\}\\\}\_\{t=1\}^\{T\}, whereqt∈ℝDq\_\{t\}\\in\\mathbb\{R\}^\{D\}denotes the humanoid joint configuration at time steptt, including whole\-body joint rotations and root motion parameters\.
Due to the lack of in\-depth research on the co speech motion of native humanoid robots\. Therefore, a more widespread approach is first generate motion in human\-centric representations such as skeletal pose spaces or parametric body models, and subsequently transfer the generated motions onto humanoid embodiments through retargeting pipelines\. This process can be summarized as
A→Stage 1H→Stage 2Q,A\\xrightarrow\{\\text\{Stage 1\}\}H\\xrightarrow\{\\text\{Stage 2\}\}Q,\(3\)whereHHdenotes motion represented in a human\-body joint space and the second stage corresponds to humanoid retargeting\.
Although this paradigm has demonstrated strong visual performance for digital humans, it implicitly assumes that human motion representations are compatible with humanoid embodiment constraints\. However, humanoid robots differ fundamentally from humans in morphology, kinematic structure, joint limits, actuation capability, and balance dynamics\[[7](https://arxiv.org/html/2606.19935#bib.bib33),[41](https://arxiv.org/html/2606.19935#bib.bib70)\]\. As a result, the feasible humanoid motion space does not coincide with the motion manifold learned from human motion data\. Shown in Fig\.[2](https://arxiv.org/html/2606.19935#S2.F2)are the feasible region of the end effectors and the joint rotation manifold between SMPL\-X and Unitree\-G1\. Bridging this gap requires mapping unconstrained human poses to constrained robot motions, typically via inverse kinematics with joint limits and collision avoidance\. We denote the human motion manifold asℳh\\mathcal\{M\}\_\{h\}and the physically executable humanoid motion manifold asℳr\\mathcal\{M\}\_\{r\}\. Existing human\-centric generation methods effectively learn distributions overℳh\\mathcal\{M\}\_\{h\}, while humanoid execution is constrained withinℳr\\mathcal\{M\}\_\{r\}\. Sinceℳh≠ℳr\\mathcal\{M\}\_\{h\}\\neq\\mathcal\{M\}\_\{r\}, retargeting inevitably becomes a projection process from human motion space into the feasible humanoid motion subspace:
ℛ:ℳh→ℳr\.\\mathcal\{R\}:\\mathcal\{M\}\_\{h\}\\rightarrow\\mathcal\{M\}\_\{r\}\.\(4\)This discrepancy forms what we define as theembodiment gapin humanoid co\-speech motion generation\.
Importantly, the embodiment gap does not necessarily manifest as severe kinematic failure\. Modern retargeting methods can often maintain basic physical feasibility and preserve coarse motion semantics\. However, the projection fromℳh\\mathcal\{M\}\_\{h\}toℳr\\mathcal\{M\}\_\{r\}inevitably compresses expressive motion distributions and alters fine\-grained temporal dynamics\. In co\-speech interaction, where gesture rhythm, motion energy, and speech prosody are tightly coupled, such distortions accumulate into perceptually significant degradation of conversational naturalness\. This phenomenon is particularly evident in expressive gestures\. Human co\-speech motions often rely on subtle upper\-body coordination, asymmetric arm dynamics, and temporally localized motion emphasis synchronized with speech rhythm\[[37](https://arxiv.org/html/2606.19935#bib.bib71),[15](https://arxiv.org/html/2606.19935#bib.bib72)\]\. During retargeting, these high\-frequency expressive components are frequently smoothed or projected into more conservative feasible robot motions, reducing motion diversity and weakening prosody\-motion synchronization even when the resulting motions remain physically executable\.
Our experimental analysis further confirms this observation\. While retargeted motions retain relatively similar semantic alignment scores, motion diversity decreases substantially after retargeting across multiple generation methods\. These results suggest that the primary limitation of human\-centric pipelines is not merely physical feasibility, but the loss of embodiment consistency between expressive speech\-driven motion dynamics and humanoid execution constraints\.
A natural solution is therefore to directly model humanoid co\-speech motion within robot\-native action spaces\[[21](https://arxiv.org/html/2606.19935#bib.bib37)\]\. Instead of learning human motion distributions followed by embodiment transfer, we aim to directly learn the conditional distributionp\(Q\|A\)p\(Q\|A\), where motion generation and embodiment constraints are jointly modeled within humanoid joint space\. This formulation motivates the proposed PhysDrift framework, which combines embodiment\-aware motion curation and robot\-native speech\-conditioned motion generation to preserve both expressive dynamics and physical executability\.
## IVMethod
### IV\-AOverview
The embodiment mismatch manifests in two aspects\. First, human\-centric representations contain large regions that are unreachable for humanoid embodiments, leading to unstable or over\-smoothed motions after retargeting\. Second, frame\-wise inverse kinematics and motion smoothing operations distort subtle temporal patterns synchronized with speech rhythm and emphasis, which are essential for natural co\-speech behavior\.
To address these issues, we reformulate humanoid co\-speech motion generation as a direct robot\-space generation problem\. Instead of treating robotic embodiment as a post\-processing constraint, embodiment consistency is explicitly modeled during both dataset construction and motion generation\. Our framework consists of two tightly coupled components\. First, we propose IK\-EER, an embodiment\-aware humanoid motion construction framework that converts human co\-speech motions into robot\-native supervision while preserving speech\-motion synchronization and physical feasibility\. Second, we introduce PhysDrift, a robot\-native speech\-driven motion generation framework that directly learns executable humanoid motion distributions in joint space without relying on intermediate human\-body representations\.
### IV\-BEmbodiment\-Aware Motion Construction
Figure 3:Pipeline of IK\-EER\. “Joint DoF” and “Root Position” are optimized through gradient\-based motion refinement, while the root orientation is directly inherited from the human motion sequence\.Constructing a high\-quality humanoid co\-speech motion dataset is critical for robot\-native motion generation\. However, due to the lack of available robot co\-speech motion, a common method is to retarget from human motion data to the robot body\. Existing retargeting pipelines\[[1](https://arxiv.org/html/2606.19935#bib.bib52),[49](https://arxiv.org/html/2606.19935#bib.bib53),[34](https://arxiv.org/html/2606.19935#bib.bib60)\]primarily optimize spatial pose reconstruction fidelity, implicitly assuming that physically feasible reconstruction is sufficient for downstream learning\. However, for co\-speech motions, preserving temporal synchronization between speech prosody and motion dynamics is equally important\. Motions that are physically executable but temporally inconsistent with speech often appear socially unnatural during interaction\.
To address this issue, we propose IK\-EER, an embodiment\-aware motion construction framework designed not merely for pose transfer, but for generating robot\-native supervision suitable for humanoid co\-speech learning\. The details of IK\-EER is shown in Fig\.[3](https://arxiv.org/html/2606.19935#S4.F3)\.
Starting from the BEAT2 dataset, We removed physically infeasible sequences from the source data \(such as floating, twisted limbs, ground penetration\)\. Instead of directly adopting SMPL\-X parameters as training targets, all motions are transformed into the native joint space of the humanoid robot\. For each frame, robot joint configurations are initialized through sparse human\-to\-robot keypoint correspondences using inverse kinematics\. Let𝐪\\mathbf\{q\}denote robot joint angles\. The initialization objective is formulated as
min𝐪∑iwip‖𝐩itarget−𝐩i\(𝐪\)‖2\+wir‖𝐑itarget−𝐑i\(𝐪\)‖F2,\\min\_\{\\mathbf\{q\}\}\\sum\_\{i\}w\_\{i\}^\{p\}\\\|\\mathbf\{p\}\_\{i\}^\{target\}\-\\mathbf\{p\}\_\{i\}\(\\mathbf\{q\}\)\\\|^\{2\}\+w\_\{i\}^\{r\}\\\|\\mathbf\{R\}\_\{i\}^\{target\}\-\\mathbf\{R\}\_\{i\}\(\\mathbf\{q\}\)\\\|\_\{F\}^\{2\},\(5\)where𝐩i\\mathbf\{p\}\_\{i\}and𝐑i\\mathbf\{R\}\_\{i\}denote the position and orientation of the corresponding robot jointi\.wiw\_\{i\}is the weighting factor andtargetdenotes the target human motion position or orientation\.
Although inverse kinematics provides feasible initialization, conventional retargeting pipelines often weaken the intrinsic coupling between speech rhythm and motion dynamics due to smoothing and kinematic corrections\. Human co\-speech motions naturally exhibit strong correlations between acoustic emphasis and motion intensity\. However, this relationship is highly sensitive to temporal distortion during embodiment projection\.
To preserve this synchronization, we introduce an Energy Envelope \(EE\) objective\. The synchronization objective is defined using normalized cross\-correlation \(NCC\)\[[26](https://arxiv.org/html/2606.19935#bib.bib58)\]bewteen motion energyEmE\_\{m\}and audio energyEaE\_\{a\}:
ℒEE=1−NCC\(Em,Ea\)\.\\mathcal\{L\}\_\{EE\}=1\-\\mathrm\{NCC\}\(E\_\{m\},E\_\{a\}\)\.\(6\)
Due to the fact that movement speed is an important manifestation of motion rhythm, given a motion sequence𝐦\\mathbf\{m\}, motion energy is estimated from joint velocities:
Em\(t\)=‖𝐦t\+1−𝐦t−12Δt‖2\.E\_\{m\}\(t\)=\\left\\\|\\frac\{\\mathbf\{m\}\_\{t\+1\}\-\\mathbf\{m\}\_\{t\-1\}\}\{2\\Delta t\}\\right\\\|\_\{2\}\.\(7\)
Similary, the raw waveform be denoted as𝐚∈ℝB×L\\mathbf\{a\}\\in\\mathbb\{R\}^\{B\\times L\}\. We first compute an8080\-bin log\-Mel spectrogram using an short\-time Fourier transform, yielding𝐒∈ℝB×80×T\\mathbf\{S\}\\in\\mathbb\{R\}^\{B\\times 80\\times T\}\. The frame\-level log\-Mel energy is then defined as
Ea\(t\)=log\(∑m=180𝐒:,m,t\+ϵ\),E\_\{a\}\(t\)=\\log\\left\(\\sum\_\{m=1\}^\{80\}\\mathbf\{S\}\_\{:,m,t\}\+\\epsilon\\right\),\(8\)whereϵ=10−6\\epsilon=10^\{\-6\}\. Both signals are normalized before alignment\.
In addition, physical regularization termsℒphys\\mathcal\{L\}\_\{phys\}\[[25](https://arxiv.org/html/2606.19935#bib.bib57)\]are incorporated to enforce embodiment feasibility, including joint\-limit penalties, foot\-contact consistency, and skating suppression:
ℒmotion=ℒEE\+ℒphys\.\\mathcal\{L\}\_\{motion\}=\\mathcal\{L\}\_\{EE\}\+\\mathcal\{L\}\_\{phys\}\.\(9\)
Unlike conventional retargeting methods that prioritize geometric pose reconstruction, IK\-EER explicitly prioritizes embodiment\-consistent temporal dynamics\. More importantly, IK\-EER serves as a humanoid co\-speech motion*curation framework*rather than merely a retargeting algorithm\. Since robot\-native generation requires large\-scale physically executable humanoid motion supervision, preserving speech\-motion synchronization during humanoid reconstruction becomes essential for downstream learning\. The resulting dataset provides robot\-native motion distributions that simultaneously preserve embodiment feasibility and conversational expressiveness\.
### IV\-CRobot\-Native Motion Representation
Existing co\-speech motion generation methods typically operate on human\-centric representations such as SMPL\-X parameters or discretized latent motion tokens\. Although these representations are suitable for digital human animation, they introduce substantial embodiment mismatch for humanoid robots\.
From the perspective of motion manifolds, human\-centric joint spaces contain large regions that are unreachable for robot embodiments due to differences in kinematic topology, actuation structure, and joint constraints\. Consequently, learning in human representation space inevitably introduces representation\-level embodiment inconsistency\.
To eliminate this mismatch, we directly represent motions in the native control space of the humanoid robot\. For each frame, the motion state is defined as
𝐱\(t\)=\[𝐫root\(t\),θ1\(t\),…,θN\(t\),𝐩root\(t\)\],\\mathbf\{x\}^\{\(t\)\}=\[\\mathbf\{r\}\_\{root\}^\{\(t\)\},\\theta\_\{1\}^\{\(t\)\},\\dots,\\theta\_\{N\}^\{\(t\)\},\\mathbf\{p\}\_\{root\}^\{\(t\)\}\],\(10\)where𝐫root\\mathbf\{r\}\_\{root\}denotes the root orientation represented using continuous 6D rotation\[[51](https://arxiv.org/html/2606.19935#bib.bib61)\],θi\\theta\_\{i\}are robot joint angles,𝐩root\\mathbf\{p\}\_\{root\}represents the global root position, and superscript\(t\)denotes the timestept\.
This formulation establishes a one\-to\-one correspondence between generated motion trajectories and executable motor commands, removing the need for intermediate human\-body reconstruction or post\-hoc retargeting\. More importantly, robot\-native representation constrains motion learning directly within the feasible humanoid motion manifold, substantially reducing embodiment inconsistency during generation\.
### IV\-DPhysDrift
Figure 4:The proposed PhysDrift pipline\. Left: the model architecture\. Right: the loss function used during the training phase\. Speech and text will be separately represented by modal encoders, and then fused with noisy motion sequences through cross\-attention\. Finally, the model outputs the corresponding predicted motion sequences\.Building upon the robot\-native motion representation, we propose PhysDrift, a direct speech\-to\-motion generation framework that operates entirely in humanoid joint space\. Unlike conventional co\-speech generation pipelines that first model human motions and subsequently adapt them to robotic embodiments, PhysDrift directly learns the distribution of executable humanoid motions, thereby eliminating the embodiment gap introduced by intermediate human\-body representations\. Fig\.[4](https://arxiv.org/html/2606.19935#S4.F4)shows the pipeline of PhysDrift\.
Recent advances in generative modeling have demonstrated the effectiveness of diffusion models and Flow Matching for motion synthesis\. In particular, Flow Matching learns a continuous transport process by estimating a velocity field
d𝐱tdt=vθ\(𝐱t,t\),\\frac\{d\\mathbf\{x\}\_\{t\}\}\{dt\}=v\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\),\(11\)which progressively transforms samples from a simple prior distribution into samples from the target motion distribution\[[28](https://arxiv.org/html/2606.19935#bib.bib62)\]\. Compared with diffusion models, Flow Matching \(FM\) significantly reduces inference cost because the generation process can be formulated as deterministic trajectory transport\.
However, we identify that directly learning velocity fields introduces a previously overlooked form of embodiment gap when applied to humanoid co\-speech motion generation\. Velocity\-field learning implicitly assumes that generated trajectories will be decoded or smoothed before execution\. This assumption is generally valid for digital human animation, where latent motions are reconstructed through VQVAE\[[43](https://arxiv.org/html/2606.19935#bib.bib66)\]Decoder or rendering pipelines that naturally suppress high\-frequency artifacts\. In contrast, humanoid robots directly execute generated joint trajectories through motor controllers\. Consequently, small velocity estimation errors are accumulated through temporal integration:
𝐱t\+Δt=𝐱t\+∫tt\+Δtvθ\(𝐱,τ\)𝑑τ,\\mathbf\{x\}\_\{t\+\\Delta t\}=\\mathbf\{x\}\_\{t\}\+\\int\_\{t\}^\{t\+\\Delta t\}v\_\{\\theta\}\(\\mathbf\{x\},\\tau\)d\\tau,\(12\)causing local oscillations to propagate into large acceleration fluctuations\. As a result, velocity\-based transport often produces excessive motioaan jerk, unstable joint accelerations, and physically undesirable behaviors\. This issue becomes particularly severe for co\-speech motions, where subtle rhythmic variations synchronized with speech prosody create high\-frequency motion components that are easily amplified during velocity integration\.
From an embodiment perspective, the root cause is that velocity fields are defined in the tangent space \(the vector space formed by all possible “directions” or “velocities” at a certain point on a manifold\) of the motion manifold
vθ:𝒳→T𝒳,v\_\{\\theta\}:\\mathcal\{X\}\\rightarrow T\\mathcal\{X\},\(13\)whereT𝒳T\\mathcal\{X\}denotes the tangent bundle of the motion manifold\. Whereas humanoid execution is performed directly in configuration space \(Eq\. \([10](https://arxiv.org/html/2606.19935#S4.E10)\)\)\. Therefore, minimizing velocity transport error does not necessarily minimize the execution error of robot motions since robot execution is ultimately performed in configuration space rather than velocity space\. This mismatch constitutes a generation\-level embodiment gap that remains even after robot\-native motion representations are adopted\.
We address this issue by reformulating motion generation as a manifold attraction problem rather than a velocity transport problem, which leads to the proposed Motion Drift Field \(MDF\)\. To better understand its relationship with FM, we revisit the transport dynamics learned by velocity\-based generative models\.
The continuous dynamics in Eq\. \([11](https://arxiv.org/html/2606.19935#S4.E11)\) can be discretized as
𝐱t\+Δt=𝐱t\+Δtvθ\(𝐱t,t\)\.\\mathbf\{x\}\_\{t\+\\Delta t\}=\\mathbf\{x\}\_\{t\}\+\\Delta t\\,v\_\{\\theta\}\(\\mathbf\{x\}\_\{t\},t\)\.\(14\)Therefore, FM can be interpreted as repeatedly predicting infinitesimal transport directions and accumulating them through numerical integration\. While this formulation provides an accurate approximation of continuous transport trajectories, it requires the model to learn dynamics in tangent space rather than the robot execution space\.
Instead of learning infinitesimal transport directions, PhysDrift directly models the accumulated displacement between a generated sample and the feasible humanoid motion manifold\. Let
Δ𝐱=𝐱∗−𝐱,\\Delta\\mathbf\{x\}=\\mathbf\{x\}^\{\*\}\-\\mathbf\{x\},\(15\)where𝐱∗\\mathbf\{x\}^\{\*\}denotes an equilibrium state on the target humanoid motion manifold\. The proposed MDF directly estimates this displacement:
𝐃\(𝐱\)≈𝐱∗−𝐱\.\\mathbf\{D\}\(\\mathbf\{x\}\)\\approx\\mathbf\{x\}^\{\*\}\-\\mathbf\{x\}\.\(16\)From this perspective, MDF can be viewed as a one\-step relaxation formulation of velocity transport\. Rather than explicitly learning tangent\-space dynamics and integrating them over time, PhysDrift directly learns the correction required to move a sample toward the target motion manifold\. Consequently, the transport process is implicitly absorbed into the network parameters during training\.
Letppdenote the distribution of robot\-native motions constructed by IK\-EER and letqqdenote the generated motion distribution\. For a motion sample𝐱\\mathbf\{x\}, the MDF is formally defined as
𝐃p,q\(𝐱\)=1ZpZq𝔼p,q\[k~\(𝐱,𝐲\+\)k~\(𝐱,𝐲−\)\(𝐲\+−𝐲−\)\],\\mathbf\{D\}\_\{p,q\}\(\\mathbf\{x\}\)=\\dfrac\{1\}\{Z\_\{p\}Z\_\{q\}\}\\mathbb\{E\}\_\{p,q\}\\left\[\\tilde\{k\}\(\\mathbf\{x\},\\mathbf\{y\}^\{\+\}\)\\tilde\{k\}\(\\mathbf\{x\},\\mathbf\{y\}^\{\-\}\)\(\\mathbf\{y\}^\{\+\}\-\\mathbf\{y\}^\{\-\}\)\\right\],\(17\)
Zp\(𝐱\)=𝔼p\[k~\(𝐱,𝐲\+\)\]\],Z\_\{p\}\(\\mathbf\{x\}\)=\\mathbb\{E\}\_\{p\}\\left\[\\tilde\{k\}\(\\mathbf\{x\},\\mathbf\{y\}^\{\+\}\)\\right\]\],\(18\)
Zq\(𝐱\)=𝔼q\[k~\(𝐱,𝐲−\)\],Z\_\{q\}\(\\mathbf\{x\}\)=\\mathbb\{E\}\_\{q\}\\left\[\\tilde\{k\}\(\\mathbf\{x\},\\mathbf\{y\}^\{\-\}\)\\right\],\(19\)
k~\(𝐱,𝐲\)=exp\(−1τ‖𝐱−𝐲‖\),\\tilde\{k\}\(\\mathbf\{x\},\\mathbf\{y\}\)=\\exp\\left\(\-\\frac\{1\}\{\\tau\}\\\|\\mathbf\{x\}\-\\mathbf\{y\}\\\|\\right\),\(20\)where𝐲\+∼p\\mathbf\{y\}^\{\+\}\\sim p,𝐲−∼q\\mathbf\{y\}^\{\-\}\\sim q\. In order to better adapt to the robot’s motion space, we have redefined∥⋅∥\\\|\\cdot\\\|as a set of structured distances to accommodate different motion representations in Eq\. \([10](https://arxiv.org/html/2606.19935#S4.E10)\), rather thanL2L\_\{2\}distance used in\[[10](https://arxiv.org/html/2606.19935#bib.bib59)\]\. Specifically, for the root orientation𝐫root\\mathbf\{r\}\_\{root\}, we reconstruct the rotation using the 6D representation and map it toSO\(3\)SO\(3\)\. The distance between two orientationsdrotd\_\{\\text\{rot\}\}is measured using the geodesic distance onSO\(3\)SO\(3\)\.
drot=arccos\(tr\(\(𝐑\(1\)\)⊤𝐑\(2\)\)−12\),d\_\{\\text\{rot\}\}=\\arccos\\Bigg\(\\frac\{\\mathrm\{tr\}\\Big\(\(\\mathbf\{R\}^\{\(1\)\}\)^\{\\top\}\\mathbf\{R\}^\{\(2\)\}\\Big\)\-1\}\{2\}\\Bigg\),\(21\)where𝐑\(1\),𝐑\(2\)∈SO\(3\)\\mathbf\{R\}^\{\(1\)\},\\mathbf\{R\}^\{\(2\)\}\\in SO\(3\)are the corresponding root rotation matrices\.tr\(⋅\)\\mathrm\{tr\}\(\\cdot\)represents the trace of the matrix\. For joint angleθi\\theta\_\{i\}, we compute rotational differences using the periodic distance between joint angles\.
djoint=∑i=1n\(min\(\|θi\(1\)−θi\(2\)\|,2π−\|θi\(1\)−θi\(2\)\|\)\)2d\_\{\\text\{joint\}\}=\\sqrt\{\\sum\_\{i=1\}^\{n\}\\left\(\\min\\left\(\\bigl\|\\theta\_\{i\}^\{\(1\)\}\-\\theta\_\{i\}^\{\(2\)\}\\bigr\|,\\;2\\pi\-\\bigl\|\\theta\_\{i\}^\{\(1\)\}\-\\theta\_\{i\}^\{\(2\)\}\\bigr\|\\right\)\\right\)^\{2\}\}\(22\)For the root position𝐩root\\mathbf\{p\}\_\{root\}, we measure positional errorsdposd\_\{\\text\{pos\}\}using theL2L\_\{2\}distance\. In summary,
‖𝐱−𝐲‖=drot\+djoint\+dpos\.\\\|\\mathbf\{x\}\-\\mathbf\{y\}\\\|=d\_\{\\text\{rot\}\}\+d\_\{\\text\{joint\}\}\+d\_\{\\text\{pos\}\}\.\(23\)
The corresponding training objective minimizes the residual magnitude of the drift field:
ℒdrift=‖𝐱−sg\(𝐱\+𝐃p,q\(𝐱\)\)‖22,\\mathcal\{L\}\_\{drift\}=\\left\\\|\\mathbf\{x\}\-\\mathrm\{sg\}\\left\(\\mathbf\{x\}\+\\mathbf\{D\}\_\{p,q\}\(\\mathbf\{x\}\)\\right\)\\right\\\|\_\{2\}^\{2\},\(24\)wheresg\(⋅\)\\mathrm\{sg\}\(\\cdot\)denotes the stop\-gradient operator\. The proposed formulation naturally possesses several desirable theoretical properties\.
Proposition 1 \(Equilibrium Property\)\.When the generated distribution matches the target distribution, the expected drift vanishes:
𝐃p,q\(𝐱\)=0\.\\mathbf\{D\}\_\{p,q\}\(\\mathbf\{x\}\)=0\.\(25\)This indicates that the target humanoid motion manifold corresponds to a stationary equilibrium of the learning dynamics\. Furthermore, the kernelized formulation induces a local attraction process\. Since the similarity kernel suppresses long\-range interactions and emphasizes nearby motion samples, the resulting dynamics behave similarly to a contraction mapping in local neighborhoods\. Consequently, generated motions are progressively attracted toward feasible humanoid motion regions while avoiding abrupt trajectory jumps\.
An important consequence of the proposed attraction dynamics is that motion corrections are predicted directly in the robot execution space\. Consequently, PhysDrift avoids derivative amplification caused by high\-frequency motion components, substantially reducing motion jerk and improving temporal smoothness\. This property is particularly important for physically embodied co\-speech motion, where motion stability strongly influences interaction quality\.
To further preserve embodiment consistency during generation, we incorporate physical feasibility and speech\-motion synchronization constraints\. The final training objective is defined as
ℒ=ℒdrift\+ℒJoint\_limit\+ℒEE,\\mathcal\{L\}=\\mathcal\{L\}\_\{drift\}\+\\mathcal\{L\}\_\{Joint\\\_limit\}\+\\mathcal\{L\}\_\{EE\},\(26\)ℒJoint\_limit=ReLU\(qmin−𝐱\)\+ReLU\(𝐱−qmax\)\.\\mathcal\{L\}\_\{Joint\\\_limit\}=\\mathrm\{ReLU\}\(q^\{min\}\-\\mathbf\{x\}\)\+\\mathrm\{ReLU\}\(\\mathbf\{x\}\-q^\{max\}\)\.\(27\)
### IV\-EInference
The proposed formulation naturally enables one\-step generation\. Conventional diffusion models require iterative denoising, while FM requires numerical integration of learned velocity trajectories\. In contrast, PhysDrift directly learns the equilibrium correction that moves a sample toward the target humanoid motion manifold\.
Because the transport dynamics are absorbed into the learned MDF during training, the multi\-step transport process is compressed into a single network evaluation\. Therefore, inference does not require iterative denoising, trajectory rollout, or numerical ordinary differential equation integration\.
Given speech features𝐀\\mathbf\{A\}and Gaussian noiseϵ\\bm\{\\epsilon\}, humanoid motion is generated through a single forward pass:
𝐱^=fθ\(ϵ,𝐀\)\.\\hat\{\\mathbf\{x\}\}=f\_\{\\theta\}\(\\bm\{\\epsilon\},\\mathbf\{A\}\)\.\(28\)
By directly learning motion\-space attraction dynamics in robot configuration space, PhysDrift eliminates the generation\-level embodiment gap introduced by velocity transport, produces smoother and more physically executable trajectories, and enables real\-time one\-step humanoid co\-speech motion generation\.
## VExperiment
### V\-AImplemented Details\.
We refer to the experimental settings in Syntalker\[[5](https://arxiv.org/html/2606.19935#bib.bib55)\]and GestureLSM\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\]\. All models used in the experiment were trained for 1000 epochs and optimized using Adam with a learning rate of 1e\-4\. The training process were conducted on single NVIDIA A100\.
### V\-BEvaluation Metrics
To comprehensively evaluate humanoid co\-speech motion generation, we consider multiple metrics covering speech\-motion synchronization, motion distribution quality, physical feasibility, motion smoothness, and real\-time inference capability\.
Alignment\(Align\.\)\[[29](https://arxiv.org/html/2606.19935#bib.bib9)\]measures the temporal alignment between speech and generated motion\. Specifically, it evaluates whether motion rhythm and motion energy changes are synchronized with speech prosody and acoustic emphasis\. Higher values indicate better speech\-motion temporal consistency\.
Diversity\[[29](https://arxiv.org/html/2606.19935#bib.bib9)\]measures the variability of generated motions across different samples\. Higher diversity indicates richer expressive motion distributions and reduced mode collapse\.
FMD\(Fréchet Motion Distance\) evaluates the distributional similarity between generated motions and ground\-truth motions\. Similar to Fréchet Inception Distance \(FID\)\[[18](https://arxiv.org/html/2606.19935#bib.bib63)\]in image generation and Fréchet Gesture Distance \(FGD\)\[[29](https://arxiv.org/html/2606.19935#bib.bib9)\]in gesture generation, FMD computes the Fréchet distance between feature distributions extracted from generated and real motion sequences\. Lower values indicate that the generated motions are statistically closer to real humanoid motion distributions\.
G\_MPJPE\(Global Mean Per Joint Position Error\)\[[44](https://arxiv.org/html/2606.19935#bib.bib64)\]is defined as the average Euclidean distance between generated joint positions and reference joint positions in global coordinate space, measured in meters \(m\)\. Unlike local pose reconstruction metrics, G\_MPJPE additionally evaluates global spatial consistency of the reconstructed humanoid motion\. Lower values indicate more accurate pose reconstruction\.
Joint Violation\[[44](https://arxiv.org/html/2606.19935#bib.bib64)\]is defined as the joint\-angle violation beyond predefined humanoid joint limits during motion execution\.
Foot Contact Distance\[[46](https://arxiv.org/html/2606.19935#bib.bib65)\]evaluates the average height of the feet above the ground during frames that are expected to maintain foot contact, measured in meters \(m\)\. Lower values indicate more stable foot\-ground contact behavior\.
Skating Velocity\[[46](https://arxiv.org/html/2606.19935#bib.bib65)\]quantifies the sliding velocity of the feet during contact phases, with the unit meters per second \(m/s\)\. Lower values indicate reduced foot skating artifacts and more physically plausible humanoid motion\.
Jerkis defined as the third\-order temporal derivative of end\-effector \(wrist and foot\) position trajectories, corresponding to the rate of change of acceleration, measured in meters per second cubed \(m/s3m/s^\{3\}\)\. Lower jerk values indicate smoother and more physically stable humanoid motion dynamics\. FMD, Align\., and Diversity alone cannot accurately reflect high\-frequency jitter; in fact, stronger jitter may even improve these metrics\. Therefore, Jerk should be added as a core reference\. Only when Jerk and the other three metrics are all within reasonable ranges can motion naturalness be assessed\.
APS\(Actions Per Second\) measures the number of motion frames generated per second during inference, reflecting end\-to\-end latency from speech input to motion output\. Higher APS indicates faster inference speed and stronger real\-time interaction capability\.
### V\-CEmbodiment Gap in Human\-Centric Co\-Speech Generation
We first analyze how retargeting affects speech\-driven motion distributions in existing human\-centric co\-speech generation pipelines\. Table[II](https://arxiv.org/html/2606.19935#S5.T2)compares motion quality before and after humanoid retargeting across both ground\-truth \(GT\) motions and generated motions from different co\-speech generation models\.
A consistent phenomenon can be observed across all methods: retargeting causes degradation in both speech\-motion alignment and motion diversity\. For example, the alignment score of GT motion decreases from 0\.6897 to 0\.6600 after retargeting and the diversity score drops substantially from 12\.76 to 7\.319\. Similar trends can also be observed for Syntalker and GestureLSM\.
These results suggest that modern retargeting pipelines are generally capable of preserving coarse semantic correspondence between speech and motion\. However, the projection from human motion space into feasible humanoid motion space substantially compresses expressive motion distributions\. Instead of causing severe kinematic failure, the embodiment gap primarily manifests as a loss of expressive variability and temporal dynamics\. This observation is consistent with the formulation introduced in Section 3, where retargeting acts as a projection from the human motion manifoldℳh\\mathcal\{M\}\_\{h\}into the feasible humanoid motion manifoldℳr\\mathcal\{M\}\_\{r\}\.
Importantly, co\-speech interaction is highly sensitive to such distribution compression\. Human co\-speech gestures rely heavily on subtle rhythmic variations, asymmetric arm dynamics, and localized motion emphasis synchronized with speech prosody\. During retargeting, these expressive components are often smoothed into more conservative feasible motions, resulting in weaker conversational expressiveness despite relatively preserved semantic alignment\.
TABLE II:The impact of retargeting on speech motion temporal alignment and motion diversity\. “Align\.\*” refers to the alignment of only upper body and hand joint movements with speech\. “Syntalker” and “GestureLSM” means data generated by Syntalker and GestureLSM model\.DataRetargetingAlign\.\*↑\\uparrowDiversity↑\\uparrowGT×\\times0\.689712\.76GT✓\\checkmark0\.66007\.319Syntalker\[[5](https://arxiv.org/html/2606.19935#bib.bib55)\]×\\times0\.735912\.31Syntalker✓\\checkmark0\.72919\.527GestureLSMDiffusion\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\]×\\times0\.738412\.57GestureLSMDiffusion✓\\checkmark0\.73707\.571GestureLSMShortcutFlow\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\]×\\times0\.749012\.46GestureLSMShortcutFlow✓\\checkmark0\.74477\.525
### V\-DRetargeting for Humanoid Co\-Speech Motion Curation
We further evaluate different humanoid retargeting strategies in Table[III](https://arxiv.org/html/2606.19935#S5.T3)\. Unlike conventional retargeting evaluation that focuses primarily on pose reconstruction accuracy, we additionally evaluate speech\-motion alignment to analyze whether retargeted motions preserve conversational dynamics required for co\-speech interaction\.
Existing retargeting baselines such as GMR, Mink, and PHC exhibit relatively poor physical quality and alignment performance\. In particular, Mink and PHC produce severe skating artifacts with infinite skating velocity, indicating unstable contact behavior during humanoid execution\. Their alignment scores also remain relatively low, ranging from 0\.50 to 0\.53\.
Optimization\-based methods substantially improve physical feasibility\. Gradient\-Based optimization reduces G\_MPJPE from 0\.72 to 0\.040\. Introducing inverse\-kinematics initialization and end\-effector constraints further stabilizes motion quality\. The slight increase in G\_MPJPE can be attributed to motion matching with high\-energy audio signals, where the increased motion amplitude leads to minor violations\.
However, the most important observation is that physical feasibility alone does not guarantee high\-quality co\-speech motion\. Although Gradient\-Based optimization and IK\-EER w/o EE Loss achieve excellent kinematic metrics, their alignment scores remain limited at 0\.55\. In contrast, the full IK\-EER framework improves alignment significantly to 0\.69 while maintaining highly stable physical behavior\. This result demonstrates that preserving speech\-motion temporal synchronization is an independent and essential objective beyond conventional retargeting feasibility\.
More importantly, these results highlight the role of IK\-EER as a humanoid co\-speech motion curation framework rather than merely a retargeting algorithm\. Since robot\-native co\-speech generation requires large\-scale physically executable humanoid motion data, preserving prosody\-motion coupling during humanoid reconstruction becomes critical for dataset quality\. The proposed IK\-EER framework enables the construction of robot\-native co\-speech datasets that simultaneously preserve embodiment feasibility and conversational dynamics\.
TABLE III:Retargeting Experiment\. Gradient\-Based refers to optimizing based solely on the remaining Physical Loss functions in Fig\.[3](https://arxiv.org/html/2606.19935#S4.F3)while without using IK init\. and EE Loss\. “NaN” means that due to the repositioning, the rear robot is in a suspended state and cannot measure the sliding speed when the foot touches the ground\.Best results inbold, second bestunderlined\.MethodG\_MPJPE↓\\downarrowFoot Contact Distance↓\\downarrowSkating Velocity↓\\downarrowAlign\.↑\\uparrowGMR\[[1](https://arxiv.org/html/2606.19935#bib.bib52)\]0\.720\.0240\.190\.59Mink\[[49](https://arxiv.org/html/2606.19935#bib.bib53)\]0\.850\.32NaN0\.53PHC\[[34](https://arxiv.org/html/2606.19935#bib.bib60)\]0\.880\.35NaN0\.50Gradient\-Based Optimization0\.0400\.00410\.0720\.55IK\-EER \(Ours\)0\.0460\.00190\.0630\.69IK\-EER w/o EE Loss0\.0450\.00190\.0630\.55
### V\-EEffect of Motion Representation and Embodiment\-Aware Generation
Table[IV](https://arxiv.org/html/2606.19935#S5.T4)investigates the influence of motion representation and generation framework design on humanoid co\-speech generation\.
Using human\-centric representations such as “6D w\. VQVAE” leads to relatively poor distribution quality, achieving an FMD of 4\.183\. Removing VQVAE slightly improves diversity from 7\.254 to 9\.612, but the overall motion distribution remains significantly inferior to robot\-native representations\. In contrast, directly modeling humanoid motion using “6D Root&\\&Joint Angle” dramatically improves motion distribution quality, reducing FMD to 0\.6244\.
These results indicate that motion representation itself strongly influences embodiment consistency\. Human\-centric latent representations introduce structural mismatch between learned motion distributions and feasible humanoid motion spaces, whereas robot\-native joint\-space representations provide a more suitable representation for humanoid co\-speech behavior\.
Flow\-based generation further improves distribution learning capability\. FM achieves the best FMD score of 0\.3799 together with extremely high diversity of 29\.41, demonstrating the strong expressive capability of flow matching models\. However, as shown later in Table[V](https://arxiv.org/html/2606.19935#S5.T5), such unconstrained expressive generation also introduces severe jerk during humanoid execution\.
Compared with purely generative flow matching, PhysDrift achieves a more balanced trade\-off between motion quality, alignment, and physical stability\. Although its diversity is lower than unconstrained flow generation, PhysDrift maintains substantially better embodiment consistency while preserving competitive distribution quality\.
TABLE IV:Experiment on motion representation and model architecture ablation\. Best results inbold, second bestunderlined\.Motion RepresentationMethodFMD↓\\downarrowAlign\.↑\\uparrowDiversity↑\\uparrow6D w\. VQVAEDiffusion4\.1830\.75937\.2546D w/o\. VQVAEDiffusion4\.1650\.69619\.6126D Root & Joint AngleDiffusion0\.62440\.624110\.956D Root & Joint AngleFM0\.37990\.674629\.416D Root & Joint AnglePhysDrift0\.5400\.680311\.17
TABLE V:Comparative experiment between digital human with retargeting pipeline and robot native generation\. “NFE” refers to the number of forward process steps required for inference\. “†\\dagger” indicates that the model is modified to accommodate robot motion representation\. Best results inbold, second bestunderlined\.MethodNFEFMD↓\\downarrowAglin\.↑\\uparrowDiversity↑\\uparrowJoint ViolationJerk↓\\downarrowAPS↑\\uparrowGT––0\.53667\.319×\\times86\.59–SMPL\-X&\\&RetargetingSyntalker\[[5](https://arxiv.org/html/2606.19935#bib.bib55)\]&\\&GMR\[[1](https://arxiv.org/html/2606.19935#bib.bib52)\]10000\.56330\.69789\.527×\\times133\.911\.74GestureLSMShortcutFlow\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\]&\\&GMR\[[1](https://arxiv.org/html/2606.19935#bib.bib52)\]20\.70120\.67817\.769×\\times141\.640\.51GestureLSMMeanFlow\[[32](https://arxiv.org/html/2606.19935#bib.bib54)\]&\\&GMR\[[1](https://arxiv.org/html/2606.19935#bib.bib52)\]10\.42410\.69647\.525×\\times149\.342\.32Robotic Native Motion RepresentationSyntalker†\\dagger10000\.62440\.624110\.95×\\times141\.418\.11GestureLSM†ShortcutFlow\{\}\_\{ShortcutFlow\}\\dagger20\.37990\.674729\.41✓\\checkmark975\.42010GestureLSM†MeanFlow\{\}\_\{MeanFlow\}\\dagger10\.55420\.710210\.92×\\times437\.22350PhysDrift \(Ours\)10\.54620\.685612\.20×\\times118\.02880PhysDrift w/o\. Joint Limitaion10\.53100\.690111\.64✓\\checkmark118\.62880PhysDrift w/o\. Joint Limitaion & Energy Envelope10\.54090\.680311\.17✓\\checkmark106\.92880
### V\-FOverall Comparison
Table[V](https://arxiv.org/html/2606.19935#S5.T5)presents the overall comparison between human\-centric pipelines, robot\-native generation methods, and the proposed PhysDrift framework\.
Human\-centric pipelines based on SMPL\-X generation followed by retargeting exhibit limited motion diversity and relatively slow inference speed\. For example, “Syntalker&\\&GMR” achieves only 11\.74 APS with a diversity score of 9\.527\. Although “GestureLSMMeanFlow&\\&GMR” improves generation quality and speed, the resulting motion diversity remains constrained after retargeting, consistent with the embodiment\-gap analysis in Table[II](https://arxiv.org/html/2606.19935#S5.T2)\.
Robot\-native generation substantially improves efficiency\. Flow\-based methods achieve extremely high inference throughput, with “GestureLSMMeanFlow” reaching 2350 APS\. Moreover, robot\-native generation also avoids the diversity collapse introduced by retargeting\. However, these methods often suffer from severe physical instability\. “GestureLSMShortcutFlow” achieves very high diversity \(29\.41\) and low FMD \(0\.3799\), but simultaneously produces catastrophic joint violations and extremely high jerk \(975\.4\), making the generated motions unsuitable for stable humanoid execution\. According to the definition of Diversity, strong jitter may actually lead to higher scores\. Therefore, high diversity scores alone does not necessarily indicate an advantage of the method in robot space\. The high jerk of “GestureLSMShortcutFlow” and “GestureLSMMeanFlow” also demonstrates that, without a VAE Decoder to smooth the motions, learning velocity fields is not suitable for the native motion space of robots\. PhysDrift achieves a substantially better balance among expressiveness, physical plausibility, and real\-time capability\. Compared with unconstrained flow\-based generation, PhysDrift reduces jerk from 437\.2 to 118\.0 while maintaining competitive diversity and alignment performance\.
The ablation variants further demonstrate the importance of embodiment\-aware regularization\. Removing joint constraints result in joint violations, while removing both Joint Limitation and Energy Envelope constraints reduces alignment and diversity simultaneously\. These results confirm that embodiment\-aware constraints are essential for stabilizing robot\-native co\-speech generation\.
Overall, the experimental results consistently support the central hypothesis of this work: the primary limitation of existing co\-speech pipelines lies not merely in physical feasibility, but in the embodiment inconsistency introduced by human\-centric motion representations\. By directly modeling speech\-driven motion in robot joint space while incorporating embodiment\-aware physical regularization, PhysDrift achieves more expressive, physically plausible, and real\-time humanoid co\-speech interaction\.
## VIConclusion
In this work, we investigated the problem of humanoid co\-speech motion generation from the perspective of embodiment consistency\. Unlike existing human\-centric pipelines that generate motion in intermediate human\-body representations and subsequently retarget them onto humanoid robots, we showed that such approaches introduce a fundamental embodiment gap between human motion manifolds and physically executable humanoid motion spaces\. Through extensive analysis, we demonstrated that the primary limitation of existing pipelines is not merely kinematic feasibility, but the distortion of expressive motion distributions and speech\-motion temporal dynamics during embodiment transfer\.
To address this problem, we proposed an embodiment\-aware robot\-native co\-speech generation framework composed of two complementary components\. First, IK\-EER enables prosody\-preserving humanoid motion curation by jointly optimizing kinematic feasibility and speech\-motion synchronization during retargeting\. Building upon the curated robot\-native dataset, PhysDrift directly models speech\-conditioned humanoid joint trajectories while incorporating embodiment\-aware physical regularization to stabilize motion dynamics and preserve expressive behavior\.
Experimental results consistently validated the proposed formulation of the embodiment gap\. Our analysis showed that retargeting degraded speech\-motion alignment and motion diversity, even when semantic information was largely preserved\. Furthermore, while robot\-native flow\-based generation substantially improves expressiveness and inference efficiency, unconstrained generation often leads to unstable humanoid dynamics\. By jointly modeling embodiment constraints and expressive generation within robot joint space, PhysDrift achieves a substantially better balance between motion realism, physical plausibility, speech alignment, and real\-time interaction capability\.
More broadly, this work suggests that humanoid co\-speech generation should move beyond conventional human\-centric motion representations toward embodiment\-aware generative modeling directly grounded in robot morphology and dynamics\. We believe that preserving embodiment consistency will become increasingly important for future socially interactive humanoid systems, where natural communication depends not only on semantic correctness, but also on physically grounded expressive motion\.
## References
- \[1\]\(2025\)Retargeting matters: general motion retargeting for humanoid motion tracking\.arXiv preprint arXiv:2510\.02252\.Cited by:[§IV\-B](https://arxiv.org/html/2606.19935#S4.SS2.p1.1),[TABLE III](https://arxiv.org/html/2606.19935#S5.T3.4.5.1),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.10.8.1.1.1),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.13.11.2.2.2),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.16.14.2.2.2)\.
- \[2\]L\. Bao, Y\. Pan, T\. Peng, D\. Kanoulas, and C\. Zhou\(2025\)Hierarchical intention\-aware expressive motion generation for humanoid robots\.arXiv preprint arXiv:2506\.01563\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1)\.
- \[3\]Y\. Bian, A\. Zeng, X\. Ju, X\. Liu, Z\. Zhang, W\. Liu, and Q\. Xu\(Mar, 2025\)MotionCraft: crafting whole\-body motion with plug\-and\-play multimodal controls\.InProc\. AAAI Conf\. Artif\. Intell\., \(AAAI\),T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),Philadelphia, PA, USA,pp\. 1880–1888\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[4\]W\. Budiharto, A\. D\. Cahyani, P\. C\.B\. Rumondor, and D\. Suhartono\(2017\)EduRobot: intelligent humanoid robot with natural interaction for education and entertainment\.Procedia Computer Science116,pp\. 564–570\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[5\]B\. Chen, Y\. Li, Y\. Ding, T\. Shao, and K\. Zhou\(2024\)Enabling synergistic full\-body control in prompt\-based co\-speech motion generation\.InProceedings of the 32nd ACM International Conference on Multimedia,New York, NY, USA,pp\. 10\.External Links:[Document](https://dx.doi.org/10.1145/3664647.3680847)Cited by:[§V\-A](https://arxiv.org/html/2606.19935#S5.SS1.p1.1),[TABLE II](https://arxiv.org/html/2606.19935#S5.T2.5.5.5.2),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.10.8.1.1.1)\.
- \[6\]H\. Cheng, T\. Wang, G\. Shi, Z\. Zhao, and Y\. Fu\(Jun, 2025\)HOP: heterogeneous topology\-based multimodal entanglement for co\-speech gesture generation\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Nashville, TN, USA,pp\. 906–916\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[7\]X\. Cheng, Y\. Ji, J\. Chen, R\. Yang, G\. Yang, and X\. Wang\(Jul, 2024\)Expressive whole\-body control for humanoid robots\.InRobotics: Science and Systems XX,Delft, The Netherlands\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1),[§III](https://arxiv.org/html/2606.19935#S3.p6.5)\.
- \[8\]Y\. Cheng, S\. Huang, X\. Chen, J\. Ning, and M\. Gong\(Mar, 2025\)DIDiffGes: decoupled semi\-implicit diffusion models for real\-time gesture generation from speech\.InProc\. AAAI Conf\. Artif\. Intell\., \(AAAI\),T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),Philadelphia, PA, USA,pp\. 2464–2472\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[9\]L\. Cui, Y\. Li, X\. Yang, X\. Liu, L\. Zhang, and L\. Hou\(2026\)Humanoid robot–assisted support for health care in older adults: systematic scoping review\.JMIR Aging9,pp\. e83849\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[10\]M\. Deng, H\. Li, T\. Li, Y\. Du, and K\. He\(2026\)Generative modeling via drifting\.arXiv preprint arXiv:2602\.04770\.Cited by:[§IV\-D](https://arxiv.org/html/2606.19935#S4.SS4.p11.8)\.
- \[11\]S\. Ekström and L\. Pareto\(2022\)The dual role of humanoid robots in education: as didactic tools and social actors\.Educ\. Inf\. Technol\.27\(9\),pp\. 12609–12644\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[12\]A\. Ephrat, I\. Mosseri, O\. Lang, T\. Dekel, K\. Wilson, A\. Hassidim, W\. T\. Freeman, and M\. Rubinstein\(2018\)Looking to listen at the cocktail party: a speaker\-independent audio\-visual model for speech separation\.ACM Trans\. Graph\.37\(4\),pp\. 112\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[13\]F\. Fang, S\. Yang, and W\. Yang\(Jun, 2026\)CoordSpeaker: exploiting gesture captioning for coordinated caption\-empowered co\-speech gesture generation\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Denver, Colorado, USA,pp\. 30761–30771\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[14\]S\. Feng, E\. C\. Whitman, X\. Xinjilefu, and C\. G\. Atkeson\(Nov, 2014\)Optimization based full body control for the atlas robot\.InIEEE\-RAS Int\. Conf\. Humanoid Rob\.,Madrid, Spain,pp\. 120–127\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[15\]C\. Fu, Y\. Wang, H\. He, S\. Wang, C\. Wang, Y\. Tai, Y\. Liu, and J\. Zhang\(2026\)MambaGesture2: co\-speech gesture generation via hierarchical fusion and spatiotemporal aggregation\.IEEE Trans\. MultimediaEarly Access,pp\. 1–9\.Cited by:[§III](https://arxiv.org/html/2606.19935#S3.p7.2)\.
- \[16\]S\. Ghorbani, Y\. Ferstl, D\. Holden, N\. F\. Troje, and M\. Carbonneau\(2023\)ZeroEGGS: zero\-shot example\-based gesture generation from speech\.Comput\. Graph\. Forum42\(1\),pp\. 206–216\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[17\]T\. He, W\. Xiao, T\. Lin, Z\. Luo, Z\. Xu, Z\. Jiang, J\. Kautz, C\. Liu, G\. Shi, X\. Wang, L\. J\. Fan, and Y\. Zhu\(May, 2025\)HOVER: versatile neural whole\-body controller for humanoid robots\.InProc\. IEEE Int\. Conf\. Robot\. Autom\., \(ICRA\),Atlanta, GA, USA,pp\. 9989–9996\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1)\.
- \[18\]M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. Hochreiter\(Dec, 2017\)GANs trained by a two time\-scale update rule converge to a local nash equilibrium\.InProc\. Adv\. neural inf\. proces\. syst\., \(NeurIPS\),Long Beach, California, USA,pp\. 6629–6640\.Cited by:[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p4.1)\.
- \[19\]T\. Jeong, Y\. Chai, S\. Choi, J\. Bak, C\. Kim, J\. Yoon, Y\. Lee, J\. Lee, K\. Lee, J\. Kim, and S\. Choi\(Sep, 2025\)CoRe: A hybrid approach of contact\-aware optimization and learning for humanoid robot motions\.InIEEE\-RAS Int\. Conf\. Humanoid Rob\.,Seoul, Republic of Korea,pp\. 293–300\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[20\]H\. Jia, J\. Song, Y\. Zhang, H\. Jin, Y\. Fan, W\. Chen, W\. Zhang, and Y\. Yue\(2026\)ECHO: edge\-cloud humanoid orchestration for language\-to\-motion control\.arXiv preprint arXiv:2603\.16188\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p2.1)\.
- \[21\]Z\. Jiang, Y\. Xie, J\. Li, Y\. Yuan, Y\. Zhu, and Y\. Zhu\(Nov, 2024\)Harmon: whole\-body motion generation of humanoid robots from language descriptions\.InProc\. Conf\. Robot Learning, \(CoRL\),Munich, Germany,pp\. 3015–3026\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1),[§III](https://arxiv.org/html/2606.19935#S3.p4.4),[§III](https://arxiv.org/html/2606.19935#S3.p9.1)\.
- \[22\]W\. Jutharee, B\. Kaewkamnerdpong, and T\. Maneewarn\(2023\)Joint reconfiguration after failure for performing emblematic gestures in humanoid receptionist robot\.Sensors23\(22\),pp\. 9277\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[23\]S\. Kajita, F\. Kanehiro, K\. Kaneko, K\. Fujiwara, K\. Harada, K\. Yokoi, and H\. Hirukawa\(Sep, 2003\)Biped walking pattern generation by using preview control of zero\-moment point\.InProc\. IEEE Int\. Conf\. Robot\. Autom\., \(ICRA\),Taipei, Taiwan, China,pp\. 1620–1626\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[24\]H\. T\. Kalidindi, A\. Balachandran, and S\. V\. Shah\(2019\)Optimal whole\-body motion planning of humanoids in cluttered environments\.Robotics Auton\. Syst\.118,pp\. 263–277\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[25\]K\. Lee, S\. Kim, M\. Park, H\. Kim, D\. Hwang, H\. Lee, and J\. Choo\(2025\)PHUMA: physically\-grounded humanoid locomotion dataset\.arXiv preprint arXiv:2510\.26236\.Cited by:[§IV\-B](https://arxiv.org/html/2606.19935#S4.SS2.p10.1)\.
- \[26\]G\. Li, H\. Shao, X\. Deng, and Y\. Jiang\(2025\)Adaptive convolutional network pruning through pixel\-level cross\-correlation and channel independence for enhanced model compression\.Eng\. Appl\. Artif\. Intell\.154,pp\. 110920\.External Links:ISSN 0952\-1976Cited by:[§IV\-B](https://arxiv.org/html/2606.19935#S4.SS2.p5.2)\.
- \[27\]Z\. Li, J\. Liu, D\. Li, T\. Teng, M\. Li, S\. Calinon, D\. G\. Caldwell, and F\. Chen\(Oct, 2025\)ManiDP: manipulability\-aware diffusion policy for posture\-dependent bimanual manipulation\.InProc\. IEEE/RSJ Int\. Conf\. Intell\. Robots Syst\.,\(IROS\),Hangzhou, China,pp\. 9956–9962\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1)\.
- \[28\]Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le\(May, 2023\)Flow matching for generative modeling\.InProc\. Int\. Conf\. Learn\. Represent\.,\(ICLR\),Kigali, Rwanda\.Cited by:[§IV\-D](https://arxiv.org/html/2606.19935#S4.SS4.p2.2)\.
- \[29\]H\. Liu, Z\. Zhu, G\. Becherini, Y\. Peng, M\. Su, Y\. Zhou, X\. Zhe, N\. Iwamoto, B\. Zheng, and M\. J\. Black\(Jun\. 2024\)EMAGE: towards unified holistic co\-speech gesture generation via expressive masked audio gesture modeling\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Seattle, WA, USA,pp\. 1144–1154\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1),[§III](https://arxiv.org/html/2606.19935#S3.p4.4),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p2.1),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p3.1),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p4.1)\.
- \[30\]H\. Liu, Z\. Zhu, N\. Iwamoto, Y\. Peng, Z\. Li, Y\. Zhou, E\. Bozkurt, and B\. Zheng\(Oct\. 2022\)BEAT: A large\-scale semantic and emotional multi\-modal dataset for conversational gestures synthesis\.InProc\. Eur\. Conf\. Comput\. Vis\., \(ECCV\),Tel Aviv, Israel,pp\. 612–630\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[31\]L\. Liu, E\. Ghaleb, A\. Ozyurek, and Z\. Yumak\(Oct, 2025\)SemGes: semantics\-aware co\-speech gesture generation using semantic coherence and relevance learning\.InProc\. IEEE Int\. Conf\. Comput\. Vis\., \(ICCV\),Honolulu, Hawaii, USA,pp\. 13963–13973\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[32\]P\. Liu, L\. Song, J\. Huang, H\. Liu, and C\. Xu\(Oct, 2025\)GestureLSM: latent shortcut based co\-speech gesture generation with spatial\-temporal modeling\.InProc\. IEEE Int\. Conf\. Comput\. Vis\., \(ICCV\),Honolulu, HI, USA,pp\. 10929–10939\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p4.1),[§V\-A](https://arxiv.org/html/2606.19935#S5.SS1.p1.1),[TABLE II](https://arxiv.org/html/2606.19935#S5.T2.11.11.11.1),[TABLE II](https://arxiv.org/html/2606.19935#S5.T2.7.7.7.1),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.13.11.2.2.2),[TABLE V](https://arxiv.org/html/2606.19935#S5.T5.16.14.2.2.2)\.
- \[33\]C\. Lu, X\. Cheng, J\. Li, S\. Yang, M\. Ji, C\. Yuan, G\. Yang, S\. Yi, and X\. Wang\(May, 2025\)Mobile\-television: predictive motion priors for humanoid whole\-body control\.InProc\. IEEE Int\. Conf\. Robot\. Autom\., \(ICRA\),Atlanta, GA, USA,pp\. 5364–5371\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[34\]Z\. Luo, J\. Cao, A\. Winkler, K\. Kitani, and W\. Xu\(Oct, 2023\)Perpetual humanoid control for real\-time simulated avatars\.InProc\. IEEE Int\. Conf\. Comput\. Vis\., \(ICCV\),Paris, France,,pp\. 10861–10870\.Cited by:[§IV\-B](https://arxiv.org/html/2606.19935#S4.SS2.p1.1),[TABLE III](https://arxiv.org/html/2606.19935#S5.T3.4.7.1)\.
- \[35\]J\. Mao, S\. Zhao, S\. Song, C\. Hong, T\. Shi, J\. Ye, M\. Zhang, H\. Geng, J\. Malik, V\. Guizilini, and Y\. Wang\(Sep, 2025\)Universal humanoid robot pose learning from internet human videos\.InIEEE\-RAS Int\. Conf\. Humanoid Rob\.,Seoul, Republic of Korea,pp\. 1–8\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[36\]G\. Pavlakos, V\. Choutas, N\. Ghorbani, T\. Bolkart, A\. A\. A\. Osman, D\. Tzionas, and M\. J\. Black\(Jun, 2019\)Expressive body capture: 3d hands, face, and body from a single image\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Long Beach, CA, USA,pp\. 10975–10985\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p2.1)\.
- \[37\]X\. Qi, C\. Liu, L\. Li, J\. Hou, H\. Xin, and X\. Yu\(2024\)EmotionGesture: audio\-driven diverse emotional co\-speech 3d gesture generation\.IEEE Trans\. Multimedia26\(\),pp\. 10420–10430\.Cited by:[§III](https://arxiv.org/html/2606.19935#S3.p7.2)\.
- \[38\]M\. U\. Saleem, M\. J\. Patel, E\. Pinyoanuntapong, Z\. Qin, L\. Yang, H\. Xue, A\. Helmy, C\. Chen, and P\. Wang\(Jun, 2026\)LiveGesture: streamable co\-speech gesture generation model\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Denver, Colorado, USA,pp\. 2264–2273\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[39\]M\. F\. Stollenga, L\. Pape, M\. Frank, J\. Leitner, A\. Förster, and J\. Schmidhuber\(Nov, 2013\)Task\-relevant roadmaps: A framework for humanoid motion planning\.In2013 IEEE/RSJ Int\. Conf\. Intell\. Robots Syst\., \(IROS\),Tokyo, Japan,pp\. 5772–5778\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[40\]W\. Suleiman, F\. Kanehiro, E\. Yoshida, J\. Laumond, and A\. Monin\(2010\)Time parameterization of humanoid\-robot paths\.IEEE Trans\. Robot\.26\(3\),pp\. 458–468\.Cited by:[§II\-B](https://arxiv.org/html/2606.19935#S2.SS2.p1.1)\.
- \[41\]A\. Tang, T\. Hiraoka, N\. Hiraoka, F\. Shi, K\. Kawaharazuka, K\. Kojima, K\. Okada, and M\. Inaba\(May, 2024\)Humanmimic: learning natural locomotion and transitions for humanoid robot via wasserstein adversarial imitation\.InProc\. IEEE Int\. Conf\. Robot\. Autom\., \(ICRA\),Yokohama, Japan,pp\. 13107–13114\.Cited by:[§III](https://arxiv.org/html/2606.19935#S3.p6.5)\.
- \[42\]U\. Tripathi, R\. S\. J, V\. Chamola, A\. Jolfaei, and A\. Chintanpalli\(2022\)Advancing remote healthcare using humanoid and affective systems\.IEEE Sens\. J\.22\(18\),pp\. 17606–17614\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p1.1)\.
- \[43\]A\. van den Oord, O\. Vinyals, and K\. Kavukcuoglu\(2017\)Neural discrete representation learning\.InProc\. Adv\. neural inf\. proces\. syst\., \(NeurIPS\),Long Beach, CA, USA,pp\. 6309–6318\.Cited by:[§IV\-D](https://arxiv.org/html/2606.19935#S4.SS4.p3.1)\.
- \[44\]W\. Xie, J\. Han, J\. Zheng, H\. Li, X\. Liu, J\. Shi, W\. Zhang, C\. Bai, and X\. Li\(2025\)KungfuBot: physics\-based humanoid whole\-body control for learning highly\-dynamic skills\.InProc\. Adv\. neural inf\. proces\. syst\., \(NeurIPS\),San Diego, CA, USA,pp\. 62406–62433\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p2.1),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p5.1),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p6.1)\.
- \[45\]Z\. Xu, M\. Hu, K\. Xiao, Q\. Fang, C\. Liu, and Q\. Chen\(2025\)Realizing text\-driven motion generation on NAO robot: A reinforcement learning\-optimized control pipeline\.arXiv preprint arXiv:2506\.05117\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1)\.
- \[46\]L\. Yang, X\. Huang, Z\. Wu, A\. Kanazawa, P\. Abbeel, C\. Sferrazza, C\. K\. Liu, R\. Duan, and G\. Shi\(2025\)Omniretarget: interaction\-preserving data generation for humanoid whole\-body loco\-manipulation and scene interaction\.arXiv preprint arXiv:2509\.26633\.Cited by:[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p7.1),[§V\-B](https://arxiv.org/html/2606.19935#S5.SS2.p8.1)\.
- \[47\]K\. Yin, W\. Zeng, K\. Fan, M\. Dai, Z\. Wang, Q\. Zhang, Z\. Tian, J\. Wang, J\. Pang, and W\. Zhang\(2026\)UniTracker: learning universal whole\-body motion tracker for humanoid robots\.IEEE Rob\. Autom\. Lett\.11\(7\),pp\. 8124–8131\.Cited by:[§I](https://arxiv.org/html/2606.19935#S1.p2.1)\.
- \[48\]Y\. Yoon, B\. Cha, J\. Lee, M\. Jang, J\. Lee, J\. Kim, and G\. Lee\(2020\)Speech gesture generation from the trimodal context of text, audio, and speaker identity\.ACM Trans\. Graph\.39\(6\),pp\. 222:1–222:16\.Cited by:[§II\-A](https://arxiv.org/html/2606.19935#S2.SS1.p1.1)\.
- \[49\]K\. Zakka\(2026\-02\)Mink: Python inverse kinematics based on MuJoCo\(Website\)External Links:[Link](https://github.com/kevinzakka/mink)Cited by:[§IV\-B](https://arxiv.org/html/2606.19935#S4.SS2.p1.1),[TABLE III](https://arxiv.org/html/2606.19935#S5.T3.4.6.1)\.
- \[50\]Y\. Ze, Z\. Chen, W\. Wang, T\. Chen, X\. He, Y\. Yuan, X\. B\. Peng, and J\. Wu\(Oct, 2025\)Generalizable humanoid manipulation with 3d diffusion policies\.InProc\. IEEE/RSJ Int\. Conf\. Intell\. Robots Syst\.,\(IROS\),Hangzhou, China,pp\. 2873–2880\.Cited by:[§II\-C](https://arxiv.org/html/2606.19935#S2.SS3.p1.1)\.
- \[51\]Y\. Zhou, C\. Barnes, J\. Lu, J\. Yang, and H\. Li\(Jun, 2019\)On the continuity of rotation representations in neural networks\.InProc\. IEEE Conf\. Comput\. Vis\. Pattern Recognit\., \(CVPR\),Long Beach, CA, USA,pp\. 5745–5753\.Cited by:[§IV\-C](https://arxiv.org/html/2606.19935#S4.SS3.p4.3)\.Similar Articles
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
The paper presents Puppeteer, a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, posture-aware, and physically grounded gestures.
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
OmniHumanoid is a framework that enables scalable cross-embodiment video generation by factorizing motion transfer and embodiment-specific adaptation, using unpaired data and branch-isolated attention to reduce interference.
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
PhyGenHOI is a novel framework that generates physically accurate 4D human-object interactions by coupling motion diffusion models with material point method simulations using 3D Gaussian representations.
CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments
CrossBFM proposes a unified encoder architecture to distill a shared latent behavior space across humanoid embodiments, enabling efficient training and cross-embodiment generalization for motion tracking, goal reaching, and reward optimization.
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
DRIFT is a framework that adapts pretrained vision-language models for continuous output decoding by combining coarse prediction with iterative flow matching refinement, improving performance on perception and planning tasks.