LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

arXiv cs.LG Papers

Summary

This paper introduces LUCID, a hierarchical model-based reinforcement learning framework for long-horizon humanoid loco-manipulation. It learns reusable latent skills and a macro-dynamics world model, enabling high-level planning via imagined rollouts and improving success rates in simulated multi-object rearrangement tasks.

arXiv:2608.07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via adversarial imitation and then freezes it while jointly learning a high-level policy and macro-dynamics world model. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high-level policy optimization through imagined rollouts. We evaluate our framework across various simulated multi-object rearrangement scenarios. Experimental results show that LUCID improves the full-task success and partial-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco-manipulation tasks.
Original Article
View Cached Full Text

Cached at: 08/11/26, 08:06 AM

# Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
Source: [https://arxiv.org/html/2608.07746](https://arxiv.org/html/2608.07746)
###### Abstract

Long\-horizon humanoid loco\-manipulation requires composing versatile whole\-body skills and reliable high\-level decision making\. Existing methods often coordinate pretrained skills with scripted planners, finite\-state machines or task\-specific model\-free policies, restricting their ability to handle complex task sequences\. To address this limitation, we proposeLUCID, a hierarchical model\-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model\. LUCID first trains a structured latent\-conditioned low\-level policy via adversarial imitation and then freezes it while jointly learning a high\-level policy and macro\-dynamics world model\. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high\-level policy optimization through imagined rollouts\. We evaluate our framework across various simulated multi\-object rearrangement scenarios\. Experimental results show that LUCID improves the full\-task success and partial\-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco\-manipulation tasks\.

## Introduction

Humanoids can perform individual locomotion and manipulation behaviors, but executing them reliably in a long\-horizon sequence remains a challenge\(Guet al\.[2026](https://arxiv.org/html/2608.07746#bib.bib1)\)\. In multi\-object rearrangement, for example, a humanoid needs to repeatedly navigate between objects and manipulate them within a continuous episode\. Such tasks require both versatile whole\-body skills and a high\-level policy that can coordinate them toward sequential goals\. Previous works have made substantial progress in physics\-based humanoid control through adversarial imitation and latent skill learning, which enable natural whole\-body motion and reusable motor behaviors\(Penget al\.[2021](https://arxiv.org/html/2608.07746#bib.bib5),[2022](https://arxiv.org/html/2608.07746#bib.bib2); Wenget al\.[2025](https://arxiv.org/html/2608.07746#bib.bib22); Wanget al\.[2025a](https://arxiv.org/html/2608.07746#bib.bib19)\)\. Recent humanoid\-scene interaction methods further extend such motor control to navigation and object manipulation in more complex environments\(Starkeet al\.[2019](https://arxiv.org/html/2608.07746#bib.bib16); Hassanet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib10); Panet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib14); Xiaoet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib15)\)\. Together, these works have expanded humanoid control from isolated motor skills to scene\-aware loco\-manipulation\.

However, coordinating these behaviors over a long task sequence remains difficult\. Existing approaches either consider primarily specialized controllers for isolated object interactions\(Xuet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib6); Wanget al\.[2025b](https://arxiv.org/html/2608.07746#bib.bib20)\), or connect consecutive interactions through hand\-designed mechanisms, such as a fixed navigation\-to\-manipulation state machine\(Panet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib17)\)or per\-object reference tracking with scripted handoffs between stages\(Xuet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib18)\)\. Although such systems handle individual interactions effectively, their temporal organization is largely predefined and does not account for how intermediate decisions affect later subtasks\. Hierarchical reinforcement learning addresses this structural problem by separating fast motor control from slower decision\-making over temporally extended skills\(Suttonet al\.[1999](https://arxiv.org/html/2608.07746#bib.bib23)\)\. In physics\-based character control, latent\-skill methods have produced controllable motion representations and compact categorical priors that can be reused for downstream tasks\(Tessleret al\.[2023](https://arxiv.org/html/2608.07746#bib.bib12); Zhuet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib24)\)\. High\-level policies can also combine pretrained humanoid motor primitives with limited task\-specific reward design\(Kuanget al\.[2025](https://arxiv.org/html/2608.07746#bib.bib13)\)\. Yet existing hierarchical systems still rely on predefined skill transitions or task\-specific model\-free high\-level policies trained through direct environment interaction\. Consequently, they can select reusable behaviors but cannot predict how one skill changes the conditions for subsequent interactions, which is a limitation that becomes increasingly important in long\-horizon multi\-object rearrangement\.

World models offer this predictive capability by approximating environment dynamics and using imagined trajectories to improve decision\-making\(Ha and Schmidhuber[2018](https://arxiv.org/html/2608.07746#bib.bib21); Hafneret al\.[2020](https://arxiv.org/html/2608.07746#bib.bib4)\)\. They have demonstrated strong performance across control domains and have recently been applied to robotic manipulation, navigation, and locomotion\(Wuet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib28); Liet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib29)\)\. For humanoid interaction, however, predicting every joint\-level transition over an extended horizon is difficult: high\-dimensional, contact\-rich dynamics require many autoregressive steps, amplifying model errors\. The hierarchy provides a natural temporal abstraction\. Rather than reproducing complete physical trajectories, a model can predict the task\-level changes induced by each temporally extended skill\. This connects reusable motor behaviors with predictive high\-level reasoning, motivating a framework that learns skill\-level dynamics and uses them to coordinate sequential humanoid interactions\.

To this end, we introduceLUCID, a hierarchical model\-based reinforcement learning framework for long\-horizon skill composition, illustrated in Figure[1](https://arxiv.org/html/2608.07746#Sx2.F1)\. Unlike hierarchical humanoid controllers that rely on scripted transitions or task\-specific model\-free skill sequencing, LUCID learns skill\-level dynamics and optimizes a goal\-conditioned high\-level policy through imagined macro\-transitions\. A frozen latent\-conditioned low\-level controller provides reusable whole\-body skills through a structured interface combining discrete skill anchors with continuous variation\. At the macro timescale, the high\-level policy selects latent commands, while a macro\-dynamics world model predicts their task\-level consequences for the humanoid, manipulated objects, and task progress rather than modeling every joint\-level transition\. Imagined rollouts through this model allow the policy to anticipate how current skill choices affect subsequent subtasks, while real simulator experience continually improves the model\. This combination enables LUCID to coordinate complete multi\-object rearrangement sequences without scripted handoffs\.

In summary, our contributions are threefold\.

- •We proposeLUCID, a hierarchical model\-based reinforcement learning framework for task planning from imagined dynamics\. Its world model predicts how latent decisions shape task progression, allowing a high\-level policy to optimize long\-horizon behavior in imagination\.
- •We train a low\-level controller with adversarial imitation and an interaction\-aware curriculum, using a structured latent interface that exposes whole\-body motor primitives to the high\-level policy\.
- •We evaluate our method in various long\-horizon multi\-object rearrangement tasks, showing improved full sequence completion and validating the key component through ablations\.

## Related Work

![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/lucid_framework.png)Figure 1:Overview of LUCID\. LLC is trained with adversarial imitation and task rewards, then frozen\. During training, latent z is produced with a state\-based oracle\. HLC is jointly trained with the macro\-dynamics world model\. The replay buffer is initialized with real transitions and refreshed during each update\. Dashed lines indicate detached \(stop\-gradient\) data flow\.### Physics\-Based Motion Generation

Physics\-based motion generation produces dynamically feasible behaviors by training controllers in simulation\. DeepMimic popularized reference\-tracking reinforcement learning for character animation\(Penget al\.[2018](https://arxiv.org/html/2608.07746#bib.bib40)\), while AMP acquired task\-compatible motion priors from unstructured demonstrations\(Penget al\.[2021](https://arxiv.org/html/2608.07746#bib.bib5)\)\. Building on these foundations, reusable and directable representations enabled broader downstream applications\(Penget al\.[2022](https://arxiv.org/html/2608.07746#bib.bib2); Tessleret al\.[2023](https://arxiv.org/html/2608.07746#bib.bib12),[2024](https://arxiv.org/html/2608.07746#bib.bib41)\)\. Interaction\-oriented research further expanded the scope to physical scene contacts, dynamic imitation, and object grasping\(Hassanet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib10); Wanget al\.[2023](https://arxiv.org/html/2608.07746#bib.bib42); Luoet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib43)\); general rearrangement\(Xuet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib6)\); and scalable multi\-skill execution\(Xuet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib18); Panet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib17)\)\. Recent systems have also transferred such capabilities to real humanoids\(Wanget al\.[2025a](https://arxiv.org/html/2608.07746#bib.bib19); Liet al\.[2026](https://arxiv.org/html/2608.07746#bib.bib30)\)\. Unlike these methods, LUCID organizes phase\-controllable skills with a high\-level controller trained in imagination\.

### Hierarchical Reinforcement Learning

Hierarchical reinforcement learning decomposes long\-horizon control into high\-level decisions and low\-level skills\(Suttonet al\.[1999](https://arxiv.org/html/2608.07746#bib.bib23); Baconet al\.[2017](https://arxiv.org/html/2608.07746#bib.bib31)\)\. Early systems built skill libraries and used schedulers to select or compose controllers\(Faloutsoset al\.[2001](https://arxiv.org/html/2608.07746#bib.bib33); Heesset al\.[2016](https://arxiv.org/html/2608.07746#bib.bib32)\)\. Later methods learned latent skill spaces that high\-level policies could steer to generate motions\(Merelet al\.[2018](https://arxiv.org/html/2608.07746#bib.bib35); Penget al\.[2019](https://arxiv.org/html/2608.07746#bib.bib34); Hasencleveret al\.[2020](https://arxiv.org/html/2608.07746#bib.bib36); Tessleret al\.[2023](https://arxiv.org/html/2608.07746#bib.bib12)\)\. These abstractions were extended to legged loco\-manipulation through skill sequencing and residual composition\(Jiet al\.[2022](https://arxiv.org/html/2608.07746#bib.bib37); Kumaret al\.[2023](https://arxiv.org/html/2608.07746#bib.bib38)\)\. Recent humanoid systems instead plan motion references with hierarchical world models or blend goal\-conditioned skills\(Hansenet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib39); Kuanget al\.[2025](https://arxiv.org/html/2608.07746#bib.bib13)\)\. LUCID learns discrete phase decisions over a frozen structured interface and optimizes long\-horizon skill sequencing through imagined macro\-dynamics for multi\-object rearrangement\.

### World Model for Robotics

World models learn action\-conditioned dynamics for planning and policy optimization through imagined trajectories\(Ha and Schmidhuber[2018](https://arxiv.org/html/2608.07746#bib.bib21)\)\. PlaNet introduced latent dynamics for planning\(Hafneret al\.[2019](https://arxiv.org/html/2608.07746#bib.bib25)\), while Dreamer trained actor–critic policies from latent imagination\(Hafneret al\.[2020](https://arxiv.org/html/2608.07746#bib.bib4),[2025](https://arxiv.org/html/2608.07746#bib.bib3)\)\. TD\-MPC combined learned representations with trajectory optimization\(Hansenet al\.[2022](https://arxiv.org/html/2608.07746#bib.bib27),[2024](https://arxiv.org/html/2608.07746#bib.bib26)\), and DayDreamer extended imagination\-based learning to physical robots\(Wuet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib28)\)\. More recent robotics\-specific models emphasize predictive reliability\. RWM uses dual\-autoregressive training to reduce compounding error and optimize locomotion policies in a learned neural simulator\(Liet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib29)\)\. HAIC instead predicts unobserved object dynamics from proprioception to improve feedback control during humanoid–object interaction\(Liet al\.[2026](https://arxiv.org/html/2608.07746#bib.bib30)\)\. However, these approaches primarily model joint\-level dynamics, provide state estimates, or address continuous control\. LUCID instead models skill\-level transitions induced by structured phase commands and trains a high\-level policy through imagined rollouts, enabling long\-horizon planning while keeping the physics\-based low\-level controller fixed\.

## Method

Figure[1](https://arxiv.org/html/2608.07746#Sx2.F1)presents LUCID, a scalable hierarchical framework for long\-horizon robot\-object interaction tasks\. Training proceeds in two stages\. First, we train a reusable latent\-conditioned low\-level controller \(LLC\) via adversarial imitation following ASE\(Penget al\.[2022](https://arxiv.org/html/2608.07746#bib.bib2)\)\. We then freeze the LLC and jointly train a macro\-dynamics world model \(WM\) and a goal\-conditioned high\-level controller \(HLC\) similar to DreamerV3\(Hafneret al\.[2025](https://arxiv.org/html/2608.07746#bib.bib3)\)\. The WM predicts the physical consequences of latent decisions, allowing the HLC to learn from imagined macro\-rollouts while real simulator interaction continually supplies training data for the WM\.

### Problem Formulation

We consider a simulated humanoid that must rearrange multiple objects in a prescribed order\. We formulate the task as a goal\-conditioned Markov Decision Process \(MDP\)⟨𝒮,𝒜,𝒯,ℛ,γ,𝒢⟩\\langle\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{T\},\\mathcal\{R\},\\gamma,\\mathcal\{G\}\\rangle\. At simulation steptt,st∈𝒮s\_\{t\}\\in\\mathcal\{S\}contains the humanoid and environment states,at∈𝒜a\_\{t\}\\in\\mathcal\{A\}denotes joint\-level PD targets, and𝒯\\mathcal\{T\}is induced by the physics simulator\. The episode goalg=\(g1,…,gM\)∈𝒢g=\(g^\{1\},\\ldots,g^\{M\}\)\\in\\mathcal\{G\}specifies an ordered sequence ofMMobject\-rearrangement subgoals\. The objective is to maximize𝔼​\[∑t=0T−1γt​rt\]\\mathbb\{E\}\[\\sum\_\{t=0\}^\{T\-1\}\\gamma^\{t\}r\_\{t\}\], wherertr\_\{t\}rewards subgoal completion and physically plausible motion\.

### Latent Skill Primitives

The LLCπL​\(at∣stL,zt\)\\pi\_\{L\}\(a\_\{t\}\\mid s\_\{t\}^\{L\},z\_\{t\}\)maps a low\-level statestLs\_\{t\}^\{L\}and a skill latentztz\_\{t\}to joint\-level actions\. The low\-level state contains proprioception, object state, and the current task\-guidance variables; in particular, the bounded guidance offset and waypoint\-advance information represented byptgp\_\{t\}^\{g\}are encoded intostLs\_\{t\}^\{L\}before LLC execution\. We train it to produce human\-like, interaction\-aware behaviors while remaining controllable throughztz\_\{t\}by combining an adversarial imitation rewardrtDr\_\{t\}^\{D\}with a goal\-conditioned task rewardrtGr\_\{t\}^\{G\}:

maxπL⁡𝔼τ∼p​\(τ∣πL,ψ\)​\[∑t=0T−1γt​\(rtD\+λG​rtG\)\]\.\\max\_\{\\pi\_\{L\}\}\\;\\mathbb\{E\}\_\{\\tau\\sim p\(\\tau\\mid\\pi\_\{L\},\\psi\)\}\\left\[\\sum\_\{t=0\}^\{T\-1\}\\gamma^\{t\}\\bigl\(r\_\{t\}^\{D\}\+\\lambda\_\{G\}r\_\{t\}^\{G\}\\bigr\)\\right\]\.\(1\)Here,τ=\(s0L,z0,a0,s1L,…\)\\tau=\(s\_\{0\}^\{L\},z\_\{0\},a\_\{0\},s\_\{1\}^\{L\},\\ldots\)is a low\-level trajectory,λG\\lambda\_\{G\}balances task progress and imitation, and the oracleψ​\(stL,g\)\\psi\(s\_\{t\}^\{L\},g\)selects the skill associated with the current interaction phase\. By consistently conditioning each phase on its corresponding latent during training, the LLC learns to produce distinct behaviors for different latent inputs, thereby providing latent\-level behavior controllability\.

##### Adversarial imitation

Letϕt=ϕ​\(stD,st\+1D\)\\phi\_\{t\}=\\phi\(s\_\{t\}^\{D\},s\_\{t\+1\}^\{D\}\)be a transition feature constructed from robot and object kinematics\. A discriminatorD​\(ϕ\)∈\(0,1\)D\(\\phi\)\\in\(0,1\)is trained to classify reference transitions as real and LLC transitions as generated\. The LLC therefore receives the non\-saturating reward:

rtD=−log⁡\(1−D​\(ϕt\)\),r\_\{t\}^\{D\}=\-\\log\\big\(1\-D\(\\phi\_\{t\}\)\\big\),\(2\)which increases when a generated transition is assigned a higher reference\-motion probability\. The full discriminator objective, including its gradient regularizer, is provided in the supplementary material \(*Adversarial Imitation Objective*\)\.

##### Structured latent space

Unlike the unstructured hyperspherical prior used by ASE, we reserve one anchor for each ofNNsemantic skills and use the remaining dimensions for within\-skill variation:

zt=\[𝐞ct;σz​ϵt\]‖\[𝐞ct;σz​ϵt\]‖2,ϵt∼𝒩​\(𝟎,𝐈Dz−N\),z\_\{t\}=\\frac\{\[\\,\\mathbf\{e\}\_\{c\_\{t\}\};\\;\\sigma\_\{z\}\\boldsymbol\{\\epsilon\}\_\{t\}\\,\]\}\{\\left\\lVert\[\\,\\mathbf\{e\}\_\{c\_\{t\}\};\\;\\sigma\_\{z\}\\boldsymbol\{\\epsilon\}\_\{t\}\\,\]\\right\\rVert\_\{2\}\},\\quad\\boldsymbol\{\\epsilon\}\_\{t\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\_\{D\_\{z\}\-N\}\),\(3\)whereDz≥ND\_\{z\}\\geq Nis the latent dimension,𝐞ct∈ℝN\\mathbf\{e\}\_\{c\_\{t\}\}\\in\\mathbb\{R\}^\{N\}is the one\-hot anchor for skillct∈\{1,…,N\}c\_\{t\}\\in\\\{1,\\ldots,N\\\}, andσz\\sigma\_\{z\}controls within\-skill variation\. During LLC training,ctc\_\{t\}is supplied byψ\\psi, the induced latent distribution is thus a mixturep​\(zt\)=∑c=1Np​\(ct=c\)​p​\(zt∣ct=c\)p\(z\_\{t\}\)=\\sum\_\{c=1\}^\{N\}p\(c\_\{t\}=c\)\\,p\(z\_\{t\}\\mid c\_\{t\}=c\)ofNNsuch neighborhoods rather than a uniform distribution over the sphere\. After the LLC is frozen, the HLC selects the structured latent\.

##### Interaction\-aware curriculum

To maintain behavior controllability during sequential object rearrangement, we trainπL\\pi\_\{L\}using a three\-stage curriculum: 1\)*carry*, which emphasizes approach, contact, and lifting; 2\)*rearrangement*, which enables transport and placement; and 3\)*retreat and chaining*, which adds the transit skill, a subtask\-completion bonus, and a post\-placement retreat term\. Each stage warm\-starts from the previous one to target progressive completions, with only the task rewardrGr^\{G\}and reference\-state initialization \(RSI\) changing\. At every stage, the task reward is defined as:

rt,mG=∑jwj\(m\)​rt\(j\),r\_\{t,m\}^\{G\}=\\sum\_\{j\}w\_\{j\}^\{\(m\)\}r\_\{t\}^\{\(j\)\},\(4\)wheremmindexes the curriculum stage,rt\(j\)r\_\{t\}^\{\(j\)\}is an interaction reward, andwj\(m\)w\_\{j\}^\{\(m\)\}is its stage\-dependent coefficient\. The stage\-specific objective, retreat reward, activation conditions, and initialization scheme are given in the supplementary material \(*Interaction\-Aware Curriculum*\)\.

### Macro\-Dynamics World Model

##### Macro transition

The WM predicts task\-level physical consequences overKKlow\-level control steps\. At macro steptt, the HLC choosesatH=\(zt,ptg\)a\_\{t\}^\{H\}=\(z\_\{t\},p\_\{t\}^\{g\}\), whereztz\_\{t\}is the structured skill latent andptgp\_\{t\}^\{g\}contains a bounded guidance offset and a waypoint\-advance gate\. Before LLC execution,ptgp\_\{t\}^\{g\}is encoded into the task\-guidance portion ofstLs\_\{t\}^\{L\}\. The frozen LLC then executes the selected latent for up toKKsimulator steps, producing one real macro transition\. We decompose the compact macro state asstH=\(stc,stf\)s\_\{t\}^\{H\}=\(s\_\{t\}^\{c\},s\_\{t\}^\{f\}\), wherestcs\_\{t\}^\{c\}contains continuous task variables \(e\.g\., poses, velocities, displacements, lift heights\) andstfs\_\{t\}^\{f\}contains binary task\-progress flags\. The dynamics network uses\(stH,atH\)\(s\_\{t\}^\{H\},a\_\{t\}^\{H\}\)to predict the next continuous state and progress flags, whereas a separate continuation head estimates whether the episode remains active:

s^t\+1c\\displaystyle\\hat\{s\}\_\{t\+1\}^\{c\}=stc\+fθ​\(stH,atH\),\\displaystyle=s\_\{t\}^\{c\}\+f\_\{\\theta\}\(s\_\{t\}^\{H\},a\_\{t\}^\{H\}\),\(5\)s^t\+1f\\displaystyle\\hat\{s\}\_\{t\+1\}^\{f\}=sigmoid⁡\(gθ​\(stH,atH\)\),\\displaystyle=\\operatorname\{sigmoid\}\\\!\\big\(g\_\{\\theta\}\(s\_\{t\}^\{H\},a\_\{t\}^\{H\}\)\\big\),ζ^t\\displaystyle\\hat\{\\zeta\}\_\{t\}=sigmoid⁡\(hθ​\(stH\)\)\.\\displaystyle=\\operatorname\{sigmoid\}\\\!\\big\(h\_\{\\theta\}\(s\_\{t\}^\{H\}\)\\big\)\.Here,fθf\_\{\\theta\}predicts the change in the continuous variables rather than their absolute next values\. This residual formulation focuses learning on the state change induced by one macro action\. In contrast,gθg\_\{\\theta\}directly predicts the probabilities of the binary flags because their abrupt changes, such as a placed flag switching from zero to one, are poorly represented by continuous residuals\. The continuation predictionζ^t∈\[0,1\]\\hat\{\\zeta\}\_\{t\}\\in\[0,1\]estimates the probability that the transition does not terminate because of failure; it subsequently discounts value propagation through imagined trajectories\.

##### Training process

First, we pretrain the world model with real macro\-transitions collected from the simulator in a supervised fashion\. We prefill a replay buffer with real macro transitions generated by an exploratory policyπrand\\pi\_\{\\mathrm\{rand\}\}that samples structured latents and guidance variables, and run a fixed number of single\-step regression updates\. The world model is optimized with the following loss:

ℒWM=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{WM\}\}=\{\}1dc​‖s^t\+1c−st\+1c‖22\\displaystyle\\frac\{1\}\{d\_\{c\}\}\\left\\lVert\\hat\{s\}\_\{t\+1\}^\{c\}\-s\_\{t\+1\}^\{c\}\\right\\rVert\_\{2\}^\{2\}\(6\)\+λf​BCE​\(s^t\+1f,st\+1f\)\+λζ​BCE​\(ζ^t,ζt\)\.\\displaystyle\+\\lambda\_\{f\}\\,\\mathrm\{BCE\}\\left\(\\hat\{s\}\_\{t\+1\}^\{f\},s\_\{t\+1\}^\{f\}\\right\)\+\\lambda\_\{\\zeta\}\\,\\mathrm\{BCE\}\\left\(\\hat\{\\zeta\}\_\{t\},\\zeta\_\{t\}\\right\)\.Here, the first term is the mean\-squared error on the continuous part, the second term is a binary cross\-entropy on the discrete task flags, and the last term trains the continuation head\. Afterwards, we train the world model andπH\\pi\_\{H\}concurrently using a shared replay buffer\. Each iteration includes three steps: \(i\) rolling out the current policy in the simulator to collect macro\-transitions and adding them to the buffer; \(ii\) updating the world model on mini\-batches sampled from the buffer; and \(iii\) updatingπH\\pi\_\{H\}on short rollouts imagined by the newly updated world model\. Note that the buffer is constantly refreshed with the real experience visited by the currentπH\\pi\_\{H\}at each update\. The supplementary material \(*Macro\-Dynamics World Model*\) specifies the complete loss, target construction, replay\-buffer procedure, and rollout convention\.

### High\-Level Learning in Imagined Dynamics

The goal\-conditioned actorπH,φ​\(atH∣stH,g\)\\pi\_\{H,\\varphi\}\(a\_\{t\}^\{H\}\\mid s\_\{t\}^\{H\},g\)and criticVξ​\(stH,g\)V\_\{\\xi\}\(s\_\{t\}^\{H\},g\)operate at the macro timescale\. At each update, real start states are sampled fromℬ\\mathcal\{B\}and the actor and WM are unrolled forQQimagined macro steps\. The high\-level rewardrtH=rH​\(stH,atH,s^t\+1H,g\)r\_\{t\}^\{H\}=r^\{H\}\(s\_\{t\}^\{H\},a\_\{t\}^\{H\},\\hat\{s\}\_\{t\+1\}^\{H\},g\)uses the same task\-progress definition for real and imagined transitions\. Letvt=Vξ​\(stH,g\)v\_\{t\}=V\_\{\\xi\}\(s\_\{t\}^\{H\},g\)\. We estimate imagined returns using:

Vtλ=rtH\+γ​ζ^t​\[\(1−λret\)​vt\+1\+λret​Vt\+1λ\],VQλ=vQ,V\_\{t\}^\{\\lambda\}=r\_\{t\}^\{H\}\+\\gamma\\hat\{\\zeta\}\_\{t\}\\big\[\(1\-\\lambda\_\{\\mathrm\{ret\}\}\)v\_\{t\+1\}\+\\lambda\_\{\\mathrm\{ret\}\}V\_\{t\+1\}^\{\\lambda\}\\big\],\\quad V\_\{Q\}^\{\\lambda\}=v\_\{Q\},\(7\)whereλret\\lambda\_\{\\mathrm\{ret\}\}is the return\-trace parameter andζ^t\\hat\{\\zeta\}\_\{t\}downweights returns across predicted termination\. Because the high\-level action includes sampled discrete decisions, we optimize the actor with a score\-function estimator:

ℒactor=𝔼τ^​\[∑t=0Q−1dt​ℓtactor\],\\mathcal\{L\}\_\{\\mathrm\{actor\}\}=\\mathbb\{E\}\_\{\\hat\{\\tau\}\}\\left\[\\sum\_\{t=0\}^\{Q\-1\}d\_\{t\}\\ell\_\{t\}^\{\\mathrm\{actor\}\}\\right\],\(8\)
ℓtactor=−sg⁡\(At\)​log⁡πH,φ​\(atH∣stH,g\)−ηH​ℋt,\\ell\_\{t\}^\{\\mathrm\{actor\}\}=\-\\operatorname\{sg\}\(A\_\{t\}\)\\log\\pi\_\{H,\\varphi\}\(a\_\{t\}^\{H\}\\mid s\_\{t\}^\{H\},g\)\-\\eta\_\{H\}\\mathcal\{H\}\_\{t\},\(9\)whereAt=\(Vtλ−vt\)/max⁡\(S,ε\)A\_\{t\}=\(V\_\{t\}^\{\\lambda\}\-v\_\{t\}\)/\\max\(S,\\varepsilon\)is a return\-normalized advantage,dt=∏k=0t−1γ​ζ^kd\_\{t\}=\\prod\_\{k=0\}^\{t\-1\}\\gamma\\hat\{\\zeta\}\_\{k\}is the cumulative continuation discount,SSis a running return scale,ε\>0\\varepsilon\>0prevents division by zero,ℋt\\mathcal\{H\}\_\{t\}is the entropy of the sampled high\-level action distribution, andsg\\operatorname\{sg\}denotes stop\-gradient\. The critic is trained toward detachedVtλV\_\{t\}^\{\\lambda\}targets with a distributional value loss, while an exponential\-moving\-average target critic provides a detached distributional regularizer\.

Algorithm 1LUCID training pipeline0:Motion data

ℳ\\mathcal\{M\}, simulator

Ψ\\Psi, oracle

ψ\\psi, macro period

KK
0:Frozen LLC

πL\\pi\_\{L\}and trained HLC

πH,φ\\pi\_\{H,\\varphi\}
1:

πL←arg⁡maxπ⁡𝒥L​\(π;ℳ,Ψ,ψ\)\\displaystyle\\pi\_\{L\}\\leftarrow\\arg\\max\_\{\\pi\}\\mathcal\{J\}\_\{L\}\(\\pi;\\mathcal\{M\},\\Psi,\\psi\); freeze

πL\\pi\_\{L\}
2:Initialize WM

ℳθ\\mathcal\{M\}\_\{\\theta\}, actor

πH,φ\\pi\_\{H,\\varphi\}, critic

VξV\_\{\\xi\}, target critic

Vξ¯V\_\{\\bar\{\\xi\}\}, and replay buffer

ℬ\\mathcal\{B\}
3:for

i=1,…,Nprei=1,\\ldots,N\_\{\\mathrm\{pre\}\}do

4:

τreal←𝖢𝗈𝗅𝗅𝖾𝖼𝗍K​\(πL,πrand\)\\tau^\{\\mathrm\{real\}\}\\leftarrow\\mathsf\{Collect\}\_\{K\}\(\\pi\_\{L\},\\pi\_\{\\mathrm\{rand\}\}\)
5:Label visited states with

ψ\\psiand insert

τreal\\tau^\{\\mathrm\{real\}\}into

ℬ\\mathcal\{B\}
6:endfor

7:for

i=1,…,NWMi=1,\\ldots,N\_\{\\mathrm\{WM\}\}do

8:

θ←θ−ηθ​∇θℒWM​\(θ;𝖲𝖺𝗆𝗉𝗅𝖾​\(ℬ\)\)\\theta\\leftarrow\\theta\-\\eta\_\{\\theta\}\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{WM\}\}\(\\theta;\\mathsf\{Sample\}\(\\mathcal\{B\}\)\)
9:endfor

10:repeat

11:

τreal←𝖢𝗈𝗅𝗅𝖾𝖼𝗍K​\(πL,πH,φ\)\\tau^\{\\mathrm\{real\}\}\\leftarrow\\mathsf\{Collect\}\_\{K\}\(\\pi\_\{L\},\\pi\_\{H,\\varphi\}\)
12:Label its visited states with

ψ\\psiand refresh

ℬ\\mathcal\{B\}
13:

θ←θ−ηθ​∇θℒWM​\(θ;𝖲𝖺𝗆𝗉𝗅𝖾​\(ℬ\)\)\\theta\\leftarrow\\theta\-\\eta\_\{\\theta\}\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{WM\}\}\(\\theta;\\mathsf\{Sample\}\(\\mathcal\{B\}\)\)
14:

𝒟∼ℬ\\mathcal\{D\}\\sim\\mathcal\{B\}
15:

τ^←𝖨𝗆𝖺𝗀𝗂𝗇𝖾Q​\(𝒟,πH,φ,ℳθ\)\\hat\{\\tau\}\\leftarrow\\mathsf\{Imagine\}\_\{Q\}\(\\mathcal\{D\},\\pi\_\{H,\\varphi\},\\mathcal\{M\}\_\{\\theta\}\)
16:

ξ←ξ−ηξ​∇ξℒcritic​\(ξ;τ^\)\\xi\\leftarrow\\xi\-\\eta\_\{\\xi\}\\nabla\_\{\\xi\}\\mathcal\{L\}\_\{\\mathrm\{critic\}\}\(\\xi;\\hat\{\\tau\}\)
17:

φ←φ−ηφ​∇φ\[ℒactor​\(φ;τ^\)\+βBC,k​ℒBC​\(φ;𝒟,ψ\)\]\\varphi\\leftarrow\\varphi\-\\eta\_\{\\varphi\}\\nabla\_\{\\varphi\}\\left\[\\mathcal\{L\}\_\{\\mathrm\{actor\}\}\(\\varphi;\\hat\{\\tau\}\)\+\\beta\_\{\\mathrm\{BC\},k\}\\mathcal\{L\}\_\{\\mathrm\{BC\}\}\(\\varphi;\\mathcal\{D\},\\psi\)\\right\]
18:

ξ¯←\(1−α\)​ξ¯\+α​ξ\\bar\{\\xi\}\\leftarrow\(1\-\\alpha\)\\bar\{\\xi\}\+\\alpha\\xi;

βBC,k\+1←𝖣𝖾𝖼𝖺𝗒​\(βBC,k\)\\beta\_\{\\mathrm\{BC\},k\+1\}\\leftarrow\\mathsf\{Decay\}\(\\beta\_\{\\mathrm\{BC\},k\}\)
19:untilthe training budget is exhausted

20:return

πL,πH,φ\\pi\_\{L\},\\pi\_\{H,\\varphi\}

To stabilize early skill discovery, we warm\-start the categorical skill\-selection component from a geometry\-based oracle\. Its behavior\-cloning coefficient is annealed to zero, so the oracle initializes rather than fixes the final controller\.

Algorithm[1](https://arxiv.org/html/2608.07746#alg1)summarizes the full optimization procedure for LUCID\. Here,𝒥L\\mathcal\{J\}\_\{L\}is the LLC objective in Eq\. \([1](https://arxiv.org/html/2608.07746#Sx3.E1)\),ℳθ\\mathcal\{M\}\_\{\\theta\}is the macro\-dynamics world model, andθ\\theta,φ\\varphi, andξ\\xidenote the WM, actor, and critic parameters, respectively\. The oracle\-labeled buffer is used only whileβBC,k\>0\\beta\_\{\\mathrm\{BC\},k\}\>0\. The exact joint action distribution, critic loss, oracle objective, and actor–WM update schedule are provided in the supplementary material \(*High\-Level Controller Training*\)\.

## Experiments

We evaluate LUCID on long\-horizon multi\-object rearrangement, focusing on three questions: whether it improves sequential task completion over representative baselines, whether it remains effective on held\-out layouts and longer task chains, and which components account for its performance\. We first define the shared evaluation protocol, then report comparison results for diverse settings, followed by world\-model and latent\-interface ablations\.

### Experimental Setup

#### Dataset

We construct multi\-object rearrangement tasks from HITR\(Xuet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib6)\)\. Movable objects are initialized on the ground in diverse physically feasible configurations across warehouse, kitchen, bedroom, and living\-room layouts\. We compose single\-object instances into sequential episodes in which each object must be placed before the next subtask becomes active\. The ID split contains 62 training tasks, whereas the OOD split contains 20 tasks\. The standard benchmark contains two\-object chains; a separate extension evaluates chains of up to five objects\. LLC training additionally uses reference motions from OMOMO\(Liet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib7)\)and SAMP\(Hassanet al\.[2021](https://arxiv.org/html/2608.07746#bib.bib8)\)\.

#### Metrics

We report three metrics\.Success RateSRkis the percentage of episodes in which the firstkkobjects are placed within0\.2​m0\.2\\,\\mathrm\{m\}of their targets\.Average Placement Error\(APE\) is the final object\-to\-goal distance averaged over all task objects\.Timeis the mean simulated completion time over successful episodes\. Values are mean±\\pmsample standard deviation over three evaluation seeds using the same task split and success criterion for every method\.

#### Baselines

We compare against the following baselines using task\-specific deployment adapters\.InterMimic\(Wanget al\.[2025b](https://arxiv.org/html/2608.07746#bib.bib20)\)is a motion\-tracking policy; for each subtask, we construct a kinematic human–object reference by warping motion segments to the required waypoints and object poses\.TokenHSI\(Panet al\.[2025](https://arxiv.org/html/2608.07746#bib.bib17)\)uses a scripted finite\-state adapter to switch between locomotion and object\-carrying skills\.HumanVLA\(Xuet al\.[2024](https://arxiv.org/html/2608.07746#bib.bib6)\)is an end\-to\-end single\-object rearrangement policy, which we extend by activating subtasks sequentially\. All methods use the same ID/OOD task lists, simulator horizon, placement threshold, and evaluation seeds\. Because these baselines do not model autonomous inter\-subtask transitions, their adapters reset only the humanoid to a favorable pose at each handoff\. Consequently, SR and APE compare LUCID against favorably initialized baseline deployments, while completion times are descriptive and not directly comparable\.

### Implementation Details

We use Isaac Gym\(Makoviychuket al\.[2021](https://arxiv.org/html/2608.07746#bib.bib9)\)and a humanoid with 15 rigid bodies and 28 PD\-controlled joints\(Hassanet al\.[2023](https://arxiv.org/html/2608.07746#bib.bib10)\)\. The LLC follows ASE but replaces its learned skill encoder and diversity objective with the structured latent interface\. We train it with PPO\(Schulmanet al\.[2017](https://arxiv.org/html/2608.07746#bib.bib11)\)and a staged curriculum that progresses from robust single\-object interaction to two\-object chaining\. The HLC uses a DreamerV3\-style actor–critic, and the deterministic world model is a plain MLP over the compact task state\. The simulator runs at60​Hz60\\,\\mathrm\{Hz\}, the LLC acts at30​Hz30\\,\\mathrm\{Hz\}, and the HLC selects one macro action every 20 LLC steps\. The HLC uses a 12\-step imagination horizon, sequence length 32, and batch size 64\. We train the LLC on two NVIDIA RTX 4090 GPUs and train the HLC and world model jointly with 8192 parallel environments\. Further architecture, optimization, and evaluation details are provided in the supplementary material\.

Table 1:Comparison on the shared ID and OOD task splits\. SRkdenotes completion through subtaskkk\.![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/livingroom_row6.png)

![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/bedroom_row6.png)

![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/kitchen_row6.png)

![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/warehouse_row6.png)

Figure 2:Qualitative filmstrips of LUCID in various task layouts \(top to bottom: living room, bedroom, kitchen, warehouse\)\. Colors indicate semantic latent decisions: locomotion \(blue\), transport \(green\), release \(red\), fetch \(yellow\), and transit \(purple\)\. The frames illustrate continuous navigation, object interaction, and autonomous handoff without resetting the humanoid\.![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/extended_task_comparison.png)Figure 3:Success rates against the number of objects\.![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/per_dim_rmse.png)Figure 4:World\-model one\-step RMSE on task\-relevant compact\-state dimensions\. Bars average per\-environment RMSE and whiskers denote the standard error across evaluation environments\.![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/trajectory_compact_env0.png)Figure 5:Multi\-step world\-model predictions\. Black curves show simulator states, dashed blue curves show one\-step predictions, and orange curves show imagined rollouts\.![Refer to caption](https://arxiv.org/html/2608.07746v1/Figures/e1_fair_2x2.png)Figure 6:Behavior controllability of the structured and unstructured latent interface\. Top: separately fitted t\-SNE visualizations of identical kinematic descriptors\. Bottom left: five\-fold random\-forest balanced accuracy for command decoding\. Bottom right: mean fraction of time each command lifts the object, with0\.40\.4used as the sustained threshold\.Table 2:WM ablation under sparse and dense rewards\.Table 3:Structured\-interface ablation\. Each HLC is trained and evaluated with its corresponding frozen LLC interface under the same ID/OOD protocol\.
### Comparison Experiments

#### Quantitative evaluation

Table[1](https://arxiv.org/html/2608.07746#Sx4.T1)compares all methods on the shared ID and OOD protocols\. LUCID obtains the highest mean SR1, SR2, and the lowest APE on both splits despite receiving no handoff reset\. On ID tasks, its SR2exceeds the strongest baseline by33\.633\.6percentage points \(73\.4%73\.4\\%versus39\.8%39\.8\\%\); on OOD tasks, the margin is31\.431\.4points \(68\.4%68\.4\\%versus37\.0%37\.0\\%\)\. The smaller SR1–SR2gap for LUCID \(15\.815\.8points ID and15\.515\.5points OOD\) indicates more reliable handoffs than the baselines\. From ID to OOD, LUCID decreases by5\.35\.3points in SR1and5\.05\.0points in SR2, while retaining the best mean placement accuracy\. InterMimic performs worst overall, consistent with the difficulty of constructing precise warped references for diverse object configurations\. Baseline completion times are shorter because their handoff adapters teleport the humanoid; we therefore treat Time as descriptive rather than a direct efficiency comparison\.

#### Qualitative performance

Figure[2](https://arxiv.org/html/2608.07746#Sx4.F2)shows LUCID completing multi\-object rearrangement chains in diverse layouts\. We can see that LUCID exhibits clean phase structures and well\-separated transitions across all scenes, with HLC producing interpretable skills that are faithfully executed by the frozen LLC\. The same latent repertoire generalizes across layouts without per\-scene tuning, adapting to environmental obstacles while completing the task chain fully autonomously\. These are consistent with the SR2 and placement\-accuracy gains reported in Table[1](https://arxiv.org/html/2608.07746#Sx4.T1)\. The supplementary material provides the corresponding evaluation protocol and implementation details\.

#### Extended multi\-objects tasks

To assess long\-horizon robustness beyond the two\-object training horizon, we evaluate chains of up to five objects and report prefix success SRkin Figure[3](https://arxiv.org/html/2608.07746#Sx4.F3)\. InterMimic and TokenHSI approach zero by SR4, and HumanVLA retains5%5\\%success at SR5\. This rapid decay reflects their reliance on scripted planners, which accumulate handoff errors without recovery\. However, LUCID degrades more gracefully, achieving around56%56\\%at SR3 and still21%21\\%at SR5\. We attribute this to the joint evolution of the world model and high\-level policy during training, which together yields a more reliable long\-horizon planner\.

### Ablation Studies

We investigate the key components of LUCID by addressing the following questions: 1\)*Does the world model improve HLC learning through imagined rollouts?*2\)*Does the structured latent interface improve behavior controllability?*

#### On world model effectiveness

We first evaluate predictive accuracy independently of downstream task success\. Figure[4](https://arxiv.org/html/2608.07746#Sx4.F4)reports one\-step errors for task\-critical dimensions, including object–goal displacement, lift height, hand–object distance, and placement progress\. Figure[5](https://arxiv.org/html/2608.07746#Sx4.F5)then visualizes a representative 12\-step open\-loop rollout\. The supplementary material specifies the action sequence, sampling protocol, and aggregate horizon\-wise errors used alongside this qualitative trajectory\. Table[2](https://arxiv.org/html/2608.07746#Sx4.T2)compares HLCs trained with world\-model \(WM\) imagination against model\-free \(MF\) learning under both sparse and dense reward settings\. The WM variants achieve higher mean SR2in all four split–reward comparisons:73\.573\.5versus59\.359\.3and74\.074\.0versus61\.961\.9on ID, and63\.163\.1versus42\.842\.8and50\.250\.2versus33\.133\.1on OOD\. This consistent advantage supports the use of imagined macro\-transitions for learning long\-horizon task progression\. Dense shaping yields only a0\.50\.5\-point ID SR2gain for LUCID but reduces OOD SR2by12\.912\.9points, suggesting weaker transfer than the sparse setting\.

#### On structured latent interface

Figure[6](https://arxiv.org/html/2608.07746#Sx4.F6)separates command decodability from task relevance\. We drive each frozen LLC with five commands in the same two\-object environment, retain states in which the object is initially reachable, and describe the resulting motions with identical kinematic features\. The independently fitted t\-SNE plots and five\-fold random\-forest probe show that both interfaces produce command\-decodable behaviors \(balanced accuracy0\.860\.86and0\.950\.95\)\. Decodability alone is therefore insufficient\. The manipulation panel instead shows that three structured commands exceed the sustained\-lift threshold, whereas none of the unstructured commands does\. Table[3](https://arxiv.org/html/2608.07746#Sx4.T3)further shows that the structured interface reaches74\.2%74\.2\\%ID and66\.4%66\.4\\%OOD SR2, while the unstructured interface obtains0\.0%0\.0\\%SR2on both splits\. These results indicate that the relevant property is not merely diverse behavior, but a command set aligned with the interaction stages required by the downstream task\.

## Conclusion

In this paper, we introduced LUCID, a hierarchical framework for multi\-object rearrangement with a simulated humanoid\. It coordinates a structured low\-level controller using a high\-level policy trained through imagined skill\-level transitions, improving long\-horizon completion\. Results show that task\-aligned temporal abstraction is central: a structured skill interface makes high\-level decisions meaningful, while macro\-level prediction supports long\-horizon reasoning without reproducing joint\-level dynamics\. Current limitations include privileged state inputs and limited task, object, and embodiment diversity; future work will address visual control, broader tasks, and real\-world transfer\.

## References

- P\. Bacon, J\. Harb, and D\. Precup \(2017\)The option\-critic architecture\.InProceedings of the AAAI conference on artificial intelligence,Vol\.31\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- P\. Faloutsos, M\. Van de Panne, and D\. Terzopoulos \(2001\)Composable controllers for physics\-based character animation\.InProceedings of the 28th annual conference on Computer graphics and interactive techniques,pp\. 251–260\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- Z\. Gu, J\. Li, W\. Shen, W\. Yu, Z\. Xie, S\. McCrory, X\. Cheng, A\. Shamsah, R\. Griffin, C\. K\. Liu,et al\.\(2026\)Humanoid locomotion and manipulation: current progress and challenges in control, planning, and learning\.IEEE/ASME Transactions on Mechatronics31\(2\),pp\. 2300–2330\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1)\.
- D\. Ha and J\. Schmidhuber \(2018\)World models\.arXiv preprint arXiv:1803\.101222\(3\),pp\. 440\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p3.1),[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- D\. Hafner, T\. Lillicrap, J\. Ba, and M\. Norouzi \(2019\)Dream to control: learning behaviors by latent imagination\.arXiv preprint arXiv:1912\.01603\.Cited by:[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- D\. Hafner, T\. Lillicrap, M\. Norouzi, and J\. Ba \(2020\)Mastering atari with discrete world models\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p3.1),[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- D\. Hafner, J\. Pasukonis, J\. Ba, and T\. Lillicrap \(2025\)Mastering diverse control tasks through world models\.Nature640\(8059\),pp\. 647–653\.Cited by:[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1),[Method](https://arxiv.org/html/2608.07746#Sx3.p1.1)\.
- N\. Hansen, H\. Su, and X\. Wang \(2024\)Td\-mpc2: scalable, robust world models for continuous control\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 47376–47405\.Cited by:[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- N\. Hansen, J\. SV, V\. Sobal, Y\. LeCun, X\. Wang, and H\. Su \(2025\)Hierarchical world models as visual whole\-body humanoid controllers\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 62175–62195\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- N\. Hansen, X\. Wang, and H\. Su \(2022\)Temporal difference learning for model predictive control\.arXiv preprint arXiv:2203\.04955\.Cited by:[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- L\. Hasenclever, F\. Pardo, R\. Hadsell, N\. Heess, and J\. Merel \(2020\)CoMic: complementary task learning & mimicry for reusable skills\.InInternational Conference on Machine Learning,pp\. 4105–4115\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- M\. Hassan, D\. Ceylan, R\. Villegas, J\. Saito, J\. Yang, Y\. Zhou, and M\. J\. Black \(2021\)Stochastic scene\-aware motion prediction\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 11374–11384\.Cited by:[Dataset](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx1.p1.1)\.
- M\. Hassan, Y\. Guo, T\. Wang, M\. Black, S\. Fidler, and X\. B\. Peng \(2023\)Synthesizing physical character\-scene interactions\.InACM SIGGRAPH 2023 Conference Proceedings,pp\. 1–9\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[Implementation Details](https://arxiv.org/html/2608.07746#Sx4.SSx2.p1.2)\.
- N\. Heess, G\. Wayne, Y\. Tassa, T\. Lillicrap, M\. Riedmiller, and D\. Silver \(2016\)Learning and transfer of modulated locomotor controllers\.arXiv preprint arXiv:1610\.05182\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- Y\. Ji, Z\. Li, Y\. Sun, X\. B\. Peng, S\. Levine, G\. Berseth, and K\. Sreenath \(2022\)Hierarchical reinforcement learning for precise soccer shooting skills using a quadrupedal robot\.In2022 IEEE/RSJ International Conference on Intelligent Robots and Systems \(IROS\),pp\. 1479–1486\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- Y\. Kuang, H\. Geng, A\. Elhafsi, T\. Do, P\. Abbeel, J\. Malik, M\. Pavone, and Y\. Wang \(2025\)Skillblender: towards versatile humanoid whole\-body loco\-manipulation via skill blending\.arXiv preprint arXiv:2506\.09366\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- K\. N\. Kumar, I\. Essa, and S\. Ha \(2023\)Cascaded compositional residual learning for complex interactive behaviors\.IEEE Robotics and Automation Letters8\(8\),pp\. 4601–4608\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- C\. Li, A\. Krause, and M\. Hutter \(2025\)Robotic world model: a neural network simulator for robust policy optimization in robotics\.arXiv preprint arXiv:2501\.10100\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p3.1),[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- D\. Li, X\. Chen, Q\. Wu, B\. Chen, S\. Wu, H\. Wu, G\. Zhang, L\. Li, M\. Zhou, D\. Xiang,et al\.\(2026\)Haic: humanoid agile object interaction control via dynamics\-aware world model\.arXiv preprint arXiv:2602\.11758\.Cited by:[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- J\. Li, J\. Wu, and C\. K\. Liu \(2023\)Object motion guided human motion synthesis\.ACM Transactions on Graphics \(TOG\)42\(6\),pp\. 1–11\.Cited by:[Dataset](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx1.p1.1)\.
- Z\. Luo, J\. Cao, S\. Christen, A\. Winkler, K\. Kitani, and W\. Xu \(2024\)Omnigrasp: grasping diverse objects with simulated humanoids\.Advances in Neural Information Processing Systems37,pp\. 2161–2184\.Cited by:[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- V\. Makoviychuk, L\. Wawrzyniak, Y\. Guo, M\. Lu, K\. Storey, M\. Macklin, D\. Hoeller, N\. Rudin, A\. Allshire, A\. Handa,et al\.\(2021\)Isaac gym: high performance gpu\-based physics simulation for robot learning\.arXiv preprint arXiv:2108\.10470\.Cited by:[Implementation Details](https://arxiv.org/html/2608.07746#Sx4.SSx2.p1.2)\.
- J\. Merel, L\. Hasenclever, A\. Galashov, A\. Ahuja, V\. Pham, G\. Wayne, Y\. W\. Teh, and N\. Heess \(2018\)Neural probabilistic motor primitives for humanoid control\.arXiv preprint arXiv:1811\.11711\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- L\. Pan, J\. Wang, B\. Huang, J\. Zhang, H\. Wang, X\. Tang, and Y\. Wang \(2024\)Synthesizing physically plausible human motions in 3d scenes\.In2024 International Conference on 3D Vision \(3DV\),pp\. 1498–1507\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1)\.
- L\. Pan, Z\. Yang, Z\. Dou, W\. Wang, B\. Huang, B\. Dai, T\. Komura, and J\. Wang \(2025\)Tokenhsi: unified synthesis of physical human\-scene interactions through task tokenization\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 5379–5391\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[Baselines](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx3.p1.1)\.
- X\. B\. Peng, P\. Abbeel, S\. Levine, and M\. Van de Panne \(2018\)Deepmimic: example\-guided deep reinforcement learning of physics\-based character skills\.ACM Transactions On Graphics \(TOG\)37\(4\),pp\. 1–14\.Cited by:[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- X\. B\. Peng, M\. Chang, G\. Zhang, P\. Abbeel, and S\. Levine \(2019\)Mcp: learning composable hierarchical control with multiplicative compositional policies\.Advances in neural information processing systems32\.Cited by:[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- X\. B\. Peng, Y\. Guo, L\. Halper, S\. Levine, and S\. Fidler \(2022\)Ase: large\-scale reusable adversarial skill embeddings for physically simulated characters\.ACM Transactions On Graphics \(TOG\)41\(4\),pp\. 1–17\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[Method](https://arxiv.org/html/2608.07746#Sx3.p1.1)\.
- X\. B\. Peng, Z\. Ma, P\. Abbeel, S\. Levine, and A\. Kanazawa \(2021\)Amp: adversarial motion priors for stylized physics\-based character control\.ACM Transactions on Graphics \(ToG\)40\(4\),pp\. 1–20\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[Implementation Details](https://arxiv.org/html/2608.07746#Sx4.SSx2.p1.2)\.
- S\. Starke, H\. Zhang, T\. Komura, and J\. Saito \(2019\)Neural state machine for character\-scene interactions\.ACM Transactions on Graphics38\(6\),pp\. 178\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1)\.
- R\. S\. Sutton, D\. Precup, and S\. Singh \(1999\)Between mdps and semi\-mdps: a framework for temporal abstraction in reinforcement learning\.Artificial intelligence112\(1\-2\),pp\. 181–211\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- C\. Tessler, Y\. Guo, O\. Nabati, G\. Chechik, and X\. B\. Peng \(2024\)Maskedmimic: unified physics\-based character control through masked motion inpainting\.ACM Transactions On Graphics \(TOG\)43\(6\),pp\. 1–21\.Cited by:[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- C\. Tessler, Y\. Kasten, Y\. Guo, S\. Mannor, G\. Chechik, and X\. B\. Peng \(2023\)Calm: conditional adversarial latent models for directable virtual characters\.InACM SIGGRAPH 2023 conference proceedings,pp\. 1–9\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[Hierarchical Reinforcement Learning](https://arxiv.org/html/2608.07746#Sx2.SSx2.p1.1)\.
- H\. Wang, W\. Zhang, R\. Yu, T\. Huang, J\. Ren, F\. Jia, Z\. Wang, X\. Niu, X\. Chen, J\. Chen,et al\.\(2025a\)Physhsi: towards a real\-world generalizable and natural humanoid\-scene interaction system\.arXiv preprint arXiv:2510\.11072\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- Y\. Wang, J\. Lin, A\. Zeng, Z\. Luo, J\. Zhang, and L\. Zhang \(2023\)Physhoi: physics\-based imitation of dynamic human\-object interaction\.arXiv preprint arXiv:2312\.04393\.Cited by:[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- Y\. Wang, Q\. Zhao, R\. Yu, H\. W\. Tsui, A\. Zeng, J\. Lin, Z\. Luo, J\. Yu, X\. Li, Q\. Chen,et al\.\(2025b\)Skillmimic: learning basketball interaction skills from demonstrations\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 17540–17549\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Baselines](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx3.p1.1)\.
- H\. Weng, Y\. Li, N\. Sobanbabu, Z\. Wang, Z\. Luo, T\. He, D\. Ramanan, and G\. Shi \(2025\)Hdmi: learning interactive humanoid whole\-body control from human videos\.arXiv preprint arXiv:2509\.16757\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1)\.
- P\. Wu, A\. Escontrela, D\. Hafner, P\. Abbeel, and K\. Goldberg \(2023\)Daydreamer: world models for physical robot learning\.InConference on robot learning,pp\. 2226–2240\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p3.1),[World Model for Robotics](https://arxiv.org/html/2608.07746#Sx2.SSx3.p1.1)\.
- Z\. Xiao, T\. Wang, J\. Wang, J\. Cao, W\. Zhang, B\. Dai, D\. Lin, and J\. Pang \(2024\)Unified human\-scene interaction via prompted chain\-of\-contacts\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 24450–24461\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p1.1)\.
- S\. Xu, H\. Y\. Ling, Y\. Wang, and L\. Gui \(2025\)Intermimic: towards universal whole\-body control for physics\-based human\-object interactions\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 12266–12277\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1)\.
- X\. Xu, Y\. Zhang, Y\. Li, L\. Han, and C\. Lu \(2024\)Humanvla: towards vision\-language directed object rearrangement by physical humanoid\.Advances in Neural Information Processing Systems37,pp\. 18633–18659\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1),[Physics\-Based Motion Generation](https://arxiv.org/html/2608.07746#Sx2.SSx1.p1.1),[Dataset](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx1.p1.1),[Baselines](https://arxiv.org/html/2608.07746#Sx4.SSx1.SSSx3.p1.1)\.
- Q\. Zhu, H\. Zhang, M\. Lan, and L\. Han \(2023\)Neural categorical priors for physics\-based character control\.ACM Transactions on Graphics \(TOG\)42\(6\),pp\. 1–16\.Cited by:[Introduction](https://arxiv.org/html/2608.07746#Sx1.p2.1)\.

Similar Articles

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Hugging Face Daily Papers

InterEvolve introduces test-time evolution of reward programs for humanoid loco-manipulation, using an object-aware forward-backward behavioral foundation model and an LLM agent that revises staged reward programs in-context to unlock untrained skills, with evolved behaviors deployed autonomously on a physical Unitree G1 robot.

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Hugging Face Daily Papers

PRISM is a real-to-sim-to-real framework that amplifies a few real human-object interaction videos into hundreds of diverse counterfactual videos via V2V generation, then reconstructs physically plausible motions to train a generalizable humanoid loco-manipulation policy deployed on a real robot without real-world fine-tuning.