InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
Summary
InterEvolve introduces test-time evolution of reward programs for humanoid loco-manipulation, using an object-aware forward-backward behavioral foundation model and an LLM agent that revises staged reward programs in-context to unlock untrained skills, with evolved behaviors deployed autonomously on a physical Unitree G1 robot.
View Cached Full Text
Cached at: 10/02/26, 08:28 AM
Paper page - InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
Source: https://huggingface.co/papers/2610.02196
Abstract
Westudytest-timeevolutionforhumanoidloco-manipulation:solvingtasksthatacontrollerwasnevertrainedforbyrepurposingitsexistingskills,improvingfromitsownattempts,andretainingwhatitlearns,withoutretraining.Ourkeyinsightisthatabroadcontrolleralreadyholdsmuchofthecompetenceanewtaskneeds,andthatthiscompetencebecomesaccessiblethroughaninterfacebetweenplanningandcontrolthatisexpressiveenoughtospecifycontact-rich,multi-stageinteractions,yetexecutableandmeasurableenoughthatexecutionfeedbackcanguideplanningfromexperience.InterEvolverealizesthisinterfacewithtwocomponents.First,wedevelopanobject-awareforward-backward(FB)behavioralfoundationmodel,whoseobjectresidualsonafrozenbodypriorturnanewrewardaboutthebodyorobjectsintoloco-manipulationbehaviorattesttime.Second,wespecifytasksasrewardprograms:stagedrewardswithcompletionconditionsandtunableconstants.Alargelanguagemodel(LLM)agentrevisestheprogramstructureincontext,drawingonexecutionfeedbackandaskilllibraryofverifiedprograms,whileanumericaloptimizertunesitsconstants.Witheverycandidateverifiedacrossparallelsimulationscenarios,theprogramexploresnewwaystoinduce,repurpose,andcomposethecontroller’sexistingmotorcompetenceforthetaskathand,andthusimprovesoveriterations.Experimentsshowthathuman-designedrewardsleavemuchoftheFBmodel’sloco-manipulationcompetenceuntapped,whereastheprogramsInterEvolveevolvesreleaseit,sometimesthroughnovelstrategies.Itfurtherproducesbehaviorsfordiversetasks,complexscenes,andlong-horizoncompositionsinsimulation,andevolvedskillsrunautonomouslyonaphysicalUnitreeG1fromegocentriconboardperception.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2610\.02196
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2610.02196 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2610.02196 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2610.02196 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
EvoTrainer introduces an autonomous training framework that co-evolves LLM policies and training harnesses through empirical feedback, outperforming human-engineered RL baselines on mathematical reasoning, code generation, and long-horizon software engineering tasks.
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
RoboEvolve is a framework that co-evolves a VLM planner and VGM simulator for robotic manipulation, achieving data efficiency with only 500 unlabeled seed images and robust continual learning.
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation
This paper introduces LUCID, a hierarchical model-based reinforcement learning framework for long-horizon humanoid loco-manipulation. It learns reusable latent skills and a macro-dynamics world model, enabling high-level planning via imagined rollouts and improving success rates in simulated multi-object rearrangement tasks.
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
EvoTest introduces J-TTL, a benchmark for measuring agent test-time learning capabilities, and proposes an evolutionary framework where an Actor Agent plays games while an Evolver Agent iteratively improves the system's prompts, memory, and hyperparameters without fine-tuning. The method demonstrates superior performance compared to reflection and memory-based baselines on complex text-based games.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.