InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Hugging Face Daily Papers Papers

Summary

InterEvolve introduces test-time evolution of reward programs for humanoid loco-manipulation, using an object-aware forward-backward behavioral foundation model and an LLM agent that revises staged reward programs in-context to unlock untrained skills, with evolved behaviors deployed autonomously on a physical Unitree G1 robot.

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
Original Article
View Cached Full Text

Cached at: 10/02/26, 08:28 AM

Paper page - InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Source: https://huggingface.co/papers/2610.02196

Abstract

Westudytest-timeevolutionforhumanoidloco-manipulation:solvingtasksthatacontrollerwasnevertrainedforbyrepurposingitsexistingskills,improvingfromitsownattempts,andretainingwhatitlearns,withoutretraining.Ourkeyinsightisthatabroadcontrolleralreadyholdsmuchofthecompetenceanewtaskneeds,andthatthiscompetencebecomesaccessiblethroughaninterfacebetweenplanningandcontrolthatisexpressiveenoughtospecifycontact-rich,multi-stageinteractions,yetexecutableandmeasurableenoughthatexecutionfeedbackcanguideplanningfromexperience.InterEvolverealizesthisinterfacewithtwocomponents.First,wedevelopanobject-awareforward-backward(FB)behavioralfoundationmodel,whoseobjectresidualsonafrozenbodypriorturnanewrewardaboutthebodyorobjectsintoloco-manipulationbehaviorattesttime.Second,wespecifytasksasrewardprograms:stagedrewardswithcompletionconditionsandtunableconstants.Alargelanguagemodel(LLM)agentrevisestheprogramstructureincontext,drawingonexecutionfeedbackandaskilllibraryofverifiedprograms,whileanumericaloptimizertunesitsconstants.Witheverycandidateverifiedacrossparallelsimulationscenarios,theprogramexploresnewwaystoinduce,repurpose,andcomposethecontroller’sexistingmotorcompetenceforthetaskathand,andthusimprovesoveriterations.Experimentsshowthathuman-designedrewardsleavemuchoftheFBmodel’sloco-manipulationcompetenceuntapped,whereastheprogramsInterEvolveevolvesreleaseit,sometimesthroughnovelstrategies.Itfurtherproducesbehaviorsfordiversetasks,complexscenes,andlong-horizoncompositionsinsimulation,andevolvedskillsrunautonomouslyonaphysicalUnitreeG1fromegocentriconboardperception.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2610\.02196

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.02196 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.02196 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.02196 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems

arXiv cs.CL

EvoTest introduces J-TTL, a benchmark for measuring agent test-time learning capabilities, and proposes an evolutionary framework where an Actor Agent plays games while an Evolver Agent iteratively improves the system's prompts, memory, and hyperparameters without fine-tuning. The method demonstrates superior performance compared to reflection and memory-based baselines on complex text-based games.

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

arXiv cs.CL

CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.