Discovering Reinforcement Learning Interfaces with Large Language Models

Hugging Face Daily Papers Papers

Summary

This paper introduces LIMEN, an LLM-guided evolutionary framework that automatically discovers reinforcement learning interfaces by jointly optimizing observation mappings and reward functions from raw simulator states. The approach reduces manual engineering effort and demonstrates that co-designing observations and rewards outperforms optimizing either component alone.

Reinforcement learning systems rely on environment interfaces that specify observations and reward functions, yet constructing these interfaces for new tasks often requires substantial manual effort. While recent work has automated reward design using large language models (LLMs), these approaches assume fixed observations and do not address the broader challenge of synthesizing complete task interfaces. We study RL task interface discovery from raw simulator state, where both observation mappings and reward functions must be generated. We propose LIMEN (Code available at https://github.com/Lossfunk/LIMEN), a LLM guided evolutionary framework that produces candidate interfaces as executable programs and iteratively refines them using policy training feedback. Across novel discrete gridworld tasks and continuous control domains spanning locomotion and manipulation, joint evolution of observations and rewards discovers effective interfaces given only a trajectory-level success metric, while optimizing either component alone fails on at least one domain. These results demonstrate that automatic construction of RL interfaces from raw state can substantially reduce manual engineering and that observation and reward components often benefit from co-design, as single-component optimization fails catastrophically on at least one domain in our evaluation suite.
Original Article
View Cached Full Text

Cached at: 05/11/26, 02:52 PM

Paper page - Discovering Reinforcement Learning Interfaces with Large Language Models

Source: https://huggingface.co/papers/2605.03408

Abstract

Automated reinforcement learning interface discovery using LLM-guided evolutionary algorithms that jointly optimize observation mappings and reward functions from raw simulator state.

Reinforcement learningsystems rely onenvironment interfacesthat specify observations andreward functions, yet constructing these interfaces for new tasks often requires substantial manual effort. While recent work has automated reward design usinglarge language models(LLMs), these approaches assume fixed observations and do not address the broader challenge of synthesizing complete task interfaces. We study RL task interface discovery from raw simulator state, where bothobservation mappingsandreward functionsmust be generated. We propose LIMEN (Code available at https://github.com/Lossfunk/LIMEN), a LLM guidedevolutionary frameworkthat produces candidate interfaces as executable programs and iteratively refines them usingpolicy trainingfeedback. Across novel discrete gridworld tasks and continuous control domains spanning locomotion and manipulation,joint evolutionof observations and rewards discovers effective interfaces given only atrajectory-level success metric, while optimizing either component alone fails on at least one domain. These results demonstrate that automatic construction of RL interfaces from raw state can substantially reduce manual engineering and that observation and reward components often benefit fromco-design, as single-component optimization fails catastrophically on at least one domain in our evaluation suite.

View arXiv pageView PDFProject pageGitHub3Add to collection

Get this paper in your agent:

hf papers read 2605\.03408

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.03408 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.03408 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.03408 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning

arXiv cs.CL

This paper proposes the LLM-as-Environment-Engineer framework, where a policy model analyzes failures to automatically redesign the training environment for reinforcement learning, and introduces MAPF-FrozenLake as a controllable testbed. The framework, using Qwen3-4B, outperforms larger models like GPT and Gemini, showing that policy learning improves the model's ability to diagnose weaknesses.