TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
Summary
TGRL proposes a temperature-grouped reinforcement learning method to enhance exploration in large language models, achieving faster training and improved performance across multiple benchmarks.
View Cached Full Text
Cached at: 09/30/26, 04:22 AM
Paper page - TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
Source: https://huggingface.co/papers/2609.33589 Published on Sep 27
·
Submitted byhttps://huggingface.co/lin1111987
zihanon Sep 30
Abstract
Efficientexplorationoftenremainsacentralbottleneckinreinforcementlearningwithverifiablerewards(RLVR).Althoughtemperaturecontrolandtest-timescalingstrategiescanincreaserolloutdiversityoflargelanguagemodels(LLMs),theyeitherexpandthesamplebudgetatrollouttimeorleavethebenefitofexplorationunquantified.Tothisend,weproposeTemperature-GroupedReinforcementLearning(TGRL),whichturnstemperature-induceddiversityintoanexplicittrainingsignal.Foreachprompt,TGRLpartitionsitsrolloutgroupintolow-andhigh-temperaturesubsets,estimatesexplorationgainthroughtheirrewardcontrast,andallocatesthisgroup-levelsignalastoken-levelcreditusingJensen--Shannon(JS)divergencebetweenthecorrespondingtemperature-scalednext-tokendistributionsinducedbythesamelogits.Notably,TGRLreachesequivalentaccuracyupto36%fasterthanstrongRLVRbaselineswithoutexpandingtherolloutbudget.Across11benchmarksfromdiversedomains,TGRLbroadlyimprovesoverstrongRLVRbaselines:itimprovesthesix-benchmarkmathaverageby1.6%at32B,[email protected]%,andimprovesALFWorld/WebShopsuccessratesby6.3%/4.9%.Comprehensiveablationsandwall-clockanalysisconfirmtheefficacyofallproposedcomponents.Codeisavailableathttps://github.com/1229095296/TGRL/tree/main.
View arXiv pageView PDFGitHub2Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33589 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33589 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33589 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
This paper introduces LEEPS, a latent-guided explore-exploit prompt sampler for efficient reinforcement learning with verifiable rewards (RLVR) in LLMs. It adaptively balances reuse of informative prompts and exploration of uncertain ones, improving reasoning benchmark scores by 2.6-3.7% over baselines while adding only ~2 seconds of overhead per training step.
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
This paper introduces LLM-as-Environment-Engineer, a framework where LLMs design their own training environments for reinforcement learning in multi-agent reasoning tasks, enabling self-improving training that surpasses larger proprietary models.
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
This paper proposes the LLM-as-Environment-Engineer framework, where a policy model analyzes failures to automatically redesign the training environment for reinforcement learning, and introduces MAPF-FrozenLake as a controllable testbed. The framework, using Qwen3-4B, outperforms larger models like GPT and Gemini, showing that policy learning improves the model's ability to diagnose weaknesses.
ExpRL: Exploratory RL for LLM Mid-Training
ExpRL is a new RL-based mid-training method that uses human-written reference solutions as dense reward scaffolds (never shown to the policy) to improve LLM reasoning, achieving significant gains on hard math benchmarks like AIME-2026.
Discovering Reinforcement Learning Interfaces with Large Language Models
This paper introduces LIMEN, an LLM-guided evolutionary framework that automatically discovers reinforcement learning interfaces by jointly optimizing observation mappings and reward functions from raw simulator states. The approach reduces manual engineering effort and demonstrates that co-designing observations and rewards outperforms optimizing either component alone.