Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
Summary
This paper introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method for reinforcement learning in large language models, which asymmetrically treats success and failure to enhance exploration and achieve better performance on reasoning tasks.
View Cached Full Text
Cached at: 09/29/26, 04:08 AM
Paper page - Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
Source: https://huggingface.co/papers/2609.33781
Abstract
Reinforcementlearningwithverifiablerewards(RLVR)enhancesreasoninginlargelanguagemodels(LLMs)throughoutcome-levelfeedback,yetrecentapproachestofiner-grainedcreditassignmentoftenrequireauxiliarymodels,additionalsampling,orprivilegedinformation.Althoughpolicyentropyprovidesareadilyavailablesignal,prioritizinguncertainpositionsunderbothreinforcementandpenalizationconcentratespenaltieswherefailedresponsesstillretainalternativesforrecovery,whichcansuppressopportunitiesforexploration.Toaddressthis,weintroduceEntropicAdvantagePolicyOptimization(EAPO),anentropy-guidedcreditassignmentmethodthattreatssuccessandfailureasymmetrically.Specifically,motivatedbytheobservationthatsuccessunderuncertaintyislessrepeatablewhileconfidentfailurestendtorecur,EAPOcouplesnormalizedpolicyentropywiththesignoftheresponseadvantagetoreinforcesurprisingsuccessandcorrectrepeatedfailure.Itassignsstrongerreinforcementtohigh-entropydecisionsinsuccessfulresponsesandstrongerpenaltiestolow-entropydecisionsinfailedresponses,whileattenuatingpenaltiesatuncertainpositionstopreserveopportunitiesforrecovery.Byredistributingtheresponseadvantageacrosstokens,EAPOderivestoken-levelcreditdirectlyfromexistingrolloutsignalswithoutadditionalsupervision.WevalidateEAPOonarangeofreasoningtasksacrossbothbaseandreasoningbackbones,demonstratingthatitachievesthebestoverallperformance.WefurthershowthatEAPOpromotesmoreeffectiveexploration,broadeningproblemcoverageandgeneratingmorediversecandidateanswers.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2609\.33781
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33781 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33781 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33781 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Introduces ACPO, a token-level credit assignment framework for reinforcement learning in LLMs that uses fine-grained surrogate entropy to improve reasoning performance on math and coding benchmarks, outperforming strong baselines like DAPO, GTPO, and SAPO.
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
This paper proposes an exploration-aware reinforcement learning framework that enables LLM agents to adaptively explore only when uncertainty is high, improving performance on text-based and GUI-based benchmarks.
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
This paper proposes Adaptive Entropy Regularization (AER), a framework that dynamically balances exploration and exploitation in LLM reinforcement learning by addressing policy entropy collapse through difficulty-aware coefficient allocation and initial-anchored target entropy. Experiments on mathematical reasoning benchmarks demonstrate consistent improvements in both accuracy and exploration capability.
ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
The paper proposes ERR+, a two-phase reinforcement learning framework that improves LLM reasoning by rewarding entropy drops in thinking phases, leading to better accuracy and efficiency across multiple datasets.
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs
This paper proposes E³RL, a reinforcement learning method that uses dynamic epistemic entropy thresholds to enable LLMs to excise local logical defects during generation, overcoming the autoregressive curse in long-horizon reasoning and achieving state-of-the-art results on mathematical reasoning benchmarks like AIME.