Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Hugging Face Daily Papers Papers

Summary

This paper introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method for reinforcement learning in large language models, which asymmetrically treats success and failure to enhance exploration and achieve better performance on reasoning tasks.

Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:08 AM

Paper page - Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Source: https://huggingface.co/papers/2609.33781

Abstract

Reinforcementlearningwithverifiablerewards(RLVR)enhancesreasoninginlargelanguagemodels(LLMs)throughoutcome-levelfeedback,yetrecentapproachestofiner-grainedcreditassignmentoftenrequireauxiliarymodels,additionalsampling,orprivilegedinformation.Althoughpolicyentropyprovidesareadilyavailablesignal,prioritizinguncertainpositionsunderbothreinforcementandpenalizationconcentratespenaltieswherefailedresponsesstillretainalternativesforrecovery,whichcansuppressopportunitiesforexploration.Toaddressthis,weintroduceEntropicAdvantagePolicyOptimization(EAPO),anentropy-guidedcreditassignmentmethodthattreatssuccessandfailureasymmetrically.Specifically,motivatedbytheobservationthatsuccessunderuncertaintyislessrepeatablewhileconfidentfailurestendtorecur,EAPOcouplesnormalizedpolicyentropywiththesignoftheresponseadvantagetoreinforcesurprisingsuccessandcorrectrepeatedfailure.Itassignsstrongerreinforcementtohigh-entropydecisionsinsuccessfulresponsesandstrongerpenaltiestolow-entropydecisionsinfailedresponses,whileattenuatingpenaltiesatuncertainpositionstopreserveopportunitiesforrecovery.Byredistributingtheresponseadvantageacrosstokens,EAPOderivestoken-levelcreditdirectlyfromexistingrolloutsignalswithoutadditionalsupervision.WevalidateEAPOonarangeofreasoningtasksacrossbothbaseandreasoningbackbones,demonstratingthatitachievesthebestoverallperformance.WefurthershowthatEAPOpromotesmoreeffectiveexploration,broadeningproblemcoverageandgeneratingmorediversecandidateanswers.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2609\.33781

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33781 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33781 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33781 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning

arXiv cs.CL

This paper proposes Adaptive Entropy Regularization (AER), a framework that dynamically balances exploration and exploitation in LLM reinforcement learning by addressing policy entropy collapse through difficulty-aware coefficient allocation and initial-anchored target entropy. Experiments on mathematical reasoning benchmarks demonstrate consistent improvements in both accuracy and exploration capability.