Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Summary
The paper introduces Density-Aware Reward Aggregation (DARA), which corrects reward-wise normalized multi-reward RL with an inverse-square-root density weighting based on active-group density, reducing signal imbalance between rewards. Experiments on tool calling and math reasoning show DARA reaches high format/length compliance in up to 26% and 65% fewer training steps than GDPO.
View Cached Full Text
Cached at: 10/02/26, 08:28 AM
Paper page - Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Source: https://huggingface.co/papers/2610.00574
Abstract
Multi-rewardreinforcementlearningtrainslargelanguagemodelstosatisfymultiplebehavioralobjectivessimultaneously.Reward-wisenormalization,asusedinGDPO,preservesreward-specificrelativeinformationwithinrolloutgroups,butdifferentobjectivescanstillexhibitunevenlearningprogress.Westudythisbehaviorthroughadvantageenergy,thesumofareward’ssquaredadvantagesoverabatch.UnderidealizedGDPOnormalization,weshowthatthisenergyisproportionaltoactive-groupdensity:thefractionofrolloutgroupsinwhichtherewardprovidesnonzerorelativeadvantages.Thisrevealsaresidualbatch-levelsignalimbalanceandprovidesabasisforcalibratingrewardcontributions.Basedonthisrelation,weproposeDensity-AwareRewardAggregation(DARA).Wederiveaninverse-square-rootdensitycorrectionthatgivesgreaterweighttosignalsfromlessfrequentlyactiverewards.DARAcomputesitsweightsfromeachrolloutbatch,adaptingtochangesinrewardactivitythroughouttrainingwithoutmodifyingtheunderlyingpolicyoptimizationobjective.ExperimentsontoolcallingandmathematicalreasoningshowthatDARAlearnsthetargetedbehaviorsfasterthanGDPO,reachinghighformatcomplianceinupto26%fewertrainingstepsontoolcallingandnear-saturatedlengthcomplianceinupto65%fewerstepsonmathematicalreasoning,whileremainingcompetitiveinfinalperformance.Ourcodeisavailableathttps://github.com/zhaihaotian/DARA.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2610\.00574
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2610.00574 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2610.00574 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2610.00574 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
This paper introduces CurveRL, a principled distribution-aware prompt reweighting approach for reinforcement learning with verifiable rewards (RLVR) that improves LLM reasoning by assigning weights based on the rank and density of pass rates rather than their absolute values, consistently outperforming GRPO and other baselines.
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
This paper proposes PRISM, a multi-reward RL framework that decomposes policy space rather than mixing rewards, improving multi-reward optimization and enabling inference-time controllability. Experiments on reasoning and alignment tasks show it outperforms existing baselines.
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
This paper identifies aggregation-induced reward hacking in multi-reward reinforcement learning for LLMs and proposes an adaptive projection method, AMRP, to dynamically adjust weights for better reward balance and performance.
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
This paper proposes an empirical 'sparse-to-dense' reward principle for language model post-training, arguing that scarce labeled data should be used with sparse rewards for teacher model discovery and dense rewards for student compression via distillation. The authors demonstrate that this staged approach, bridging sparse RL and on-policy distillation, outperforms direct GRPO on deployment-sized models in math benchmarks.
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation
SMOPD proposes a two-stage specialize-and-merge online policy distillation method to improve multi-reward reinforcement learning, addressing issues with sparse and dense reward signals where GDPO struggles. It outperforms GDPO across 1.5B, 3B, and 7B backbones in complementary and conflicting reward settings.