Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Hugging Face Daily Papers Papers

Summary

The paper introduces Density-Aware Reward Aggregation (DARA), which corrects reward-wise normalized multi-reward RL with an inverse-square-root density weighting based on active-group density, reducing signal imbalance between rewards. Experiments on tool calling and math reasoning show DARA reaches high format/length compliance in up to 26% and 65% fewer training steps than GDPO.

Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Original Article
View Cached Full Text

Cached at: 10/02/26, 08:28 AM

Paper page - Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Source: https://huggingface.co/papers/2610.00574

Abstract

Multi-rewardreinforcementlearningtrainslargelanguagemodelstosatisfymultiplebehavioralobjectivessimultaneously.Reward-wisenormalization,asusedinGDPO,preservesreward-specificrelativeinformationwithinrolloutgroups,butdifferentobjectivescanstillexhibitunevenlearningprogress.Westudythisbehaviorthroughadvantageenergy,thesumofareward’ssquaredadvantagesoverabatch.UnderidealizedGDPOnormalization,weshowthatthisenergyisproportionaltoactive-groupdensity:thefractionofrolloutgroupsinwhichtherewardprovidesnonzerorelativeadvantages.Thisrevealsaresidualbatch-levelsignalimbalanceandprovidesabasisforcalibratingrewardcontributions.Basedonthisrelation,weproposeDensity-AwareRewardAggregation(DARA).Wederiveaninverse-square-rootdensitycorrectionthatgivesgreaterweighttosignalsfromlessfrequentlyactiverewards.DARAcomputesitsweightsfromeachrolloutbatch,adaptingtochangesinrewardactivitythroughouttrainingwithoutmodifyingtheunderlyingpolicyoptimizationobjective.ExperimentsontoolcallingandmathematicalreasoningshowthatDARAlearnsthetargetedbehaviorsfasterthanGDPO,reachinghighformatcomplianceinupto26%fewertrainingstepsontoolcallingandnear-saturatedlengthcomplianceinupto65%fewerstepsonmathematicalreasoning,whileremainingcompetitiveinfinalperformance.Ourcodeisavailableathttps://github.com/zhaihaotian/DARA.

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2610\.00574

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.00574 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.00574 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.00574 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning

arXiv cs.LG

This paper introduces CurveRL, a principled distribution-aware prompt reweighting approach for reinforcement learning with verifiable rewards (RLVR) that improves LLM reasoning by assigning weights based on the rank and density of pass rates rather than their absolute values, consistently outperforming GRPO and other baselines.

Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training

Hugging Face Daily Papers

This paper proposes an empirical 'sparse-to-dense' reward principle for language model post-training, arguing that scarce labeled data should be used with sparse rewards for teacher model discovery and dense rewards for student compression via distillation. The authors demonstrate that this staged approach, bridging sparse RL and on-policy distillation, outperforms direct GRPO on deployment-sized models in math benchmarks.