BiasReducer: Adaptive Bias Mitigation for Reward Models

Hugging Face Daily Papers Papers

Summary

BiasReducer is a lightweight framework that edits only the linear reward head of reward models to adaptively mitigate biases toward superficial attributes like response length and confidence, using a sparse-autoencoder-style encoder to detect and rank relevant biases per dataset. Across five reward models it improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming training-based baselines and reducing downstream verbosity and sycophancy.

Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.
Original Article
View Cached Full Text

Cached at: 10/01/26, 08:24 PM

Paper page - BiasReducer: Adaptive Bias Mitigation for Reward Models

Source: https://huggingface.co/papers/2609.32720

Abstract

Rewardmodelsscoreresponsesfromlargelanguagemodels(LLMs)andguideLLMtrainingtowardhumanpreferences.However,rewardmodelscanfavorsuperficialattributessuchaslengthorconfidence,leadingLLMstoproducehigher-scoringbutnotmorecorrectresponses.Existingmitigationmethodseitherretraintherewardmodelorapplyafixedcorrectiontooneknownbias,suchasapreferenceforlongerresponses.Retrainingrequiresadditionaldataandcomputationalresources,whileexistingeditingmethodsrequirethetargetbiastobespecifiedinadvanceanduseafixededitforthatbias.Tothisend,weproposeBiasReducer,alightweightframeworkthateditsonlythelinearrewardheadandselectstherelevanteditsforeachnewdataset.First,BiasReducerusesasparseautoencoder(SAE)-styleencodertolearnwhichattributes(e.g.,lengthandconfidence)therewardmodelissensitiveto.Second,itlearnshowtoreducetherewardmodel’sdependenceoneachattributebydeterminingwhichdirectiontoadjusttherewardheadandhowmuchtoadjustit.Third,foranewdataset,itrankstheattributesbytheirinfluenceonrewardscores,selectstherelevantones,andeditstherewardmodelaccordingly.BiasReducerconsistentlyimprovesreward-modelrobustnesstobiasestowardsuperficialresponseattributes.Acrossfiverewardmodels,BiasReducer-Mimprovesthethreebenchmarksby8.3,18.0,and6.9percentagepointsonaverage,outperformingthetwotraining-basedbaselines.Thegainstransferdownstream,reducingunnecessaryverbosityandsycophancywhilemaintainingcomparablejudgedquality.

View arXiv pageView PDFGitHub1Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.32720 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32720 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32720 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Detecting and Mitigating Bias by Treating Fairness as a Symmetry Operation

arXiv cs.AI

The paper proposes treating fairness as a symmetry operation in machine learning classifiers, implementing loss-based regularization to enforce invariance under swapping of sensitive attributes while holding merit features fixed. The framework achieves over 90% bias reduction with minimal accuracy loss and requires no causal graph knowledge.

Reward Models Can Be Too Sensitive (22 minute read)

TLDR AI

This paper argues that reward models in RL are often oversensitive, assigning different scores to equally good responses, and proposes a training-free discretization algorithm using Monte Carlo dropout to reduce oversensitivity, improving policy quality.