BiasReducer: Adaptive Bias Mitigation for Reward Models
Summary
BiasReducer is a lightweight framework that edits only the linear reward head of reward models to adaptively mitigate biases toward superficial attributes like response length and confidence, using a sparse-autoencoder-style encoder to detect and rank relevant biases per dataset. Across five reward models it improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming training-based baselines and reducing downstream verbosity and sycophancy.
View Cached Full Text
Cached at: 10/01/26, 08:24 PM
Paper page - BiasReducer: Adaptive Bias Mitigation for Reward Models
Source: https://huggingface.co/papers/2609.32720
Abstract
Rewardmodelsscoreresponsesfromlargelanguagemodels(LLMs)andguideLLMtrainingtowardhumanpreferences.However,rewardmodelscanfavorsuperficialattributessuchaslengthorconfidence,leadingLLMstoproducehigher-scoringbutnotmorecorrectresponses.Existingmitigationmethodseitherretraintherewardmodelorapplyafixedcorrectiontooneknownbias,suchasapreferenceforlongerresponses.Retrainingrequiresadditionaldataandcomputationalresources,whileexistingeditingmethodsrequirethetargetbiastobespecifiedinadvanceanduseafixededitforthatbias.Tothisend,weproposeBiasReducer,alightweightframeworkthateditsonlythelinearrewardheadandselectstherelevanteditsforeachnewdataset.First,BiasReducerusesasparseautoencoder(SAE)-styleencodertolearnwhichattributes(e.g.,lengthandconfidence)therewardmodelissensitiveto.Second,itlearnshowtoreducetherewardmodel’sdependenceoneachattributebydeterminingwhichdirectiontoadjusttherewardheadandhowmuchtoadjustit.Third,foranewdataset,itrankstheattributesbytheirinfluenceonrewardscores,selectstherelevantones,andeditstherewardmodelaccordingly.BiasReducerconsistentlyimprovesreward-modelrobustnesstobiasestowardsuperficialresponseattributes.Acrossfiverewardmodels,BiasReducer-Mimprovesthethreebenchmarksby8.3,18.0,and6.9percentagepointsonaverage,outperformingthetwotraining-basedbaselines.Thegainstransferdownstream,reducingunnecessaryverbosityandsycophancywhilemaintainingcomparablejudgedquality.
View arXiv pageView PDFGitHub1Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.32720 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32720 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32720 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
BiasGRPO proposes a framework using Group Relative Policy Optimization (GRPO) to stabilize social bias mitigation in LLMs by normalizing rewards across sampled completions, outperforming DPO and PPO on multiple benchmarks. The authors also release a compute-efficient bias reward model designed for integration into multi-objective RLHF pipelines.
Detecting and Mitigating Bias by Treating Fairness as a Symmetry Operation
The paper proposes treating fairness as a symmetry operation in machine learning classifiers, implementing loss-based regularization to enforce invariance under swapping of sensitive attributes while holding merit features fixed. The framework achieves over 90% bias reduction with minimal accuracy loss and requires no causal graph knowledge.
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
Introduces PEBS, a per-rater empirical-Bayes shrinkage estimator for calibrating reward models in RLHF, reducing within-user RMSE by over 8.5% on PRISM and over 9.6% on PluriHarms.
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
This paper investigates preference instability in reward models for LLMs, where subtle input variations cause contradictory preference assignments. The authors propose two SAE-based mitigation strategies—SAE Feature Steering and SAE Residual Correction—to reduce incorrect preference assignments without retraining.
Reward Models Can Be Too Sensitive (22 minute read)
This paper argues that reward models in RL are often oversensitive, assigning different scores to equally good responses, and proposes a training-free discretization algorithm using Monte Carlo dropout to reduce oversensitivity, improving policy quality.