preference-averaging

Tag

Cards List
#preference-averaging

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv cs.LG · 2026-08-12 Cached

This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.

0 favorites 0 likes
← Back to home

Submit Feedback