Fair Reinforcement Learning

Reddit r/AI_Agents Papers

Summary

Fair Reinforcement Learning introduces Democratic Alignment to incorporate multiple competing value sets from different agents, overcoming traditional RLHF limitations, and achieves orders of magnitude faster optimization via a black-box policy wrapper.

**ICLR 2026 Publication** * ⚖️ **Democratic Alignment:** Seamlessly incorporates multiple, competing sets of values from different agents, moving past the "one-size-fits-all" limitation of traditional RLHF. * 📦 **Black-Box Policy Optimization:** Operates as a wrapper around *standard policy optimization* algorithms, removing direct dependency on the total number of states or actions. * 🚀 **Orders of Magnitude Faster:** Drastically reduces sample complexity and orders of magnitude more efficient with respect to computation compared to prior tabular methods.
Original Article

Similar Articles

RL-FAT: Reinforcement Learning for Fair Adversarial Training

arXiv cs.LG

RL-FAT is a reinforcement learning framework for fair adversarial training that improves robustness while reducing class-wise disparities. Experiments demonstrate competitive accuracy and better fairness compared to standard methods.

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv cs.LG

This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.

Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

arXiv cs.AI

This paper introduces a multi-agent environment based on the board game Fog of Love to evaluate affinity-based reinforcement learning for instilling virtuous behavior in AI agents. The authors demonstrate that localized affinities improve agent performance in both competitive and cooperative objectives, advancing machine ethics research beyond simple grid-world environments.