Fair Reinforcement Learning
Summary
Fair Reinforcement Learning introduces Democratic Alignment to incorporate multiple competing value sets from different agents, overcoming traditional RLHF limitations, and achieves orders of magnitude faster optimization via a black-box policy wrapper.
Similar Articles
RL-FAT: Reinforcement Learning for Fair Adversarial Training
RL-FAT is a reinforcement learning framework for fair adversarial training that improves robustness while reducing class-wise disparities. Experiments demonstrate competitive accuracy and better fairness compared to standard methods.
Inference-Time Policy Alignment for Fair Reinforcement Learning
This paper proposes an inference-time policy shaping framework to steer pretrained reinforcement learning policies toward welfare-based fairness objectives without retraining, inspired by inference-time alignment in LLMs.
Procedural Fairness Failures in RLHF from Preference Averaging
This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
This paper proposes PRISM, a multi-reward RL framework that decomposes policy space rather than mixing rewards, improving multi-reward optimization and enabling inference-time controllability. Experiments on reasoning and alignment tasks show it outperforms existing baselines.
Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment
This paper introduces a multi-agent environment based on the board game Fog of Love to evaluate affinity-based reinforcement learning for instilling virtuous behavior in AI agents. The authors demonstrate that localized affinities improve agent performance in both competitive and cooperative objectives, advancing machine ethics research beyond simple grid-world environments.