Variance reduction for policy gradient with action-dependent factorized baselines
Summary
OpenAI researchers derive a bias-free action-dependent baseline for variance reduction in policy gradient methods, demonstrating improved learning efficiency on high-dimensional control tasks, multi-agent, and partially observed environments.
View Cached Full Text
Cached at: 04/20/26, 02:56 PM
Similar Articles
Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation
This paper proposes a robust gradient-based algorithm for learning behavior policies that reduce variance in online reinforcement learning policy evaluation, addressing uncertainties in transition functions with theoretical guarantees and numerical validation.
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL
This paper introduces FADE (Focal Advantage with Dynamic Entropy), a self-adapting advantage function that dynamically schedules gradient weights during RL post-training of LLMs, achieving faster convergence and better accuracy-diversity trade-offs compared to static baselines.
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
BiasGRPO proposes a framework using Group Relative Policy Optimization (GRPO) to stabilize social bias mitigation in LLMs by normalizing rewards across sampled completions, outperforming DPO and PPO on multiple benchmarks. The authors also release a compute-efficient bias reward model designed for integration into multi-objective RLHF pipelines.
@msjgriffiths: Oh, neat. I've had an itch in the back of my head for the past year that GRPO (a naive method) should be related to Ste…
This paper proposes shrinkage baselines for reinforcement learning with verifiable rewards to reduce variance in policy gradient estimators and improve training stability.
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
BiPACE introduces a drop-in advantage estimator that fixes state-action credit mismatch in stepwise group-based RL for LLM agents, using bisimulation-guided state clustering and action counterfactual estimation, achieving significant performance gains on ALFWorld, WebShop, and TextCraft with Qwen2.5 models.