policy-gradient

Tag

Cards List
#policy-gradient

PACT: From Credit Assignment to Critic Alignment

Hugging Face Daily Papers ↗ · 6d ago Cached

The paper formulates regularity conditions for token-level credit in reinforcement learning for LLMs and introduces Policy Aligned Critic Training (PACT), which outperforms GRPO and PPO in mathematical reasoning and coding benchmarks.

0 favorites 0 likes
#policy-gradient

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.

0 favorites 0 likes
#policy-gradient

@msjgriffiths: Oh, neat. I've had an itch in the back of my head for the past year that GRPO (a naive method) should be related to Ste…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

This paper proposes shrinkage baselines for reinforcement learning with verifiable rewards to reduce variance in policy gradient estimators and improve training stability.

0 favorites 0 likes
#policy-gradient

Tail-Likelihood Reinforcement Learning

arXiv cs.LG ↗ · 2026-09-04 Cached

The paper proposes Tail-Likelihood Reinforcement Learning (TailRL), an optimization method that focuses on the upper tails of reward distributions to improve policy performance in generative tasks, demonstrated across various applications.

0 favorites 0 likes
#policy-gradient

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper identifies value mismatch in parallel reinforcement learning where a shared critic degrades performance by mixing environment-specific values, and proposes conditioning the critic on environment indices to improve learning stability and return in benchmarks like Procgen.

0 favorites 0 likes
#policy-gradient

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

arXiv cs.LG ↗ · 2026-08-21 Cached

This paper analyzes a single-loop, entropy-regularized Natural Actor-Critic algorithm and proves accelerated convergence rates for the unregularized objective in stochastic and deterministic regimes under linear function approximation.

0 favorites 0 likes
#policy-gradient

Vector Symbolic Policy Gradient

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper introduces Vector-Symbolic Policy Gradient (VSPG), a novel method that uses vector symbolic architecture for discrete-action policy gradients in reinforcement learning, achieving competitive performance with robust degradation under noise for edge systems.

0 favorites 0 likes
#policy-gradient

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-12 Cached

Introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method for safe reinforcement learning that actively drives the policy toward the constraint boundary, with convergence guarantees and improved reward/boundary tracking on a Safety-Gymnasium task.

0 favorites 0 likes
#policy-gradient

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-04 Cached

This paper proves that using error-penalized scoring rules with abstention as a discrete action can kill both the reward gradient and the KL anchor, causing models to collapse toward refusing everything. It proposes a structural repair — training a mandatory confidence report — and validates the mechanism with simulations and language model experiments.

0 favorites 0 likes
#policy-gradient

Policy Gradient Steering: Interventions from Behavioral Objectives

arXiv cs.LG ↗ · 2026-07-31 Cached

Introduces Policy Gradient Steering (PGS), a method that formulates activation steering as a reinforcement learning problem, using policy gradients to construct removable, composable steering vectors from behavioral objectives. Validated in gridworld, chess puzzle, and football environments.

0 favorites 0 likes
#policy-gradient

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

arXiv cs.LG ↗ · 2026-07-21 Cached

Introduces Hindsight Policy Optimization (HPO), a novel policy gradient method that uses an intent space and Wasserstein distance to reduce variance in long-horizon language agent training, showing improved stability over GRPO and PPO.

0 favorites 0 likes
#policy-gradient

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv cs.LG ↗ · 2026-07-13 Cached

SafeExplorer introduces an unbiased policy gradient estimator for reinforcement learning with recovery interventions, significantly reducing training-time falls on robot tasks while matching or exceeding standard PPO's final reward.

0 favorites 0 likes
#policy-gradient

Predictive Divergence Masks for LLM RL

Hugging Face Daily Papers ↗ · 2026-07-12 Cached

Proposes predictive divergence masks for LLM reinforcement learning that improve upon PPO's direction criterion by predicting whether the next policy gradient step will increase or decrease the divergence used by the trust region, leading to better alignment and improved RL training across model scales.

0 favorites 0 likes
#policy-gradient

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL

arXiv cs.LG ↗ · 2026-07-03 Cached

This paper introduces FADE (Focal Advantage with Dynamic Entropy), a self-adapting advantage function that dynamically schedules gradient weights during RL post-training of LLMs, achieving faster convergence and better accuracy-diversity trade-offs compared to static baselines.

0 favorites 0 likes
#policy-gradient

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

arXiv cs.LG ↗ · 2026-06-24 Cached

EMAgnet introduces parameter-space exponential moving average regularization for policy gradient self-play in large two-player zero-sum games, achieving lower exploitability compared to uniform regularization targets.

0 favorites 0 likes
#policy-gradient

KLip-PPO: A per-sample KL perspective on PPO-Clip

arXiv cs.LG ↗ · 2026-06-24 Cached

This paper shows that the gradient of the clipped surrogate in Proximal Policy Optimization (PPO) is exactly reproduced by a per-sample Kullback-Leibler penalty with a variable coefficient, revealing structural features of the clipped surrogate and suggesting new design directions.

0 favorites 0 likes
#policy-gradient

In game theory, generalists sometimes win out over specialists

MIT News — Artificial Intelligence ↗ · 2026-06-17 Cached

MIT researchers co-authored a paper showing that general-purpose policy gradient algorithms can outperform specialized game-theoretic algorithms in imperfect-information games, challenging long-held assumptions in the field.

0 favorites 0 likes
#policy-gradient

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper formalizes embedding model routing as an adversarial contextual linear bandit with low-rank experts, proposing the Hypentropy Policy Gradient (HPG) algorithm that achieves O~(s√(MT)) policy regret, avoiding the curse of dimensionality.

0 favorites 0 likes
#policy-gradient

Diffusion Policy Optimization without Drifting Apart

arXiv cs.LG ↗ · 2026-06-15 Cached

DiPOD stabilizes diffusion policy optimization by interleaving self-distillation with policy-gradient updates to maintain a tight ELBO, preventing the double-drift phenomenon and achieving higher rewards in both language and continuous control tasks.

0 favorites 0 likes
#policy-gradient

Self-Distilled Policy Gradient

arXiv cs.LG ↗ · 2026-06-04 Cached

SDPG (Self-Distilled Policy Gradient) is a new RL training framework for LLMs that combines group-relative verifier advantages with on-policy self-distillation and KL regularization to address sparse rewards and instability in RLVR training. The method uses a shared model as both student and teacher by conditioning on privileged context, showing improved stability and performance over RLVR and self-distillation baselines.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback