top-k-estimator

Tag

Cards List
#top-k-estimator

Predictive Divergence Masks for LLM RL

Hugging Face Daily Papers · 2026-07-12 Cached

Proposes predictive divergence masks for LLM reinforcement learning that improve upon PPO's direction criterion by predicting whether the next policy gradient step will increase or decrease the divergence used by the trust region, leading to better alignment and improved RL training across model scales.

0 favorites 0 likes
← Back to home

Submit Feedback