off-policy-updates

Tag

Cards List
#off-policy-updates

Soft Adaptive Policy Optimization

Papers with Code Trending · 2025-11-25 Cached

SAPO introduces a smooth, temperature-controlled gate to adaptively attenuate off-policy updates in reinforcement learning for large language models, enhancing training stability and performance compared to methods with hard clipping.

0 favorites 0 likes
← Back to home

Submit Feedback