advantage-estimation

Tag

Cards List
#advantage-estimation

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

arXiv cs.AI · 2026-06-30 Cached

BV-Blend is a critic-free reinforcement learning framework that combines prompt-local on-policy statistics with historical moments from semantic clusters to stabilize advantage estimation, improving training stability and performance for aligning large language models with verifiable rewards.

0 favorites 0 likes
← Back to home

Submit Feedback