Tag
This paper addresses advantage scale calibration in group-relative optimization under low-variance rewards, proposing methods like the Reward-Resolution Protocol and MaxNorm-AC to filter sub-resolution jitter and provide bounded recovery for credible gaps.