Tag
This paper shows that tool-result caching, even if marginally correct, can reverse the expected group-normalized policy updates in reinforcement learning, as demonstrated through mathematical analysis and experiments with a two-action model.