Tag
The paper investigates when historical action credits in tool-using agents require updating after policy changes, introducing a Decision-Sufficient Credit Gate (DSC-Gate) to efficiently reuse data and reduce new tool interactions.
This paper shows that tool-result caching, even if marginally correct, can reverse the expected group-normalized policy updates in reinforcement learning, as demonstrated through mathematical analysis and experiments with a two-action model.