Tag
This paper proposes PIVOT, a dual-level learning framework that enhances visually-grounded reasoning in large vision-language models by using self-calibrated experience replay and vision-guided advantage allocation to optimize reinforcement learning.
ReasoningBank gives LLM agents a memory layer so they continuously learn from both successes and failures, driving higher task success rates and faster execution.