Tag
The article describes testing to reduce reading cost in an AI verification pipeline while preserving high recall and minority evidence, noting challenges with current approaches and inviting community input.
This paper identifies a reward-variance collapse failure mode in GRPO for multi-turn evidence-reading agents and proposes CIGPO, which uses per-turn contextual information-gain rewards to maintain gradient signal, achieving +105% F1 improvement on HotpotQA.