Tag
The tweet describes a reasoning-intensive regression task that evaluates where a flawed reasoning trace first goes wrong, and shows that pedagogical reinforcement learning achieves the best performance with an 18% decrease in NMSE and 5% increase in CCC.