Tag
STRIDE introduces a training framework that uses learnable stepwise language feedback instead of scalar rewards to improve LLM reasoning, achieving state-of-the-art results on diverse benchmarks.
Variational Policy Distillation (VPD) formalizes learning from language feedback as a variational EM problem, co-training teacher and student networks to improve policy learning in reinforcement learning from verifiable rewards. It shows consistent improvements over baselines on code generation and scientific reasoning tasks.