Tag
STRIDE introduces a training framework that uses learnable stepwise language feedback instead of scalar rewards to improve LLM reasoning, achieving state-of-the-art results on diverse benchmarks.