Tag
The paper introduces Reflective Recovery, a self-supervised method that enhances LLM reasoning by transforming failed trajectories into training data, breaking scaling collapse and enabling emergent self-correction.