ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Summary
ReflectRL is a framework that learns from 'golden negative trajectories' (failed reasoning attempts by expert models) by reflecting on them, then transfers this reflective reasoning back to direct reasoning, improving LLM performance across benchmarks.
View Cached Full Text
Cached at: 08/05/26, 05:46 PM
Paper page - ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Source: https://huggingface.co/papers/2608.03972 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
On-policytraininghasemergedasapowerfulpost-trainingparadigmforimprovingthereasoningcapabilitiesoflargelanguagemodels,andisoftenenhancedbygoldentrajectoriesfromstrongerexpertmodels.However,whentheexpertfailsonharderproblems,existingtrajectory-guidedmethodslosetheirmainsourceofsupervision,andthesefailedtrajectoriesaretypicallydiscardedasnegativesamples.Wearguethatsuchfailures,whichwecallGoldenNegativeTrajectories,canstillprovidevaluablereasoningsignalswhentreatednotasdemonstrationstoimitate,butasflawedtrajectoriestoreflectupon.WeidentifyaReflectionAdvantage:forhardproblems,reflectingonaflawedtrajectorycanbeeasierandmoreeffectivethansolvingtheproblemdirectlyfromscratch.Motivatedbythis,weproposeReflectRL,alightweightplug-and-playframeworkthatlearnsfromGoldenNegativeTrajectoriesduringon-policytraining.ReflectRLfirstusesthesetrajectoriestoelicitReflectiveReasoning,thenappliesReflective-to-DirectPolicyTransitiontotransfertheacquiredreasoningbehaviorbacktoDirectReasoning.Experimentsacross9benchmarks,4LLMbackbones,and4on-policytrainingmethodsshowthatReflectRLconsistentlyimprovesreasoningperformancewithminimaloverhead.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2608\.03972
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.03972 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.03972 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03972 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes
The paper introduces Reflective Recovery, a self-supervised method that enhances LLM reasoning by transforming failed trajectories into training data, breaking scaling collapse and enabling emergent self-correction.
ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation
ReflectMT introduces a two-stage RL method that trains LRMs to internalize reflection, enabling single-pass high-quality translation with 94% fewer tokens than multi-step reasoning models like DeepSeek-R1.
ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning
This paper introduces ReFlect, a training-free harness system that wraps LLMs with deterministic error detection and recovery logic to improve performance on complex, long-horizon reasoning tasks.
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
This paper introduces ResRL, a method to boost LLM reasoning by decoupling semantic distributions between positive and negative responses through negative sample projection. It aims to maintain generation diversity while improving performance on various benchmarks.
Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation
The paper introduces Reflection-Enhanced Self-Distillation (Resd), a framework that transforms failure feedback into corrective supervision for LLMs, enabling efficient learning from rare successes. It outperforms standard self-distillation baselines and achieves faster early improvement than GRPO with fewer samples.