ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Summary
ThoughtFold proposes a framework using introspective preference learning to reduce redundant explorations in Chain-of-Thought reasoning for Large Reasoning Models, achieving ~56% token reduction on DeepSeek-R1-Distill-Qwen-7B without accuracy loss.
View Cached Full Text
Cached at: 06/03/26, 09:44 AM
# ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Source: [https://arxiv.org/abs/2606.03503](https://arxiv.org/abs/2606.03503) [View PDF](https://arxiv.org/pdf/2606.03503) > Abstract:Large Reasoning Models \(LRMs\) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards \(RLVR\) on Chain\-of\-Thoughts \(CoTs\)\. However, since long CoTs naturally contain trial and errors and mainstream RLVR approaches choose outcome\-correct CoT trajectories for memorization, the redundant explorations in long CoTs are inevitably reinforced, which results in the over\-thinking issues of LRMs\. Previous attempts to resolve this issue mainly give more advantage to shorter trajectories, yet their learning signals are still outcome\-based and cannot reduce the memorization of redundant explorations in long CoTs\. Therefore, we propose ThoughtFold, a framework that leverages fine\-grained preference learning to mitigate redundant explorations for efficient reasoning\. ThoughtFold employs an introspective strategy to identify redundancy within each correct trajectory, which yields a spectrum of candidate sub\-trajectories\. Leveraging this spectrum, we introduce a masked preference optimization objective that explicitly penalizes redundant explorations and encourages the model to directly bridge essential reasoning segments, effectively folding its reasoning chains into a more concise path\. Extensive experiments show that ThoughtFold significantly enhances efficiency\. It reduces the token usage of DeepSeek\-R1\-Distill\-Qwen\-7B by approximately 56% while maintaining state\-of\-the\-art accuracy\. ## Submission history From: Ziyan Liu \[[view email](https://arxiv.org/show-email/ea6547b1/2606.03503)\] **\[v1\]**Tue, 2 Jun 2026 11:21:27 UTC \(692 KB\)
Similar Articles
TabRank: Chain-of-Thought Distillation for Table Re-Rankers
TabRank introduces a framework for training reasoning rerankers for tabular retrieval by distilling chain-of-thought traces from a large reasoning model (DeepSeek-R1) into compact student models, achieving significant improvements across multiple table retrieval benchmarks.
Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
This paper introduces Semi-CoT, a semi-supervised learning framework for chain-of-thought reasoning that uses unlabeled questions with an entropy-based selection to generate reliable pseudo reasoning chains, showing promising but mixed results on math reasoning benchmarks.
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
This research paper from MediaTek and National Taiwan University challenges the assumption that reasoning chains must be dense and sequential, showing that models can extract answers from sparse, shuffled, and noisy reasoning traces. The findings suggest that answer extraction is robust and order-independent, potentially enabling more efficient, parallelized reasoning generation.
Reasoning Fine-Tuning Induces Persistent Latent Policy States
This paper models Chain-of-Thought reasoning as a switching dynamical system, showing that reasoning fine-tuning globally reorganizes latent policy states, leading to improved multi-step reasoning. The proposed framework combines time-aware contrastive learning with discrete regime discovery, and experiments demonstrate that fine-tuned models exhibit richer latent-policy organization with functional specialization.
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
This paper systematically evaluates multimodal Chain-of-Thought reasoning across 12 tasks, finding it selectively effective for reasoning tasks but detrimental for perception tasks, and identifies a 'Look Light, Think Heavy' pattern where visual introspection declines during reasoning.