COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models
Summary
COFT is a training-free decoding method that applies token-level fairness control and conformal calibration to reduce bias in chain-of-thought reasoning of large language models, achieving 30-55% bias reduction with minimal computational overhead.
View Cached Full Text
Cached at: 06/01/26, 09:26 AM
# COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models Source: [https://arxiv.org/abs/2605.30641](https://arxiv.org/abs/2605.30641) [View PDF](https://arxiv.org/pdf/2605.30641) > Abstract:Large language models \(LLMs\) can reveal and amplify societal biases during chain\-of\-thought \(CoT\) generation\. We present COFT \(Chain of Fair Thought\), a training\-free decoding method that applies token\-level fairness control at decode time, with distribution\-free marginal validity guarantees \(under exchangeability\) for any frozen causal language model\. COFT operates in three stages\. First, it creates a masked counterfactual prompt by replacing sensitive spans with neutral tokens\. Second, it compares the factual and masked logit distributions through lightweight logit fusion to attenuate attribute\-driven biases\. Third, it uses dual\-branch split\-conformal calibration to certify per\-step candidate token sets at a user\-chosen risk level\. We evaluate COFT across six models and multiple bias benchmarks\. Our method reduces standard bias metrics by 30\-55% \(median 38%\) while preserving task utility and language quality\. Reasoning accuracies remain unchanged within run\-to\-run noise margins\. The computational overhead is modest, equivalent to one additional cached forward pass \(<=11%\)\. COFT offers a clear, auditable path to safer CoT generation with significant bias reduction, negligible utility loss, and no requirement for retraining, auxiliary classifiers, or weight access\. ## Submission history From: Arya Fayyazi \[[view email](https://arxiv.org/show-email/c4a9d3d9/2605.30641)\] **\[v1\]**Thu, 28 May 2026 22:52:15 UTC \(2,107 KB\)
Similar Articles
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
Proposes ProxyCoT, a training framework that improves long-context reasoning in large language models by first obtaining chain-of-thought reasoning traces on short proxy contexts (via reinforcement learning or distillation) and then grounding them in full long contexts through supervised fine-tuning. Experiments show consistent improvements over baselines with reduced computational cost.
Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding
This paper proposes Counterfactual Ensemble Decoding (CED) to mitigate social biases in large vision-language models by constructing multi-group counterfactual perspectives and integrating them during decoding, achieving substantial bias reduction while preserving model capabilities.
OpenCoF: Learning to Reason Through Video Generation
OpenCoF introduces a reasoning video dataset and a fine-tuned video generation model that improves temporal reasoning through diverse supervision and explicit reasoning tokens, showing significant gains on four video reasoning benchmarks.
Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.
CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness
Proposes CASE, a framework combining training-time causal alignment and inference-time structural enforcement to improve faithfulness of chain-of-thought reasoning in large language models, achieving a 37% average improvement in CoT faithfulness across benchmarks.