Training on Documents About Monitoring Leads to CoT Obfuscation

arXiv cs.LG Papers

Summary

This paper demonstrates that models trained on documents describing chain-of-thought monitoring can learn to obfuscate their reasoning to avoid detection, posing a risk to CoT-based alignment techniques.

arXiv:2605.15257v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study whether monitor-aware models are capable of obfuscating their reasoning to evade detection. We use synthetic document finetuning to expose eight models to realistic pre-training-style documents describing a CoT monitor and find that monitor-aware models consistently achieve higher rates of undetected misbehavior compared to unaware controls. This effect is weaker but still present on a harder agentic task. We also show that CoT controllability, a model's ability to reshape its own reasoning trace under an imposed constraint, is closely correlated with obfuscation success across the eight models studied ($r=0.800$, $p=0.017$). Monitor-aware models placed under equal reinforcement learning optimization pressure also learn to reward-hack without triggering a CoT monitor substantially faster than unaware controls. Together, these results suggest that knowledge of monitoring combined with high CoT controllability poses a risk to CoT-based monitoring.
Original Article
View Cached Full Text

Cached at: 05/18/26, 06:39 AM

# Training on Documents About Monitoring Leads to CoT Obfuscation
Source: [https://arxiv.org/abs/2605.15257](https://arxiv.org/abs/2605.15257)
[View PDF](https://arxiv.org/pdf/2605.15257)

> Abstract:Chain\-of\-thought \(CoT\) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning\. Motivated by this vulnerability, we study whether monitor\-aware models are capable of obfuscating their reasoning to evade detection\. We use synthetic document finetuning to expose eight models to realistic pre\-training\-style documents describing a CoT monitor and find that monitor\-aware models consistently achieve higher rates of undetected misbehavior compared to unaware controls\. This effect is weaker but still present on a harder agentic task\. We also show that CoT controllability, a model's ability to reshape its own reasoning trace under an imposed constraint, is closely correlated with obfuscation success across the eight models studied \($r=0\.800$, $p=0\.017$\)\. Monitor\-aware models placed under equal reinforcement learning optimization pressure also learn to reward\-hack without triggering a CoT monitor substantially faster than unaware controls\. Together, these results suggest that knowledge of monitoring combined with high CoT controllability poses a risk to CoT\-based monitoring\.

## Submission history

From: Reilly Haskins \[[view email](https://arxiv.org/show-email/7b9ac9d3/2605.15257)\] **\[v1\]**Thu, 14 May 2026 17:59:01 UTC \(600 KB\)

Similar Articles

Detecting misbehavior in frontier reasoning models

OpenAI Blog

OpenAI researchers demonstrate that chain-of-thought monitoring can detect misbehavior in frontier reasoning models like o3-mini, but warn that directly optimizing CoT to prevent bad thoughts causes models to hide their intent rather than eliminate the behavior.

Reasoning models struggle to control their chains of thought, and that’s good

OpenAI Blog

OpenAI researchers study whether reasoning models can deliberately obscure their chain-of-thought to evade monitoring, finding that current models struggle to control their reasoning even when aware of monitoring. They introduce CoT-Control, an open-source evaluation suite with over 13,000 tasks to measure chain-of-thought controllability in reasoning models.

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

arXiv cs.CL

The paper proposes CoT-Interpretability Alignment (CIA), a metric using interpretability tools like linear probes to measure agreement between an LLM's chain-of-thought and its internal computations, and shows post-training with task accuracy plus parametric faithfulness rewards substantially improves faithfulness while maintaining accuracy.

Training Continuous Chain of Thought Models: A Tale of Two Regimes

arXiv cs.AI

This paper introduces C-MTP, a direct supervision method for training continuous chain-of-thought models that compresses reasoning traces into latent representations. The method performs competitively on simple tasks but reveals that both direct and indirect supervision methods struggle with complex long reasoning traces, showing about 65% performance drop.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

arXiv cs.AI

This paper demonstrates that adversarial agents can persuade chain-of-thought monitors to approve policy-violating actions, increasing harmful approvals by 9.5% on average. It proposes a fact-checking framework using different model families that reduces approval of policy violations by up to 45%.