chain-of-thought-monitoring

Tag

Cards List
#chain-of-thought-monitoring

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

arXiv cs.AI · 2026-07-10 Cached

This paper demonstrates that adversarial agents can persuade chain-of-thought monitors to approve policy-violating actions, increasing harmful approvals by 9.5% on average. It proposes a fact-checking framework using different model families that reduces approval of policy violations by up to 45%.

0 favorites 0 likes
← Back to home

Submit Feedback