unfaithfulness

Tag

Cards List
#unfaithfulness

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

arXiv cs.AI · 2026-07-09 Cached

This paper introduces reasoning consistency scanning, a method to audit whether chain-of-thought reasoning is logically consistent with the final answer in AI safety evaluations, distinguishing it from faithfulness. The authors formalize inconsistency subtypes, build a benchmark, implement a scanner, and report findings across models and tasks.

0 favorites 0 likes
← Back to home

Submit Feedback