Tag
CheckRLM is a framework that uses retrieval-augmented generation to detect and correct factual errors in the reasoning chains of reasoning language models, improving coherence and reducing error accumulation.
This paper introduces a white-box diagnostic framework that localizes instruction hierarchy failures in reasoning language models into identification, conflict resolution, and response realization stages. It evaluates several models and proposes two training-free self-monitoring mechanisms that reduce non-compliance by 81–99%.