Tag
This paper investigates why LLMs fail to apply CBT principles effectively despite high theoretical accuracy, introducing a knowledge-guided framework and a behavioral metric (Protocol Leverage Force) to measure intervention shifts. Experiments show that even with multi-chain-of-thought prompting, models remain biased toward validation and reflection.