Tag
This paper studies behavioral detection of unfaithful chain-of-thought reasoning in LLMs, finding that answer correctness structures detection performance: on incorrect answers, where most unfaithfulness occurs, behavioral signals are at chance, while on correct answers they offer modest separation.
This paper characterizes backdoors in LoRA adapters that activate at the token feature level, and proposes behavioral and weight-level detection methods. The backdoor generalizes across related token patterns but not structurally identical ones, and detection methods show strong separation.