Tag
This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.
This paper introduces a method for ensuring LLMs report their true beliefs by using counterfactual report coordinates that resist pressure but remain responsive to genuine evidence. The approach achieves high performance on a benchmark, demonstrating a causal certificate for internal incentive compatibility.