mechanism-interpretability

Tag

Cards List
#mechanism-interpretability

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.

0 favorites 0 likes
#mechanism-interpretability

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper introduces a method for ensuring LLMs report their true beliefs by using counterfactual report coordinates that resist pressure but remain responsive to genuine evidence. The approach achieves high performance on a benchmark, demonstrating a causal certificate for internal incentive compatibility.

0 favorites 0 likes
← Back to home

Submit Feedback