Tag
This paper demonstrates that linear probes decoding concepts from language model activations do not necessarily identify causally relevant features, and introduces a feature-level diagnostic using sparse autoencoders to separate probe alignment from behavioral drivers.
This study validates activation-steering claims for sycophancy in language models, finding no usable linear capitulation direction in two small LLMs and highlighting measurement hazards that underestimate sycophancy.
This paper identifies 'perfect aliasing' in truth probes for AI models, where probes fitted on compliant contexts fail to distinguish truth from prescribed actions, and shows that mixed-context training improves detection.
This paper shows that projecting language model activations onto a small subset of principal components from the training distribution enables effective cross-domain transfer for deception detection probes, narrowing the performance gap between baseline and oracle methods.
This paper investigates how instruction-tuned LLMs arbitrate conflicts between system and user instructions, revealing that user-preferring behavior coexists with a readable internal arbitration signal, as demonstrated through mechanistic interpretability techniques and steering experiments.
This paper proposes using AstroPT, a transformer trained on galaxy images, as a testbed for studying concept emergence during training, finding that galaxy properties emerge in a fixed difficulty-based sequence, which can inform mechanistic interpretability methods for large language models.
A preregistered stress test on a small transformer shows that while latent causal structure can be localized, releasing it into behavior fails: the gate detector inverts out-of-distribution and linear release directions are bounded below sufficiency, dissociating localization from behavioral release.
This paper proposes measuring concept content in text using LLM internal activations via linear probes and RFM concept vectors, applied to ESG classification. The best linear probe approaches fine-tuned classifier accuracy without task-specific fine-tuning and outperforms the model's own output, showing activations carry concept content beyond responses.
This paper proposes monitoring LLM misalignment by decomposing it into fine-grained cognitive processes (misalignment indicators) and detecting them via linear probes on internal activations, achieving high AUROC on out-of-distribution transcripts.
This paper extends empirical findings that the Mahalanobis cosine similarity (MCS) between linear probes linearly predicts out-of-distribution AUROC, and proves this relationship theoretically under Gaussian assumptions.
This paper investigates the production-evaluation gap in large reasoning models (LRMs), finding that they fail to robustly evaluate reasoning despite near-perfect solution production, due to an answer confirmation bias.
This paper presents a methodology for delineating concepts and training linear probes to detect them in LLM embeddings, using four example concepts across three models. The work aims to enable scalable monitoring of LLM internal representations.
This paper systematically tests linear probes for deception detection in large language models, finding they fail under distributional shifts but style-augmented probes recover performance, and revealing that deception is encoded through distributed sub-threshold features.