linear-probes

Tag

Cards List
#linear-probes

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

arXiv cs.AI · yesterday Cached

This paper demonstrates that linear probes decoding concepts from language model activations do not necessarily identify causally relevant features, and introduces a feature-level diagnostic using sparse autoencoders to separate probe alignment from behavioral drivers.

0 favorites 0 likes
#linear-probes

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

arXiv cs.CL · yesterday Cached

This study validates activation-steering claims for sycophancy in language models, finding no usable linear capitulation direction in two small LLMs and highlighting measurement hazards that underestimate sycophancy.

0 favorites 0 likes
#linear-probes

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

arXiv cs.LG · 2026-09-11 Cached

This paper identifies 'perfect aliasing' in truth probes for AI models, where probes fitted on compliant contexts fail to distinguish truth from prescribed actions, and shows that mixed-context training improves detection.

0 favorites 0 likes
#linear-probes

Probe Generalization as Subspace Selection for OOD Deception Detection

arXiv cs.CL · 2026-09-04 Cached

This paper shows that projecting language model activations onto a small subset of principal components from the training distribution enables effective cross-domain transfer for deception detection probes, narrowing the performance gap between baseline and oracle methods.

0 favorites 0 likes
#linear-probes

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

arXiv cs.AI · 2026-09-01 Cached

This paper investigates how instruction-tuned LLMs arbitrate conflicts between system and user instructions, revealing that user-preferring behavior coexists with a readable internal arbitration signal, as demonstrated through mechanistic interpretability techniques and steering experiments.

0 favorites 0 likes
#linear-probes

What AstroPT knows about galaxies, and what that can teach us about LLMs

Hugging Face Daily Papers · 2026-08-23 Cached

This paper proposes using AstroPT, a transformer trained on galaxy images, as a testbed for studying concept emergence during training, finding that galaxy properties emerge in a fixed difficulty-based sequence, which can inform mechanistic interpretability methods for large language models.

0 favorites 0 likes
#linear-probes

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

arXiv cs.CL · 2026-08-13 Cached

A preregistered stress test on a small transformer shows that while latent causal structure can be localized, releasing it into behavior fails: the gate detector inverts out-of-distribution and linear release directions are bounded below sufficiency, dissociating localization from behavioral release.

0 favorites 0 likes
#linear-probes

Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes

arXiv cs.CL · 2026-08-10 Cached

This paper proposes measuring concept content in text using LLM internal activations via linear probes and RFM concept vectors, applied to ESG classification. The best linear probe approaches fine-tuned classifier accuracy without task-specific fine-tuning and outperforms the model's own output, showing activations carry concept content beyond responses.

0 favorites 0 likes
#linear-probes

Probing the Misaligned Thinking Process of Language Models

arXiv cs.AI · 2026-06-24 Cached

This paper proposes monitoring LLM misalignment by decomposing it into fine-grained cognitive processes (misalignment indicators) and detecting them via linear probes on internal activations, achieving high AUROC on out-of-distribution transcripts.

0 favorites 0 likes
#linear-probes

Comparing Linear Probes with Mahalanobis Cosine Similarity

Hugging Face Daily Papers · 2026-06-17 Cached

This paper extends empirical findings that the Mahalanobis cosine similarity (MCS) between linear probes linearly predicts out-of-distribution AUROC, and proves this relationship theoretically under Gaussian assumptions.

0 favorites 0 likes
#linear-probes

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Hugging Face Daily Papers · 2026-05-31 Cached

This paper investigates the production-evaluation gap in large reasoning models (LRMs), finding that they fail to robustly evaluate reasoning despite near-perfect solution production, due to an answer confirmation bias.

0 favorites 0 likes
#linear-probes

What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs

arXiv cs.CL · 2026-05-29 Cached

This paper presents a methodology for delineating concepts and training linear probes to detect them in LLM embeddings, using four example concepts across three models. The work aims to enable scalable monitoring of LLM internal representations.

0 favorites 0 likes
#linear-probes

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

Hugging Face Daily Papers · 2026-05-27 Cached

This paper systematically tests linear probes for deception detection in large language models, finding they fail under distributional shifts but style-augmented probes recover performance, and revealing that deception is encoded through distributed sub-threshold features.

0 favorites 0 likes
← Back to home

Submit Feedback