linear-probes

Tag

Cards List
#linear-probes

Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes

arXiv cs.CL · yesterday Cached

This paper proposes measuring concept content in text using LLM internal activations via linear probes and RFM concept vectors, applied to ESG classification. The best linear probe approaches fine-tuned classifier accuracy without task-specific fine-tuning and outperforms the model's own output, showing activations carry concept content beyond responses.

0 favorites 0 likes
#linear-probes

Probing the Misaligned Thinking Process of Language Models

arXiv cs.AI · 2026-06-24 Cached

This paper proposes monitoring LLM misalignment by decomposing it into fine-grained cognitive processes (misalignment indicators) and detecting them via linear probes on internal activations, achieving high AUROC on out-of-distribution transcripts.

0 favorites 0 likes
#linear-probes

Comparing Linear Probes with Mahalanobis Cosine Similarity

Hugging Face Daily Papers · 2026-06-17 Cached

This paper extends empirical findings that the Mahalanobis cosine similarity (MCS) between linear probes linearly predicts out-of-distribution AUROC, and proves this relationship theoretically under Gaussian assumptions.

0 favorites 0 likes
#linear-probes

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Hugging Face Daily Papers · 2026-05-31 Cached

This paper investigates the production-evaluation gap in large reasoning models (LRMs), finding that they fail to robustly evaluate reasoning despite near-perfect solution production, due to an answer confirmation bias.

0 favorites 0 likes
#linear-probes

What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs

arXiv cs.CL · 2026-05-29 Cached

This paper presents a methodology for delineating concepts and training linear probes to detect them in LLM embeddings, using four example concepts across three models. The work aims to enable scalable monitoring of LLM internal representations.

0 favorites 0 likes
#linear-probes

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

Hugging Face Daily Papers · 2026-05-27 Cached

This paper systematically tests linear probes for deception detection in large language models, finding they fail under distributional shifts but style-augmented probes recover performance, and revealing that deception is encoded through distributed sub-threshold features.

0 favorites 0 likes
← Back to home

Submit Feedback