llm-interpretability

Tag

Cards List
#llm-interpretability

Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

arXiv cs.CL ↗ · 2026-09-18 Cached

This paper introduces an unsupervised, training-free approach to discover prompt-conditional stylistic axes in LLM hidden activations using sampling and PCA, validated through human studies.

0 favorites 0 likes
#llm-interpretability

Interpreting and Steering LLM Agents for Social Simulations

arXiv cs.LG ↗ · 2026-09-16 Cached

This paper explores methods to interpret and steer large language model agents in social simulations, comparing prompt-based, SAE-based, and probe-based techniques, and finds that SAE and probe methods often outperform basic prompting for control and interpretability.

0 favorites 0 likes
#llm-interpretability

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

arXiv cs.LG ↗ · 2026-08-26 Cached

This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.

0 favorites 0 likes
#llm-interpretability

Locating and Controlling Implicit Personalization in Large Language Models

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper investigates how large language models implicitly personalize outputs based on demographic cues, locating an internal activation signal that tracks these shifts and showing that removing this signal can suppress the behavior.

0 favorites 0 likes
#llm-interpretability

Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv cs.AI ↗ · 2026-08-12 Cached

This paper introduces Feature Nonlocality (FNL), an entropy-based metric for measuring the semantic abstractness of SAE features in LLMs, and demonstrates applications in auditing jailbreak mitigation features and improving MATH-500 accuracy via steering.

0 favorites 0 likes
#llm-interpretability

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper asks whether lesion parameters in LLaVA-Vicuna 13B can be recovered from aphasic picture-naming error profiles. The authors find that perturbation intensity is recoverable while layer index is only approximate, with 81.4% counterfactual fidelity and syndrome-discriminative generalization to stroke survivors, suggesting functional redundancy across transformer layers.

0 favorites 0 likes
#llm-interpretability

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper proposes a three-stream detector that combines residual-stream motion with coarse regions and fine directions to better identify reasoning errors in LLMs, improving selection accuracy by up to 12% over state-of-the-art displacement-only methods.

0 favorites 0 likes
#llm-interpretability

A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper applies graph signal processing to analyze how LLMs internally represent numerical sequences during in-context learning, finding that attention-induced token graphs and hidden-state signals show systematic, context-dependent signatures related to input complexity.

0 favorites 0 likes
#llm-interpretability

Inverted Detection and Control in Steering Vectors

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper identifies an 'inverted detection-control' phenomenon where some discriminative steering vectors, despite aligning with positive concept representations, consistently promote the opposite behavior. The authors propose a method to detect such inverted steering vectors without generation, enabling sign flips that improve a detection-based steering pipeline across multiple LLMs and concepts.

0 favorites 0 likes
#llm-interpretability

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

arXiv cs.AI ↗ · 2026-07-31 Cached

This paper investigates internal representational differences between RL and SFT fine-tuned models on mathematical reasoning, finding that RL models exhibit more linearly separable hidden states and hierarchical layer importance. Token allocation variability under repeated sampling suggests training pipeline dependence rather than RL vs SFT alone.

0 favorites 0 likes
#llm-interpretability

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

arXiv cs.CL ↗ · 2026-07-28 Cached

This paper studies behavioral detection of unfaithful chain-of-thought reasoning in LLMs, finding that answer correctness structures detection performance: on incorrect answers, where most unfaithfulness occurs, behavioral signals are at chance, while on correct answers they offer modest separation.

0 favorites 0 likes
#llm-interpretability

Truth is not a direction: a Tarski attack on LLM probes

Hacker News Top ↗ · 2026-07-27 Cached

This article presents a Tarski-inspired diagonal argument showing that no linear probe on an LLM's embedding space can reliably detect truth, drawing parallels to Gödel's incompleteness and Turing's halting problem. It critiques the linear representation hypothesis for truth in language models.

0 favorites 0 likes
#llm-interpretability

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper introduces a framework using Jungian cognitive functions for activation steering in LLMs, demonstrating effective monotonic control over eight functions and revealing structured geometric relationships in activation space.

0 favorites 0 likes
#llm-interpretability

For What Reason? Interpreting Models' Encoding of Causation and Antithesis

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper investigates how instruction-tuned Transformer models (LLaMA and Mistral) encode the discourse relations of causation and antithesis using interpretability techniques on next-token prediction.

0 favorites 0 likes
#llm-interpretability

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper investigates how alignment tuning introduces cue-induced biases such as sycophancy in LLMs, finding that biases are installed by alignment rather than pretraining and can be decoded and steered via hidden state directions.

0 favorites 0 likes
#llm-interpretability

Diagnosing Correctness Probes under Self-Judgement Confounding

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper investigates whether neural network probes that predict correctness of language model outputs actually capture objective correctness or the model's own self-judgement, using conflict cases where the two disagree. The authors find that transferable directions predominantly preserve self-judgement polarity, challenging the interpretation of correctness readouts.

0 favorites 0 likes
#llm-interpretability

Belief-reality separation lives in routing over a shared value slot in language models

arXiv cs.CL ↗ · 2026-07-15 Cached

This paper investigates how language models separate a character's belief from reality, finding that they use a shared value slot for attributed values and a router at the query position to select the frame (belief or reality) to read out. It identifies two routes for asserted and derived beliefs, and shows that the slot itself carries no belief-reality tag; the separation lies in dissociated routing subspaces.

0 favorites 0 likes
#llm-interpretability

What Anthropic’s latest AI discovery does—and doesn’t—show

MIT Technology Review ↗ · 2026-07-13 Cached

Anthropic discovered a hidden internal space (J-space) in LLMs like Claude that contains words influencing reasoning, advancing understanding of AI model internals.

0 favorites 0 likes
#llm-interpretability

The Download: Claude’s inner workings and OpenAI’s “super app”

MIT Technology Review ↗ · 2026-07-10 Cached

A digest covering Anthropic's discovery of a hidden 'J-space' inside Claude's LLM, OpenAI's launch of the ChatGPT Work super app with GPT 5.6 models, humanoid robot surgery on live pigs, SK Hynix's record US listing, and Tencent's deal to acquire Manus.

0 favorites 0 likes
#llm-interpretability

Unsupervised Features Mining via Activation Geometry

arXiv cs.AI ↗ · 2026-07-07 Cached

This paper introduces Mining via Activation Geometry (MAG), an unsupervised framework that extracts reasoning features from LLM activations using natural-language instructions, enabling activation steering and effective training data selection for classifier probes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback