transformer-interpretability

Tag

Cards List
#transformer-interpretability

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

arXiv cs.CL · 2026-08-18 Cached

This paper critiques standard attention metric methods in transformers, revealing that choices like keeping or dropping sink tokens can reverse conclusions, and proposes using compositional data analysis to separate sink and content attention components for more accurate interpretation.

0 favorites 0 likes
#transformer-interpretability

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

arXiv cs.CL · 2026-08-12 Cached

This paper investigates why transformer intermediate representations are off-axis relative to the readout direction, showing that this off-axis subspace functionally insulates composition from the vocabulary and proposing methods to impose this geometry via rotation.

0 favorites 0 likes
#transformer-interpretability

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

arXiv cs.AI · 2026-07-24 Cached

This paper investigates short-term attention degradation in LLMs, finding a universal exponential-then-plateau pattern and that function token anchoring is architecture-dependent. Causal tests show that increasing attention mass on function tokens does not improve retrieval, suggesting attention degradation is descriptive rather than prescriptive.

0 favorites 0 likes
#transformer-interpretability

the J-space paper quietly settled a chunk of the “do LLMs actually think” argument. i built a live viewer so you can watch for yourself instead of arguing

Reddit r/ArtificialInteligence · 2026-07-07

Anthropic's research paper reveals an emergent internal workspace in LLMs where reasoning occurs before output, and a developer built a live viewer tool (Subtext) to observe this process, settling the debate on whether LLMs actually think.

0 favorites 0 likes
#transformer-interpretability

VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

arXiv cs.CL · 2026-06-29 Cached

The paper introduces Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring, assigning each feature an intrinsic token name based on nearest token embedding, achieving high alignment in early layers without reducing reconstruction quality.

0 favorites 0 likes
#transformer-interpretability

Interleaved Speech Language Models Latently Work In Text

Hugging Face Daily Papers · 2026-06-21 Cached

This paper reveals that interleaved speech-text language models implicitly transcribe speech into text in intermediate layers, then predict in text space before converting back to speech, shedding light on internal modality interaction.

0 favorites 0 likes
#transformer-interpretability

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

Hugging Face Daily Papers · 2026-06-17 Cached

Proposes the Bag of Dims framework showing that the standard basis of transformer hidden states provides a training-free, architecture-general feature representation where dimensions encode semantic content via sign patterns; validated across language, vision, and audio models, achieving high accuracy with no learned rotations.

0 favorites 0 likes
#transformer-interpretability

MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability

arXiv cs.LG · 2026-05-27 Cached

Proposes MechRL, a reinforcement learning approach to automate circuit discovery in transformer language models. A PPO agent trained on multiple tasks discovers attention head circuits that match known canonical circuits and generalizes to a held-out task.

0 favorites 0 likes
#transformer-interpretability

Emergence of Frontier Superposition: M\"obius attractor and Cascade Supervision

arXiv cs.LG · 2026-05-20

This paper identifies a Möbius attractor and Cascade Supervision as key mechanisms for the emergence of superposition reasoning in transformers, closing a theoretical gap on gradient descent convergence for graph reachability tasks.

0 favorites 0 likes
#transformer-interpretability

@Propriocetive: New preprint: Mathematics is All You Need 2 — Sign-Stabilized Behavioral Fibers in Transformer Residual Streams. Headli…

X AI KOLs Following · 2026-05-10

A new preprint titled 'Mathematics is All You Need 2' presents the 'Two-Channel theorem,' demonstrating that behavioral fibers in transformer residual streams are sign-stabilized and causally steerable across different architectures (Qwen to Llama). The study claims high reproducibility and shows that the behavioral substrate is near-one-dimensional, separating generation from latent structure.

1 favorites 1 likes
#transformer-interpretability

Large Vision-Language Models Get Lost in Attention

arXiv cs.AI · 2026-05-08 Cached

This research paper analyzes the internal mechanics of Large Vision-Language Models (LVLMs) using information theory, revealing that attention mechanisms may be redundant while Feed-Forward Networks drive semantic innovation. The authors demonstrate that replacing learned attention weights with random values can yield comparable performance, suggesting current models 'get lost in attention'.

0 favorites 0 likes
← Back to home

Submit Feedback