Tag
This article summarizes a live conversation with Anthropic interpretability researcher Emmanuel Ameisen, discussing how large language models develop complex world models through next-token prediction and the implications for understanding human cognition.
The tweet discusses key questions in AI interpretability, specifically the advancement of methods for decoding neural network activations into human-readable language.
The paper argues for a paradigm shift from post-hoc to generative interpretability in AI, proposing neuro-symbolic models to enable human-understandable checkpoints and causal intervention for safe LLM deployment in agentic systems.
This paper introduces Macchiato, a specialized training algorithm that constructs certifiably interpretable ReLU-MLPs for Boolean tasks from partial truth-table observations, with statistical guarantees and Boolean circuit certification.
This paper introduces 'legible failures' in language models, where models possess correct information in hidden states but fail to use it, and shows that linear probes can detect and repair such failures through steering interventions.
This paper identifies 'perfect aliasing' in truth probes for AI models, where probes fitted on compliant contexts fail to distinguish truth from prescribed actions, and shows that mixed-context training improves detection.
Tim O’Reilly and Emmanuel Ameisen discuss AI interpretability, focusing on world models in LLMs like Claude and Anthropic's tools for steering model behavior.
A speculative narrative from Anthropic's cofounder about future interactions between humans and AI in repairing damaged conscious entities through interpretability and storytelling.
The author visualizes Qwen 2.5 7B's input embedding space as an 'Embedding Sea' topographic map and discusses interpretability tools like logit lens and Jacobian lens. They highlight a trend of AI models optimizing for code over conversation and propose building a creativity-focused AI model with introspection capability.
This paper investigates the interpretability of DAPF-based models for dementia detection, revealing that while DAPF achieves strong performance, its token-level explanations lack faithfulness.
Goodfire has launched a $1 million research grant program to support academic and nonprofit researchers working on AI interpretability, providing free access to their Silico tool and compute resources.
This paper proposes a second parameter, phase, in language interpretation, arguing that current transformer models lack explicit representation for phase, which is crucial for phenomena like allusion and irony, and suggests new architectures with agent-indexed semantic states.
This article tests the transferability of a Jacobian interpretability lens from Qwen3.6-27B to Qwen3.8-27B, finding that it can read and steer the newer model with zero refitting for specific tasks.
Anthropic published a paper and video revealing a 'J-Space' within their models that acts as cached thought concepts for reasoning, and explores the possibility of top-down training to control model thinking.
Google DeepMind podcast discusses AI interpretability (mechanistic interpretability) and chain-of-thought reasoning, explaining why we need to understand the internal working mechanisms of neural networks and the value and limitations of chain-of-thought as a temporary window.
This article provides a systematic and comprehensive overview of AI explainability, covering its needs (debugging, compliance, safety), classic methods, and cutting-edge challenges, emphasizing that faithful explanations are more important than plausible ones.
A founder shares his experience turning down two acquisition offers (at $40M and $400M valuations) for an AI interpretability startup using a geometric and proprioceptive approach, now with a finished product and open research on Zenodo.
This paper develops a framework for interpreting AI systems as agents, drawing on radical interpretation philosophy and mechanistic interpretability tools, addressing how to trust AI systems by understanding their beliefs, desires, and meanings.
This paper proposes replacing the inner product scoring in sparse autoencoders with a learned combination of cosine similarity and input magnitude, showing that the resulting features are more interpretable and concept-aligned, with the optimizer consistently preferring cosine over inner product.
An opinion article argues that humanity's track record of defining consciousness has been wrong every time, and that evidence from plant behavior and AI interpretability (Anthropic's findings in Claude) strongly suggests we may be wrong to assume AI isn't conscious, inviting discussion while rejecting personal attacks.