Interpretability Can Be Actionable

arXiv cs.LG Papers

Summary

This position paper argues that interpretability research should be evaluated based on actionability—the extent to which insights enable concrete decisions and interventions. The authors propose a framework with evaluation criteria aligned with practical outcomes to address the lack of real-world impact in current interpretability work.

arXiv:2605.11161v1 Announce Type: new Abstract: Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impact, raising questions about its relevance and utility. This position paper argues that the central missing ingredient is not new methods, but evaluation criteria: interpretability should be evaluated by actionability--the extent to which insights enable concrete decisions and interventions beyond interpretability research itself. We define actionable interpretability along two dimensions--concreteness and validation--and analyze the barriers currently preventing real-world impact. To address these barriers, we identify five domains where interpretability offers unique leverage and present a framework for actionable interpretability with evaluation criteria aligned with practical outcomes. Our goal is not to downplay exploratory research, but to establish actionability as a core objective of interpretability research.
Original Article
View Cached Full Text

Cached at: 05/13/26, 06:32 AM

# Interpretability Can Be Actionable
Source: [https://arxiv.org/html/2605.11161](https://arxiv.org/html/2605.11161)
Fazl BarezTal HaklayIsabelle LeeMarius MosbachAnja ReuschNaomi SaphraByron C WallaceSarah WiegreffeEric WongIan TenneyMor Geva

###### Abstract

Interpretability aims to explain the behavior of deep neural networks\. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impact, raising questions about its relevance and utility\. This position paper argues that the central missing ingredient is not new methods, but evaluation criteria: interpretability should be evaluated by*actionability*—the extent to which insights enable concrete decisions and interventions beyond interpretability research itself\. We define actionable interpretability along two dimensions—concreteness and validation—and analyze the barriers currently preventing real\-world impact\. To address these barriers, we identify five domains where interpretability offers unique leverage and present a framework for actionable interpretability with evaluation criteria aligned with practical outcomes\. Our goal is not to downplay exploratory research, but to establish actionability as a core objective of interpretability research\.

Machine Learning, ICML

![Refer to caption](https://arxiv.org/html/2605.11161v1/x1.png)Figure 1:Actionability checklist for interpretability research\.## 1Introduction

Interpretability research seeks to explain modern machine learning systems\. In recent years, it has grown into a large and active research area\(Mosbachet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib94); Maslejet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib110)\), driven by the intuition that understanding models should help make them more reliable, efficient, safer, and aligned with human values\(Bereska and Gavves,[2024](https://arxiv.org/html/2605.11161#bib.bib111)\)\.

Despite its growth, interpretability work is often seen as lacking practical impact such as informing changes to models, training practices, deployment decisions, or policy\(Krishnan,[2020](https://arxiv.org/html/2605.11161#bib.bib97); Greenblattet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib77); Potts,[2025](https://arxiv.org/html/2605.11161#bib.bib95)\), motivating calls to focus on clearly demonstrable outcomes beyond “understanding” itself\(Haklayet al\.,[2025b](https://arxiv.org/html/2605.11161#bib.bib106); Upadhyay and Barez,[2025](https://arxiv.org/html/2605.11161#bib.bib28); Nandaet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib78); Barez,[2026](https://arxiv.org/html/2605.11161#bib.bib2)\)\. Our framing draws in part on discussions from the ICML 2025 workshop on actionable interpretability, which aimed to foster dialogue on leveraging interpretability insights to drive tangible advancements in AI\.

In this paper, we argue that interpretability research should be evaluated not only by how well it explains models, but by what those explanations enable us to do\.That is, interpretability should be held to a standard of*actionability*\.

We contend that the field’s impact will be strengthened if it explicitly tracks not only what we understand, but what that understanding enables us to do\. We do not, however, argue that all interpretability work must immediately yield actionable outcomes, nor that purely exploratory contributions lack value\. Indeed, methodological novelty and demonstrated applications are not at odds—grounding findings in real\-world actions holds methods to a higher standard, providing evidence that insights reflect genuine model behavior rather than artifacts of a particular analysis\. What is missing in interpretability research is notmethods, butevaluation criteria: a shared framework for determining when interpretability research is successful from a practical, decision\-oriented perspective\. We therefore advance a framework for actionable interpretability: analyzing current limitations, identifying opportunities for impact, and suggesting practical tools to increase actionability\.[Figure1](https://arxiv.org/html/2605.11161#S0.F1)translates our central thesis into concrete steps for researchers\.

Scope\.We consider interpretability in modern machine learning, focusing on deep learning and foundation models\. Although we draw many examples from LLMs, our arguments apply broadly to any deep neural networks domain which may require explanation\. This is a position paper: rather than an exhaustive survey, we propose actionability as a unifying lens for evaluating interpretability work\.

The paper is organized as follows: Section[2](https://arxiv.org/html/2605.11161#S2)defines actionable interpretability\.[Section3](https://arxiv.org/html/2605.11161#S3)diagnoses the barriers that currently prevent interpretability from achieving real\-world impact\. The rest of the paper paves the way towards more actionable interpretability\.[Section4](https://arxiv.org/html/2605.11161#S4)identifies opportunities for actionability\.[Section5](https://arxiv.org/html/2605.11161#S5)presents a framework for categorizing actions and[Section6](https://arxiv.org/html/2605.11161#S6)discusses evaluation criteria aligned with actionability\.[Section7](https://arxiv.org/html/2605.11161#S7)addresses counter\-arguments\.[Section8](https://arxiv.org/html/2605.11161#S8)reviews related work\.[Section9](https://arxiv.org/html/2605.11161#S9)concludes with an actionable checklist for researchers, summarized in[Figure1](https://arxiv.org/html/2605.11161#S0.F1)\.

## 2Defining Actionable Interpretability

We consider a work111By “work”, we refer broadly to research or engineering contributions, including methods, models, analyses, benchmarks, and empirical studies\.to beinterpretability\-orientedif it aims to explain or analyze an AI model—for example, works that analyze model representations, explain specific behaviors or capabilities, or discover internal mechanisms\. Having this distinction, we provide the following definition:

Actionable InterpretabilityAn interpretability\-oriented work is actionable if it producesinsightsabout an AI model that inform or guideactionstoward non\-interpretability objectives\.

Insightsare outputs of interpretability work: findings about how models represent or process inputs, explanations of internal mechanisms, or methods that clarify behavior\.

Actions\(toward non\-interpretability objectives\) are decisions made by humans in response to interpretability insights that would not have been taken otherwise\. These fall outside the scope of interpretability itself and ideally lead to concrete improvements such as enhanced performance, better\-calibrated trust, or improved safety\.

### 2\.1Dimensions of Actionability

In practice, actionability is more fine\-grained and not binary\. Interpretability\-oriented work can support different levels of actionability, which we characterize along two key dimensions:concretenessandvalidation\.

Concretenesscaptures how precisely an action is articulated\. At the low end are vague suggestions \(“could inform safety research”\) or no suggestions at all; at the high end are exact specifications with implementation details\.

Validationcaptures empirical support for an action’s utility\. At the low end, actions are untested hypotheses; at the high end, they are systematically evaluated with quantitative or qualitative evidence of meaningful outcomes beyond interpretability research itself\.

Together, these dimensions span a space for situating interpretability work \(illustrated in[Figure3](https://arxiv.org/html/2605.11161#A1.F3)in the Appendix\):

*Low concreteness, low validation*: Work in this region recommends no specific actions to validate\. The insights from this work may, however, inform future work by providing a starting point that others can build upon and test\. For example,Gevaet al\.\([2021](https://arxiv.org/html/2605.11161#bib.bib66)\)’s key\-value memory view of MLPs directionally motivated subsequent work on knowledge localization and model editing\.Wanget al\.\([2023](https://arxiv.org/html/2605.11161#bib.bib71)\),Conmyet al\.\([2023](https://arxiv.org/html/2605.11161#bib.bib153)\)and others laid groundwork for circuit\-based analysis\. While not the emphasis of this paper, such exploratory work is imperative to drive the field forward\.

*High concreteness, low validation*: Concrete actions proposed but not empirically validated—e\.g\., approaches for verifying scientific models to build trust in their predictions\(Kinget al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib59);[Liet al\.,](https://arxiv.org/html/2605.11161#bib.bib91); Ferreiraet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib54)\)or optimizing model deployment and training\(Zhaoet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib92); Chenet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib148)\)\.

High concreteness, high validation\.Precise specifications with demonstrated utility, informed by interpretability insights drawn either from the work itself or prior work\. Examples include model editing methods leveraging the MLP key\-value store view\(Menget al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib88); Wanget al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib71); Aradet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib150); Fanget al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib115)\), are based on sparse\-auto\-encoders\(Gur\-Ariehet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib60); Ashuachet al\.,[2025a](https://arxiv.org/html/2605.11161#bib.bib61)\)or insights into the role of cross\-attention layers\(Orgadet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib151); Gandikotaet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib96)\)\. Representation finetuning\(Wuet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib87)\), an alternative to LoRA\-based methods, was inspired by interpretability findings\.Schutet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib62)\)use concept vectors to uncover novel chess concepts transferable to human players\.Anthropic \([2025](https://arxiv.org/html/2605.11161#bib.bib53)\)analyzed internal activations during a safety audit of Claude\.

## 3Why Interpretability Isn’t \(Yet\) Actionable

Despite growing interest, several barriers limit interpretability’s real\-world impact: misaligned incentives, methodological limitations, and deployment challenges\. These reinforce a cycle where actionability is not prioritized, methods lack validation, and deployment yields little feedback\. The rest of the paper discusses how to advance actionable interpetability despite these limitations\.

### 3\.1Misaligned Incentives

The interpretability community does not sufficiently reward work for demonstrating practical value\. Without a strong incentive to prove that interpretability methods deliver real\-world value, researchers are less likely to conduct or show interest in actionable interpretability work\.

Publication standards do not require actionability\.Papers can be accepted based purely on methodological novelty, with no requirement to demonstrate applications\. Meanwhile,application\-focused work is under\-rewarded,Practical demonstrations may be dismissed as “merely engineering” despite their greater potential impact\. We argue that methodological novelty and application demonstration are not at odds—demonstrating applications holds interpretability methods to a higher standard, providing evidence that findings are grounded in reality\. This asymmetry—low requirements for actionability combined with low rewards for demonstrating it—substantially reduces researchers’ incentive to pursue practical applications\.

These issues are not unique to the interpretability field, and also exist in mainstream machine learning \(ML\) research\. However, unlike applied ML, where benchmark performance provides immediate feedback,interpretability lacks clear signals of success\. Mainstream ML research has a forcing function interpretability lacks: new methods must demonstrate gains on established benchmarks\. The field has matured by moving from toy problems to real\-world tasks—MNIST to ImageNet, Penn Treebank to diverse downstream tasks\. However, the interpretability field has yet to fully mature, lacking agreed\-upon standards\.

### 3\.2Methodological Limitations

These incentive gaps often manifest as concrete technical problems that prevent interpretability insights from translating into action\. In this section, we outline such technical problems and associated methods\.

Lack of actionable insights\.Interpretability work often fails to articulate how findings can inspire concrete actions\. This limitation was reflected in the ICML 2025 workshop on Actionable Interpretability, where in 21\.8% of the submitted papers, at least one reviewer explicitly flagged the work as insufficiently actionable\.Mosbachet al\.\([2024](https://arxiv.org/html/2605.11161#bib.bib94)\)showed that although interpretability papers are cited, their impact is predominantly conceptual—most citations do not credit changes to training, architecture, or evaluation\. While foundational work may eventually drive actionability\(Bau,[2025](https://arxiv.org/html/2605.11161#bib.bib65)\), the field should explicitly reflect on how insights matter beyond its boundaries\.

Oversimplified setups\.Much research uses simplified tasks and small models\. For instance, many mechanistic studies on LLMs focus on single next\-token predictions\(Muelleret al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib27)\), whereas real usage involves multi\-token generation\. These settings are valuable as controlled testbeds, but their insights may not transfer to realistic settings\. Recent work byHaklayet al\.\([2025a](https://arxiv.org/html/2605.11161#bib.bib181)\)has begun addressing these limitations with circuit discovery that handles variable\-length inputs\.

Insufficient comparative analysis\.Many works lack rigorous comparisons against alternative approaches and fail to evaluate robustness across architectures, datasets, and tasks\. AsCasper \([2023](https://arxiv.org/html/2605.11161#bib.bib74)\)argues, weak evaluation hinders progress toward practical tools\. Recent benchmarks have begun to address this limitation, highlighting the importance of empirical comparisons\. AxBench\(Wuet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib86)\)showed that prompting and finetuning often outperform interpretability methods for LLM steering\. MIB\(Muelleret al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib27)\)evaluates both circuit localization and causal variable localization—two widely studied directions that previously lacked a means to compare methods\.

![Refer to caption](https://arxiv.org/html/2605.11161v1/x2.png)Figure 2:Five domains where interpretability offers unique leverage to drive concrete improvements\.
### 3\.3Deployment Challenges

Even when interpretability methods offer practical value, several barriers hinder their adoption\.

Technical complexity\.To employ interpretability techniques, a user must deeply understand model internals and be familiar with specialized libraries\(Nanda and Bloom,[2022](https://arxiv.org/html/2605.11161#bib.bib25); Fiotto\-Kaufmanet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib26)\)\. Those outside the community often lack the expertise required\(Ashtariet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib183)\)and so rarely adopt these methods, especially when simpler alternatives exist\.

The open\-weights assumption\.Most methods require direct access to weights and activations, restricting applicability to open\-weight models\. This creates a tension: interpretability is often motivated by safety concerns around powerful frontier models, yet these same models are often proprietary and therefore resistant to such analysis\.

## 4Making Interpretability Actionable

In[Section3](https://arxiv.org/html/2605.11161#S3), we identified limitations that prevent interpretability research from delivering sustained practical impact\. Here, we turn to solutions: we identify opportunities where interpretability is uniquely positioned to drive concrete improvements\. We identify five domains in which interpretability offers unique leverage—where there is a fundamental advantage from answeringwhyquestions about the model\. In Sections[5](https://arxiv.org/html/2605.11161#S5)and[6](https://arxiv.org/html/2605.11161#S6)we discuss the implementation: a framework for actionable interpretability, and its evaluation\.

Problems scaling does not solve\.The scaling hypothesis—the claim that many capabilities improve predictably with increased model size—has proven remarkably successful\. Yet certain failure modes persist or even worsen with scale, including hallucinations, catastrophic forgetting, biases and adversarial brittleness\. The persistence of these failures across model scales suggests they are fundamental to our current modeling paradigm rather than due to limited capacity\. Interpretability offers a path forward precisely because it can identify why models fail\. Standard evaluations detect failures but cannot explain their origins or suggest principled interventions\. By contrast, interpretability enables sharper hypotheses about underlying mechanisms and reasoning about potential fixes\. Even partial insights can rule out hypotheses and guide the design of solutions\.

Alignment\.As AI systems become more capable, ensuring they behave as intended becomes more critical and more difficult\. Alignment today still relies on fine\-tuning and data curation rather than understanding\-driven interventions, but as AI progresses, verifying that AI goals match human goals will shift from aspiration to necessity\. Can we credibly claim a model has no deceptive capabilities without understanding its decision\-making? Can we audit for backdoors through black\-box testing alone? Since alignment concerns what a model optimizes for, it cannot be fully established without interpretability\.

Surgical interventions\.Retraining a flawed model is expensive and risks introducing other unexpected outcomes\. Interpretability enables targeted modifications; identifying components responsible for unwanted behaviors allows surgical fixes while preserving other functionality\. Though not yet fully practical, this is among interpretability research’s most actionable outcomes—it’s efficient and affordable\. These techniques can enable post\-hoc maintenance: bug fixes, policy updates, and rapid responses to new failure modes\.

Architectural design\.Current improvements emerge largely through trial and error—an inefficient, opaque process where success may not scale or transfer to new domains\. Interpretability can transform this paradigm by linking specific design choices \(data curation, architecture, optimization\) to their effects on model behavior\. This approach could accelerate progress by narrowing the space of plausible architecture modifications, reducing both labor and compute required\.

Translation of explanation to meaningful concepts\.The most natural role of interpretability is explaining model behavior, yet translating internal signals into meaningful concepts remains a critical bottleneck\. In high\-stakes domains like healthcare, a radiologist needs to know if an AI\-assisted diagnosis depends on clinically relevant features, not which pixels activate; A developer debugging failures needs specific, legible insights than what current circuit discovery methods provide \(“layers 7 and 9 interact together”\)\. Automated methods that translate technical explanations into domain\-appropriate,actionableconcepts could unlock interpretability’s core promise\. This also includes methods that scale interpretability methods beyond a single input or template into a more natural, diverse setting\.

## 5A Framework for Actionable Interpretability

We now present a framework for the actions interpretability enables and the actors who carry them out\. This framework is intended to help researchers identify and articulate the actionability potential of their own work\. The examples we present throughout this section demonstrate successful cases of actionable interpretability, yet they represent a relatively small fraction of the broader literature\.

Table 1:Different audiences of interpretability and their actionable outputs\.Who takes the “action” in “actionable”?Different stakeholders have different capabilities and motivation, as illustrated in[Table1](https://arxiv.org/html/2605.11161#S5.T1)\.AI developersmay use mechanistic insights to inform model design\.Deployment engineersmay focus on controlling behavior in specific applications\.Domain expertslike clinicians need feature\-level rationales to justify diagnoses\.A policymakerrelies on system\-level summaries of fairness or compliance\. These actors rarely operate in isolation—interpretability outputs should serve as communication interfaces across roles\. A clinician’s feedback about unreliable explanations may reveal failure modes to engineers\. Similarly, a policy maker’s compliance requirements may drive developers toward mitigation\.

An interpretability work will become more actionableif it is explicit about its intended audience and the decisions it aims to support\.At the same time,[Section3\.3](https://arxiv.org/html/2605.11161#S3.SS3)emphasizes technical complexity as a major barrier to deployment\. Taken together, these observations suggest that effective actionable interpretability work must do more than produce insights or analyses: it must specify who can act on its findings, what actions are enabled and how \(by providing code or explicit instructions\), and where those actions plausibly apply\.

We next describe the different types of actions that interpretability insights can drive\. We classify them according to what each action affects\. Rather than providing an exhaustive taxonomy, we provide representative examples for each\. Additional examples are listed in[AppendixB](https://arxiv.org/html/2605.11161#A2)\.

### 5\.1Actions that Modify Model Output

Interpretability can inform decisions that directly modify model behavior—changes in training, inputs, weights, or internal computations\. These decisions are primarily made by developers and researchers with access to model internals\.

Data curation\.Influence functions can help identify training examples that help or harm model performance\.Koh and Liang \([2017](https://arxiv.org/html/2605.11161#bib.bib9)\)used them to detect mislabeled examples and improve accuracy\.Hanet al\.\([2020](https://arxiv.org/html/2605.11161#bib.bib68)\)used them to expose artifacts in training data\. More recently,Agiaet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib156)\)applied influence functions to robot learning, identifying detrimental demonstrations and achieving state\-of\-the\-art results with only 33% of the original training data\.

Model input\.Interpretability can inform decisions about what inputs to provide to models\.Zhouet al\.\([2024](https://arxiv.org/html/2605.11161#bib.bib158)\)built on the insight that transformers implement an internal optimizer for in\-context learning\(Akyüreket al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib157)\), using influence functions to identify which in\-context demonstrations help versus harm performance\.

Training decisions\.Casperet al\.\([2024a](https://arxiv.org/html/2605.11161#bib.bib168)\)built on insights about how models internally represent concepts and used latent adversarial training—perturbing internal representations—to defend against unforeseen vulnerabilities, thereby removing backdoors and improving robustness\.

Direct control\.Interpretability can identify components responsible for specific behaviors, enabling targeted interventions\. Model editing modifies weights to insert, remove, or correct behaviors without full retraining\(Menget al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib88),[2023](https://arxiv.org/html/2605.11161#bib.bib114); Orgadet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib151)\)\. Runtime interventions steer activations along interpretable directions at inference time\(Liet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib154); Turneret al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib1)\)\. Concept bottleneck models introduce human\-defined concepts as intermediate representations, enabling expert\-guided control\(Kohet al\.,[2020b](https://arxiv.org/html/2605.11161#bib.bib4); Yuksekgonulet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib5); Oikarinenet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib6)\)\.

Safety\.Interpretability can support efforts to remove or suppress unsafe behaviors encoded in model weights\. Techniques such as concept erasure\(Ravfogelet al\.,[2020](https://arxiv.org/html/2605.11161#bib.bib17); Elazaret al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib18)\)and machine unlearning\(Gandikotaet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib96); Ashuachet al\.,[2025b](https://arxiv.org/html/2605.11161#bib.bib31); Bourtouleet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib20); Cao and Yang,[2015](https://arxiv.org/html/2605.11161#bib.bib19)\)provide principled approaches for mitigating privacy risks and removing unwanted behaviors by identifying and neutralizing specific learned associations\.

### 5\.2Actions about Deployment and Use

Interpretability can inform decisions that end users—including domain experts \(e\.g\., clinicians\) and other practitioners—make when interacting with model outputs\. Unlike decisions that modify the model’s final output, these actions changewhat humans dowith model predictions: when to trust them, when to override them, and how to integrate them into their workflows\.

End user decisions\.Prenosilet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib155)\)developed a neuro\-symbolic system combining GPT\-4 with rule\-based expert systems for clinical data extraction, providing the transparency and auditability that enabled radiologists to confidently use AI while maintaining oversight\. Activation patching may reveal when models are \(overly\) relying on patient demographic information when making clinical predictions\(Ahsanet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib67)\)\.

Work on uncertainty estimation from internal representations\(Kadavathet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib162), inter alia\)enables users to detect potential errors and make informed decisions about when to trust model outputs\.

Deployment decisions\.Interpretability enables predictions about where models will fail\.Huanget al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib16)\)use internal mechanisms to identify out\-of\-distribution failures during inference, whileLiet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib15)\)use them to predict errors on unseen distributions\. Such insights support routing decisions—whether to return a model’s answer or escalate to alternative methods\. Work that detects model errors based on internal representations\(Kadavathet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib162), inter alia\)can also be used in this context\.Chenet al\.\([2024a](https://arxiv.org/html/2605.11161#bib.bib167)\)demonstrate the potential value, achieving up to 98% cost reduction while matching GPT\-4 performance through uncertainty\-based routing\.

### 5\.3Shaping Future Practice

Beyond immediate interventions, interpretability may offer insights that inform how the field builds and governs future systems\. This has longer\-term, broader impact\.

Policy and regulation\.Interpretability requirements are increasingly embedded in regulations\. The EU AI Act mandates explainability for high\-risk AI systems\(ArtificialIntelligenceAct\.eu,[2025](https://arxiv.org/html/2605.11161#bib.bib170)\)\. GDPR’s Article 22\(European Union,[2016](https://arxiv.org/html/2605.11161#bib.bib171)\)restricts automated decision\-making with significant effects and requires safeguards, including the right for human intervention\. A central challenge to regulation is verification: whether interpretability enables credible claims about the absence of dangerous mechanisms, even with limited access to proprietary model details\. This is currently largely unsolved, even for public or open\-weight models\.

Learning from superhuman models\.When models exceed human expertise, interpretability becomes a mechanism not just for trust or safety, but for transferring knowledge from machines back to humans\.Schutet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib62)\)show that interpreting AlphaZero’s\(Silveret al\.,[2018](https://arxiv.org/html/2605.11161#bib.bib177)\)strategies surfaces novel chess concepts that can teach human grandmasters—demonstrating interpretability’s potential for knowledge transfer from AI to humans\.

Development of future models\.Interpretability can shift model design from trial\-and\-error toward principled engineering guided by an understanding of the computations performed\. For example, induction heads in Transformers\(Elhageet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib69); Olssonet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib172)\)provided a mechanism for in\-context learning that traditional state\-space models lacked, directly influencing the design of the Mamba architecture’s selective state updates\(Gu and Dao,[2024](https://arxiv.org/html/2605.11161#bib.bib176)\)\.

## 6Evaluating Actionability

How do we know if interpretability work is actionable? Current practice often evaluates methods against other interpretability techniques or relies on intuitive notions of “understanding”\. This is insufficient for actionable interpretability\. Here, we require metrics that measure whether insights actually enable better decisions and outcomes\. We propose evaluation criteria for insights that can enable each of the three action categories from[Section5](https://arxiv.org/html/2605.11161#S5)\. All of these criteria share a common principle:interacting with the world beyond the field of interpretabilityrather than solely comparing methods within the field\.

### 6\.1Evaluating Actions that Modify Outputs

Comparative utility against standard baselines\.A major limitation of current evaluation is “grading on a curve”—comparing interpretability methods only against each other\. Instead, performance should be measured against standard, pragmatic ML baselines, such as prompting or fine\-tuning, and using standard metrics such as accuracy, or benchmark\-specific measures\. For example, does steering with SAEs improve refusal behavior more than targeted prompting or small LoRA\(Huet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib184)\)adapters? This defines actionability as marginal leverage gained over simpler methods that do not require deep mechanistic understanding\.

Mechanistic faithfulness\.This measures whether an explanation correctly identifies model components causally involved in a specific computation\. Evaluation uses intervention\-based verification on tasks with well\-defined semantics: an explanation is faithful if intervening on identified components produces predicted changes—altering target computations while leaving unrelated behaviors intact\. For example, when reverse engineering an LLM’s sorting algorithm, one can identify the comparison component and intervene to reliably swap two specific items in the output\.

Generalization\.To address whether an insight holds beyond a specific setting, our metrics must evaluate whether it generalizes, e\.g\., across seeds, input perturbations, architectures, and scales, without requiring rediscovery\. Concrete evaluations may include transferring identified circuits, features, or mechanisms between models of different sizes, or from toy settings to frontier models\. Successful transfer indicates the method captures robust, reusable structure\.

Specificity\.Next, we consider whether an interpretability claim identifies a component that is specifically linked to a distinctive target, rather than a broad correlation\. This is evaluated in two ways\. First, does the proposed component explain the behavior better than alternatives? This establishes that the finding is genuinely informative rather than arbitrary\. Second, when intervening on the component to modify target behavior, do unrelated behaviors remain unchanged? This should be evaluated on standard benchmarks that quantify model capabilities\. For example, if a neuronspecificallycontrols review sentiment, an intervention may affect the tone while preserving the factual content and performance on unrelated tasks\. Interventions that reveal broad effects suggest the component plays a generalized, entangled role in model behavior\.

### 6\.2Evaluating Actions about Deployment

Task\-enhancement\.The most direct user\-facing metric is whether explanations improve performance on the task the model supports—not the model itself, but human decision\-making, speed, or reliance on outputs\. This typically requires human\-subject evaluations, which are critical since prior work suggests explanations do not reliably improve performance\(Spillneret al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib141)\)\.

Understandability\.Incomprehensible explanations are unlikely to influence user decisions, even if technically correct\. This is especially pronounced in high\-stakes settings where practitioners face severe penalties for errors\. For example, in medicine, clinical training requirements and legal liability create high barriers to AI adoption, even for superhuman systems\. Importantly, understandability is orthogonal to faithfulness—an explanation may accurately reflect model behavior while failing to be usable\. We expand on these evaluations in[SectionC\.1](https://arxiv.org/html/2605.11161#A3.SS1)\.

Reliability\.Even when explanations improve task performance and are understandable, they may fail to be actionable if perceived as unreliable\. Explanations that vary substantially across random seeds or minor perturbations introduce uncertainty that undermines trust\. While related to generalization \(Cf\.[Section6\.1](https://arxiv.org/html/2605.11161#S6.SS1)\), reliability focuses on within\-task stability—whether a user can expect consistent explanations across repeated or slightly varied contexts\. This framing is especially important in high\-stakes domains where explanations guide interventions and brittle explanations are often viewed as unsafe or uninformative\(Ghassemiet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib173); Arunet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib174)\)\. We provide examples for measuring reliability in[SectionC\.2](https://arxiv.org/html/2605.11161#A3.SS2)\.

### 6\.3Evaluating Actions Shaping Future Practices

In policy contexts, interpretability acts as an institutional lever rather than a scientific diagnostic\. Actionability should be measured by whether methods enable feasible AI governance by regulators and safety teams\(Upadhyay and Barez,[2025](https://arxiv.org/html/2605.11161#bib.bib28)\), and not by depth of insight for a handful of researchers staring at neuron visualizations\. From this perspective, interpretability is policy\-actionable to the extent that it expands feasible governance interventions: supporting safety audits, interpretable proxy models, or verifying the absence of dangerous mechanisms\. Practically, interpretability should reduce monitoring and mitigation costs relative to blunt instruments like pausing deployment, supporting concrete policy tools \(e\.g\., risk audits, model cards, licensing regimes\) while remaining legible to non\-experts\.

## 7Alternative Viewpoints

Is actionability the right goal?Some defend interpretability as basic science regardless of actionability\.Bau \([2025](https://arxiv.org/html/2605.11161#bib.bib65)\)argues for curiosity\-driven research since we do not yet know which techniques may permit actionable insights\. We do not disagree—we argue that impact will increase by tracking actionability as a yardstick, not that all work must be \(immediately\) actionable\.

How should actionability be defined and measured?Many believe interpretability’s value lies in AI safety\(Nanda,[2022](https://arxiv.org/html/2605.11161#bib.bib76); Olah,[2023](https://arxiv.org/html/2605.11161#bib.bib99); Anwaret al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib82); Amodei,[2025](https://arxiv.org/html/2605.11161#bib.bib79); Markset al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib80); Shahet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib83)\)\. Some argue safety is the*singular*actionable goal\(Greenblattet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib77); Nanda,[2022](https://arxiv.org/html/2605.11161#bib.bib76); Nandaet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib78); Hendrycks and Hiscott,[2025](https://arxiv.org/html/2605.11161#bib.bib98); Marks,[2025](https://arxiv.org/html/2605.11161#bib.bib103)\), and performance improvements represent “dual use” problem\(Segerie,[2023](https://arxiv.org/html/2605.11161#bib.bib75); So8res,[2023](https://arxiv.org/html/2605.11161#bib.bib112); Shovelain and Mckernon,[2023](https://arxiv.org/html/2605.11161#bib.bib113)\)\. We argue for a broader framework centered on human users, encompassing both safety and performance improvements\.

Is actionability achievable?For practitioners whose priority is building better models, there must be decisive evidence that interpretability methods outperform alternatives with minimal additional effort\. We currently lack this evidence for many research lines, contributing to distrust within the broader ML community\. However, we can reduce skepticism by ensuring baselines include non\-interpretability methods, as we discuss in[Section6](https://arxiv.org/html/2605.11161#S6)\. Recent community efforts aim to unify the discussion around actionability\(Haklayet al\.,[2025b](https://arxiv.org/html/2605.11161#bib.bib106)\)\. Many remain optimistic, and the community is actively pivoting\(Gao,[2025](https://arxiv.org/html/2605.11161#bib.bib102); Ho,[2025](https://arxiv.org/html/2605.11161#bib.bib101); Steinhardt and Schwettmann,[2024](https://arxiv.org/html/2605.11161#bib.bib100); Marks,[2025](https://arxiv.org/html/2605.11161#bib.bib103)\)\. We have additionally laid out arguments in this paper for why interpretability*is already actionable*in many scenarios \([Section5](https://arxiv.org/html/2605.11161#S5)\)\.

It would be premature to discount actionable interpretability when the field is still at an early\-stage compared to to other scientific disciplines, and our objects of study have only emerged in their current form in the past five years\.

## 8Previous Work

Previous conceptual and position work\.Lipton \([2018](https://arxiv.org/html/2605.11161#bib.bib37)\)pointed out that interpretability is an overloaded term, and distinguishes betweentransparency\(understanding how a model works\) andpost\-hoc interpretability\(explaining its decisions after the fact\)\.Miller \([2019](https://arxiv.org/html/2605.11161#bib.bib39)\)argues that interpretability requires attention to user and social context because much work neglects decades of findings from from philosophy, psychology, and cognitive science research which highlight how explanations should be grounded\. Similarly,Jacovi and Goldberg \([2021](https://arxiv.org/html/2605.11161#bib.bib38)\)emphasize the role of*social attribution*in explanations, namely the implicit attribution of intent to models\.Rudin \([2019](https://arxiv.org/html/2605.11161#bib.bib43)\)advances this perspective, suggesting that researchers should abandon post\-hoc explanations of models entirely and instead focus on inherently interpretable models\.

Others\(Doshi\-Velez and Kim,[2017](https://arxiv.org/html/2605.11161#bib.bib42)\)argue that without clear criteria, interpretability research may prioritize intuitively appealing methods over practically valuable ones\. While their emphasis aligns with actionable interpretability, we contend evaluation should focus on the specific interventions and decisions the insights enable\.Calderon and Reichart \([2025](https://arxiv.org/html/2605.11161#bib.bib175)\)note that NLP interpretability often fail to generalize beyond their initial domains and stressed the importance of defining stakeholders\. While complementing to our view, our focus is on translating insights into actionable outcomes\.

Most recently,Nandaet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib78)\)advocate a “pragmatic” approach to interpretability that focuses on solving specific problems rather than solely reverse\-engineering models, using meaningful “proxy tasks” to drive rapid iteration\. On the other hand,Bau \([2025](https://arxiv.org/html/2605.11161#bib.bib65)\)argues for the importance of curiosity\-driven research, noting that we cannot yet predict which interpretability techniques may yield actionable insights in the future\.

Actionable Explainability\.The field of actionable explainability originates primarily from human\-AI interaction \(HCI\) and algorithmic recourse research, focusing on enabling individuals to act on model outputs\. An explanation is considered “actionable” if it helps a person understand the changes needed to receive a different outcome in the future\(Joshiet al\.,[2019](https://arxiv.org/html/2605.11161#bib.bib50); Ustunet al\.,[2019](https://arxiv.org/html/2605.11161#bib.bib32); Karimiet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib24); Singhet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib49)\)\. For instance, increasing income to obtain a loan approval\. Other approaches\(Singhet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib41); Poyiadziet al\.,[2020](https://arxiv.org/html/2605.11161#bib.bib47)\)define actionability as the ability to translate explanations into feasible behavioral changes, while the idea was also extended to human\-in\-the\-loop settings\(Sarantiet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib48)\), allowing domain experts to directly adjust model parameters\. While actionable explainability centers on enabling human action, actionable interpretability reframes interpretability as a way to also drive concrete improvements in model performance and reliability, not only human understanding\.

## 9Conclusion

In this position paper, we argue that interpretability can have greater real\-world impact if actionability is incorporated as a core evaluation criterion\. This is not to say that conceptual or theoretical work without immediate practical utility has no place in the field; such research remains valuable and necessary\. Rather, making actionability a common evaluation criterion and explicitly tracking what insights make possible can accelerate progress for exploratory work\.

We conclude by offering an actionable checklist for interpretability researchers\.

1. 1\.Define a clear goal\.Identify a specific problem that your interpretability question aims to eventually solve\.
2. 2\.Identify your audience\.Although academic papers are primarily read by researchers, their insights may be acted upon by different stakeholders \(e\.g\., developers, practitioners, policymakers\), each of whom may require different framing, language, or levels of abstraction\.
3. 3\.Propose concrete actions\.Articulate what decisions or interventions your insights enable\.
4. 4\.Validate empirically\.Where possible, implement the proposed action yourself and demonstrate its effects\.
5. 5\.Evaluate in realistic settings\.Apply your methods to realistic scenarios, including large\-scale models and non\-synthetic datasets\.
6. 6\.Use actionable metrics, as descibed in[Section6](https://arxiv.org/html/2605.11161#S6)\. Especially, ask whether your contribution: - •Surpasses standard baselines \(e\.g\., prompting, fine\-tuning\) on target metrics\. - •Generalizes across models and other variations of the setting\. - •Produces targeted effects without degrading unrelated capabilities\. - •Yields explanations that are useful for the target audience—and if not, whether they can be translated into a more accessible form\.

The burden now falls on the research community: to reward actionable contributions alongside explanatory depth, to establish evaluation criteria that track the utility of interpretability insights, and to build infrastructure that connects understanding to impact\.

## Acknowledgments

This work has been made possible in part by a gift from the Chan Zuckerberg Initiative Foundation to establish the Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University\. A\.R\. was funded through the Azrieli international postdoctoral fellowship and the Ali Kaufman postdoctoral fellowship\. B\.W\. is supported by a grant from Coefficient Giving and the National Science Foundation \(NSF\), RI 2211954\. I\.L\. and N\.S\. are supported by a Technical AI Safety Research Grant from Coefficient Giving via Berkeley Existential Risk Initiative\. M\.M\. is supported by the Mila P2v5 grant and the Mila\-Samsung grant\. E\.W\. is supported by a grant from the National Science Foundation \(NSF\), CCF 2442421 and ARPA\-H program on Safe and Explainable AI under the award D24AC00253\-00\.

## References

- A\. Achara, E\. P\. Anton, A\. Hammers, and A\. P\. King \(2025\)Invisible attributes, visible biases: exploring demographic shortcuts in mri\-based alzheimer’s disease classification\.arXiv preprint arXiv:2509\.09558\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p2.1)\.
- C\. Agarwal, S\. H\. Tanneru, and H\. Lakkaraju \(2024\)Faithfulness vs\. plausibility: on the \(un\) reliability of explanations from large language models\.arXiv preprint arXiv:2402\.04614\.Cited by:[§C\.1](https://arxiv.org/html/2605.11161#A3.SS1.p1.1)\.
- R\. Agarwal, L\. Melnick, N\. Frosst, X\. Zhang, B\. Lengerich, R\. Caruana, and G\. Hinton \(2021\)Neural additive models: interpretable machine learning with neural nets\.InAdvances in Neural Information Processing Systems,A\. Beygelzimer, Y\. Dauphin, P\. Liang, and J\. W\. Vaughan \(Eds\.\),Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- C\. Agia, R\. Sinha, J\. Yang, R\. Antonova, M\. Pavone, H\. Nishimura, M\. Itkina, and J\. Bohg \(2025\)CUPID: curating data your robot loves with influence functions\.InProceedings of The 9th Conference on Robot LearningThe Eleventh International Conference on Learning RepresentationsThe Thirty\-eighth Annual Conference on Neural Information Processing SystemsAdvances in Neural Information Processing SystemsMachine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2020, Ghent, Belgium, September 14–18, 2020, Proceedings, Part IIThe Thirteenth International Conference on Learning RepresentationsProceedings of the 2024 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)First conference on language modelingFindings of the Association for Computational Linguistics: EMNLP 2023Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)Proceedings of the 2023 Conference on Empirical Methods in Natural Language ProcessingProceedings of the 2023 ACM Designing Interactive Systems Conference,J\. Lim, S\. Song, H\. Park, Y\. Al\-Onaizan, M\. Bansal, Y\. Chen, L\. Chiruzzo, A\. Ritter, L\. Wang, H\. Bouamor, J\. Pino, K\. Bali, W\. Che, J\. Nabende, E\. Shutova, M\. T\. Pilehvar, H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Proceedings of Machine Learning Research, Vol\.30533,pp\. 2907–2932\.External Links:[Link](https://proceedings.mlr.press/v305/agia25a.html)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p2.1)\.
- H\. Ahsan, A\. S\. Sharma, S\. Amir, D\. Bau, and B\. C\. Wallace \(2025\)Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare\.InProceedings of the Findings of Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p2.1)\.
- E\. Akyürek, D\. Schuurmans, J\. Andreas, T\. Ma, and D\. Zhou \(2023\)What learning algorithm is in\-context learning? investigations with linear models\.External Links:[Link](https://openreview.net/forum?id=0g0X4H8yN4I)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p3.1)\.
- D\. Alvarez\-Melis and T\. S\. Jaakkola \(2018\)On the robustness of interpretability methods\.arXiv preprint arXiv:1806\.08049\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- D\. Amodei \(2025\)The urgency of interpretability\.\(en\)\.Note:BlogpostExternal Links:[Link](https://www.darioamodei.com/post/the-urgency-of-interpretability)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- Anthropic \(2025\)System card: claude sonnet 4\.5\.Note:https://assets\.anthropic\.com/m/12f214efcc2f457a/original/Claude\-Sonnet\-4\-5\-System\-Card\.pdfCited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- U\. Anwar, A\. Saparov, J\. Rando, D\. Paleka, M\. Turpin, P\. Hase, E\. S\. Lubana, E\. Jenner, S\. Casper, O\. Sourbut,et al\.\(2024\)Foundational challenges in assuring alignment and safety of large language models\.Transactions on Machine Learning Research\.Note:arXiv:2404\.09932 \[cs\]External Links:[Link](http://arxiv.org/abs/2404.09932),[Document](https://dx.doi.org/10.48550/arXiv.2404.09932)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- D\. Arad, H\. Orgad, and Y\. Belinkov \(2024\)ReFACT: updating text\-to\-image models by editing the text encoder\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 2537–2558\.External Links:[Link](https://aclanthology.org/2024.naacl-long.140/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.140)Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- ArtificialIntelligenceAct\.eu \(2025\)Article 86: right to explanation of individual decision\-making\.ArtificialIntelligenceAct\.eu\.Note:Accessed: 2025\-12\-28OnlineExternal Links:[Link](https://artificialintelligenceact.eu/article/86/)Cited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p2.1)\.
- N\. Arun, N\. Gaw, P\. Singh, K\. Chang, M\. Aggarwal, B\. Chen, K\. Hoebel, S\. Gupta, J\. Patel, M\. Gidwani,et al\.\(2021\)Assessing the trustworthiness of saliency maps for localizing abnormalities in medical imaging\.Radiology: Artificial Intelligence3\(6\),pp\. e200267\.Cited by:[§6\.2](https://arxiv.org/html/2605.11161#S6.SS2.p3.1)\.
- N\. Ashtari, R\. Mullins, C\. Qian, J\. Wexler, I\. Tenney, and M\. Pushkarna \(2023\)From discovery to adoption: understanding the ml practitioners’ interpretability journey\.pp\. 2304–2325\.Cited by:[§3\.3](https://arxiv.org/html/2605.11161#S3.SS3.p2.1)\.
- T\. Ashuach, D\. Arad, A\. Mueller, M\. Tutek, and Y\. Belinkov \(2025a\)CRISP: persistent concept unlearning via sparse autoencoders\.arXiv preprint arXiv:2508\.13650\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- T\. Ashuach, M\. Tutek, and Y\. Belinkov \(2025b\)REVS: unlearning sensitive information in language models via rank editing in the vocabulary space\.External Links:2406\.09325,[Link](https://arxiv.org/abs/2406.09325)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- A\. Azaria and T\. Mitchell \(2023\)The internal state of an LLM knows when it’s lying\.Singapore,pp\. 967–976\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.68/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.68)Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1)\.
- F\. Barez \(2026\)Automated interpretability\-driven model auditing and control: a research agenda\.AI Governance Initiative, University of Oxford\.Note:Working paper\. Research agenda dated January 8, 2026External Links:[Link](https://aigi.ox.ac.uk/publications/automated-interpretability-driven-model-auditing-and-control-a-research-agenda/)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1)\.
- S\. Bassan and G\. Katz \(2023\)Towards formal xai: formally approximate minimal explanations of neural networks\.InInternational Conference on Tools and Algorithms for the Construction and Analysis of Systems,pp\. 187–207\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- D\. Bau \(2025\)In defense of curiosity\.Note:[https://davidbau\.com/archives/2025/12/09/in\_defense\_of\_curiosity\.html](https://davidbau.com/archives/2025/12/09/in_defense_of_curiosity.html)Accessed: 2025\-12\-09Cited by:[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p2.1),[§7](https://arxiv.org/html/2605.11161#S7.p1.1),[§8](https://arxiv.org/html/2605.11161#S8.p3.1)\.
- L\. Bereska and S\. Gavves \(2024\)Mechanistic interpretability for AI safety \- a review\.Transactions on Machine Learning Research\.Note:Survey Certification, Expert CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=ePUVetPKu6)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p1.1)\.
- G\. Blanc, J\. Lange, and L\. Tan \(2021\)Provably efficient, succinct, and precise explanations\.Advances in Neural Information Processing Systems34,pp\. 6129–6141\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- L\. Bourtoule, V\. Chandrasekaran, C\. A\. Choquette\-Choo, H\. Jia, A\. Travers, B\. Zhang, D\. Lie, and N\. Papernot \(2021\)Machine unlearning\.InProceedings of the 42nd IEEE Symposium on Security and Privacy \(SP\),pp\. 141–159\.External Links:[Document](https://dx.doi.org/10.1109/SP40001.2021.00019),[Link](https://doi.org/10.1109/SP40001.2021.00019)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- W\. Brendel and M\. Bethge \(2019\)Approximating cnns with bag\-of\-local\-features models works surprisingly well on imagenet\.International Conference on Learning Representations\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- N\. Calderon and R\. Reichart \(2025\)On behalf of the stakeholders: trends in NLP model interpretability in the era of LLMs\.Albuquerque, New Mexico,pp\. 656–693\.External Links:[Link](https://aclanthology.org/2025.naacl-long.29/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.29),ISBN 979\-8\-89176\-189\-6Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p2.1)\.
- Y\. Cao and J\. Yang \(2015\)Towards making systems forget with machine unlearning\.InProceedings of the IEEE Symposium on Security and Privacy,pp\. 463–480\.External Links:[Document](https://dx.doi.org/10.1109/SP.2015.35),[Link](https://doi.org/10.1109/SP.2015.35)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- S\. Casper, L\. Schulze, O\. Patel, and D\. Hadfield\-Menell \(2024a\)Defending against unforeseen failure modes with latent adversarial training\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=mVPPhQ8cAd)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p4.1)\.
- S\. Casper, J\. Yun, J\. Baek, Y\. Jung, M\. Kim, K\. Kwon, S\. Park, H\. Moore, D\. Shriver, M\. Connor,et al\.\(2024b\)The satml’24 cnn interpretability competition: new innovations for concept\-level interpretability\.arXiv preprint arXiv:2404\.02949\.Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px2.p1.1)\.
- S\. Casper \(2023\)Broad critiques of interpretability research\.Note:The Engineer’s Interpretability Sequence \(blogpost 3 of 12\)External Links:[Link](https://www.lesswrong.com/posts/gwG9uqw255gafjYN4/eis-iii-broad-critiques-of-interpretability-research)Cited by:[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p4.1)\.
- T\. Chakraborti, C\. R\. Banerji, A\. Marandon, V\. Hellon, R\. Mitra, B\. Lehmann, L\. Bräuninger, S\. McGough, C\. Turkay, A\. F\. Frangi,et al\.\(2025\)Personalized uncertainty quantification in artificial intelligence\.Nature Machine Intelligence7\(4\),pp\. 522–530\.Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1)\.
- L\. Chen, M\. Zaharia, and J\. Zou \(2024a\)FrugalGPT: how to use large language models while reducing cost and improving performance\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=cSimKw5p6R)Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p4.1)\.
- R\. Chen, A\. Arditi, H\. Sleight, O\. Evans, and J\. Lindsey \(2025\)Persona vectors: monitoring and controlling character traits in language models\.arXiv preprint arXiv:2507\.21509\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p6.1)\.
- Y\. Chen, A\. Wu, T\. DePodesta, C\. Yeh, K\. Li, N\. C\. Marin, O\. Patel, J\. Riecke, S\. Raval, O\. Seow, M\. Wattenberg, and F\. Viégas \(2024b\)Designing a dashboard for transparency and control of conversational ai\.External Links:2406\.07882,[Link](https://arxiv.org/abs/2406.07882)Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p2.1)\.
- A\. Conmy, A\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-Alonso \(2023\)Towards automated circuit discovery for mechanistic interpretability\.Advances in Neural Information Processing Systems36,pp\. 16318–16352\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p5.1)\.
- J\. Crabbé and M\. van der Schaar \(2023\)Evaluating the robustness of interpretability methods through explanation invariance and equivariance\.Advances in Neural Information Processing Systems36,pp\. 71393–71429\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- F\. Doshi\-Velez and B\. Kim \(2017\)Towards a rigorous science of interpretable machine learning\.arXiv preprint arXiv:1702\.08608\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p2.1)\.
- Y\. Elazar, S\. Ravfogel, A\. Jacovi, and Y\. Goldberg \(2021\)Amnesic probing: behavioral explanation with amnesic counterfactuals\.InTransactions of the Association for Computational Linguistics \(TACL\),Vol\.9,pp\. 147–163\.External Links:[Link](https://aclanthology.org/2021.tacl-1.13/)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah \(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.Note:https://transformer\-circuits\.pub/2021/framework/index\.htmlCited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p4.1)\.
- European Union \(2016\)General data protection regulation \(gdpr\), article 22: automated individual decision\-making, including profiling\.Note:Regulation \(EU\) 2016/679Accessed: 2025\-12\-28External Links:[Link](https://gdpr.eu/article-22-automated-individual-decision-making-including-profiling/)Cited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p2.1)\.
- J\. Fang, H\. Jiang, K\. Wang, Y\. Ma, J\. Shi, X\. Wang, X\. He, and T\. Chua \(2025\)AlphaEdit: null\-space constrained model editing for language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HvSytvg3Jh)Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- P\. Ferreira, W\. Aziz, and I\. Titov \(2025\)Truthful or fabricated? using causal attribution to mitigate reward hacking in explanations\.arXiv preprint arXiv:2504\.05294\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p6.1)\.
- J\. F\. Fiotto\-Kaufman, A\. R\. Loftus, E\. Todd, J\. Brinkmann, K\. Pal, D\. Troitskii, M\. Ripa, A\. Belfki, C\. Rager, C\. Juang, A\. Mueller, S\. Marks, A\. S\. Sharma, F\. Lucchetti, N\. Prakash, C\. E\. Brodley, A\. Guha, J\. Bell, B\. C\. Wallace, and D\. Bau \(2025\)NNsight and NDIF: democratizing access to open\-weight foundation model internals\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,External Links:[Link](https://openreview.net/forum?id=MxbEiFRf39)Cited by:[§3\.3](https://arxiv.org/html/2605.11161#S3.SS3.p2.1)\.
- R\. Gandikota, H\. Orgad, Y\. Belinkov, J\. Materzyńska, and D\. Bau \(2024\)Unified concept editing in diffusion models\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,Note:arXiv:2308\.14761Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1),[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- L\. Gao \(2025\)An ambitious vision for interpretability\.Note:Alignment ForumExternal Links:[Link](https://www.alignmentforum.org/posts/Hy6PX43HGgmfiTaKu/an-ambitious-vision-for-interpretability)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p3.1)\.
- M\. Geva, R\. Schuster, J\. Berant, and O\. Levy \(2021\)Transformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5484–5495\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p5.1)\.
- M\. Ghassemi, L\. Oakden\-Rayner, and A\. L\. Beam \(2021\)The false hope of current approaches to explainable artificial intelligence in health care\.The lancet digital health3\(11\),pp\. e745–e750\.Cited by:[§6\.2](https://arxiv.org/html/2605.11161#S6.SS2.p3.1)\.
- L\. Gorton, N\. Wang, N\. Nguyen, M\. Deng, E\. Ho, D\. Balsam, and T\. McGrath \(2025\)Interpreting evo 2: arc institute’s next\-generation genomic foundation model\.Goodfire Research\.External Links:[Link](https://www.goodfire.ai/research/interpreting-evo-2)Cited by:[§B\.3](https://arxiv.org/html/2605.11161#A2.SS3.SSS0.Px1.p1.1)\.
- D\. Gottesman and M\. Geva \(2024\)Estimating knowledge in large language models without generating a single token\.Miami, Florida, USA,pp\. 3994–4019\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.232/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.232)Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1)\.
- R\. Greenblatt, N\. Nanda, Buck, and habryka \(2023\)How useful is mechanistic interpretability?\.Note:LessWrong BlogpostExternal Links:[Link](https://www.lesswrong.com/posts/tEPHGZAb63dfq2v8n/how-useful-is-mechanistic-interpretability)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1),[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- A\. Gu and T\. Dao \(2024\)Mamba: linear\-time sequence modeling with selective state spaces\.Cited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p4.1)\.
- Y\. Gur\-Arieh, C\. H\. Suslik, Y\. Hong, F\. Barez, and M\. Geva \(2025\)Precise in\-parameter concept erasure in large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 18997–19017\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- T\. Haklay, H\. Orgad, D\. Bau, A\. Mueller, and Y\. Belinkov \(2025a\)Position\-aware automatic circuit discovery\.Vienna, Austria,pp\. 2792–2817\.External Links:[Link](https://aclanthology.org/2025.acl-long.141/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.141),ISBN 979\-8\-89176\-251\-0Cited by:[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p3.1)\.
- T\. Haklay, H\. Orgad, A\. Reusch, M\. Mosbach, S\. Wiegreffe, I\. Tenney, and M\. Geva \(2025b\)1st Actionable Interpretability Workshop at ICML\.\(en\)\.External Links:[Link](https://icml.cc/virtual/2025/workshop/39962)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1),[§7](https://arxiv.org/html/2605.11161#S7.p3.1)\.
- X\. Han, B\. C\. Wallace, and Y\. Tsvetkov \(2020\)Explaining black box predictions and unveiling data artifacts through influence functions\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5553–5563\.External Links:[Link](https://aclanthology.org/2020.acl-main.492/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.492)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p2.1)\.
- S\. Havaldar, H\. Jin, C\. Kim, A\. Xue, W\. You, M\. Gatti, B\. Jain, H\. Qu, D\. A\. Hashimoto, A\. Madani,et al\.\(2025\)T\-fix: text\-based explanations with features interpretable to experts\.arXiv preprint arXiv:2511\.04070\.Cited by:[§C\.1](https://arxiv.org/html/2605.11161#A3.SS1.p1.1)\.
- D\. Hendrycks and L\. Hiscott \(2025\)The misguided quest for mechanistic ai interpretability\.AI Frontiers\.External Links:[Link](https://ai-frontiers.org/articles/the-misguided-quest-for-mechanistic-ai-interpretability)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- E\. Ho \(2025\)On optimism for interpretability\.\(en\)\.Note:Goodfire AI BlogExternal Links:[Link](https://www.goodfire.ai/blog/on-optimism-for-interpretability)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p3.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)LoRA: low\-rank adaptation of large language models\.\.ICLR1\(2\),pp\. 3\.Cited by:[§6\.1](https://arxiv.org/html/2605.11161#S6.SS1.p1.1)\.
- J\. Huang, J\. Tao, T\. Icard, D\. Yang, and C\. Potts \(2025\)Internal causal mechanisms robustly predict language model out\-of\-distribution behaviors\.External Links:2505\.11770,[Link](https://arxiv.org/abs/2505.11770)Cited by:[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p4.1)\.
- A\. Jacovi and Y\. Goldberg \(2021\)Aligning faithful interpretations with their social attribution\.Transactions of the Association for Computational Linguistics9,pp\. 294–310\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p1.1)\.
- S\. Jain, S\. Wiegreffe, Y\. Pinter, and B\. C\. Wallace \(2020\)Learning to faithfully rationalize by construction\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- H\. Jin, S\. Havaldar, C\. Kim, A\. Xue, W\. You, H\. Qu, M\. Gatti, D\. A\. Hashimoto, B\. Jain, A\. Madani,et al\.\(2024\)The fix benchmark: extracting features interpretable to experts\.arXiv preprint arXiv:2409\.13684\.Cited by:[§C\.1](https://arxiv.org/html/2605.11161#A3.SS1.p1.1)\.
- H\. Jin, A\. Xue, W\. You, S\. Goel, and E\. Wong \(2025\)Probabilistic stability guarantees for feature attributions\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p2.1)\.
- S\. Joshi, O\. Koyejo, W\. Vijitbenjaronk, B\. Kim, and J\. Ghosh \(2019\)Towards realistic individual recourse and actionable explanations in black\-box decision making systems\.arXiv preprint arXiv:1907\.09615\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell, T\. Henighan, D\. Drain, E\. Perez, N\. Schiefer, Z\. Hatfield\-Dodds, N\. DasSarma, E\. Tran\-Johnson,et al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p3.1),[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p4.1)\.
- A\. Karimi, B\. Schölkopf, and I\. Valera \(2021\)Algorithmic recourse: from counterfactual explanations to interventions\.InProceedings of the 2021 ACM conference on fairness, accountability, and transparency,pp\. 353–362\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- C\. Kim, W\. You, S\. Havaldar, and E\. Wong \(2024\)Evaluating groups of features via consistency, contiguity, and stability\.InThe Second Tiny Papers Track at ICLR 2024,Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p2.1)\.
- P\. Kindermans, S\. Hooker, J\. Adebayo, M\. Alber, K\. T\. Schütt, S\. Dähne, D\. Erhan, and B\. Kim \(2019\)The \(un\) reliability of saliency methods\.InExplainable AI: Interpreting, explaining and visualizing deep learning,pp\. 267–280\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- F\. King, C\. Pettersen, D\. Posselt, S\. Ringerud, and Y\. Xie \(2025\)Leveraging sparse autoencoders to reveal interpretable features in geophysical models\.Journal of Geophysical Research: Machine Learning and Computation2\(4\),pp\. e2025JH000769\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p6.1)\.
- P\. W\. Koh and P\. Liang \(2017\)Understanding black\-box predictions via influence functions\.InProceedings of the 34th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.70,pp\. 1885–1894\.External Links:[Link](https://proceedings.mlr.press/v70/koh17a.html)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p2.1)\.
- P\. W\. Koh, T\. Nguyen, Y\. S\. Tang, S\. Mussmann, E\. Pierson, B\. Kim, and P\. Liang \(2020a\)Concept bottleneck models\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.119,pp\. 5338–5348\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- P\. W\. Koh, T\. Nguyen, Y\. S\. Tang, S\. Mussmann, E\. Pierson, B\. Kim, and P\. Liang \(2020b\)Concept bottleneck models\.InProceedings of the 37th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.119,pp\. 5338–5348\.External Links:[Link](https://arxiv.org/abs/2007.04612)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- M\. Krishnan \(2020\)Against interpretability: a critical examination of the interpretability problem in machine learning\.Philosophy & Technology33\(3\),pp\. 487–502\.Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1)\.
- S\. Lai, L\. Hu, J\. Wang, L\. Berti\-Equille, and D\. Wang \(2024\)Faithful vision\-language interpretation via concept bottleneck models\.InThe Twelfth International Conference on Learning Representations,Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. Wattenberg \(2023\)Inference\-time intervention: eliciting truthful answers from a language model\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=aLLuYpn83y)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- V\. R\. Li, J\. Kaufmann, M\. Wattenberg, D\. Alvarez\-Melis, and N\. Saphra \(2025\)Can interpretation predict behavior on unseen data?\.External Links:2507\.06445,[Link](https://arxiv.org/abs/2507.06445)Cited by:[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p4.1)\.
- \[77\]Z\. Li, J\. Ji, and Y\. ZhangFrom kepler to newton: explainable ai for science discovery\.InICML 2022 2nd AI for Science Workshop,Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p6.1)\.
- Z\. C\. Lipton \(2018\)The mythos of model interpretability: in machine learning, the concept of interpretability is both important and slippery\.Queue16\(3\),pp\. 31–57\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p1.1)\.
- K\. Liu, S\. Casper, D\. Hadfield\-Menell, and J\. Andreas \(2023\)Cognitive dissonance: why do language model outputs disagree with internal representations of truthfulness?\.Singapore,pp\. 4791–4797\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.291/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.291)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px3.p1.1)\.
- Q\. Lyu, S\. Havaldar, A\. Stein, L\. Zhang, D\. Rao, E\. Wong, M\. Apidianaki, and C\. Callison\-Burch \(2023\)Faithful chain\-of\-thought reasoning\.InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),J\. C\. Park, Y\. Arase, B\. Hu, W\. Lu, D\. Wijaya, A\. Purwarianti, and A\. A\. Krisnadhi \(Eds\.\),Nusa Dua, Bali,pp\. 305–329\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.ijcnlp-main.20)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- C\. Ma, J\. Donnelly, W\. Liu, S\. Vosoughi, C\. Rudin, and C\. Chen \(2024\)Interpretable image classification with adaptive prototype\-based vision transformers\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- I\. Magnusson, N\. Tai, B\. Bogin, D\. Heineman, J\. D\. Hwang, L\. Soldaini, A\. Bhagia, J\. Liu, D\. Groeneveld, O\. Tafjord, N\. A\. Smith, P\. W\. Koh, and J\. Dodge \(2025\)DataDecide: how to predict best pretraining data with small experiments\.External Links:2504\.11393,[Link](https://arxiv.org/abs/2504.11393)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px1.p1.1)\.
- P\. Mangla, V\. Singh, and V\. N\. Balasubramanian \(2020\)On saliency maps and adversarial robustness\.Berlin, Heidelberg,pp\. 272–288\.External Links:ISBN 978\-3\-030\-67660\-5,[Link](https://doi.org/10.1007/978-3-030-67661-2_17),[Document](https://dx.doi.org/10.1007/978-3-030-67661-2%5F17)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px1.p1.1)\.
- S\. Marks, P\. Hase, and Alignment Science Team \(2025\)Recommendations for technical ai safety research directions\.Note:Anthropic Alignment Science BlogExternal Links:[Link](https://alignment.anthropic.com/2025/recommended-directions/)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- S\. Marks \(2025\)Principles for picking practical interpretability projects\.Note:Alignment ForumExternal Links:[Link](https://www.lesswrong.com/posts/DqaoPNqhQhwBFqWue/principles-for-picking-practical-interpretability-projects)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1),[§7](https://arxiv.org/html/2605.11161#S7.p3.1)\.
- N\. Maslej, L\. Fattorini, R\. Perrault, Y\. Gil, V\. Parli, N\. Kariuki, E\. Capstick, A\. Reuel, E\. Brynjolfsson, J\. Etchemendy,et al\.\(2025\)Artificial intelligence index report 2025\.InArtificial Intelligence Index Report 2025,Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p1.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1),[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- K\. Meng, A\. S\. Sharma, A\. J\. Andonian, Y\. Belinkov, and D\. Bau \(2023\)Mass\-editing memory in a transformer\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MkbcAHIYgyS)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- T\. Miller \(2019\)Explanation in artificial intelligence: insights from the social sciences\.Artificial intelligence267,pp\. 1–38\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p1.1)\.
- M\. Mosbach, V\. Gautam, T\. Vergara Browne, D\. Klakow, and M\. Geva \(2024\)From insights to actions: the impact of interpretability and analysis research on NLP\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3078–3105\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.181/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.181)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p1.1),[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p2.1)\.
- A\. Mueller, A\. Geiger, S\. Wiegreffe, D\. Arad, I\. Arcuschin, A\. Belfki, Y\. S\. Chan, J\. F\. Fiotto\-Kaufman, T\. Haklay, M\. Hanna, J\. Huang, R\. Gupta, Y\. Nikankin, H\. Orgad, N\. Prakash, A\. Reusch, A\. Sankaranarayanan, S\. Shao, A\. Stolfo, M\. Tutek, A\. Zur, D\. Bau, and Y\. Belinkov \(2025\)MIB: A mechanistic interpretability benchmark\.InForty\-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13\-19, 2025,External Links:[Link](https://openreview.net/forum?id=sSrOwve6vb)Cited by:[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p4.1)\.
- N\. Nanda and J\. Bloom \(2022\)TransformerLens\.Note:[https://github\.com/TransformerLensOrg/TransformerLens](https://github.com/TransformerLensOrg/TransformerLens)Cited by:[§3\.3](https://arxiv.org/html/2605.11161#S3.SS3.p2.1)\.
- N\. Nanda, J\. Engels, A\. Conmy, S\. Rajamanoharan, bilalchughtai, C\. McDougall, J\. Kramár, and L\. Smith \(2025\)A pragmatic vision for interpretability\.Note:[https://www\.alignmentforum\.org/posts/StENzDcD3kpfGJssR/a\-pragmatic\-vision\-for\-interpretability](https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability)Accessed: 2025\-12\-02Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1),[§7](https://arxiv.org/html/2605.11161#S7.p2.1),[§8](https://arxiv.org/html/2605.11161#S8.p3.1)\.
- N\. Nanda \(2022\)A longlist of theories of impact for interpretability\.Note:LessWrong Blogpost\. Accessed: 2025\-11\-24\.External Links:[Link](https://www.alignmentforum.org/posts/uK6sQCNMw8WKzJeCQ/a-longlist-of-theories-of-impact-for-interpretability)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- B\. Newman, A\. Ravichander, J\. Jung, R\. Xin, H\. Ivison, Y\. Kuznetsov, P\. W\. Koh, and Y\. Choi \(2025\)The curious case of factuality finetuning: models’ internal beliefs can improve factuality\.External Links:2507\.08371,[Link](https://arxiv.org/abs/2507.08371)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px3.p1.1)\.
- O\. Obeso, A\. Arditi, J\. Ferrando, J\. Freeman, C\. Holmes, and N\. Nanda \(2025\)Real\-time detection of hallucinated entities in long\-form generation\.External Links:2509\.03531,[Link](https://arxiv.org/abs/2509.03531)Cited by:[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1)\.
- T\. Oikarinen, S\. Das, L\. M\. Nguyen, and T\. Weng \(2023\)Label\-free concept bottleneck models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=FlCg47MNvBA)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- C\. Olah \(2023\)Interpretability dreams\.Note:Transformer Circuits Thread BlogpostExternal Links:[Link](https://transformer-circuits.pub/2023/interpretability-dreams/index.html)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. DasSarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen,et al\.\(2022\)In\-context learning and induction heads\.arXiv preprint arXiv:2209\.11895\.Cited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p4.1)\.
- H\. Orgad, B\. Kawar, and Y\. Belinkov \(2023\)Editing implicit assumptions in text\-to\-image diffusion models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1),[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- H\. Orgad, M\. Toker, Z\. Gekhman, R\. Reichart, I\. Szpektor, H\. Kotek, and Y\. Belinkov \(2025\)LLMs know more than they show: on the intrinsic representation of LLM hallucinations\.External Links:[Link](https://openreview.net/forum?id=KRnsX5Em3W)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px3.p1.1),[§B\.2](https://arxiv.org/html/2605.11161#A2.SS2.SSS0.Px1.p1.1)\.
- A\. Peysakhovich and A\. Lerer \(2023\)Attention sorting combats recency bias in long context language models\.arXiv preprint arXiv:2310\.01427\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px2.p1.1)\.
- C\. Potts \(2025\)Assessing skeptical views of interpretability research\.Note:BlogpostExternal Links:[Link](https://web.stanford.edu/~cgpotts/blog/interp/)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1)\.
- R\. Poyiadzi, K\. Sokol, R\. Santos\-Rodriguez, T\. De Bie, and P\. Flach \(2020\)FACE: feasible and actionable counterfactual explanations\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,pp\. 344–350\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- G\. A\. Prenosil, T\. K\. Weitzel, S\. C\. Bello, C\. Mingels, G\. Manzini, L\. P\. Meier, K\. Shi, A\. Rominger, and A\. Afshar\-Oromieh \(2025\)Neuro\-symbolic ai for auditable cognitive information extraction from medical reports\.Communications Medicine5\(1\),pp\. 491\.Cited by:[§5\.2](https://arxiv.org/html/2605.11161#S5.SS2.p2.1)\.
- G\. Pruthi, F\. Liu, S\. Kale, and M\. Sundararajan \(2020\)Estimating training data influence by tracing gradient descent\.pp\. 19920–19930\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px1.p1.1)\.
- S\. Ravfogel, Y\. Elazar, H\. Gonen, M\. Twiton, and Y\. Goldberg \(2020\)Null it out: guarding protected attributes by iterative nullspace projection\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 7237–7256\.External Links:[Link](https://aclanthology.org/2020.acl-main.647/)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p6.1)\.
- C\. Rudin \(2019\)Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead\.Nature machine intelligence1\(5\),pp\. 206–215\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p1.1)\.
- A\. Saranti, M\. Hudec, E\. Mináriková, Z\. Takáč, U\. Großschedl, C\. Koch, B\. Pfeifer, A\. Angerschmid, and A\. Holzinger \(2022\)Actionable explainable ai \(axai\): a practical example with aggregation functions for adaptive classification and textual explanations for interpretable machine learning\.Machine Learning and Knowledge Extraction4\(4\),pp\. 924–953\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- L\. Schut, N\. Tomašev, T\. McGrath, D\. Hassabis, U\. Paquet, and B\. Kim \(2025\)Bridging the human–ai knowledge gap through concept discovery and transfer in alphazero\.Proceedings of the National Academy of Sciences122\(13\),pp\. e2406675122\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2406675122),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2406675122),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2406675122Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1),[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p3.1)\.
- C\. Segerie \(2023\)Against almost every theory of impact of interpretability\.Note:LessWrong Blogpost\. Accessed: 2024\-08\-17\.External Links:[Link](https://www.lesswrong.com/posts/LNA8mubrByG7SFacm/against-almost-every-theory-of-impact-of-interpretability-1)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- R\. Shah, A\. Irpan, A\. M\. Turner, A\. Wang, A\. Conmy, D\. Lindner, J\. Brown\-Cohen, L\. Ho, N\. Nanda, R\. A\. Popa,et al\.\(2025\)An approach to technical agi safety and security\.External Links:2504\.01849,[Link](https://arxiv.org/abs/2504.01849)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- J\. Shovelain and E\. Mckernon \(2023\)The risk\-reward tradeoff of interpretability research\.Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/HdqdqNC3MyABHzSqf/the-risk-reward-tradeoff-of-interpretability-research)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- D\. Silver, T\. Hubert, J\. Schrittwieser, I\. Antonoglou, M\. Lai, A\. Guez, M\. Lanctot, L\. Sifre, D\. Kumaran, T\. Graepel,et al\.\(2018\)A general reinforcement learning algorithm that masters chess, shogi, and go through self\-play\.Science362\(6419\),pp\. 1140–1144\.Cited by:[§5\.3](https://arxiv.org/html/2605.11161#S5.SS3.p3.1)\.
- R\. Singh, T\. Miller, H\. Lyons, L\. Sonenberg, E\. Velloso, F\. Vetere, P\. Howe, and P\. Dourish \(2023\)Directive explanations for actionable explainability in machine learning applications\.ACM Trans\. Interact\. Intell\. Syst\.13\(4\)\.External Links:ISSN 2160\-6455,[Link](https://doi.org/10.1145/3579363),[Document](https://dx.doi.org/10.1145/3579363)Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- R\. Singh, T\. Miller, L\. Sonenberg, E\. Velloso, F\. Vetere, P\. Howe, and P\. Dourish \(2024\)An actionability assessment tool for explainable ai\.arXiv preprint arXiv:2407\.09516\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- So8res \(2023\)If interpretability research goes well, it may get dangerous\.Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/BinkknLBYxskMXuME/if-interpretability-research-goes-well-it-may-get-dangerous)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p2.1)\.
- L\. Spillner, R\. Ringe, R\. Porzel, and R\. Malaka \(2025\)Can ai explanations make you change your mind?\.arXiv preprint arXiv:2508\.08158\.Cited by:[§6\.2](https://arxiv.org/html/2605.11161#S6.SS2.p1.1)\.
- J\. Steinhardt and S\. Schwettmann \(2024\)Introducing transluce\.Note:Transluce AI BlogExternal Links:[Link](https://transluce.org/introducing-transluce)Cited by:[§7](https://arxiv.org/html/2605.11161#S7.p3.1)\.
- A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid \(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.External Links:[Link](https://arxiv.org/abs/2308.10248)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- S\. Upadhyay and F\. Barez \(2025\)Martian interpretability challenge, part 2: the core problems in interpretability\.Note:Martian BlogAccessed: 2025\-12\-08External Links:[Link](https://withmartian.com/post/interpretability-prize-part2)Cited by:[§1](https://arxiv.org/html/2605.11161#S1.p2.1),[§6\.3](https://arxiv.org/html/2605.11161#S6.SS3.p1.1)\.
- B\. Ustun, A\. Spangher, and Y\. Liu \(2019\)Actionable recourse in linear classification\.InProceedings of the Conference on Fairness, Accountability, and Transparency \(FAT\*\),pp\. 10–19\.Cited by:[§8](https://arxiv.org/html/2605.11161#S8.p4.1)\.
- J\. Vincent, R\. Moreno, J\. Takala, S\. Willatts, A\. De Mendonça, H\. Bruining, C\. K\. Reinhart, P\. Suter, and L\. G\. Thijs \(1996\)The sofa \(sepsis\-related organ failure assessment\) score to describe organ dysfunction/failure: on behalf of the working group on sepsis\-related problems of the european society of intensive care medicine \(see contributors to the project in the appendix\)\.Intensive care medicine22\(7\),pp\. 707–710\.Cited by:[§C\.1](https://arxiv.org/html/2605.11161#A3.SS1.p1.1)\.
- K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt \(2023\)Interpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NpsVSN6o4ul)Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p5.1),[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- X\. Wen, W\. Tan, and R\. O\. Weber \(2024\)GAProtoNet: a multi\-head graph attention\-based prototypical network for interpretable text classification\.External Links:2409\.13312Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- Z\. Wu, A\. Arora, A\. Geiger, Z\. Wang, J\. Huang, D\. Jurafsky, C\. D\. Manning, and C\. Potts \(2025\)AxBench: steering llms? even simple baselines outperform sparse autoencoders\.\(en\)\.External Links:[Link](https://openreview.net/forum?id=K2CckZjNy0)Cited by:[§3\.2](https://arxiv.org/html/2605.11161#S3.SS2.p4.1)\.
- Z\. Wu, A\. Arora, Z\. Wang, A\. Geiger, D\. Jurafsky, C\. D\. Manning, and C\. Potts \(2024\)Reft: representation finetuning for language models\.Advances in Neural Information Processing Systems37,pp\. 63908–63962\.Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p7.1)\.
- A\. Xue, R\. Alur, and E\. Wong \(2023\)Stability guarantees for feature attributions with multiplicative smoothing\.Advances in Neural Information Processing Systems36,pp\. 62388–62413\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p2.1)\.
- Y\. Yang, M\. Gandhi, Y\. Wang, Y\. Wu, M\. S\. Yao, C\. Callison\-Burch, J\. C\. Gee, and M\. Yatskar \(2024\)A textbook remedy for domain shifts: knowledge priors for medical image analysis\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 90683–90713\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- Y\. Yang, A\. Panagopoulou, S\. Zhou, D\. Jin, C\. Callison\-Burch, and M\. Yatskar \(2023\)Language in a bottle: language model guided concept bottlenecks for interpretable image classification\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 19187–19197\.Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- C\. Yeh, C\. Hsieh, A\. Suggala, D\. I\. Inouye, and P\. K\. Ravikumar \(2019\)On the \(in\) fidelity and sensitivity of explanations\.Advances in neural information processing systems32\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p1.1)\.
- W\. You, H\. Qu, M\. Gatti, B\. Jain, and E\. Wong \(2025a\)Sum\-of\-parts: self\-attributing neural networks with end\-to\-end learning of feature groups\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=r6y9TEdLMh)Cited by:[§B\.1](https://arxiv.org/html/2605.11161#A2.SS1.SSS0.Px4.p1.1)\.
- W\. You, A\. Xue, S\. Havaldar, D\. Rao, H\. Jin, C\. Callison\-Burch, and E\. Wong \(2025b\)Probabilistic soundness guarantees in llm reasoning chains\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 7517–7536\.Cited by:[§C\.2](https://arxiv.org/html/2605.11161#A3.SS2.p2.1)\.
- M\. Yuksekgonul, M\. Wang, and J\. Zou \(2023\)Post\-hoc concept bottleneck models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=nA5AZ8CEyow)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p5.1)\.
- X\. Zhao, F\. Yin, and G\. Durrett \(2025\)Understanding synthetic context extension via retrieval heads\.InForty\-second International Conference on Machine Learning,Cited by:[§2\.1](https://arxiv.org/html/2605.11161#S2.SS1.p6.1)\.
- Z\. Zhou, X\. Lin, X\. Xu, A\. Prakash, D\. Rus, and B\. K\. H\. Low \(2024\)DETAIL: task DEmonstration attribution for interpretable in\-context learning\.External Links:[Link](https://openreview.net/forum?id=4jRNkAH15k)Cited by:[§5\.1](https://arxiv.org/html/2605.11161#S5.SS1.p3.1)\.

## Appendix AVisualizing the space of Actionable Interpetability work

[Figure3](https://arxiv.org/html/2605.11161#A1.F3)demonstrates the different types of actionable interpretability work as spanned by the dimensions of actionability\.

CONCRETENESSVALIDATIONLOWHIGHLOWHIGHHigh concreteness, high validationPrecise specifications with validated utility and robust evidence\.Low concreteness, low validationDirectional insights that motivate future work but lack steps\.High concreteness, low validationConcrete actions or plans waiting for empirical validation\.Figure 3:The Space of Actionable Interpretability Work\.
## Appendix BAdditional Examples of Actionable Work

### B\.1Actions That Modify Model’s Output

This appendix provides additional examples of actionable interpretability work that modifies model’s output\.

#### Data curation\.

Pruthiet al\.\([2020](https://arxiv.org/html/2605.11161#bib.bib159)\)estimate the training data influence by tracing how the loss on a test example changes due to each training example during gradient descent\.Manglaet al\.\([2020](https://arxiv.org/html/2605.11161#bib.bib160)\)use saliency maps to guide adversarial training\. They demonstrated that they can identify mislabeled training examples and data poisoning attacks, and that removing the most negatively influential training examples improves model performance\.Magnussonet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib36)\)show that strategic data filtering enables small\-scale evaluations to forecast large\-scale benchmark performance at orders\-of\-magnitude lower compute cost\.

#### Model input\.

Peysakhovich and Lerer \([2023](https://arxiv.org/html/2605.11161#bib.bib161)\)used attention analysis to discover that models attend more to relevant documents even when not using them, then developed “attention sorting”—reordering documents at inference time to improve retrieval\-augmented generation performance\.

#### Training decisions\.

Newmanet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib178)\)build on findings that models’ internal representations of truthfulness can contradict their outputs\(Liuet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib182); Orgadet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib163)\)\. They use these internal representations to curate post\-training data, selecting the model’s own generations that align with its internal signals of truthfulness, and demonstrate that this approach reduces hallucinations\.

#### Self\-explaining models\.

Models that generate explanations as an integral part of their prediction process offer a distinct form of actionability: users can inspect and potentially intervene on the intermediate reasoning\. This paradigm includes self\-attribution architectures\(Agarwalet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib116); Brendel and Bethge,[2019](https://arxiv.org/html/2605.11161#bib.bib117); Jainet al\.,[2020](https://arxiv.org/html/2605.11161#bib.bib118)\), interpretability wrappers for foundation models\(Youet al\.,[2025a](https://arxiv.org/html/2605.11161#bib.bib127)\), selecting prototypes\(Maet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib119); Wenet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib120)\), predicting directly off of concepts\(Kohet al\.,[2020a](https://arxiv.org/html/2605.11161#bib.bib122); Yanget al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib123); Laiet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib124); Yanget al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib125)\), or generating programs as explanations that calculate the outcome\(Lyuet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib126)\)\.

### B\.2Actions About Deployment and Use

#### End user decisions\.

Recent work demonstrates how the internal representations of LLMs can provide uncertainty estimations about their outputs, enabling users to detect errors and make informed decisions\(Kadavathet al\.,[2022](https://arxiv.org/html/2605.11161#bib.bib162); Azaria and Mitchell,[2023](https://arxiv.org/html/2605.11161#bib.bib180); Gottesman and Geva,[2024](https://arxiv.org/html/2605.11161#bib.bib164); Orgadet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib163), inter alia\)\. Building on this foundation,Chakrabortiet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib165)\)argue that AI systems in high\-stakes domains like healthcare must provide personalized uncertainty estimates to support decision\-making: when a clinical decision support system indicates high uncertainty, clinicians can choose to override recommendations rather than following potentially unreliable outputs\.Obesoet al\.\([2025](https://arxiv.org/html/2605.11161#bib.bib179)\)presented a method for real\-time identification of hallucinated tokens in long\-form generations\.

Chenet al\.\([2024b](https://arxiv.org/html/2605.11161#bib.bib166)\)developed a dashboard that exposes the model’s internal “user model”” in real time; user studies showed participants valued this transparency for identifying biased behavior\.

#### Deployment decisions\.

AlthoughChenet al\.\([2024a](https://arxiv.org/html/2605.11161#bib.bib167)\)use a separate scoring function rather than interpretability techniques directly, it illustrates how uncertainty estimates enable benefits\.Casperet al\.\([2024b](https://arxiv.org/html/2605.11161#bib.bib169)\)organized a competition to evaluate whether interpretability tools could help humans detect backdoors implanted in ImageNet\-scale CNNs using feature synthesis methods inspired by interpretability research\. They achieved 49% human detection rates, significantly outperforming dataset\-based attribution methods\.

### B\.3Shaping Future Practice

#### Learning from superhuman models\.

Goodfire’s interpretation of Arc Institute’s biological foundation model Evo 2\(Gortonet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib30)\)identifies biologically relevant structure in model representations, demonstrating interpretability’s potential to guide scientific investigation\.

## Appendix CAdditional Examples on Evaluating Actionability

### C\.1Evaluating Understandability

In high\-stake decisions contexts, interpretability must present model behavior in a form that aligns with users’ existing conceptual frameworks in order to be acted upon\. Understandability can be evaluated by measuring how well explanations align with a user’s domain\-specific reasoning\. Benchmarks such as Features Interpretable to eXperts \(FIX\)\(Jinet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib130)\)and its textual extension T\-FIX\(Havaldaret al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib129)\)operationalize this notion by assessing whether explanations correspond to established concepts in domains such as astrophysics or medicine \(e\.g\., cosmological structures or clinical scoring systems like SOFA\(Vincentet al\.,[1996](https://arxiv.org/html/2605.11161#bib.bib131)\)\)\. For non\-expert users, understandability is often captured through plausibility metrics\(Agarwalet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib132)\), which evaluate whether explanations appear coherent and reasonable given common\-sense expectations\.

### C\.2Evaluating Reliability

A large body of prior work proposes metrics for evaluating the robustness of explanations, including sensitivity of feature attributions\(Alvarez\-Melis and Jaakkola,[2018](https://arxiv.org/html/2605.11161#bib.bib143); Yehet al\.,[2019](https://arxiv.org/html/2605.11161#bib.bib144); Kindermanset al\.,[2019](https://arxiv.org/html/2605.11161#bib.bib145)\), explanation invariance\(Crabbé and van der Schaar,[2023](https://arxiv.org/html/2605.11161#bib.bib142)\), and provable guarantees on explanation behavior\(Blancet al\.,[2021](https://arxiv.org/html/2605.11161#bib.bib138); Bassan and Katz,[2023](https://arxiv.org/html/2605.11161#bib.bib140)\)\.

One example is work on robustness guarantees in the form of*stability certificates*\(Xueet al\.,[2023](https://arxiv.org/html/2605.11161#bib.bib133); Kimet al\.,[2024](https://arxiv.org/html/2605.11161#bib.bib134)\)\. These certificates explicitly quantify how sensitive a model’s predictions are to changes implied by an explanation, such as removing or altering explanatory features\. More recent work has extended such guarantees to large\-scale foundation models\(Jinet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib135)\), chain\-of\-thought explanations\(Youet al\.,[2025b](https://arxiv.org/html/2605.11161#bib.bib137)\), and clinical applications such as Alzheimer’s disease\(Acharaet al\.,[2025](https://arxiv.org/html/2605.11161#bib.bib136)\)\.

Similar Articles

Interpretability

Anthropic Research

Anthropic's Interpretability team focuses on understanding large language models internally to enhance AI safety and positive outcomes, utilizing a multidisciplinary approach.

@fnruji316625: Agentic interpretability is becoming a research direction of its own. Instead of one-shot labeling, AI agents can: form…

X AI KOLs Timeline

Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.

Radical AI Interpretability

arXiv cs.AI

This paper develops a framework for interpreting AI systems as agents, drawing on radical interpretation philosophy and mechanistic interpretability tools, addressing how to trust AI systems by understanding their beliefs, desires, and meanings.

Interactive Evaluation Requires a Design Science

Hugging Face Daily Papers

This position paper argues that interactive AI evaluation should be treated as a design science paradigm, proposing a two-axis taxonomy and reporting standards for assessing dynamic system behavior through trajectories.