From Plausible to Actionable: A Position on LLM Self-Explanations
Summary
This position paper argues that LLM self-explanations can be plausible, questionably faithful, but highly actionable, and proposes evaluation guidelines beyond traditional metrics.
View Cached Full Text
Cached at: 07/20/26, 09:36 AM
# From Plausible to Actionable: A Position on LLM Self-Explanations Source: [https://arxiv.org/html/2607.15957](https://arxiv.org/html/2607.15957) Benedetta Muscato Scuola Normale Superiore, Pisa, Italy University of Pisa, Italy benedetta\.muscato@sns\.it Elize Herrewijnen1,2,Benedetta Muscato3,4,Gizem Gezici3,Fosca Giannotti3 1University of Utrecht, Utrecht, Netherlands 2National Police Lab AI, Netherlands Police, Driebergen, Netherlands 3Scuola Normale Superiore, Pisa, Italy 4University of Pisa, Pisa, Italy Corresponding Author:[e\.herrewijnen@uu\.nl](https://arxiv.org/html/2607.15957v1/mailto:email@domain) ###### Abstract Large Language Models \(LLMs\) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self\-explanations\. Such explanations have emerged as a promising direction for explainable artificial intelligence \(XAI\), particularly for interpreting LLM behavior\. However, while self\-explanations often appear plausible, whether they faithfully reflect a model’s underlying reasoning process remains an open question\. In this opinion paper, we argue that self\-explanations can be highly plausible, questionably faithful, and yet highly actionable\. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM\-generated self\-explanations and propose practical guidelines for assessing their plausibility and faithfulness\. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision\-making and appropriate action across diverse stakeholders\. From Plausible to Actionable: A Position on LLM Self\-Explanations ## 1Introduction The emergence of Large Language Models \(LLMs\) has transformed the landscape of eXplainable Artificial Intelligence \(XAI\), creating new opportunities while introducing significant challenges\. As LLMs are increasingly integrated into real\-world applications, their explanations should accommodate the diverse goals, expertise, and technical backgrounds of stakeholders, thereby fostering informed trust\(Calderon and Reichart,[2025](https://arxiv.org/html/2607.15957#bib.bib112)\)\. Traditional post\-hoc XAI methods generate feature attributions, often using surrogate models such as LIME\(Ribeiroet al\.,[2016](https://arxiv.org/html/2607.15957#bib.bib71)\)\. These approaches typically produce intuitive explanations, for example by highlighting the input tokens that influenced a prediction, making them relatively accessible to stakeholders\. However, their computational cost limits scalability to modern LLMs with hundreds of billions of parameters\(Jean\-Quartieret al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib106)\)\. Mechanistic interpretability offers a complementary approach by reverse\-engineering the internal representations and computational mechanisms of models\(Somvanshiet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib56); Bereska,[2022](https://arxiv.org/html/2607.15957#bib.bib55); Ranaldi,[2025](https://arxiv.org/html/2607.15957#bib.bib54)\), but its complexity limits accessibility to stakeholders with advanced expertise\(Wilde and Rugolon,[2025](https://arxiv.org/html/2607.15957#bib.bib4)\)\. LLM\-generated self\-explanations \(hereafter, self\-explanations\) have emerged as a promising direction for bridging this gap\. Instead of relying on external explanation methods or probing internal mechanisms, LLMs can generate natural language explanations for their outputs\(Fayyazet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib27); Paltaet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib31); Ajwaniet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib83)\), or produce intermediate reasoning through chain\-of\-thought \(CoT\) prompting\(Weiet al\.,[2022](https://arxiv.org/html/2607.15957#bib.bib39)\)\. The natural\-language format of self\-explanations often creates a strong impression of transparency and alignment with human reasoning\(Chenet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib34); Barezet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib35)\), making them particularly appealing in high\-stakes applications\. For example, in healthcare, where LLMs are increasingly explored for clinical decision support and patient\-facing applications, understandable explanations can help build trust among clinicians and patients\(Afolabiet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib105)\)\. Although self\-explanations appear promising for explaining LLM decisions, whilst remaining accessible to many stakeholders, this does not necessarily imply that self\-explanations align with human expectations or truly reflect the model’s underlying decision\-making process\. This distinction is captured by two central concepts in XAI: plausibility and faithfulness\.*Plausibility*refers to how convincing or aligned with human reasoning an explanation appears\(Jacovi and Goldberg,[2021](https://arxiv.org/html/2607.15957#bib.bib17); Paltaet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib31); Lyuet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib82)\), whereas*Faithfulness*refers to the extent to which it accurately reflects the model’s underlying decision process\(DeYounget al\.,[2020](https://arxiv.org/html/2607.15957#bib.bib91); Lyuet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib82)\)\. #### Contributions In this position paper, we argue that self\-explanations can be highly plausible but questionably faithful\. We support our position by highlighting the limitations of current evaluation protocols for self\-explanations and provide practical guidelines for assessing both plausibility and faithfulness \(§[2](https://arxiv.org/html/2607.15957#S2)\)\. Given these limitations, we advocate shifting the focus towards the*Actionability*of self\-explanations, that is their ability to help stakeholders make informed decisions and take concrete actions\(Orgadet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib103)\)To this end, we outline a research agenda that repositions self\-explanations as communicative interfaces for stakeholders without technical expertise, as tools for supporting human decision\-making and facilitating deliberation from diverse perspectives \(§[3](https://arxiv.org/html/2607.15957#S3)\)\. Overall, we call for moving beyond plausibility and faithfulness as the primary objectives of self\-explanations, and instead emphasize their actionable value in supporting informed decision\-making\. ## 2Self\-Explanations: Highly Plausible, Questionably Faithful Explanations generated by XAI methods are commonly evaluated along two key dimensions: plausibility and faithfulness\(Doshi\-Velez and Kim,[2017](https://arxiv.org/html/2607.15957#bib.bib72); Hase and Bansal,[2020](https://arxiv.org/html/2607.15957#bib.bib68)\)\. In the following, we examine the limitations of applying standard XAI evaluation protocols to self\-explanations and propose practical guidelines for evaluating their plausibility and faithfulness\. ### 2\.1Plausibility Prior work has shown that self\-explanations are often highly plausible and well aligned with human expectations\(Paltaet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib31); Ajwaniet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib83); Fanet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib14)\)\. One factor contributing to this high level of plausibility is the sycophantic behavior of LLMs, whereby models tend to agree with or flatter stakeholders\(Malmqvist,[2025](https://arxiv.org/html/2607.15957#bib.bib22); Wanget al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib24)\)\. As a result, stakeholders may be more likely to accept explanations that align with their expectations or existing beliefs\. In addition,Palodet al\.\([2026](https://arxiv.org/html/2607.15957#bib.bib90)\); Agarwalet al\.\([2024](https://arxiv.org/html/2607.15957#bib.bib15)\)note that LLMs are optimized to generate coherent and helpful responses, which together with the accessibility of natural\-language explanations, can increase their perceived interpretability\(Madsenet al\.,[2024a](https://arxiv.org/html/2607.15957#bib.bib110); Kunz and Kuhlmann,[2024](https://arxiv.org/html/2607.15957#bib.bib111)\)\. The high plausibility of self\-explanations can also be attributed to how plausibility is evaluated\. In most studies, plausibility is assessed either through human ratings\(Paltaet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib31); Ajwaniet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib83); Fanet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib14)\)or by comparing generated explanations with human explanations \(e\.g\.,Fayyazet al\.\([2024](https://arxiv.org/html/2607.15957#bib.bib27)\); DeYounget al\.\([2020](https://arxiv.org/html/2607.15957#bib.bib91)\); Di Bonaventuraet al\.\([2024](https://arxiv.org/html/2607.15957#bib.bib42)\)\)\. While these approaches capture important aspects of plausibility, we argue that they overlook two key elements\. #### Plausibility evaluation does not consider the stakeholder’s expertise\. Doshi\-Velez and Kim \([2017](https://arxiv.org/html/2607.15957#bib.bib72)\)propose explanation evaluation that includes both task experts \(e\.g\., doctors\) and lay users, where expert evaluations focus on explanation quality within the task context, and lay evaluations capture more general notions of explanation quality\. This distinction is crucial for assessing plausibility, as a stakeholder’s ability to judge an explanation depends on their expertise\. For example, in a medical diagnosis task, a non\-expert stakeholder may find an explanation with technical terminology convincing, whereas a doctor may recognize flawed reasoning or hallucinated content and judge it as less plausible\. Thus,plausibility evaluation depends on the stakeholder’s expertise and background, and it should therefore be collected and assessed accordingly\. #### Plausibility evaluation does not embrace variation\. Generally, plausibility is measured by comparing a model’s explanation with a human explanation, where closer alignment is interpreted as greater plausibility\. However, this approach overlooks variation in how humans explain the same decision, particularly for ambiguous or subjective tasks\. For example, in hate speech detection, annotators may interpret the same content differently and provide distinct labels and explanations based on their backgrounds and values\. This variation naturally arises from instance difficulty, task ambiguity, and subjective interpretation\(Plank,[2022](https://arxiv.org/html/2607.15957#bib.bib115); Kulmizevet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib109)\)\. Consequently, relying on a single reference explanation risks collapsing diverse but valid perspectives into one standard\. To better reflect human reasoning, plausibility evaluation should accommodate multiple valid explanations by aligning evaluation metrics with the task and validating them with domain experts\(Muscatoet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib93)\)\. Therefore,plausibility evaluation should embrace variation in ambiguous or subjective tasks, where multiple valid explanations may exist, rather than relying on a single reference explanation as the definitive standard\. ### 2\.2Faithfulness A growing body of work has questioned whether self\-explanations satisfy faithfulness requirements\(Paulet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib6); Lanhamet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib5); Turpinet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib7); Lewis\-Limet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib89); Madsenet al\.,[2024b](https://arxiv.org/html/2607.15957#bib.bib29); Colegadoet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib30)\)\. We argue that self\-explanations are questionably faithful for three main reasons\. First, LLMs lack access to the internal processes that generates their outputs\. As argued bySarkar \([2024b](https://arxiv.org/html/2607.15957#bib.bib18)\), self\-explanations are generated through the same process as prediction, without access to meta\-information about the prediction process itself\. Therefore, it is unlikely that self\-explanations reflect genuine introspective analysis\. Second, faithful explanations require separating decision\-making from explanation generation\. However, prompting models to explain their predictions may influence the decision process itself, as suggested by the observed impact of self\-explanations on prediction accuracy\(Liet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib116)\)\. Third, LLMs may provide coherent but misleading rationales due to sycophantic behavior, hallucinations, or biases\(Fenget al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib21)\), raising doubts about whether self\-explanations reflect the true causes of their decisions\. As a result, when models produce incorrect or hallucinated outputs\(Linet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib25)\), their explanations may still present coherent but misleading justifications, raising concerns about whether the reported rationale truly corresponds to the underlying cause of the decision\. Despite the above concerns, several studies have investigated the faithfulness of self\-explanations\(Paulet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib6); Huanget al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib74); Mayneet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib13); Turpinet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib7)\)\. While we encourage investigating the faithfulness of self\-explanations,*we argue that existing faithfulness evaluation methods are not suitable for self\-explanations*\. Many of these methods are perturbation\-based, where model inputs are perturbed \(i\.e\., removing explanation words\) and model outputs \(i\.e\., predictions\) for these perturbed inputs are compared\(Ribeiroet al\.,[2016](https://arxiv.org/html/2607.15957#bib.bib71); Huanget al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib74); Fayyazet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib27); DeYounget al\.,[2020](https://arxiv.org/html/2607.15957#bib.bib91)\)\. These methods rest on three assumptions identified byJacovi and Goldberg \([2020](https://arxiv.org/html/2607.15957#bib.bib26)\): the model assumption, the prediction assumption, and the linearity assumption\. When it comes to LLMs and self\-explanations, these assumptions do not hold for three main reasons\. #### Faithfulness evaluation overlooks nondeterminism in LLMs\. The first assumption identified byJacovi and Goldberg \([2020](https://arxiv.org/html/2607.15957#bib.bib26)\)is the*model assumption*, which states that ‘two models will make the same predictions if and only if they use the same reasoning process\.’ This assumption is problematic for LLMs, as their nondeterministic nature means that identical inputs do not necessarily produce identical outputs across multiple runs, and the same variability extends to their self\-explanations\(Astekinet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib75); Atılet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib70)\)\. Moreover,Bogaertet al\.\([2025](https://arxiv.org/html/2607.15957#bib.bib84)\)provide empirical evidence that models with equivalent predictive performance can nonetheless generate substantially different explanations for their decisions\. Therefore,faithfulness evaluation should account for the nondeterministic behavior of LLMs, rather than relying on single\-run outputs as definitive evidence of faithfulness\. #### Faithfulness evaluation does not account for prompt sensitivity in LLMs\. The second assumption, the*prediction assumption*, states that ‘on similar inputs, the model makes similar decisions if and only if its reasoning is similar\.’ When it comes to LLMs, the notion of a ‘similar input’ is very narrow; LLMs are highly sensitive to subtle changes \(e\.g\., spacing, capitalization, punctuation\) in prompts\(Heet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib12); Zhuoet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib8); Loyaet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib10); Erricaet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib9); Ganet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib96)\), making it possible that the LLM produces different outputs for similar inputs\. For example, removing a comma from the input can unexpectedly change the LLM’s prediction\. This is also the case for modifications to prompts that seem unrelated to the decision; for instance,Huanget al\.\([2023](https://arxiv.org/html/2607.15957#bib.bib74)\)note that asking the LLM to provide explanations negatively affects predictive performance\. Therefore,faithfulness evaluation should not rely on an assumed \(dis\)similarity between inputs, but account for prompt sensitivity in LLMs\. #### Faithfulness evaluation based on perturbations is unreliable for LLMs\. Finally, the*linearity assumption*poses that ‘certain parts of the input are more important to the model reasoning than others\. Moreover, the contributions of different parts of the input are independent from each other’\(Jacovi and Goldberg,[2020](https://arxiv.org/html/2607.15957#bib.bib26)\)\. Relying on this assumption, faithfulness metrics commonly remove or mask parts of the input and assess the change in the model’s predictions\(DeYounget al\.,[2020](https://arxiv.org/html/2607.15957#bib.bib91); Ribeiroet al\.,[2016](https://arxiv.org/html/2607.15957#bib.bib71)\)\. This approach is not suitable for LLMs, because an LLM may disregard the task input and instead rely solely on its internal knowledge to complete the task\. For example, an LLM may still make a sentiment prediction even after all sentiment\-bearing words have been removed from the input\.Fayyazet al\.\([2024](https://arxiv.org/html/2607.15957#bib.bib27)\)report that masking all words in the task input does not change the predicted label, attributing this behavior to label bias in pre\-trained LLMs\. Moreover, perturbing inputs may result in unwarranted model behaviour \(e\.g\. refusal\(Huanget al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib74)\)\), especially when the perturbation increases the unnaturalness of the text\. This also applies to perturbation techniques that rely on synthetic inputs, such as counterfactuals\(Colegadoet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib30); Madsenet al\.,[2024b](https://arxiv.org/html/2607.15957#bib.bib29)\)\. Such synthetic inputs may contain hidden patterns that LLMs can exploit\(Panicksseryet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib66); Wataokaet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib52)\), potentially enabling the model to cheat the faithfulness test\. Thus,faithfulness evaluation should account for two key concerns: the model’s reliance on prior knowledge over the task input, and the possibility that perturbed inputs introduce exploitable artifacts instead of meaningful signal\. ### 2\.3Implications of High Plausibility but Questionable Faithfulness Explanations that are highly plausible, but not faithful can be misleading\.Paltaet al\.\([2026](https://arxiv.org/html/2607.15957#bib.bib31)\)andFanet al\.\([2026](https://arxiv.org/html/2607.15957#bib.bib14)\)show that humans often find LLM self\-explanations persuasive even when the model’s prediction is incorrect\.Ajwaniet al\.\([2025](https://arxiv.org/html/2607.15957#bib.bib83)\)refer to this phenomenon as*adversarial helpfulness*, where explanations make incorrect predictions appear correct\. Such effects may reduce users’ ability to detect erroneous model outputs\(Fanet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib14)\), potentially leading to unwarranted trust in the LLM’s capabilities\. This trust may foster automation bias, i\.e\., the tendency of users to over\-rely on automated recommendations\(Romeo and Conti,[2026](https://arxiv.org/html/2607.15957#bib.bib62)\)\. Such over\-reliance is particularly concerning in high\-stakes domain\(Jacoviet al\.,[2021](https://arxiv.org/html/2607.15957#bib.bib61)\)\. ## 3Towards Actionable Self\-Explanations In the previous section, we discussed the challenges of evaluating the plausibility and faithfulness of self\-explanations\. Although these explanations are often highly plausible, they are not necessarily faithful, as they may not reflect the model’s underlying decision\-making process \(§[2](https://arxiv.org/html/2607.15957#S2)\)\. Given the limitations of current faithfulness evaluation methods, we argue that their value extends beyond faithfully reflecting this process\. Thus, we advocate viewing self\-explanations through the lens of*actionability*: their ability to help diverse stakeholders in making better\-informed decisions and take effective actions\(Orgadet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib103)\)\. Here, we consider self\-explanations as outputs of an LLM’s rationalization capability, i\.e\., the ability to generate justifications for decisions\. From this perspective, we highlight three practical applications of LLM rationalization capabilities: \(i\) serving as a communicative interface that makes XAI more accessible to non\-expert stakeholders, and \(ii\) supporting human decision\-making \(iii\) facilitating deliberation through diverse perspectives\. #### Making XAI accessible to non\-expert stakeholders A key challenge in the practical adoption of XAI is that the most faithful explanations are often the least accessible to non\-expert stakeholders\. Self\-explanations can bridge this gap by translating faithful but complex explanations into natural language rationales that are easier to understand and act upon\(Serafimet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib107)\)\. They can also adapt explanations to different stakeholders by adjusting terminology, formulation, and level of detail to the expertise and information needs of stakeholders\(Albertet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib50); Serafimet al\.,[2025](https://arxiv.org/html/2607.15957#bib.bib107)\)\.Self\-explanations should therefore be viewed not as standalone explanations, but as communicative interfaces that make faithful XAI methods more accessible\. Self\-explanations could also communicate uncertainty and potential hallucinations in LLM outputs\. Although a growing body of work has focused on detecting hallucinations in LLMs\(Farquharet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib64); Sriramananet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib65)\), self\-explanations can complement these methods by conveying their signals in natural language, helping users better assess the reliability of model outputs\. This perspective aligns withUlmeret al\.\([2026](https://arxiv.org/html/2607.15957#bib.bib108)\), who argue that LLMs should communicate uncertainty in a human\-like manner\. #### LLMs in support of human decision\-making In high\-stakes domains, AI systems should be designed to support rather than replace human judgments\(Kostick\-Quenet and Gerke,[2022](https://arxiv.org/html/2607.15957#bib.bib48); Lazaroset al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib46); Paniguttiet al\.,[2023](https://arxiv.org/html/2607.15957#bib.bib45)\)\. Rather than treating self\-explanations as faithful accounts of model reasoning,*we argue they should be viewed as arguments that human decision\-makers can critically evaluate*\. Their value lies not in revealing how a model reached a conclusion, but in surfacing assumptions, uncertainties, and alternative perspectives that help decision\-makers critically evaluate the output\(Sarkar,[2024a](https://arxiv.org/html/2607.15957#bib.bib11)\)\. This perspective can be operationalized through safety protocols \(e\.g\., guardrails\(Jalanet al\.,[2026](https://arxiv.org/html/2607.15957#bib.bib33)\)\), that constrain LLMs to generate arguments rather than final decisions\. For example, in the medical domain, an LLM could present the symptoms and clinical evidence that support or challenge candidate diagnoses, leaving the final decision to the doctor\. #### LLMs as advocates to facilitate deliberation Viewing self\-explanations as a means of supporting human decision\-making, we argue that such explanations can foster deliberation through diverse perspectives\(Vijjiniet al\.,[2024](https://arxiv.org/html/2607.15957#bib.bib114)\)\. Under this framing, variation is essential: different LLMs can surface complementary arguments and counterarguments that enrich the decision process\. We call this paradigm*LLMs as advocates, where models provide diverse viewpoints while humans remain the final decision\-makers*\. By expanding the range of considerations available, these systems have the potential to enhance human judgment and foster reflective deliberation\. ## 4Conclusion In this opinion paper, we have highlighted the limitations of existing XAI evaluation frameworks for measuring the plausibility and faithfulness of LLM\-generated self\-explanations\. Nevertheless, we argue that self\-explanations derive their value from their actionability: their ability to support informed decision\-making and enable stakeholders with diverse goals, expertise, and backgrounds to take appropriate actions\. We position self\-explanations not as literal accounts of a model’s internal reasoning, but as communicative interfaces for non\-experts, decision\-support tools, and facilitators of deliberation by presenting diverse perspectives\. Overall, we advocate moving beyond treating plausibility and faithfulness as the primary objectives of self\-explanations, and instead emphasizing their actionability in enabling informed decision\-making across diverse stakeholders\. ## References - H\. Afolabi, Z\. Afolabi, E\. Friel, J\. Roberts, A\. Ji\-Xu, L\. Chen, E\. Ogbomo, E\. Imevbore, P\. Eneje, W\. El Ouahidi, A\. Sohal, A\. Kennan, S\. Srivastava, A\. Vairavan, L\. Napitu, and K\. McClure \(2026\)Faithful or Just Plausible? Evaluating the Faithfulness of Closed\-Source LLMs in Medical Reasoning\.InProceedings of the Fifth Machine Learning for Health Symposium,P\. Argaw, H\. Zhang, S\. Jabbour, P\. Chandak, J\. Ji, S\. Mukherjee, O\. Salaudeen, T\. Chang, E\. Healey, F\. Gröger, A\. Adibi, S\. Hegselmann, B\. Wild, and A\. Noori \(Eds\.\),Proceedings of Machine Learning Research, Vol\.297,pp\. 1562–1591\.External Links:[Link](https://proceedings.mlr.press/v297/afolabi26a.html)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1)\. - C\. Agarwal, S\. H\. Tanneru, and H\. Lakkaraju \(2024\)Faithfulness vs\. Plausibility: On the \(Un\)Reliability of Explanations from Large Language Models\.arXiv\.External Links:2402\.04614,[Document](https://dx.doi.org/10.48550/arXiv.2402.04614)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - R\. D\. Ajwani, S\. R\. Javaji, F\. Rudzicz, and Z\. Zhu \(2025\)LLM\-Generated Black\-box Explanations Can Be Adversarially Helpful\.InNeurIPS 2024 Workshop on Regulatable ML,External Links:[Link](https://openreview.net/forum?id=F0j4PPyQzt)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2607.15957#S2.SS3.p1.1)\. - J\. Albert, B\. Martin, M\. Doh, J\. Bogaert, L\. De Vos, B\. Renard, V\. Stragier, E\. Jean,et al\.\(2024\)User Preferences for Large Language Model versus Template\-Based Explanations of Movie Recommendations: A Pilot Study\.InDutch\-Belgian Workshop on Recommender Systems,External Links:[Link](https://orbi.umons.ac.be/bitstream/20.500.12907/48327/1/TRAIL23___RECLLM___DBWRS2023.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px1.p1.1)\. - M\. Astekin, M\. Hort, and L\. Moonen \(2024\)An Exploratory Study on How Non\-Determinism in Large Language Models Affects Log Parsing\.InProceedings of the ACM/IEEE 2nd International Workshop on Interpretability, Robustness, and Benchmarking in Neural Software Engineering,Lisbon Portugal,pp\. 13–18\.External Links:[Document](https://dx.doi.org/10.1145/3643661.3643952),ISBN 979\-8\-4007\-0564\-9Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px1.p1.1)\. - B\. Atıl, S\. Aykent, A\. Chittams, L\. Fu, R\. J\. Passonneau, E\. Radcliffe, G\. R\. Rajagopal, A\. Sloan, T\. Tudrej, F\. Türe,et al\.\(2025\)Non\-Determinism of “Deterministic” LLM System Settings in Hosted Environments\.pp\. 135–148\.External Links:[Link](https://arxiv.org/abs/2408.04667)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px1.p1.1)\. - F\. Barez, T\. Wu, I\. Arcuschin, M\. Lan, V\. Wang, N\. Siegel, N\. Collignon, C\. Neo, I\. Lee, A\. Paren,et al\.\(2025\)Chain\-of\-Thought Is Not Explainability\.Preprint, alphaXiv\.External Links:[Link](https://fazlbarez.com/assets/pdf/Cot_Is_Not_Explainability.pdf)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1)\. - L\. F\. Bereska \(2022\)Mechanistic Interpretability for AI Safety—A Review\.InProceedings of The 1st Conference on Lifelong Learning Agents,External Links:[Link](https://arxiv.org/abs/2404.14082)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1)\. - J\. Bogaert, A\. Descampe, and F\. Standaert \(2025\)Consolidating explanation stability metrics\.InWorld Conference on Explainable Artificial Intelligence,pp\. 310–323\.External Links:[Link](https://link.springer.com/content/pdf/10.1007/978-3-032-08327-2_15.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px1.p1.1)\. - N\. Calderon and R\. Reichart \(2025\)On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 656–693\.External Links:[Link](https://aclanthology.org/2025.naacl-long.29.pdf)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p1.1)\. - X\. Chen, A\. Plaat, and N\. van Stein \(2026\)How does Chain of Thought Think? Mechanistic Interpretability of Chain\-of\-Thought Reasoning with Sparse Autoencoding\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 30297–30305\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/download/40281/44242)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1)\. - S\. Colegado, J\. Jin, and Y\. Hou \(2026\)Faithful or Plausible? A Counterfactual Analysis of LLM\-Generated Explanations for Machine Learning Models\.In2026 International Conference on Semantic Computing \(ICSC\),Laguna Hills, CA, USA,pp\. 83–90\.External Links:[Document](https://dx.doi.org/10.1109/ICSC67292.2026.00018),ISBN 979\-8\-3315\-4748\-6Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1)\. - J\. DeYoung, S\. Jain, N\. F\. Rajani, E\. Lehman, C\. Xiong, R\. Socher, and B\. C\. Wallace \(2020\)ERASER: A benchmark to evaluate rationalized NLP models\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 4443–4458\.External Links:[Link](https://aclanthology.org/2020.acl-main.408.pdf)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p4.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - C\. Di Bonaventura, L\. Siciliani, P\. Basile, A\. Meroño\-Peñuela, and B\. McGillivray \(2024\)Is explanation all you need? an expert survey on llm\-generated explanations for abusive language detection\.InProceedings of the 10th italian conference on computational linguistics \(CLiC\-it 2024\),pp\. 280–288\.External Links:[Link](https://aclanthology.org/2024.clicit-1.34.pdf)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1)\. - F\. Doshi\-Velez and B\. Kim \(2017\)Towards A Rigorous Science of Interpretable Machine Learning\.arXiv\.External Links:1702\.08608,[Document](https://dx.doi.org/10.48550/arXiv.1702.08608)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2607.15957#S2.p1.1)\. - F\. Errica, D\. Sanvito, G\. Siracusano, and R\. Bifulco \(2025\)What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 1543–1558\.External Links:[Link](https://aclanthology.org/2025.naacl-long.73.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1)\. - S\. Fan, L\. Zhang, and X\. Yuan \(2026\)When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI\-Assisted Decision Making\.arXiv\.External Links:2602\.04003,[Document](https://dx.doi.org/10.48550/arXiv.2602.04003)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2607.15957#S2.SS3.p1.1)\. - S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. Gal \(2024\)Detecting hallucinations in large language models using semantic entropy\.Nature630\(8017\),pp\. 625–630\.External Links:[Link](https://www.nature.com/articles/s41586-024-07421-0.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px1.p1.1)\. - M\. Fayyaz, F\. Yin, J\. Sun, and N\. Peng \(2024\)Evaluating Human Alignment and Model Faithfulness of LLM Rationale\.arXiv\.External Links:2407\.00219,[Document](https://dx.doi.org/10.48550/arXiv.2407.00219)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - Z\. Feng, Z\. Chen, J\. Ma, Y\. T\. Po, E\. Chersoni, and B\. Li \(2026\)Good Arguments Against the People Pleasers: How Reasoning Mitigates \(Yet Masks\) LLM Sycophancy\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 24536–24570\.External Links:[Link](https://aclanthology.org/2026.acl-long.1126.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p2.1)\. - E\. Gan, Y\. Zhao, L\. Cheng, M\. Yancan, A\. Goyal, K\. Kawaguchi, M\. Kan, and M\. Shieh \(2024\)Reasoning Robustness of LLMs to Adversarial Typographical Errors\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 10449–10459\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.584.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1)\. - P\. Hase and M\. Bansal \(2020\)Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5540–5552\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.491)Cited by:[§2](https://arxiv.org/html/2607.15957#S2.p1.1)\. - J\. He, M\. Rungta, D\. Koleczek, A\. Sekhon, F\. X\. Wang, and S\. Hasan \(2024\)Does prompt formatting have any impact on LLM performance?\.arXiv preprint arXiv:2411\.10541\.External Links:[Link](https://arxiv.org/abs/2411.10541)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1)\. - S\. Huang, S\. Mamidanna, S\. Jangam, Y\. Zhou, and L\. H\. Gilpin \(2023\)Can Large Language Models Explain Themselves? A Study of LLM\-Generated Self\-Explanations\.arXiv\.External Links:2310\.11207,[Document](https://dx.doi.org/10.48550/arXiv.2310.11207)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - A\. Jacovi and Y\. Goldberg \(2020\)Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,Online,pp\. 4198–4205\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.386)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - A\. Jacovi and Y\. Goldberg \(2021\)Aligning Faithful Interpretations with Their Social Attribution\.Transactions of the Association for Computational Linguistics9,pp\. 294–310\.External Links:https://direct\.mit\.edu/tacl/article\-pdf/doi/10\.1162/tacl\_a\_00367/1923972/tacl\_a\_00367\.pdf,ISSN 2307\-387X,[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00367)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p4.1)\. - A\. Jacovi, A\. Marasović, T\. Miller, and Y\. Goldberg \(2021\)Formalizing trust in artificial intelligence: prerequisites, causes and goals of human trust in ai\.InProceedings of the 2021 ACM conference on fairness, accountability, and transparency,pp\. 624–635\.External Links:[Link](https://dl.acm.org/doi/pdf/10.1145/3442188.3445923)Cited by:[§2\.3](https://arxiv.org/html/2607.15957#S2.SS3.p1.1)\. - P\. Jalan, V\. Abishethvarman, B\. Chandna, and U\. Naseem \(2026\)Survey on LLM safety: Attacks, defenses, alignment, metrics, and guardrails\.Machine Learning115\(6\),pp\. 130\.External Links:[Link](https://link.springer.com/content/pdf/10.1007/s10994-026-07060-8.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px2.p1.1)\. - C\. Jean\-Quartier, K\. Bein, L\. Hejny, E\. Hofer, A\. Holzinger, and F\. Jeanquartier \(2023\)The cost of understanding—XAI algorithms towards sustainable ML in the view of computational cost\.Computation11\(5\),pp\. 92\.External Links:[Link](https://www.mdpi.com/2079-3197/11/5/92)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1)\. - K\. M\. Kostick\-Quenet and S\. Gerke \(2022\)AI in the hands of imperfect users\.NPJ digital medicine5\(1\),pp\. 197\.External Links:[Link](https://pubmed.ncbi.nlm.nih.gov/36577851/)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px2.p1.1)\. - A\. Kulmizev, E\. Lombart, P\. Watrin, and M\. de Marneffe \(2026\)Label and Explanation Variation in LLM\-Based Annotation: a Case Study in Natural Language Inference\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16526–16543\.External Links:[Link](https://aclanthology.org/2026.acl-long.752/)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.SSS0.Px2.p1.1)\. - J\. Kunz and M\. Kuhlmann \(2024\)Properties and Challenges of LLM\-Generated Explanations\.InProceedings of the Third Workshop on Bridging Human–Computer Interaction and Natural Language Processing,pp\. 13–27\.External Links:[Link](https://aclanthology.org/2024.hcinlp-1.2/)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - T\. Lanham, A\. Chen, A\. Radhakrishnan, B\. Steiner, C\. Denison, D\. Hernandez, D\. Li, E\. Durmus, E\. Hubinger, J\. Kernion, K\. Lukošiūtė, K\. Nguyen, N\. Cheng, N\. Joseph, N\. Schiefer, O\. Rausch, R\. Larson, S\. McCandlish, S\. Kundu, S\. Kadavath, S\. Yang, T\. Henighan, T\. Maxwell, T\. Telleen\-Lawton, T\. Hume, Z\. Hatfield\-Dodds, J\. Kaplan, J\. Brauner, S\. R\. Bowman, and E\. Perez \(2023\)Measuring Faithfulness in Chain\-of\-Thought Reasoning\.arXiv\.External Links:2307\.13702,[Document](https://dx.doi.org/10.48550/arXiv.2307.13702),[Link](https://arxiv.org/abs/2307.13702)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1)\. - K\. Lazaros, A\. G\. Vrahatis, and S\. Kotsiantis \(2026\)Human\-in\-the\-loop artificial intelligence: a systematic review of concepts, methods, and applications\.Entropy28\(4\),pp\. 377\.External Links:[Link](https://www.mdpi.com/1099-4300/28/4/377)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px2.p1.1)\. - S\. Lewis\-Lim, X\. Tan, Z\. Zhao, and N\. Aletras \(2025\)Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post\-hoc Rationalisation?\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 29826–29841\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1516.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1)\. - D\. Li, B\. Hu, Q\. Chen, and S\. He \(2023\)Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training\.InProceedings of the 3rd Workshop on Trustworthy Natural Language Processing \(TrustNLP 2023\),pp\. 1–14\.External Links:[Link](https://aclanthology.org/2023.trustnlp-1.1.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p2.1)\. - L\. Lin, L\. Wang, J\. Guo, and K\. Wong \(2025\)Investigating bias in LLM\-based bias detection: Disparities between LLMs and human perception\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 10634–10649\.External Links:[Link](https://aclanthology.org/2025.coling-main.709.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p2.1)\. - M\. Loya, D\. Sinha, and R\. Futrell \(2023\)Exploring the Sensitivity of LLMs’ Decision\-Making Capabilities: Insights from Prompt Variation and Hyperparameters\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 3711–3716\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.241.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1)\. - Q\. Lyu, M\. Apidianaki, and C\. Callison\-Burch \(2024\)Towards Faithful Model Explanation in NLP: A Survey\.Computational Linguistics50\(2\),pp\. 657–723\.External Links:[Link](https://aclanthology.org/2024.cl-2.6.pdf)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p4.1)\. - A\. Madsen, S\. Chandar, and S\. Reddy \(2024a\)Are self\-explanations from Large Language Models faithful?\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 295–337\.External Links:[Link](https://aclanthology.org/2024.findings-acl.19.pdf)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - A\. Madsen, S\. Chandar, and S\. Reddy \(2024b\)Are Self\-Explanations from Large Language Models Faithful?\.InFindings of the Association for Computational Linguistics ACL 2024,Bangkok, Thailand and virtual meeting,pp\. 295–337\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.19)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1)\. - L\. Malmqvist \(2025\)Sycophancy in large language models: Causes and mitigations\.InIntelligent Computing\-Proceedings of the Computing Conference,pp\. 61–74\.External Links:[Link](https://link.springer.com/chapter/10.1007/978-3-031-92611-2_5)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - H\. Mayne, J\. S\. Kang, D\. Gould, K\. Ramchandran, A\. Mahdi, and N\. Y\. Siegel \(2026\)A Positive Case for Faithfulness: LLM Self\-Explanations Help Predict Model Behavior\.arXiv\.External Links:2602\.02639,[Document](https://dx.doi.org/10.48550/arXiv.2602.02639)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - B\. Muscato, B\. Chen, G\. Gezici, B\. Plank, and F\. Giannotti \(2026\)Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection\.arXiv preprint arXiv:2605\.31563\.External Links:[Link](https://arxiv.org/abs/2605.31563)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.SSS0.Px2.p1.1)\. - H\. Orgad, F\. Barez, T\. Haklay, I\. Lee, M\. Mosbach, A\. Reusch, N\. Saphra, B\. Wallace, S\. Wiegreffe, E\. Wong,et al\.\(2026\)Interpretability Can Be Actionable\.arXiv preprint arXiv:2605\.11161\.External Links:[Link](https://arxiv.org/pdf/2605.11161)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.15957#S3.p1.1)\. - V\. Palod, U\. Biswas, and S\. Kambhampati \(2026\)Evaluating the False Trust engendered by LLM Explanations\.arXiv preprint arXiv:2605\.10930\.External Links:[Link](https://arxiv.org/abs/2605.10930)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - S\. Palta, P\. Rankel, S\. Wiegreffe, and R\. Rudinger \(2026\)Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility\.arXiv\.External Links:2510\.08091,[Document](https://dx.doi.org/10.48550/arXiv.2510.08091)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1),[§1](https://arxiv.org/html/2607.15957#S1.p4.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2607.15957#S2.SS3.p1.1)\. - A\. Panickssery, S\. R\. Bowman, and S\. Feng \(2024\)LLM evaluators recognize and favor their own generations\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385,[Link](https://dl.acm.org/doi/10.5555/3737916.3740113)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1)\. - C\. Panigutti, R\. Hamon, I\. Hupont, D\. Fernandez Llorca, D\. Fano Yela, H\. Junklewitz, S\. Scalzo, G\. Mazzini, I\. Sanchez, J\. Soler Garrido,et al\.\(2023\)The role of explainable AI in the context of the AI Act\.InProceedings of the 2023 ACM conference on fairness, accountability, and transparency,pp\. 1139–1150\.External Links:[Link](https://dl.acm.org/doi/10.1145/3593013.3594069)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px2.p1.1)\. - D\. Paul, R\. West, A\. Bosselut, and B\. Faltings \(2024\)Making Reasoning Matter: Measuring and Improving Faithfulness of Chain\-of\-Thought Reasoning\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Miami, Florida, USA,pp\. 15012–15032\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.882)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - B\. Plank \(2022\)The “problem” of human label variation: on ground truth in data, modeling and evaluation\.InProceedings of the 2022 conference on empirical methods in natural language processing,pp\. 10671–10682\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.731/)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.SSS0.Px2.p1.1)\. - L\. Ranaldi \(2025\)Survey on the role of mechanistic interpretability in generative AI\.Big Data and Cognitive Computing9\(8\),pp\. 193\.External Links:[Link](https://www.mdpi.com/2504-2289/9/8/193)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1)\. - M\. T\. Ribeiro, S\. Singh, and C\. Guestrin \(2016\)"Why should I trust you?" Explaining the predictions of any classifier\.InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 1135–1144\.External Links:[Link](https://dl.acm.org/doi/pdf/10.1145/2939672.2939778?.)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - G\. Romeo and D\. Conti \(2026\)Exploring automation bias in human–AI collaboration: a review and implications for explainable AI\.Ai & Society41\(1\),pp\. 259–278\.External Links:[Link](https://dl.acm.org/doi/10.1007/s00146-025-02422-7)Cited by:[§2\.3](https://arxiv.org/html/2607.15957#S2.SS3.p1.1)\. - A\. Sarkar \(2024a\)AI should challenge, not obey\.Communications of the ACM67\(10\),pp\. 18–21\.External Links:[Link](https://dl.acm.org/doi/10.1145/3649404)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px2.p1.1)\. - A\. Sarkar \(2024b\)Large language models cannot explain themselves\.InHCXAI workshop,CHI EA ’24,New York, NY, USA\.External Links:ISBN 9798400703317,[Link](https://arxiv.org/abs/2405.04382)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p2.1)\. - P\. B\. Serafim, R\. F\. Filho, S\. Freitas, G\. Gezici, F\. Giannotti, F\. Raimondi, and A\. Santos \(2025\)MAINLE: A Multi\-Agent, Interactive, Natural Language Local Explainer of Classification Tasks\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 149–165\.External Links:[Link](https://ecmlpkdd-storage.s3.eu-central-1.amazonaws.com/preprints/2025/research/preprint_ecml_pkdd_2025_research_1246.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px1.p1.1)\. - S\. Somvanshi, M\. M\. Islam, A\. Rafe, A\. G\. Tusti, A\. Chakraborty, A\. Baitullah, T\. I\. Chowdhury, N\. Alnawmasi, A\. Dutta, and S\. Das \(2026\)Bridging the Black Box: A Survey on Mechanistic Interpretability in AI\.ACM Computing Surveys58\(8\),pp\. 1–35\.External Links:[Link](https://dl.acm.org/doi/pdf/10.1145/3787104)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1)\. - G\. Sriramanan, S\. Bharti, V\. S\. Sadasivan, S\. Saha, P\. Kattakinda, and S\. Feizi \(2024\)LLM\-Check: Investigating Detection of Hallucinations in Large Language Models\.Advances in Neural Information Processing Systems37,pp\. 34188–34216\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/3c1e1fdf305195cd620c118aaa9717ad-Paper-Conference.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px1.p1.1)\. - M\. Turpin, J\. Michael, E\. Perez, and S\. R\. Bowman \(2023\)Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain\-of\-Thought Prompting\.Advances in Neural Information Processing Systems36,pp\. 74952–74965\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.p3.1)\. - D\. Ulmer, A\. Lorson, I\. Titov, and C\. Hardmeier \(2026\)Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing\.Transactions of the Association for Computational Linguistics14,pp\. 1505–1540\.External Links:[Link](https://direct.mit.edu/tacl/article/doi/10.1162/TACL.a.739/137440)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px1.p1.1)\. - A\. R\. Vijjini, R\. R\. Menon, J\. Fu, S\. Srivastava, and S\. Chaturvedi \(2024\)SOCIALGAZE: Improving the Integration of Human Social Norms in Large Language Models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 16487–16506\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.962.pdf)Cited by:[§3](https://arxiv.org/html/2607.15957#S3.SS0.SSS0.Px3.p1.1)\. - K\. Wang, J\. Li, S\. Yang, Z\. Zhang, and D\. Wang \(2026\)When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 33566–33574\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40645/44606)Cited by:[§2\.1](https://arxiv.org/html/2607.15957#S2.SS1.p1.1)\. - K\. Wataoka, T\. Takahashi, and R\. Ri \(2024\)Self\-Preference Bias in LLM\-as\-a\-Judge\.InNeurips Safe Generative AI Workshop 2024,External Links:[Link](https://arxiv.org/abs/2410.21819)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px3.p1.1)\. - J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\.Advances in neural information processing systems35,pp\. 24824–24837\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p3.1)\. - K\. M\. Wilde and F\. Rugolon \(2025\)Evaluating Explainability Techniques for Machine Learning in Healthcare\-A Human\-Centered Approach Through Expert Interviews\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 219–234\.External Links:[Link](https://link.springer.com/chapter/10.1007/978-3-032-19099-4_16)Cited by:[§1](https://arxiv.org/html/2607.15957#S1.p2.1)\. - J\. Zhuo, S\. Zhang, X\. Fang, H\. Duan, D\. Lin, and K\. Chen \(2024\)ProSA: Assessing and understanding the prompt sensitivity of LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 1950–1976\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.108.pdf)Cited by:[§2\.2](https://arxiv.org/html/2607.15957#S2.SS2.SSS0.Px2.p1.1)\.
Similar Articles
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
This paper proposes a definition of good explanations based on counterfactuals and prior beliefs, and discusses the inherent difficulties in explaining LLM outputs under this definition.
What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
This paper formalizes the concept of explanation sufficiency for LLMs, proposes a new metric called SCSuff to evaluate free-text explanations using the model's own input beliefs, and demonstrates that current LLM explanations are generally insufficient.
LLMs are not the black box you were promised
An article summarizing Anthropic's 2025 paper on mechanistic interpretability, showing that LLMs are not black boxes and that circuit tracing can reveal multi-step reasoning and human-identifiable concepts.
Position: It's Time to Optimize LLMs for Self-Consistency
This position paper argues that many LLM failures stem from evaluating outputs independently and proposes a self-consistency framework that treats diverse techniques as special cases of consistency optimization.
LLMs know when they are wrong. I made a fix relating to Anthropic's new "global workspace" paper [R]
The author presents a method to make LLMs verbalize calibrated confidence by using a linear probe on mid-layer states and a small trained bridge to confidence logits, requiring only 200 labeled examples and no weight modification. This is linked to Anthropic's global workspace paper explaining the know-say gap.