Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures
Summary
A comprehensive survey reviewing recent advances in intrinsic interpretability for Large Language Models, categorizing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. The paper addresses the challenge of building transparency directly into model architectures rather than relying on post-hoc explanation methods.
View Cached Full Text
Cached at: 04/20/26, 08:30 AM
# Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures Source: https://arxiv.org/html/2604.16042 Yutong Gao1,4,*, Qinglin Meng5,*, Yuan Zhou5, Liangming Pan1,2,3,† 1 MOE Key Laboratory of Computational Linguistics, Peking University 2 School of Computer Science, Peking University 3 Beijing Academy of Artificial Intelligence, Beijing, China 4 Nanjing University of Science and Technology 5 Purdue University [email protected], {meng160, zhou1475}@purdue.edu, [email protected] ## Abstract While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment. Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations. In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative. This paper presents a systematic review of recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. We further discuss open challenges and outline future research directions in this emerging field. The paper list is available at: [Survey-Intrinsic-Interpretability-of-LLMs](https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs) ## 1 Introduction Large Language Models have achieved remarkable success across diverse tasks (Brown et al., 2020; Raffle et al., 2020; Chowdhery et al., 2022; Team et al., 2025). However, their complexity often makes them "black boxes" (Bommasani et al., 2022), hiding their internal decision-making. This lack of transparency creates trust and safety risks, especially in high-stakes fields like healthcare and law (Rudin, 2019; Pawar et al., 2020). To address these concerns, interpretability research is often divided into two paradigms: post-hoc explanation and intrinsic design. Post-hoc methods analyze trained, fixed models using external tools such as LIME, SHAP, sparse autoencoders, or causal interventions (Ribeiro et al., 2016; Lundberg and Lee, 2017; Huben et al., 2024; Meng et al., 2022). Many rely on surrogate models or statistical attributions, resulting in a well-known fidelity gap between the explanation and the model's true computation (Jacovi and Goldberg, 2020). Causal-based post-hoc methods partially address this issue by intervening directly on internal components, yielding stronger local faithfulness (Meng et al., 2022; Wang et al., 2023). However, their explanations remain highly fine-grained and are difficult to aggregate into coherent, high-level accounts of overall model behavior. In contrast, intrinsic interpretability builds transparency directly into the model architecture and training process (Fedus et al., 2022; Gao et al., 2025). By ensuring that the model's internal computation is itself interpretable, these approaches aim to achieve structural fidelity—namely, a direct correspondence between model behavior and its explanation, without relying on external surrogates or post-hoc aggregation. Historically, however, intrinsic methods were constrained by a severe trade-off: models that were transparent by construction typically lacked the expressive power required for complex language tasks (Linaardatos et al., 2021). Recent advances demonstrate that interpretability and performance need not be mutually exclusive, showing that large-scale models can be designed with interpretable internal structure while retaining competitive task performance (Rudin, 2019; Sharkey et al., 2025). By incorporating inductive biases such as modularity, sparsity, disentanglement, and structured representations directly into modern architectures and training objectives (Shazeer et al., 2017; Louizos et al., 2018; Fedus et al., 2022; Gao et al., 2025), these methods enable interpretability to emerge as a property of the model itself rather than as an after-the-fact analysis. Despite this rapid progress, the literature on intrinsic interpretability remains fragmented, spanning disparate model classes, architectural choices, and training principles. Unlike post-hoc explanation methods whose taxonomy and limitations have been extensively surveyed (Molnar, 2025; Madsen et al., 2022; Zhao et al., 2024a; Palikhe et al., 2025), there remains a need for a unified framework that organizes intrinsic approaches around shared design principles or clarifies how different mechanisms contribute to transparency in LLMs. This survey aims to fill this gap by systematically reviewing intrinsic interpretability methods for LLMs, distilling common design principles, and highlighting open challenges and promising future directions. Our contributions are threefold. First, we distinguish post-hoc explanation from intrinsic interpretability, clarifying their differences in faithfulness, scope, and design philosophy. Second, we introduce a structured taxonomy of intrinsic interpretability methods organized around five core design principles: **Functional Transparency**, **Concept Alignment**, **Representational Decomposability**, **Explicit Modularity**, and **Latent Sparsity Induction**. Finally, we synthesize existing work within this framework, analyze methodological strengths and limitations, and identify key open challenges and future research directions. ## References - Towards robust interpretability with self-explaining neural networks. Advances in Neural Information Processing Systems 31. Cited by: Table 1. - M. Böhle, M. Fritz, and B. Schiele (2022) B-cos networks: alignment is all we need for interpretability. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10329–10338. Cited by: Table 1. - R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang (2022) On the opportunities and risks of foundation models. External Links: 2108.07258. Cited by: §1. - T. B. Brown, B. Mann, N. Ryder, S. Subbiah, J. Kaplan, P. Dhariwal, N. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020) Language models are few-shot learners. External Links: 2005.14165. Cited by: §1. - C. Chang, R. Caruana, and A. Goldenberg (2022) NODE-GAM: neural generalized additive model for interpretable deep learning. External Links: 2106.01613. Cited by: Table 1. - A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel (2022) PaLM: Scaling language modeling with pathways. External Links: 2204.02311. Cited by: §1. - Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier (2017) Language modeling with gated convolutional networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6–11 August 2017, D. Precup and Y. W. Teh (Eds.), Proceedings of Machine Learning Research, Vol. 70, pp. 933–941. Cited by: Table 1. - G. Do, H. Le, and T. Tran (2025) Unified sparse mixture of experts. External Links: 2503.22996. Cited by: Table 1. - W. Fedus, B. Zoph, and N. Shazeer (2022) Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23, pp. 120:1–120:39. Cited by: §1. - L. Gao, A. Rajaram, J. Coxon, S. V. Govande, B. Baker, and D. Mossing (2025) Weight-sparse Transformers have interpretable circuits. External Links: 2511.13653. Cited by: Table 1, §1. - Z. Gao, P. Liu, W. X. Zhao, Z. Lu, and J. Wen (2022) Parameter-efficient mixture-of-experts architecture for pre-trained language models. In Proceedings of the 29th International Conference on Computational Linguistics, N. Calzolari, C. Huang, H. Kim, J. Pustejovsky, L. Wanner, K. Choi, P. Ryu, H. Chen, L. Donatelli, H. Ji, S. Kurohashi, P. Paggio, N. Xue, S. Kim, Y. Hahm, Z. He, T. K. Lee, E. Santus, F. Bond, and S. Na (Eds.), Gyeongju, Republic of Korea, pp. 3263–3273. Cited by: Table 1. - H. Guo, H. Lu, G. Nan, B. Chu, J. Zhuang, Y. Yang, W. Che, S. Leng, Q. Cui, and X. Jiang (2025) Advancing expert specialization for better MoE. External Links: 2505.22323. Cited by: Table 1. - T. Hastie and R. Tibshirani (1986) Generalized Additive Models. Statistical Science, 1(3), pp. 297–310. Cited by: Table 1. - M. Havasi, S. Parbhoo, and F. Doshi-Velez (2022) Addressing leakage in concept bottleneck models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 – December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.). Cited by: Table 1. - J. Hewitt, J. Thickstun, C. D. Manning, and P. Liang (2023) Backpack language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9–14, 2023, A. Rogers, J. L. Boyd-Graber, and N. Okazaki (Eds.), pp. 9103–9125. Cited by: Table 1. - R. Huben, H. Cunningham, L. R. Smith, A. Ewart, and L. Sharkey (2024) Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7–11, 2024. Cited by: §1. - A. Jacovi and Y. Goldberg (2020) Towards faithfully interpretable NLP systems: how should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5–10, 2020, D. Jurafsky, J. Chai, N. Schluter, and J. R. Tetreault (Eds.), pp. 4198–4205. Cited by: §1.
Similar Articles
Can We Understand How Large Language Models Reason?
This article explores the ongoing efforts and challenges in understanding how large language models reason, focusing on interpretability research.
Applied Explainability for Large Language Models: A Comparative Study
A comparative study evaluating three explainability techniques (Integrated Gradients, Attention Rollout, SHAP) on fine-tuned DistilBERT for sentiment classification, highlighting trade-offs between gradient-based, attention-based, and model-agnostic approaches for LLM interpretability.
Explainable artificial intelligence (XAI): From inherent explainability to large language models
This paper examines the progression from inherent explainability in artificial intelligence to the development and application of explainable methods for large language models.
Understanding Large Language Models
This chapter reviews current understanding of Large Language Models, discussing their Transformer architecture, emergent capabilities resembling human cognition, and debates about whether LLMs genuinely understand or merely simulate understanding.
Memory for Large Language Models
This survey presents a systematic taxonomy of memory mechanisms in large language models, classifying along axes of representation, update dynamics, and persistence, and formalizing the underlying mechanistic components.