@DanKornas: LLM interpretability is a rabbit hole. This repo gives you a map. Awesome LLM Interpretability is a curated GitHub list…
Summary
A curated GitHub list of tools, papers, and communities for LLM interpretability, helping researchers navigate the field efficiently.
View Cached Full Text
Cached at: 06/22/26, 07:40 AM
LLM interpretability is a rabbit hole. This repo gives you a map.
Awesome LLM Interpretability is a curated GitHub list of tools, papers, articles, groups, and a survey paper focused on understanding large language models.
It helps you study the field faster by grouping practical tools, research papers, explainers, and communities into scan-friendly sections instead of chasing random bookmarks.
Key features:
• Tool index – links to LLM interpretability and analysis tools like LIT, TransformerLens, Inseq, ecco, Pythia, and Automated Interpretability • Paper list – collects academic and industry papers on topics like sparse probing, copy suppression, monosemanticity, causal tracing, and model editing • Article section – includes explainers and interactive resources on grokking, the logit lens, activation patching, causal scrubbing, and evaluation pitfalls • Community map – points to interpretability and alignment groups including PAIR, Alignment Lab AI, Nous Research, and EleutherAI • Contribution path – includes fork, branch, PR, review, and merge guidelines so the list can keep improving
Free public GitHub repo.
Link in the reply
Similar Articles
@DanKornas: Reading LLM research gets slow when every topic starts with another search. Awesome-LLM-Survey is a topic-organized col…
Awesome-LLM-Survey is a topic-organized GitHub repository collecting LLM survey papers and project links to help researchers quickly find overviews across training, prompting, modalities, and applications.
@DanKornas: Keeping up with LLM systems research is messy when papers, reports, frameworks, and course links are scattered everywhe…
LLMSys-PaperList is a curated reading list on GitHub that organizes LLM systems research papers and resources into practical categories such as training systems, serving systems, and multi-modal coverage, helping AI/ML engineers and researchers stay updated.
LLMs are not the black box you were promised
An article summarizing Anthropic's 2025 paper on mechanistic interpretability, showing that LLMs are not black boxes and that circuit tracing can reveal multi-step reasoning and human-identifiable concepts.
@DanKornas: LLM eval is where most AI demos start becoming real systems. LLM-Evaluation is a public GitHub resource with workshop s…
A tweet announces LLM-Evaluation, a public GitHub repository containing workshop slides, sample notebooks, prompts, and reference links for evaluating LLMs, generative AI, and RAG systems, aiming to provide a practical map of evaluation workflows.
I kept a doc of every LLM term that confused me while building. Cleaned it up and open sourced it.
The author compiled a glossary of confusing LLM terms with production-oriented explanations, cleaned it up, and open-sourced it as a browsable UI on GitHub.