@DanKornas: LLM interpretability is a rabbit hole. This repo gives you a map. Awesome LLM Interpretability is a curated GitHub list…

X AI KOLs Timeline Tools

Summary

A curated GitHub list of tools, papers, and communities for LLM interpretability, helping researchers navigate the field efficiently.

LLM interpretability is a rabbit hole. This repo gives you a map. Awesome LLM Interpretability is a curated GitHub list of tools, papers, articles, groups, and a survey paper focused on understanding large language models. It helps you study the field faster by grouping practical tools, research papers, explainers, and communities into scan-friendly sections instead of chasing random bookmarks. Key features: • Tool index – links to LLM interpretability and analysis tools like LIT, TransformerLens, Inseq, ecco, Pythia, and Automated Interpretability • Paper list – collects academic and industry papers on topics like sparse probing, copy suppression, monosemanticity, causal tracing, and model editing • Article section – includes explainers and interactive resources on grokking, the logit lens, activation patching, causal scrubbing, and evaluation pitfalls • Community map – points to interpretability and alignment groups including PAIR, Alignment Lab AI, Nous Research, and EleutherAI • Contribution path – includes fork, branch, PR, review, and merge guidelines so the list can keep improving Free public GitHub repo. Link in the reply
Original Article
View Cached Full Text

Cached at: 06/22/26, 07:40 AM

LLM interpretability is a rabbit hole. This repo gives you a map.

Awesome LLM Interpretability is a curated GitHub list of tools, papers, articles, groups, and a survey paper focused on understanding large language models.

It helps you study the field faster by grouping practical tools, research papers, explainers, and communities into scan-friendly sections instead of chasing random bookmarks.

Key features:

• Tool index – links to LLM interpretability and analysis tools like LIT, TransformerLens, Inseq, ecco, Pythia, and Automated Interpretability • Paper list – collects academic and industry papers on topics like sparse probing, copy suppression, monosemanticity, causal tracing, and model editing • Article section – includes explainers and interactive resources on grokking, the logit lens, activation patching, causal scrubbing, and evaluation pitfalls • Community map – points to interpretability and alignment groups including PAIR, Alignment Lab AI, Nous Research, and EleutherAI • Contribution path – includes fork, branch, PR, review, and merge guidelines so the list can keep improving

Free public GitHub repo.

Link in the reply

Similar Articles

LLMs are not the black box you were promised

Hacker News Top

An article summarizing Anthropic's 2025 paper on mechanistic interpretability, showing that LLMs are not black boxes and that circuit tracing can reveal multi-step reasoning and human-identifiable concepts.