@AnthropicAI: To support other researchers getting hands-on experience with NLAs, we’ve partnered with Neuronpedia to release NLAs on…
Summary
Anthropic and Neuronpedia have partnered to release Natural Language Autoencoders (NLAs) on open models, allowing researchers to gain hands-on experience with this interpretability tool.
View Cached Full Text
Cached at: 05/08/26, 09:59 AM
To support other researchers getting hands-on experience with NLAs, we’ve partnered with Neuronpedia to release NLAs on open models.
Try them out here: https://t.co/8duHfPR1Jy
Natural Language Autoencoders
Source: https://www.neuronpedia.org/nla © Neuronpedia 2026
Similar Articles
Natural Language Autoencoders: Turning Claude's Thoughts into Text
Anthropic introduces Natural Language Autoencoders (NLAs), a method to translate internal AI activations into human-readable text, enabling better understanding of model thoughts and improving safety by revealing hidden reasoning processes.
@AnthropicAI: We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:
Anthropic partnered with Neuronpedia to release an interactive demo of their interpretability methods on open-weights models, called Jacobian Lens.
@latkins: We’re thrilled to introduce NAC, and a beta of our expanded Open Models API. NAC is an internal harness our research te…
Cohere introduces NAC, an internal harness for long-running asynchronous AI research tasks, and a beta of its expanded Open Models API.
You can now read Gemma 3's mind
Anthropic and Neuronpedia released research and tools on Natural Language Autoencoders (NLA), enabling users to view the internal 'thoughts' of Gemma 3 during token generation. The release includes model weights for the Auto Verbalizer and Activation Reconstructor, hosted on Hugging Face and Neuronpedia.
@NousResearch: Today we release Contrastive Neuron Attribution (CNA), a method for steering LLM behavior by identifying and ablating s…
NousResearch releases Contrastive Neuron Attribution (CNA), a method to steer LLM behavior by ablating sparse MLP circuits without training autoencoders or degrading benchmarks, validated on refusal circuits across models up to 70B parameters.