Tag
Anthropic releases open source interactive demos built on top of Qwen models, in partnership with Neuronpedia.
Anthropic partnered with Neuronpedia to release an interactive demo of their interpretability methods on open-weights models, called Jacobian Lens.
Anthropic and Neuronpedia released research and tools on Natural Language Autoencoders (NLA), enabling users to view the internal 'thoughts' of Gemma 3 during token generation. The release includes model weights for the Auto Verbalizer and Activation Reconstructor, hosted on Hugging Face and Neuronpedia.
Anthropic and Neuronpedia have partnered to release Natural Language Autoencoders (NLAs) on open models, allowing researchers to gain hands-on experience with this interpretability tool.