We removed an LM's ability to speak German (3 minute read)
Summary
GoodfireAI releases a research agenda on understanding neural geometry in language models, demonstrating the ability to precisely control a model's capabilities, such as removing its ability to speak German.
View Cached Full Text
Cached at: 06/26/26, 05:10 PM
Similar Articles
@GoodfireAI: Neural networks might speak English, but they think in shapes. Understanding their rich *neural geometry* is key to und…
Goodfire AI announces a new research agenda focused on neural geometry to improve the understanding, debugging, and control of neural networks.
LLM Neuroanatomy III - LLMs seem to think in geometry, not language
Researcher analyzes LLM internal representations across 8 languages and multiple models, finding that concept thinking occurs in geometric space in middle transformer layers independent of input language, supporting a universal deep structure hypothesis similar to Chomsky's theory rather than Sapir-Whorf linguistic relativism.
Natively Unlearnable Large Language Models
The paper proposes NULLs (Natively Unlearnable LLMs), a model class that isolates source-specific contributions in sparsely activated sinks while sharing backbone neurons, enabling clean unlearning of individual data sources without retraining and preserving general language capabilities.
You can now read Gemma 3's mind
Anthropic and Neuronpedia released research and tools on Natural Language Autoencoders (NLA), enabling users to view the internal 'thoughts' of Gemma 3 during token generation. The release includes model weights for the Auto Verbalizer and Activation Reconstructor, hosted on Hugging Face and Neuronpedia.
@mtschannen: For the past years my research focus was on unifying models and training paradigms across modalities. Today I'm excited…
Google DeepMind researcher announces the release of Gemma 4 12B, a dense encoder-free model that processes text, image, and audio inputs, continuing work on unifying models across modalities.