Tag
Author Alexey Fateev dissected the Qwen3.8-27B AI model to carve out a MoE structure through zero-training weight surgery, finding only two neurons active on over 90% of tokens.
This paper investigates whether repetition loops in long factual enumeration tasks by Gemma 4 models can be fixed by editing a single neuron. It finds that targeted weight edits on a small set of MLP neurons can significantly reduce loop failures, though not completely eliminate doom looping in larger models.