Tag
An analysis shows that Gemma 4 models' benchmark performance is heavily influenced by chat templates rather than model weights, with template changes causing behavioral shifts without altering any parameters; notably, all sizes fail a crisis-signal scenario.