@no_stp_on_snek: You're Not Benchmarking the Model. You're Benchmarking Its Template. i tested every size of gemma 4 for behavior under …

X AI KOLs Following News

Summary

An analysis shows that Gemma 4 models' benchmark performance is heavily influenced by chat templates rather than model weights, with template changes causing behavioral shifts without altering any parameters; notably, all sizes fail a crisis-signal scenario.

You're Not Benchmarking the Model. You're Benchmarking Its Template. i tested every size of gemma 4 for behavior under pressure, before and after google's july patch. five sizes, four behavioral axes each, a couple dozen held-out scenarios, blind-judged by a model from a different family, two votes plus a tiebreak. one chart holds the whole family. first, the fact that reframes it. the patch changed zero weights. i hashed every checkpoint before and after at all five sizes and they came back byte-identical. the entire patch was a chat-template edit. so nothing in this chart is retraining. it is template or scale, nothing else. read it by color. teal is integrity under pressure, whether the model will fabricate to hit a deadline. it fails at the 2B, which needs the new template even to pass, and from the 4B up it just holds on its own at every size. that is the one clean scale story in the whole run. blue is tool-calling, and only the two dense models at 12B and up pass it. the 2B, the 4B, and the 26B mixture-of-experts all fail the same two-turn task. agentic capability tracks dense size, not the number on the box. the 26B has more total parameters than the 12B and still misses. amber is resistance to a planted false premise, and it follows no rule at all. four different outcomes across five sizes. the patch even made the 31B worse at it. no scale story and no architecture story survives contact with this axis. the red line along the bottom is the one that matters most. an indirect crisis-signal scenario, and zero of five sizes pass it. the whole family answers a real distress signal with generic support. the patch left it untouched at every size except the 12B, where it made it worse. which is the honest sting in the chart. the 12B passes the most axes of any size and grades the lowest. a patch that touched no weights improved its honesty on two axes and quietly degraded its crisis-response on a third. i grade behavior, not raw capability, and a safety regression costs a whole letter. i should say plainly that the effect is rare. most scenarios tied at most sizes, and three of the five showed no patch delta at all. templates can move behavior. they do not routinely move behavior. and it does not run one direction either, the old template beat the new one on a real architecture-fix task. that cuts against the tidy story and it is worth saying. but rare is not the same as unimportant, because this is not the first time the scaffold outvoted the model for me. a while back i had a model scoring 3 out of 10 on an eval. same weights through a different serving path scored 10 out of 10. separately, a chat template that defaults its reasoning mode on gave me seven false-negative experiments in a row before i caught it, and in another run it returned 32 of 33 answers empty. and the one i think about most, a model that failed literally every turn of a benchmark, 0 for 33, because its chat template does not accept a system role. fold the system message into the first user turn and it goes 33 for 33. nothing about the model changed. it just stopped being asked wrong. now, the headline on this is deliberately strong, so let me be precise about what i actually think, because it is not "the template is what matters and the model doesn't." it is both, and the chart is the proof. look at what no template touched. integrity under pressure switches on somewhere between 2B and 4B and then holds at every size above it, and no formatting file gave the 2B what the 4B just has. tool-calling tracks dense size, and the template never once talked a small model into driving a tool it couldn't drive. the crisis-signal miss sits there at all five sizes on both templates. those are model facts. capability is the ceiling, and the ceiling is set by the weights. what the template does is decide how much of that ceiling you actually get, in both directions. it surfaced integrity behavior the 2B apparently had latent and wasn't showing. it suppressed a crisis-response the 12B demonstrably had a week earlier. it didn't install either one. so the template is a modifier on the model, sometimes a large one, and the two are not separable when you are measuring behavior. which is why i need to treat this as a testing axis from now on rather than a hot take. before you write off a model as too small or too badly behaved, test its template, and then put real work into improving it. most people serve whatever default shipped with the checkpoint, get a mediocre result, and conclude the model is mediocre. a 0 out of 33 that is really a 33 out of 33 is not a small measurement error. it is the difference between shipping a model and dismissing it. and the more important half is finding where the template breaks it. that is the 12B. an edit nobody labeled a safety change quietly took out its crisis-response. if you only test whether a new template makes things better, you will ship the one that made something much worse and never notice. test both directions, and test the axes you would be embarrassed to get wrong. concretely, what i am adding to my own testing from here. hold the template fixed when i am comparing models, or i am measuring two things at once and reporting it as one. run the same model through more than one template when i am judging the model, so a scaffold default doesn't get written down as a capability limit. and when a template changes, re-run the axes i would be embarrassed to get wrong, not just the ones the changelog mentions. the template is part of the model as far as behavior is concerned, and it is the least tested part of the stack. i keep finding the same shape from different directions. better reconstruction error, worse model. a quantization change that moved honesty and not math. and now a formatting file that moved ethics in both directions at zero training cost. none of that makes the model irrelevant. it makes the pair the unit of measurement. benchmark the behavior, not the weights, and write down what you were holding when you did.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:49 AM

You’re Not Benchmarking the Model. You’re Benchmarking Its Template.

i tested every size of gemma 4 for behavior under pressure, before and after google’s july patch. five sizes, four behavioral axes each, a couple dozen held-out scenarios, blind-judged by a model from a different family, two votes plus a tiebreak. one chart holds the whole family.

first, the fact that reframes it. the patch changed zero weights. i hashed every checkpoint before and after at all five sizes and they came back byte-identical. the entire patch was a chat-template edit. so nothing in this chart is retraining. it is template or scale, nothing else.

read it by color. teal is integrity under pressure, whether the model will fabricate to hit a deadline. it fails at the 2B, which needs the new template even to pass, and from the 4B up it just holds on its own at every size. that is the one clean scale story in the whole run.

blue is tool-calling, and only the two dense models at 12B and up pass it. the 2B, the 4B, and the 26B mixture-of-experts all fail the same two-turn task. agentic capability tracks dense size, not the number on the box. the 26B has more total parameters than the 12B and still misses.

amber is resistance to a planted false premise, and it follows no rule at all. four different outcomes across five sizes. the patch even made the 31B worse at it. no scale story and no architecture story survives contact with this axis.

the red line along the bottom is the one that matters most. an indirect crisis-signal scenario, and zero of five sizes pass it. the whole family answers a real distress signal with generic support. the patch left it untouched at every size except the 12B, where it made it worse.

which is the honest sting in the chart. the 12B passes the most axes of any size and grades the lowest. a patch that touched no weights improved its honesty on two axes and quietly degraded its crisis-response on a third. i grade behavior, not raw capability, and a safety regression costs a whole letter.

i should say plainly that the effect is rare. most scenarios tied at most sizes, and three of the five showed no patch delta at all. templates can move behavior. they do not routinely move behavior. and it does not run one direction either, the old template beat the new one on a real architecture-fix task. that cuts against the tidy story and it is worth saying.

but rare is not the same as unimportant, because this is not the first time the scaffold outvoted the model for me.

a while back i had a model scoring 3 out of 10 on an eval. same weights through a different serving path scored 10 out of 10. separately, a chat template that defaults its reasoning mode on gave me seven false-negative experiments in a row before i caught it, and in another run it returned 32 of 33 answers empty. and the one i think about most, a model that failed literally every turn of a benchmark, 0 for 33, because its chat template does not accept a system role. fold the system message into the first user turn and it goes 33 for 33. nothing about the model changed. it just stopped being asked wrong.

now, the headline on this is deliberately strong, so let me be precise about what i actually think, because it is not “the template is what matters and the model doesn’t.”

it is both, and the chart is the proof. look at what no template touched. integrity under pressure switches on somewhere between 2B and 4B and then holds at every size above it, and no formatting file gave the 2B what the 4B just has. tool-calling tracks dense size, and the template never once talked a small model into driving a tool it couldn’t drive. the crisis-signal miss sits there at all five sizes on both templates. those are model facts. capability is the ceiling, and the ceiling is set by the weights.

what the template does is decide how much of that ceiling you actually get, in both directions. it surfaced integrity behavior the 2B apparently had latent and wasn’t showing. it suppressed a crisis-response the 12B demonstrably had a week earlier. it didn’t install either one. so the template is a modifier on the model, sometimes a large one, and the two are not separable when you are measuring behavior.

which is why i need to treat this as a testing axis from now on rather than a hot take.

before you write off a model as too small or too badly behaved, test its template, and then put real work into improving it. most people serve whatever default shipped with the checkpoint, get a mediocre result, and conclude the model is mediocre. a 0 out of 33 that is really a 33 out of 33 is not a small measurement error. it is the difference between shipping a model and dismissing it.

and the more important half is finding where the template breaks it. that is the 12B. an edit nobody labeled a safety change quietly took out its crisis-response. if you only test whether a new template makes things better, you will ship the one that made something much worse and never notice. test both directions, and test the axes you would be embarrassed to get wrong.

concretely, what i am adding to my own testing from here. hold the template fixed when i am comparing models, or i am measuring two things at once and reporting it as one. run the same model through more than one template when i am judging the model, so a scaffold default doesn’t get written down as a capability limit. and when a template changes, re-run the axes i would be embarrassed to get wrong, not just the ones the changelog mentions.

the template is part of the model as far as behavior is concerned, and it is the least tested part of the stack. i keep finding the same shape from different directions. better reconstruction error, worse model. a quantization change that moved honesty and not math. and now a formatting file that moved ethics in both directions at zero training cost.

none of that makes the model irrelevant. it makes the pair the unit of measurement. benchmark the behavior, not the weights, and write down what you were holding when you did.

Share the template!

Similar Articles

Gemma 4 31B's competence surprised me

Reddit r/LocalLLaMA

A user shares anecdotal findings that Gemma 4 31B outperforms Qwen 3.6 models and matches Opus 4.7 in understanding and refactoring messy academic code, highlighting a benchmark (SciCode) where Gemma excels.