PSA: Gemma 4 12B is NOT completely broken for coding and tool calling, you need a special chat template
Summary
Gemma 4 12B has a known issue with tool calling and coding, but using a custom chat template in llama.cpp resolves the bugs. Users should compile llama.cpp from source and apply the fix before evaluating the model's coding ability.
Similar Articles
Tested (the updated) Gemma 4 locally on coding with OpenCode
Tested the updated Gemma 4 locally using llama.cpp on an M5 Pro, achieving 60 tokens/s for coding tasks with OpenCode; good for backend but poor for UI/UX.
Gemma 4 Chat Template now has preserve thinking
Google's Gemma 4 31B IT model now has a chat template fix that preserves thinking and improves null handling, reasoning preservation, and input validation.
Experimental "Preserve Thinking" Jinja Template for Gemma4 31B in llama.cpp
An experimental Jinja template for Gemma4 31B in llama.cpp that improves stability for multi-turn tool calls by fixing common thinking tag issues. Community feedback is welcome, but this is not recommended by Google.
@no_stp_on_snek: You're Not Benchmarking the Model. You're Benchmarking Its Template. i tested every size of gemma 4 for behavior under …
An analysis shows that Gemma 4 models' benchmark performance is heavily influenced by chat templates rather than model weights, with template changes causing behavioral shifts without altering any parameters; notably, all sizes fail a crisis-signal scenario.
Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)
Gemma-4-26B-a4B, with an updated chat template, outperforms Qwen3.6-MoE and Qwen3.5-MoE in fine-tuned instruct mode and reasoning efficiency.