@googlegemma: We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is …

X AI KOLs Timeline Models

Summary

Google Gemma is rolling out significant improvements to Gemma 4, driven by community feedback and contributions, as detailed in a thread.

We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: 🧵👇 https://t.co/SMIGbJaUZg
Original Article
View Cached Full Text

Cached at: 07/16/26, 02:19 PM

We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions!

Here is a breakdown of what’s being fixed and updated in this release:

Flash Attention: We’ve enabled uniform Flash Attention 4 (FA4) support on NVIDIA Hopper GPUs!

Expect prefill throughput to jump by 25-70% and time-to-first-token (TTFT) to drop by up to 31%. (2/5)

Chat Template: Smoother conversational formatting.

Tool Calling: Patched issues for accurate, consistent tool execution. (3/5)

Vision Options: Want to make Gemma see even better?

The default vision bucket is 280 for token efficiency. To capture maximum detail (like sharp OCR and 2.51MP resolution), manually bump max_soft_tokens to 1120!

Try our new interactive Space to see how it works: https://huggingface.co/spaces/google/gemma4_vision_token_budget…

Less Laziness: We’ve significantly reduced edge cases where the model was holding back or cutting answers short, leading to more complete responses. (4/5)

A huge shoutout to the community for submitting fixes and finding new ways to make Gemma even better. We couldn’t do this without you!

Ready to test the speedup? Download the latest Gemma 4 updates now on Hugging Face: https://huggingface.co/collections/google/gemma-4… (5/5)

Similar Articles

Do you want new Gemma?

Reddit r/LocalLLaMA

A teaser or inquiry about a new version of Google's Gemma model, suggesting an upcoming release.