@googlegemma: We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is …
Summary
Google Gemma is rolling out significant improvements to Gemma 4, driven by community feedback and contributions, as detailed in a thread.
View Cached Full Text
Cached at: 07/16/26, 02:19 PM
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions!
Here is a breakdown of what’s being fixed and updated in this release:
Flash Attention: We’ve enabled uniform Flash Attention 4 (FA4) support on NVIDIA Hopper GPUs!
Expect prefill throughput to jump by 25-70% and time-to-first-token (TTFT) to drop by up to 31%. (2/5)
Chat Template: Smoother conversational formatting.
Tool Calling: Patched issues for accurate, consistent tool execution. (3/5)
Vision Options: Want to make Gemma see even better?
The default vision bucket is 280 for token efficiency. To capture maximum detail (like sharp OCR and 2.51MP resolution), manually bump max_soft_tokens to 1120!
Try our new interactive Space to see how it works: https://huggingface.co/spaces/google/gemma4_vision_token_budget…
Less Laziness: We’ve significantly reduced edge cases where the model was holding back or cutting answers short, leading to more complete responses. (4/5)
A huge shoutout to the community for submitting fixes and finding new ways to make Gemma even better. We couldn’t do this without you!
Ready to test the speedup? Download the latest Gemma 4 updates now on Hugging Face: https://huggingface.co/collections/google/gemma-4… (5/5)
Similar Articles
Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!
Google updated Gemma 4's chat templates with major fixes to tool calling, reduced laziness, enabled Flash Attention 4 on Hopper GPUs, and released an interactive vision guide. The updates are available on Hugging Face.
Talking with Gemma 4 31B!
Announcing Gemma 4 31B, a new large language model from Google.
Do you want new Gemma?
A teaser or inquiry about a new version of Google's Gemma model, suggesting an upcoming release.
@googlegemma: Gemma 4 up to 3x faster, directly in your phone! Check out the difference Speculative Decoding makes! Multi-Token Predi…
Google's Gemma 4 achieves up to 3x faster inference speeds through speculative decoding and multi-token prediction, enabling efficient on-device deployment.
@lmstudio: Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Avail…
Google released Gemma 4 12B, a dense mid-sized model that runs on laptops, under Apache 2.0, now available in LM Studio.