@_philschmid: We just launched a Gemma 4 12B! Our first mid-sized model with native audio inputs. Gemma 4 12 B is a unified, encoder-…

X AI KOLs Following Models

Summary

We just launched Gemma 4 12B, a mid-sized multimodal model with native audio inputs, requiring only 16GB memory and released under Apache 2.0.

We just launched a Gemma 4 12B! Our first mid-sized model with native audio inputs. Gemma 4 12 B is a unified, encoder-free multimodal model. 🧠 vision and audio directly into the LLM. 💻 Just need 16GB of memory. 📊 Benchmark nearing 26B. 📄 Apache 2.0. https://t.co/o7sKQBHoWx
Original Article
View Cached Full Text

Cached at: 06/03/26, 05:53 PM

We just launched a Gemma 4 12B! Our first mid-sized model with native audio inputs. Gemma 4 12 B is a unified, encoder-free multimodal model.

🧠 vision and audio directly into the LLM. 💻 Just need 16GB of memory. 📊 Benchmark nearing 26B. 📄 Apache 2.0. https://t.co/o7sKQBHoWx

Similar Articles

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Google DeepMind Blog

Google DeepMind announces Gemma 4 12B, a novel encoder-free multimodal AI model that integrates vision and audio directly into the LLM backbone, delivering advanced reasoning and agentic capabilities on laptops with 16GB of RAM, released under Apache 2.0 license.