Tag
A practical benchmark comparing Gemma dense (31B) and MoE (26B) models shows MoE is 25.5% faster and 20% cheaper per query with identical quality, validating theoretical cost savings.
Google released QAT (4-bit) versions of their Gemma 4 model series, including the 31B Dense and 26B MoE models, furthering open-source AI.