gemma

Tag

Cards List
#gemma

No wonder Qwen and Gemma are so different

Reddit r/LocalLLaMA · 8h ago

A user shares an observation that Qwen and Gemma tokenize code very differently, with Qwen using far fewer tokens for the same HTML/JS input, which may explain differences in coding and language performance. They also note a potential retraining project by LiquidAI using a more efficient tokenizer.

0 favorites 0 likes
#gemma

Scotoma-2: Gemma4, but with less annoying slop and better writing.

Reddit r/LocalLLaMA · 2d ago Cached

Scotoma-2 is an updated fine-tune of Gemma-4-31B-it that reduces repetitive writing tics via targeted preference training while keeping the base model's intelligence. It is not uncensored but aims to produce cleaner, less annoying prose for roleplay.

0 favorites 0 likes
#gemma

10% faster decode with Q4_K MTP draft model with Gemma 4 31b

Reddit r/LocalLLaMA · 2d ago

A user reports that quantising the f16 MTP draft model to Q4_K for Gemma 4 31b gives roughly 10% faster decode (65 to 72 TPS) on dual 3090s compared to the default Q4_0, while Q2_K performs worse.

0 favorites 0 likes
#gemma

Gemma 4 31b AttnRes Project

Reddit r/LocalLLaMA · 3d ago

An independent developer updates the AttnRes project: replacing standard residual stream with attention-based routing, distilling from Gemma 4 31b via a weaning schedule and top-K logits, with plans for an Apache 2.0 community model.

0 favorites 0 likes
#gemma

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Reddit r/LocalLLaMA · 4d ago

User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.

0 favorites 0 likes
#gemma

Fast Gemma's Verified Inference Optimization Recipe (7 minute read)

TLDR AI · 5d ago Cached

The VIDRAFT team shares their verified state-of-the-art inference optimization recipe for the Fast Gemma Challenge, achieving 510.58 TPS on a single A10G with PPL 2.39 using a fully public vLLM-based config.

0 favorites 0 likes
#gemma

@rohanpaul_ai: New Google DeepMind Paper. Model weights do not have to remain static artifacts that can only be fine-tuned or averaged…

X AI KOLs Following · 6d ago Cached

A new Google DeepMind paper, SkillSmith, treats prefix key-value caches as an input modality, composing textual knowledge and existing weights into a fresh prefix cache for a frozen Gemma 3 4B model at inference time, improving adaptation without a target-specific training run.

0 favorites 0 likes
#gemma

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

Reddit r/LocalLLaMA · 6d ago

A user lets a small LLM (Gemma4-31b) run on a laptop for a day to analyze r/LocalLLaMA, concluding that brilliant open-weight research exists but is buried under benchmark drama and hardware flexes.

0 favorites 0 likes
#gemma

Coding Diffusion Gemma from scratch

Reddit r/ArtificialInteligence · 2026-07-29

Tutorial on implementing a diffusion model based on Google's Gemma architecture from scratch.

0 favorites 0 likes
#gemma

Do you want new Gemma?

Reddit r/LocalLLaMA · 2026-07-26

A teaser or inquiry about a new version of Google's Gemma model, suggesting an upcoming release.

0 favorites 0 likes
#gemma

@sundarpichai: 1B next!

X AI KOLs Following · 2026-07-25 Cached

Sundar Pichai celebrates Gemma model family reaching 900M downloads, anticipating 1 billion.

0 favorites 0 likes
#gemma

@_philschmid: This week, Gemma surpassed 900 million downloads!

X AI KOLs Following · 2026-07-25 Cached

Gemma has surpassed 900 million downloads, marking a significant milestone for the AI model.

0 favorites 0 likes
#gemma

Getting the most out of MTP

Reddit r/LocalLLaMA · 2026-07-24

A guide on optimizing MTP (Multi-Token Prediction) performance by tuning n_max parameter, with benchmark results for various models like Gemma-31b and Qwen on P100 and V100 GPUs.

0 favorites 0 likes
#gemma

Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

arXiv cs.LG · 2026-07-24 Cached

This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.

0 favorites 0 likes
#gemma

When a translation model starts solving the problem instead of translating it (small rant)

Reddit r/LocalLLaMA · 2026-07-22

The author recounts issues when using Gemma models to translate reasoning traces, where the model executes instructions in the text instead of translating, highlighting a boundary failure between instruction and payload.

0 favorites 0 likes
#gemma

@googledevs: Building trustable AI for environments where failure isn't an option. At Sonoma Raceway, GDEs deployed an AI Race Coach…

X AI KOLs Following · 2026-07-22 Cached

Google Developer Experts deployed an AI Race Coach using Antigravity, Gemini, and Gemma on Pixel 10 at Sonoma Raceway, delivering split-second telemetry coaching at 100 mph to demonstrate trustable AI for high-stakes environments.

0 favorites 0 likes
#gemma

Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)

Reddit r/LocalLLaMA · 2026-07-21

Gemma-4-26B-a4B, with an updated chat template, outperforms Qwen3.6-MoE and Qwen3.5-MoE in fine-tuned instruct mode and reasoning efficiency.

0 favorites 0 likes
#gemma

Benchmarked Dense gemma-4-31b-it vs MoE gemma-4-26b-a4b-it to see if the cost reduction holds up in practice

Reddit r/AI_Agents · 2026-07-21

A practical benchmark comparing Gemma dense (31B) and MoE (26B) models shows MoE is 25.5% faster and 20% cheaper per query with identical quality, validating theoretical cost savings.

0 favorites 0 likes
#gemma

@heyshrutimishra: Sundar Pichai just reminded everyone that Google was built on open source. He personally worked on Chromium, Android, a…

X AI KOLs Following · 2026-07-17 Cached

Sundar Pichai reminded that Google was built on open source and applies the same philosophy to AI, with Gemma models designed for edge devices, while noting that frontier models require massive capital investment.

0 favorites 0 likes
#gemma

@h100envy: Google engineer explained how to fine-tune a tiny LLM from 46% to 90% accuracy on your phone in 21 minutes - better tha…

X AI KOLs Timeline · 2026-07-16 Cached

A Google engineer shares a method to fine-tune a Gemma 270M model from 46% to 90% accuracy in 21 minutes on a phone, using synthetic data, LoRA, int4 quantization, achieving 2000 tokens per second offline.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback