gemma

Tag

Cards List
#gemma

Scotoma-2: Gemma4, but with less annoying slop and better writing.

Reddit r/LocalLLaMA · 2026-08-06 Cached

Scotoma-2 is an updated fine-tune of Gemma-4-31B-it that reduces repetitive writing tics via targeted preference training while keeping the base model's intelligence. It is not uncensored but aims to produce cleaner, less annoying prose for roleplay.

0 favorites 0 likes
#gemma

10% faster decode with Q4_K MTP draft model with Gemma 4 31b

Reddit r/LocalLLaMA · 2026-08-06

A user reports that quantising the f16 MTP draft model to Q4_K for Gemma 4 31b gives roughly 10% faster decode (65 to 72 TPS) on dual 3090s compared to the default Q4_0, while Q2_K performs worse.

0 favorites 0 likes
#gemma

Gemma 4 31b AttnRes Project

Reddit r/LocalLLaMA · 2026-08-05

An independent developer updates the AttnRes project: replacing standard residual stream with attention-based routing, distilling from Gemma 4 31b via a weaning schedule and top-K logits, with plans for an Apache 2.0 community model.

0 favorites 0 likes
#gemma

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Reddit r/LocalLLaMA · 2026-08-04

User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.

0 favorites 0 likes
#gemma

Fast Gemma's Verified Inference Optimization Recipe (7 minute read)

TLDR AI · 2026-08-04 Cached

The VIDRAFT team shares their verified state-of-the-art inference optimization recipe for the Fast Gemma Challenge, achieving 510.58 TPS on a single A10G with PPL 2.39 using a fully public vLLM-based config.

0 favorites 0 likes
#gemma

@rohanpaul_ai: New Google DeepMind Paper. Model weights do not have to remain static artifacts that can only be fine-tuned or averaged…

X AI KOLs Following · 2026-08-02 Cached

A new Google DeepMind paper, SkillSmith, treats prefix key-value caches as an input modality, composing textual knowledge and existing weights into a fresh prefix cache for a frozen Gemma 3 4B model at inference time, improving adaptation without a target-specific training run.

0 favorites 0 likes
#gemma

Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes.

Reddit r/LocalLLaMA · 2026-08-02

A user lets a small LLM (Gemma4-31b) run on a laptop for a day to analyze r/LocalLLaMA, concluding that brilliant open-weight research exists but is buried under benchmark drama and hardware flexes.

0 favorites 0 likes
#gemma

Coding Diffusion Gemma from scratch

Reddit r/ArtificialInteligence · 2026-07-29

Tutorial on implementing a diffusion model based on Google's Gemma architecture from scratch.

0 favorites 0 likes
#gemma

Do you want new Gemma?

Reddit r/LocalLLaMA · 2026-07-26

A teaser or inquiry about a new version of Google's Gemma model, suggesting an upcoming release.

0 favorites 0 likes
#gemma

@sundarpichai: 1B next!

X AI KOLs Following · 2026-07-25 Cached

Sundar Pichai celebrates Gemma model family reaching 900M downloads, anticipating 1 billion.

0 favorites 0 likes
#gemma

@_philschmid: This week, Gemma surpassed 900 million downloads!

X AI KOLs Following · 2026-07-25 Cached

Gemma has surpassed 900 million downloads, marking a significant milestone for the AI model.

0 favorites 0 likes
#gemma

Getting the most out of MTP

Reddit r/LocalLLaMA · 2026-07-24

A guide on optimizing MTP (Multi-Token Prediction) performance by tuning n_max parameter, with benchmark results for various models like Gemma-31b and Qwen on P100 and V100 GPUs.

0 favorites 0 likes
#gemma

Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

arXiv cs.LG · 2026-07-24 Cached

This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.

0 favorites 0 likes
#gemma

When a translation model starts solving the problem instead of translating it (small rant)

Reddit r/LocalLLaMA · 2026-07-22

The author recounts issues when using Gemma models to translate reasoning traces, where the model executes instructions in the text instead of translating, highlighting a boundary failure between instruction and payload.

0 favorites 0 likes
#gemma

@googledevs: Building trustable AI for environments where failure isn't an option. At Sonoma Raceway, GDEs deployed an AI Race Coach…

X AI KOLs Following · 2026-07-22 Cached

Google Developer Experts deployed an AI Race Coach using Antigravity, Gemini, and Gemma on Pixel 10 at Sonoma Raceway, delivering split-second telemetry coaching at 100 mph to demonstrate trustable AI for high-stakes environments.

0 favorites 0 likes
#gemma

Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)

Reddit r/LocalLLaMA · 2026-07-21

Gemma-4-26B-a4B, with an updated chat template, outperforms Qwen3.6-MoE and Qwen3.5-MoE in fine-tuned instruct mode and reasoning efficiency.

0 favorites 0 likes
#gemma

Benchmarked Dense gemma-4-31b-it vs MoE gemma-4-26b-a4b-it to see if the cost reduction holds up in practice

Reddit r/AI_Agents · 2026-07-21

A practical benchmark comparing Gemma dense (31B) and MoE (26B) models shows MoE is 25.5% faster and 20% cheaper per query with identical quality, validating theoretical cost savings.

0 favorites 0 likes
#gemma

@heyshrutimishra: Sundar Pichai just reminded everyone that Google was built on open source. He personally worked on Chromium, Android, a…

X AI KOLs Following · 2026-07-17 Cached

Sundar Pichai reminded that Google was built on open source and applies the same philosophy to AI, with Gemma models designed for edge devices, while noting that frontier models require massive capital investment.

0 favorites 0 likes
#gemma

@h100envy: Google engineer explained how to fine-tune a tiny LLM from 46% to 90% accuracy on your phone in 21 minutes - better tha…

X AI KOLs Timeline · 2026-07-16 Cached

A Google engineer shares a method to fine-tune a Gemma 270M model from 46% to 90% accuracy in 21 minutes on a phone, using synthetic data, LoRA, int4 quantization, achieving 2000 tokens per second offline.

0 favorites 0 likes
#gemma

I got Gemma 4 running directly inside Godot using only GDScript and Vulkan compute shaders

Reddit r/LocalLLaMA · 2026-07-13

A developer successfully integrated Gemma 4 AI model into the Godot game engine using only GDScript and Vulkan compute shaders, enabling local AI inference within games.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback