@analogalok: my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 …

X AI KOLs Timeline News

Summary

User runs Gemma 4 31B dense model on 8GB VRAM gaming laptop at ~3 tokens/sec using llama.cpp with MTP speculative decoding, demonstrating feasibility of running a 31B dense model on consumer hardware and proposing agentic workflows where a fast MoE model routes to this slower dense model for hard tasks.

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.
Original Article
View Cached Full Text

Cached at: 06/17/26, 01:45 AM

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it.

ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM

last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp’s new MTP support. smooth. snappy.

but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that’s why it flies.

so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine?

Hardware:

GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest.

The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model’s unsloth huggingface link in the comments)

This is Google DeepMind’s flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded.

the flags I used:

-m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv –spec-type draft-mtp –spec-draft-model mtp-gemma-4-31B-it.gguf –spec-draft-n-max 8 –spec-draft-p-min 0.6 -c 6000 -v

Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup.

Results:

→ Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes

so is 3 tps actually usable? For real time back and forth chat? Not ideal. You’re not having a fluid conversation at 3 tps.

but slow ≠ useless. And this is where it gets genuinely interesting.

think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior.

That’s exactly the local AI agent architecture this unlocks:

→ Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev.

→ Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That’s a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus)

other workflows where 3 tps is completely fine:

  • overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results.
  • One shot deep reasoning
  • Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints
  • Any workflow where output quality > output speed

A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we’re doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting.

Google DeepMind said the Gemma 4 31B targets “consumer GPUs and workstations.” They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping.

the tools are here. the models are here. you just have to be willing to abuse your laptop a little.

what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.

unsloth/gemma-4-31B-it-qat-GGUF · Hugging Face

How to Run MTP Models: Multi-Token Prediction Guide | Unsloth Documentation

exactly this. robotics, animation rigs, prosthetics motion planning, any domain where the solver is iterative and correctness matters more than latency. the overnight batch paradigm completely reframes what 3 tps means. it’s not slow, it’s just async. sleep is free compute time and a 31B dense model doesn’t care that you’re not watching.

with llama.cpp: run llama-server instead of llama-cli. each model gets its own server instance pointed at a different GGUF via -m. your app or agent hits whichever model it needs.

on a single 8GB VRAM card, you can only have one loaded at a time, so the swap means killing one server and spinning up the other.

the fast model signals delegation via a tool call, your host script/harness intercepts it, swaps the server, passes relevant context to the heavier model, gets the result back. this will need some tuning with a single card but would work flawlessly with 2 8 gb cards.

Similar Articles

24+ tok/s from ~30B MoE models on an old GTX 1080 (8 GB VRAM, 128k context)

Reddit r/LocalLLaMA

A developer demonstrates running MoE models like Qwen 3.6 35B-A3B and Gemma 4 26B-A4B at 24+ tok/s on an old GTX 1080 (8GB VRAM) with 128k context using llama.cpp with MoE offloading and TurboQuant KV cache quantization, revealing optimization tricks for Gemma's MTP speculative decoding.

You don't need a GPU to run gemma-4-26B-A4B

Reddit r/LocalLLaMA

The author demonstrates that the Gemma-4-26B-A4B model runs efficiently on a CPU-only system using Koboldcpp, achieving 7 tokens per second on an old desktop, suggesting that powerful GPUs may not be necessary for local LLM inference.