diffusiongemma

Tag

Cards List
#diffusiongemma

Auditing DiffusionGemma Transparency (9 minute read)

TLDR AI · 2026-06-22 Cached

An analysis of how transparent Google's DiffusionGemma model release is, discussing the implications for AI safety and accountability.

0 favorites 0 likes
#diffusiongemma

DiffusionGemma 26b on a 4090 at up to 475t/s... and some thoughts...

Reddit r/LocalLLaMA · 2026-06-18

A user shares their experience running DiffusionGemma 26B on a 4090 GPU via vLLM, achieving up to 475t/s but noting drawbacks like single-user limitation, lower accuracy, and short context, concluding it's not worth using over the regular 26B model.

0 favorites 0 likes
#diffusiongemma

Can we stop dunking on DiffusionGemma and hack it instead?

Reddit r/LocalLLaMA · 2026-06-14

Discusses various methods to optimize DiffusionGemma inference, reduce hallucination, and improve performance for tool use and agents, including entropy-bounded sampling, schema scaffolding, and retrieval during denoising.

0 favorites 0 likes
#diffusiongemma

DifussionGemma 4 on 4x7900xtx

Reddit r/LocalLLaMA · 2026-06-11

Reports running DiffusionGemma 26B on four AMD 7900 XTX GPUs using vllm, achieving 100 tps generation with overall 45-60 t/s, sharing performance metrics and setup commands.

0 favorites 0 likes
#diffusiongemma

DiffusionGemma 26B A4B results on my 5090

Reddit r/LocalLLaMA · 2026-06-11

This post presents benchmark results and tuning parameters for running DiffusionGemma 26B A4B GGUF models on an RTX 5090 GPU, showing up to 44% speedup via optimized temperature settings and quantization choices.

0 favorites 0 likes
#diffusiongemma

@HuggingPapers: NVIDIA just released an NVFP4-quantized DiffusionGemma on Hugging Face A 26B MoE multimodal model generating text via p…

X AI KOLs Following · 2026-06-10 Cached

NVIDIA released a 26B MoE multimodal model called DiffusionGemma on Hugging Face, using NVFP4 quantization and achieving over 1,100 tokens per second on Hopper hardware.

0 favorites 0 likes
#diffusiongemma

NVIDIA Accelerates Google DeepMind’s DiffusionGemma for Local AI

NVIDIA Blog · 2026-06-10 Cached

NVIDIA optimizes Google DeepMind's DiffusionGemma, an open model that generates text in parallel 256-token blocks, achieving up to 4x faster performance on local RTX GPUs, DGX Spark, and DGX Station systems.

0 favorites 0 likes
#diffusiongemma

DiffusionGemma: The Developer Guide- Google Developers Blog

Reddit r/LocalLLaMA · 2026-06-10 Cached

DiffusionGemma is a new experimental model from Google DeepMind that uses parallel generation on a 256-token canvas, achieving up to 4x faster token generation on GPUs. This developer guide explains its architecture, bidirectional context, and includes a fine-tuning recipe for solving Sudoku.

0 favorites 0 likes
#diffusiongemma

unsloth/diffusiongemma-26B-A4B-it-GGUF

Hugging Face Models Trending · 2026-06-10 Cached

Unsloth releases GGUF quantizations of Google DeepMind's DiffusionGemma (26B-A4B), a new block-diffusion architecture for faster text generation, ready for llama.cpp.

0 favorites 0 likes
← Back to home

Submit Feedback