RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA News

Summary

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

No content available
Original Article

Similar Articles

DiffusionGemma 26B A4B results on my 5090

Reddit r/LocalLLaMA

This post presents benchmark results and tuning parameters for running DiffusionGemma 26B A4B GGUF models on an RTX 5090 GPU, showing up to 44% speedup via optimized temperature settings and quantization choices.