rtx5090

Tag

Cards List
#rtx5090

Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

Reddit r/LocalLLaMA · 20h ago

Nifer is a tool that achieves 700 tokens per second inference on Qwen 3.6 35B without thinking, specifically optimized for the RTX 5090, with support for full 250k context.

0 favorites 0 likes
#rtx5090

DiffusionGemma 26B A4B results on my 5090

Reddit r/LocalLLaMA · 2026-06-11

This post presents benchmark results and tuning parameters for running DiffusionGemma 26B A4B GGUF models on an RTX 5090 GPU, showing up to 44% speedup via optimized temperature settings and quantization choices.

0 favorites 0 likes
#rtx5090

@witcheer: everyone says NVFP4 makes blackwell cards "faster." I benchmarked Qwen3.6-27B three ways on my 5090: >NVFP4 >plain Q4_K…

X AI KOLs Timeline · 2026-06-05 Cached

A benchmark of NVFP4 on an RTX 5090 with Qwen3.6-27B shows prefill speed gains of 32-42% over equal-bit Q4_K_M and 52-68% over Q6_K, but decode gains are modest (+9% vs Q4) as decode is memory-bandwidth bound. The quality loss compared to Q6 is minimal (-0.8 average), making NVFP4 a good choice for local inference.

0 favorites 0 likes
#rtx5090

I tested Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.5-27B and Gemma 4 on the same real architecture-writing task on an RTX 5090

Reddit r/LocalLLaMA · 2026-04-23

A hands-on benchmark of four local LLMs—Qwen3.6-27B, Qwen3.6-35B, Qwen3.5-27B and Gemma 4—on a 20k-token architecture-writing task shows Qwen3.6-27B delivering the best overall balance of clarity, completeness and usefulness on an RTX 5090.

1 favorites 1 likes
← Back to home

Submit Feedback