Glimmer: 233.4 tps on 5090 with Dflash
Summary
Glimmer reportedly hits 233.4 tps on an RTX 5090 using Dflash, with 256k context fitting on 24GB VRAM, sparking excitement among users.
Similar Articles
Achievable 253 t/s - unsloth/Muse Glimmer 30B UD-Q5_K_M on a 5090
Benchmarks unsloth's Muse Glimmer 30B on an RTX 5090 with speculative decoding, achieving up to 253 t/s using a DFlash draft model and a GPU-based argmax PR, though the PR is still a draft.
Gemma 4 26B Hits 600 Tok/s on One RTX 5090
A benchmark shows that using vLLM with DFlash speculative decoding boosts Gemma 4 26B inference to ~578 tokens per second on a single RTX 5090, achieving a 2.56x speedup over baseline.
Muse Glimmer ACTUALLY fits on a single RTX 3090
User reports that Muse Glimmer, a 30B model, fits on a single RTX 3090 with full 256k context using Q4_K_XL quantization and DFlash, achieving 64-124 tok/s and perfect long-context retrieval, unlike comparable models.
BeeLlama v0.2.0 – major DFlash update. Single RTX 3090: Qwen 3.6 27B up to 164 tps (4.40x), Gemma 4 31B up to 177.8 tps (4.93x). Prompt processing speed near baseline.
BeeLlama v0.2.0 introduces major DFlash speculative decoding improvements, achieving up to 4.93x speedup on single RTX 3090 for Gemma 4 31B and 4.40x for Qwen 3.6 27B, with prompt processing near baseline.
@0xSero: GLM-5.1-478B-NVFP4 Running on: - 4x RTX Pro 6000 - Sglang - 370,000 max tokens (1.75x full context) - p10 27.7 | p90 45…
A quantized 478B-parameter GLM-5.1 model runs on 4×RTX Pro 6000 GPUs via SGLang, delivering 370k-token context at up to 45 tok/s decode and 1340 tok/s prefill, and is demoed driving Figma.