Glimmer: 233.4 tps on 5090 with Dflash

Reddit r/LocalLLaMA Models

Summary

Glimmer reportedly hits 233.4 tps on an RTX 5090 using Dflash, with 256k context fitting on 24GB VRAM, sparking excitement among users.

That's insane yo! I haven't had a chance yet to test it on my 4090 at home but it sounds so promising. And read here that 256k CTX is easily reachable on 24gb unlike Qwen. Super excited!
Original Article

Similar Articles

Gemma 4 26B Hits 600 Tok/s on One RTX 5090

Reddit r/LocalLLaMA

A benchmark shows that using vLLM with DFlash speculative decoding boosts Gemma 4 26B inference to ~578 tokens per second on a single RTX 5090, achieving a 2.56x speedup over baseline.

Muse Glimmer ACTUALLY fits on a single RTX 3090

Reddit r/LocalLLaMA

User reports that Muse Glimmer, a 30B model, fits on a single RTX 3090 with full 256k context using Q4_K_XL quantization and DFlash, achieving 64-124 tok/s and perfect long-context retrieval, unlike comparable models.