NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM
Summary
The author has updated their NInfer fork with rk2v4-e8 quantization for the KV cache, enabling up to 250-350K tokens context on a single RTX 4090 without system RAM spill and with optimizations for faster generation speeds.
Similar Articles
NInfer day0 support for Qwen3.8 27b: ~200 tok/s generation, with tons of engine improvments
NInfer adds day-0 support for the Qwen3.8-27B model, achieving around 200 tokens per second on a single RTX 5090 with speculative decoding, and includes engine improvements like concurrent requests and kernel optimizations.
This is amazing. Token speed doubled + kv cache now need low vram - qwen 27b
A new KV cache optimization called kvflash doubles generation speed and reduces VRAM usage for Qwen 3.6-27B on a single RTX 3090 while maintaining accuracy.
Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
Nifer is a tool that achieves 700 tokens per second inference on Qwen 3.6 35B without thinking, specifically optimized for the RTX 5090, with support for full 250k context.
@DeepTechTR: Qwen 3.6 27B is incredibly fast with 16 GB VRAM! The impact of Pure Quant The era of the 27B model that runs seamlessly…
Qwen 3.6 27B runs fast on 16 GB VRAM thanks to 'Pure Quant' technology, achieving 40 tokens/s with MTP and supporting 64k contexts, enabling local AI on consumer GPUs like RTX 4060 Ti.
2 old RTX 2080 Ti with 22GB vram each Qwen3.6 27B at 38 token/s with f16 kv cache
A user shares their setup using two modded RTX 2080 Ti GPUs with 22GB VRAM each to run Qwen 3.6 27B at 38 tokens/s with llama.cpp, including tips on power limiting, tensor split mode, and KV cache settings.