qwen-ai

Tag

Cards List
#qwen-ai

NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM

Reddit r/LocalLLaMA · 7h ago

The author has updated their NInfer fork with rk2v4-e8 quantization for the KV cache, enabling up to 250-350K tokens context on a single RTX 4090 without system RAM spill and with optimizations for faster generation speeds.

0 favorites 0 likes
← Back to home

Submit Feedback