Tag
The post announces new quantized versions of the Qwen3.8 Flash model that outperform other high-quality quants through modified KLD measurement and optimized prefill performance.
The author compares various GPUs for LLM inference, critiquing common benchmarks and emphasizing the importance of prefill performance over generation speed, offering recommendations for different budgets and use cases.