Tag
The author shares a technical lesson on using BF16 KV cache instead of turboquant for AI model optimization and implements CPU offloading in SlimServe for Qwen models to manage VRAM consumption.
QuixiAI reports running DeepSeek v4 Flash 0731 on 4x A100 with SlimServe, achieving 175 tok/s for single requests and 1k tok/s for 64 concurrent requests.