slimserve

Tag

Cards List
#slimserve

@QuixiAI: LESSON LEARNED: Always use BF16 kv cache. I was using turboquant. Yeah the VRAM consumption sucks - so another trick is…

X AI KOLs Timeline · 4d ago Cached

The author shares a technical lesson on using BF16 KV cache instead of turboquant for AI model optimization and implements CPU offloading in SlimServe for Qwen models to manage VRAM consumption.

0 favorites 0 likes
#slimserve

@QuixiAI: I got DeepSeek v4 Flash 0731 running on 4x A100 with SlimServe. 175 tok/s for single-request 1k tok/s for 64 concurrent…

X AI KOLs Timeline · 2026-08-10 Cached

QuixiAI reports running DeepSeek v4 Flash 0731 on 4x A100 with SlimServe, achieving 175 tok/s for single requests and 1k tok/s for 64 concurrent requests.

0 favorites 0 likes
← Back to home

Submit Feedback