Tag
User reports improved tokens per second (tps) with ThinkingCap-Qwen3.6-27B compared to Qwen3.5-27B, with no quality loss, recommending it as a daily driver until the next Qwen release.
DeepSeek has reduced inference costs by approximately 5x while maintaining 50 tokens per second throughput.
Achieved 1000 tokens per second generation on Qwen3.6 27B using V100 GPUs with 128 concurrent requests, and 80 t/s for single user.