ThinkingCap-Qwen3.6-27B warrants a look
Summary
User reports improved tokens per second (tps) with ThinkingCap-Qwen3.6-27B compared to Qwen3.5-27B, with no quality loss, recommending it as a daily driver until the next Qwen release.
Similar Articles
@_akhaliq: bottlecapai/ThinkingCap-Qwen3.6-27B Capability of Qwen3.6-27B with 50% less thinking tokens on average, and over 90% le…
bottlecapai releases ThinkingCap-Qwen3.6-27B, a finetuned version of Qwen3.6-27B that achieves 50% fewer thinking tokens on average and over 90% fewer in best cases, improving efficiency.
ThinkingCap 3.8-27B vs. Swift 3.8-27B vs. Qwen 3.8-27B Benchmarks
The article benchmarks ThinkingCap-Qwen3.8-27B and Swift-Qwen3.8-27B against the original Qwen3.8-27B, showing both fine-tunes reduce reasoning tokens by ~40% with minimal performance loss, though with differences in language-specific results and token usage patterns.
Qwen 3.8 27B is faster than expected
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Qwen 3.8 27B is a powerful open-source 27B parameter vision-capable LLM from Alibaba's Qwen research lab, praised for its benchmarks but criticized for defaulting to excessive reasoning effort, which slows down performance on consumer hardware.
Qwen3.8-27B: slower tokens, faster and better results
Qwen3.8-27B is a new AI model that emphasizes wall-clock time over token speed, offering superior intelligence for local deployment on consumer hardware with 32 GB VRAM.