@witcheer: this is the first Qwen3.6-27B coding tune I've measured that improves real bug-fixing (!!!). - quality (MMLU/ARC/HellaS…
Summary
A community fine-tune of Qwen3.6-27B improves real bug-fixing on SWE-bench while maintaining quality, unlike synthetic distillations that regress.
View Cached Full Text
Cached at: 06/17/26, 04:00 PM
this is the first Qwen3.6-27B coding tune I’ve measured that improves real bug-fixing (!!!).
-
quality (MMLU/ARC/HellaSwag/GSM8K/HumanEval): 93.3 vs base 94.0. flat.
-
agentic score (native tool-calling, 40 tasks): 98.0 vs base 98.6. flat.
-
real bugs (SWE-bench Verified, 30, official harness): 20/30 vs base 19/30. up. it solves 2 the base can’t and gives up less (6 empty patches vs 8).
-
MTP drafter: 2.0 to 2.4x vs base 1.8 to 2.2x. the fine-tune kept its drafter.
this is the third Qwen3.6-27B coding tune I’ve benched. the other two were distilled on synthetic agent traces and both regressed on real bugs.
across all three the synthetic agentic score is a 2.4pt band (97.6 to 100) while real SWE spans 11 to 20.
the cheap axis can't tell them apart.
pi-tune even has the lowest quality of the group and the best real resolve. real capability tracks the training data, not the agentic coder label: real traces improved it, synthetic distill narrowed it.
only the reality anchor could see the difference.
> **Tongyi Lab (@Ali_TongyiLab):**
> We are pleased to highlight an excellent community model from developer : Qwen3.6-27B-MTP-pi-reasoning-GGUF.
>
> Built on our Qwen3.6-27B base model, this release focuses on optimizing automated programming and debugging workflows for local coding agents.
>
> If you are exploring local
Similar Articles
Alright, We got Qwen3.8-27B. Now it's community's turn to make it more better & faster
Community discussion on the release of Qwen3.8-27B, focusing on performance comparisons, memory usage, and creative writing capabilities with previous versions like Qwen3.6-27B and Qwen3.5-27B.
@malikwas1f: The first ever fine tune of @Alibaba_Qwen qwen 3.6 27b to lift up its quality. @migtissera https://github.com/noonghunn…
Announcement of the first fine-tune of Alibaba Qwen 3.6 27B to improve its quality, along with a GitHub repository (club-3090) providing recipes for running LLMs locally on RTX 3090s.
Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable
Fixed three bugs in a qMLX fork for running Qwen3.5-122B on Mac Studio, reducing prefill time from minutes to sub-seconds for long-context inference; open-sourced the fork and benchmark script.
Qwen 3.8 27b is strong even at Q3_xxs
The user finds Qwen 3.8 27b in Q3 quantization highly effective for coding tasks with fast inference speeds, outperforming previous models, despite minor issues in general conversations.
Qwen3.8-27B is identical to Qwen3.6-27B!
Qwen3.8-27B is identical in architecture to Qwen3.6-27B, with all capability gains attributed to training improvements.