Qwen3.8 27B vs Qwen3.6 27B vs Qwen3.5 27B, a slight improvement in oneshotting ability across generations.

Reddit r/LocalLLaMA News

Summary

Comparison of Qwen 27B models across generations shows slight improvement in oneshot prompting ability, with average ratings increasing from 2.46 to 3.00.

Ran Qwen3.8 27B through all 35 oneshot prompts on oneshotlm and compared them to Qwen3.6 27B and Qwen3.5 27B. Here are the results. 3.6 was not able to create a working 2048 game but 3.8 did. 3.8's pelican is accurate 3.8 was able to get Wolfenstein mostly correct whereas 3.6 did not draw anything. I used Sonnet 5 to evaluate the outputs of all these models and looks like the average rating increased across the 27B generations 2.46 -> 2.74 -> 3.00.
Original Article

Similar Articles

Qwen 3.6 35B A3B vs Qwen 3.5 122B A10B

Reddit r/LocalLLaMA

User reports Qwen 3.5 122B significantly outperforms Qwen 3.6 35B on multi-step tasks despite benchmark claims, questioning if quantization or setup issues are to blame.

Qwen 3.6 27B on DeepSWE

Reddit r/LocalLLaMA

Qwen 3.6 27B scored 2% on the DeepSWE benchmark, placing 18/20 above Haiku 4.5 and Minimax M2.7, highlighting the gap between local and leading-edge models.