Qwen 3.6 35B A3B vs Qwen 3.5 122B A10B

Reddit r/LocalLLaMA Models

Summary

User reports Qwen 3.5 122B significantly outperforms Qwen 3.6 35B on multi-step tasks despite benchmark claims, questioning if quantization or setup issues are to blame.

Does anyone else have the same experience comparing these two - for me 3.5 122B outperforms 3.6 by a large margin. 3.6 gets lost as long as the task requires a couple of more steps. I'm asking because I got the impression that it overperforms in some benchmarks, and I'm thinking that maybe I'm doing something wrong? My experience shows quite the contrary. Would be great to benefit from the speed if I can fix it, so if you have any advice to share let me know. EDIT: I'm using Qwen3.5 122b UD-Q5\_K\_XL and Qwen3.6 35b UD-Q8\_K\_XL. Maybe I should try the full BF16, but I don't think it should be too different. CUDA runtime is also 13.1, I'm aware of the issues with 13.2 and smaller quants.
Original Article

Similar Articles

Qwen/Qwen3.6-35B-A3B-FP8

Hugging Face Models Trending

Alibaba releases Qwen3.6-35B-A3B-FP8, an open-weight quantized variant of Qwen3.6 with 35B parameters and 3B activated via MoE, featuring improved agentic coding capabilities and thinking preservation for iterative development.

The Qwen 3.6 35B A3B hype is real!!!

Reddit r/LocalLLaMA

The author benchmarks small local LLMs, highlighting Qwen 3.6 35B A3B for its superior ability to map academic code to research papers compared to models like Gemma 4 and Nemotron 3 Nano.