Qwen3.7 Max scored by Artificial Analysis, 27B/35B waiting room
Summary
Qwen3.7 Max ranks 5th on Artificial Analysis benchmarks, matching GPT-5.4 and outperforming Gemini 3.5 Flash, while Qwen3.6 27B trails significantly.
Similar Articles
Qwen 3.6 27B on DeepSWE
Qwen 3.6 27B scored 2% on the DeepSWE benchmark, placing 18/20 above Haiku 4.5 and Minimax M2.7, highlighting the gap between local and leading-edge models.
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
The Qwen 3.8 27B AI model achieved a score of 52 on the Artificial Analysis Intelligence Index, as highlighted in a blog post by Simon Willison.
Qwen 3.8 27B Aider score
A user benchmarks Qwen 3.8 27B using Aider and finds it scores 72.9, matching Gemini 2.5 Pro and outperforming other state-of-the-art models on a MacBook with local inference via vLLM.
Qwen 3.8 27B in 9th position on code arena. Gemma 4 31B is 80th.
Qwen3.8-27B by Alibaba Qwen achieves 9th place on the Code Arena benchmark with 1595 points, outperforming larger models like Gemma 4-31B and reshaping the Pareto Frontier in coding performance.
Qwen3.6-35B-A3B and 9B are officially on the public Terminal-Bench 2.0 leaderboard!
Qwen3.6-35B-A3B and Qwen3.5-9B models are officially on the Terminal-Bench 2.0 leaderboard, with little-coder achieving 24.6% on the 35B variant, surpassing Gemini 2.5 Pro and Qwen3-Coder-480B, while the 9B model shows that sub-10B local models can compete on hard agentic benchmarks.