Qwen3.7 Max scored by Artificial Analysis, 27B/35B waiting room

Reddit r/LocalLLaMA Models

Summary

Qwen3.7 Max ranks 5th on Artificial Analysis benchmarks, matching GPT-5.4 and outperforming Gemini 3.5 Flash, while Qwen3.6 27B trails significantly.

https://preview.redd.it/42ak5qmus82h1.png?width=1133&format=png&auto=webp&s=744ea3dfc06c83d0c4d8aa128c39b3238b17d7be Qwen 3.7 Max sitting at 5th, pretty much on par with GPT 5.4 (xhigh) and a notch above the just released Gemini 3.5 Flash. On the other end, we see DSV4 Flash and Qwen3.6 27B which is exactly 6 points behind its max counter part. Let's hope Qwen3.7 can get in the same ballpark of its max big bro as well.
Original Article

Similar Articles

Qwen 3.6 27B on DeepSWE

Reddit r/LocalLLaMA

Qwen 3.6 27B scored 2% on the DeepSWE benchmark, placing 18/20 above Haiku 4.5 and Minimax M2.7, highlighting the gap between local and leading-edge models.

Qwen 3.8 27B Aider score

Reddit r/LocalLLaMA

A user benchmarks Qwen 3.8 27B using Aider and finds it scores 72.9, matching Gemini 2.5 Pro and outperforming other state-of-the-art models on a MacBook with local inference via vLLM.