All Qwen model oneshots: 1109 outputs to look at and compare!

Reddit r/LocalLLaMA Tools

Summary

A weekend project generated and compared one-shot outputs from all 33 Qwen models on OpenRouter across 35 prompts, with 1109 total outputs available to view on oneshotlm.com.

I've been busy this weekend generating oneshots for all the cheapest models on the openrouter and ended up going through all 33 qwen models across 35 prompts (there were some failures and only 1109 made out of 33*35 matrix). Here they are https://oneshotlm.com/model/?q=qwen Qwen 3.7: qwen3.7-plus, qwen3.7-flash Qwen 3.6: qwen3.6-plus, qwen3.6-35b-a3b, qwen3.6-27b, qwen3.6-flash Qwen 3.5: qwen3.5-plus (02-15) · (20260420), qwen3.5-397b-a17b, qwen3.5-122b-a10b, qwen3.5-35b-a3b, qwen3.5-27b, qwen3.5-9b, qwen3.5-flash Qwen 3 (+2507 refresh): qwen3-235b-a22b (+2507, +thinking), qwen3-next-80b-a3b, qwen3-30b-a3b (+instruct-2507), qwen3-32b, qwen3-14b, qwen3-8b Qwen3-Coder: qwen3-coder, qwen3-coder-next, qwen3-coder-30b-a3b, qwen3-coder-flash Qwen3-VL: qwen3-vl-235b-a22b, qwen3-vl-32b, qwen3-vl-30b-a3b, qwen3-vl-8b Qwen 2.5: qwen2.5-72b, qwen2.5-7b
Original Article

Similar Articles

8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics

Reddit r/LocalLLaMA

This article presents a comprehensive benchmark of 8 abliterated variants of the Qwen 3.8 27B model against the base model, using weight analysis, KL divergence, 13 benchmarks, and HarmBench refusal tests over 167 GPU hours. The analysis reveals that surgical edits significantly outperform heavy modifications, with aggressive abliteration causing thinking loops in up to 45% of adversarial responses and chat template manipulation detected in some variants.

Qwen3.8-Flash-Next

Simon Willison's Blog

Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal MoE model with 125B tokens but only 6B active parameters, providing a performance boost and serving as an early preview of the Qwen4 architecture.