Hey guys, Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000: RadixArk/Qwen3.8-27B-NVFP4 (dense) orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune) RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE) Each model got the same prompts and its own model card's sampler, with one attempt per task. Video with the battles, the castles and the ball run: https://youtu.be/VOtfja_Toj4 Short version Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three). Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1. Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8×. SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146). Tool use (BFCL subset, thinking off): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%. Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking. Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole ~111K-token answer budget thinking and never placed a part. CAPTCHA (40 puzzles, local copy): 27B 24/40, Flash-Next 21/40, Uncensored 19/40. Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric). https://preview.redd.it/4e22yukvi4th1.png?width=1484&format=png&auto=webp&s=e1f88f95089bff77765bdfb6847f956ad9108da3 Setup GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB CPU: AMD Ryzen 9 9950X System RAM: 96GB DDR5 27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs. Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings. The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard. One workstation, one model server at a time, and every number comes from a saved run. 1. Speed: drafters, SGLang vs vLLM, and long prompts For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0. Which drafter (tok/s, one user): Drafter SGLang · RadixArk NVFP4 vLLM · Inferact NVFP4 none 75 59 MTP (built into the model) 160 113 DSpark 174 137 DFlash2 210 160 DFlash2 + torch.compile 214 not run DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP. Which build, on which engine (tok/s): Build · engine DFlash2, 1 user No drafter, 1 user DFlash2, 4 users (total) RadixArk NVFP4 · SGLang 210 75 607 Inferact NVFP4 · vLLM 160 59 517 Uncensored NVFP4 · SGLang 146 46 473 Uncensored NVFP4 · vLLM 169 63 538 BF16 · SGLang (full precision) 97 29 291 The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster. The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146. 4-bit vs full precision: NVFP4 with DFlash2 is 2.2× the BF16 speed. Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s. https://preview.redd.it/vnd83h5zi4th1.png?width=1484&format=png&auto=webp&s=949cb3353a299d0d81cc953cce4e1a0b442a6a46 https://preview.redd.it/0v1unh5zi4th1.png?width=1484&format=png&auto=webp&s=a735ba12f3c20662b35439706686dcfab8335595 The three models: Metric Qwen3.8-27B 27B-Uncensored Flash-Next Prefill, full window 97s 99s 22.4s Decode, Spec-Bench, one user 210 tok/s 146 tok/s not run Drafter vs no drafter 2.8× 3.2× n/a The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds. One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps https://preview.redd.it/jx28sde1j4th1.png?width=1484&format=png&auto=webp&s=eb0abbfb00f1058a71865cb5ba9d3e2e02de0513 2. Tool use: the dense 27B leads This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top_p 0.8 from the card. Metric Qwen3.8-27B 27B-Uncensored Flash-Next BFCL core 73.3% 70.8% 64.5% Tool accuracy 87.8% 88.0% 82.4% Abstention 79.5% 69.5% 68.5% Multi-turn 52.5% 55.0% 42.5% Malformed calls 0.08% 0.27% 0.28% These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0. https://preview.redd.it/yqdc2j32j4th1.png?width=3396&format=png&auto=webp&s=243a206611eab4609251c7a8fa149465bbc05c42 3. Long context: perfect for all three I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request. 27B: 27/27 Uncensored: 27/27 Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning) The largest prompt was about 259.5K tokens. https://preview.redd.it/fgeuq733j4th1.png?width=3396&format=png&auto=webp&s=7b221fa5f9f2b14224d3370c3a5b41bfb3fadd7a 4. Battle arena: the local 27B beat Claude Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 × its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped). Rank Entry Score 1 RadixArk/Qwen3.8-27B-NVFP4 700 2 Claude Fable 5.1 (chat, max thinking) 678 3 Qwen3.8-27B-Uncensored 603 4 Qwen3.8-Flash-Next 535 5 GPT-5.6 (chat, ultra thinking) 473 In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked. Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes. https://preview.redd.it/1wbb6p28j4th1.png?width=3396&format=png&auto=webp&s=a372f6254ef889aa56e0e149c154886735d47fc2 https://preview.redd.it/o2yccq77j4th1.png?width=1484&format=png&auto=webp&s=1a59ffd33b733094abc19a4e2404742773103398 5. The SVG test is also a fact check Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched. 27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5_K_M, 21.73 GB, the real file size. Q6_K at 25.09 GB is correctly marked as not fitting. Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card. Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4_K_M, and its "12 tok/s" isn't in anything it fetched. All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look. https://preview.redd.it/oijy66u9j4th1.png?width=1920&format=png&auto=webp&s=7ff5d4a19dbeb73374060241b395fc99ed0d26e6 6. Video editing, voxel and design Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10: Flash-Next 9 + 10 = 19 27B 8 + 8 = 16 Uncensored 5 + 8 = 13 Voxel (Wawel Castle in three.js), ranked by eye: Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill. 27B built a clean but generic castle. Uncensored placed the camera inside its own build. Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B. Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules. https://preview.redd.it/ue511ppaj4th1.png?width=1920&format=png&auto=webp&s=1ffb968e03b4a003d51517b90cf873c673404f7c 7. Rube Goldberg: only one machine The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup? Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times. 27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet. https://preview.redd.it/px8mlhhbj4th1.png?width=1280&format=png&auto=webp&s=1229ddd46d6d3a1cb94f4b5d451a7e152cdcd16d 8. CAPTCHA: local models in a real browser I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes. Model Solved Median time Thinking tokens, all 40 Qwen3.8-27B 24/40 55s 456K Flash-Next 21/40 145s 1.0M 27B-Uncensored 19/40 36s 401K With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal. Which one should you run? RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA. RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine. orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG. Resources Configs, Docker setup and reports The test harness is still private while it's changing. Full video with the battles, the castles and the ball run: https://youtu.be/VOtfja_Toj4 I abused AI to help write this up and to check it against the report. Every number above comes from a saved run. Which test do you find most interesting and maybe you have some other creative ideas how to test models?
A user benchmarks the newly merged DFlash speculative decoding method in llama.cpp on Qwen 3.6 27B, achieving up to 4.44x speedup at 36K context compared to baseline, with detailed leaderboard and quality tests.
This guide explains how to configure DFlash2 speculative decoding with llama.cpp to achieve up to 1.72x faster inference for the Qwen3.8-27B model on an RTX 4080 GPU with 16GB VRAM.
User benchmarks and details a GPU-resident expert cache PR in llama.cpp that boosts decode speed for Qwen3.8-Flash-Next on a dual RTX 3090 system from 17 to 25-29 tokens per second.
The article reports that the Qwen3.8-Flash-Next model achieves 120 tokens/second generation speed and 12k tokens/second prefill on a system with 4x AMD R9700 GPUs using optimized vLLM and a custom Docker image.
A same-harness benchmark on a Mac Studio M5 Ultra 256GB compares Qwen3.8-Flash-Next (182GB oQ8e) vs Laguna-S-2.1 GGUF at up to 262K context. Qwen sustains ~4,200 tok/s linear prefill and stable decode, while Laguna degrades superlinearly; quality is a 4/4 draw with very different answer styles.