Are you running Qwen 3.8 27b or Qwen Flash Next?

Reddit r/LocalLLaMA News

Summary

The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.

Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx? Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/ Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.
Original Article

Similar Articles

Qwen3.8-Flash-Next optimised for Macs

Reddit r/LocalLLaMA

The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.

Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results

Reddit r/LocalLLaMA

This article presents detailed test results comparing the performance of Qwen3.8-Flash-Next-NVFP4 and Qwen3.8-27B-FP8 AI models across various tasks, highlighting that Flash-Next is faster with fewer failures but struggles with multi-step symbolic work.