Are you running Qwen 3.8 27b or Qwen Flash Next?
Summary
The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.
Similar Articles
Qwen3.8-27B vs Qwen3.8-Flash-Next smaller quant?
A user compares Qwen3.8-27B and Qwen3.8-Flash-Next models for intelligence and coding performance with 128GB RAM, seeking advice on which is better.
Qwen3.8-Flash-Next optimised for Macs
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results
This article presents detailed test results comparing the performance of Qwen3.8-Flash-Next-NVFP4 and Qwen3.8-27B-FP8 AI models across various tasks, highlighting that Flash-Next is faster with fewer failures but struggles with multi-step symbolic work.
Benchmark results: what is the best and fastest engine to run Qwen3.8-27B on macOS
This article benchmarks various engines for running the Qwen 3.8 27B model on macOS, comparing their speed and performance in agentic coding tasks, and recommends MTPLX or llama.cpp with MTP for best results.
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
This article reports on running the Qwen3.8-Flash-Next model on a MacBook Pro M5 Max, benchmarking speed versus context depth over 100 turns, with insights into performance and issues like role confusion at long contexts.