Is it just me or is Qwen3.8-Flash-Next ... really buggy?

Reddit r/LocalLLaMA News

Summary

A user questions if others have noticed hallucinations and weird reasoning with the Qwen3.8-Flash-Next model on Mac, reporting issues with AppImage installation and tensor metadata from HuggingFace despite high quantization levels.

I mean, this is on a Mac, why is a 8 years old Ubuntu AppImage being halu-installed...? And this message is in the middle of pulling some tensor metadata from HF. Never even heard of OpenD before this ... totally hallucinated stuff. And this is not a low quant - it's a 5bpw quant, with Q4 the lowest of any tensors. EDIT: I'm not looking for a solution -> I'm genuinely asking if other people have noticed hallucinations and weird reasoning.
Original Article

Similar Articles

Are you running Qwen 3.8 27b or Qwen Flash Next?

Reddit r/LocalLLaMA

The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.

TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

Reddit r/LocalLLaMA

A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.

Qwen3.8-Flash-Next optimised for Macs

Reddit r/LocalLLaMA

The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.

orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF

Hugging Face Models Trending

This article releases GGUF quantizations of the uncensored Qwen3.8-Flash-Next model, a Mixture-of-Experts preview of the Qwen4 architecture designed for llama.cpp with vision support, requiring a custom build for compatibility.