Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s

Reddit r/LocalLLaMA Models

Summary

Demonstrates that Qwen 3.8 Flash Next can run locally on a mid-range Android phone at 3.5 tokens per second with optimizations and low quantization.

Qwen 3.8 Flash Next (80gb) now at 3.5 tok/s on 12gb mid range android phone thanks to some optimizations and with a low quantization on dense part. I don't want to promote the project, but simply show that it's possible on a $400–$500 phone
Original Article

Similar Articles

TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

Reddit r/LocalLLaMA

A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.