Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?
Summary
User shares performance results and setup details for running the Qwen3.8-Flash-Next model with Llama.cpp on a system featuring an NVIDIA 5090 GPU and 64GB RAM, noting unexpected low RAM usage during inference.
Similar Articles
TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram
A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.
Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors
A Reddit user shares how they successfully ran the Qwen 3.8 Next MoE model on a system with 16GB VRAM and 32GB RAM using aggressive quantization and specific llama.cpp settings, achieving usable performance for large models on limited hardware.
Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
The article benchmarks llama.cpp and ik_llama.cpp for running the Qwen3.8-Flash-Next model on dual RTX 3060 hardware, showing that optimizing tensor splitting and ubatch settings can dramatically improve prefill speed from 36 t/s to 400 t/s, with comparisons of RAM usage.
Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
This article provides a detailed guide on running the Qwen3.8-Flash AI model on an RTX 3090 with 64GB RAM, discussing performance metrics, quantization settings, and deployment steps using tools like llamacpp.
Maxing out 64GB of RAM - Qwen3.5 122B A10B at UD-Q2_K_XL w/ MTP fully replaced Qwen3 Next 80B at UD-Q4_K_XL for me
A user compares running quantized Qwen3 Next 80B and Qwen3.5 122B on a 64GB RAM system, noting the trade-offs in speed, quality, and memory usage for local LLM inference.