qwen3-flash-next

Tag

Cards List
#qwen3-flash-next

Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM

Reddit r/LocalLLaMA · yesterday

The article reports that the Qwen3.8-Flash-Next model achieves 120 tokens/second generation speed and 12k tokens/second prefill on a system with 4x AMD R9700 GPUs using optimized vLLM and a custom Docker image.

0 favorites 0 likes
#qwen3-flash-next

llama.cpp support for Qwen3.8-Flash-Next has been merged

Reddit r/LocalLLaMA · 3d ago Cached

Support for the Qwen3.8-Flash-Next model has been merged into llama.cpp, enhancing its capabilities for local LLM inference in C/C++.

0 favorites 0 likes
← Back to home

Submit Feedback