@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…
Summary
A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.
View Cached Full Text
Cached at: 09/26/26, 07:04 PM
Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable.
So instead of buying a bigger GPU, he built an inference engine.
That’s how Strata happened.
On the same machine:
llama.cpp: ~15 tok/s Strata: up to 65.1 tok/s
And the hardware isn’t some giant AI workstation.
It’s:
RTX 5070 12GB 64GB DDR5-5600 Ryzen 5 7600 Windows
The interesting part is that Strata isn’t trying to be a generic inference engine.
It was built specifically around Qwen3.8-Flash-Next and the kind of setup where the model has to work with a relatively small GPU alongside system RAM.
At 128K context, the reported numbers look like this:
Q2_0 65.1 tok/s decode 543 tok/s prompt processing
IQ2_XS 52.0 tok/s decode 472 tok/s prompt processing
IQ3_XXS 44.8 tok/s decode 414 tok/s prompt processing
That’s a pretty dramatic difference from the original ~15 tok/s result.
And this is exactly why I find local AI inference so interesting.
People often compare models as if the model itself determines the experience.
It doesn’t.
The stack underneath matters.
Model architecture.
Quantization.
KV cache.
GPU memory.
System RAM.
CPU/GPU transfers.
CUDA kernels.
Speculative decoding.
And eventually, the inference engine tying all of it together.
A model that feels painfully slow in one runtime can become surprisingly usable after someone spends the time optimizing the actual workload.
In this case, Strata was designed around Qwen3.8-Flash-Next, CUDA, system-RAM offloading and RCO-GSQ quantized builds.
And the best part?
It’s open source.
This is the kind of local AI project I love seeing.
Someone hits a performance wall, decides the runtime is the problem, builds a better solution for their exact hardware, and then gives the result back to everyone else.
You don’t always need a bigger GPU.
Sometimes you need a better engine.
FHILY👑 (@Oluwaphilemon1): Qwen3.8 Flash is starting to look very interesting on a pair of RTX 3090s.
After testing different quantization levels, 3.5bpw looks like a pretty sweet spot between model quality and inference speed.
The setup:
• Qwen3.8 Flash • 2× RTX 3090 • 3.5bpw • ExLlamaV3 backend
Similar Articles
@Oluwaphilemon1: Qwen3.8-27B at 56 tok/s on a 9-year-old GPU. Let that sink in. The GPU? NVIDIA V100 32GB. A card that launched at aroun…
Achieves 56 tokens per second inference speed for the Qwen3.8-27B model on an NVIDIA V100 GPU, demonstrating cost-effective local AI deployment on older hardware using speculative decoding techniques.
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
The article reports that the Qwen3.8-Flash-Next model achieves 120 tokens/second generation speed and 12k tokens/second prefill on a system with 4x AMD R9700 GPUs using optimized vLLM and a custom Docker image.
2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.
Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.
The article describes a method to run the Qwen3.8-Flash-Next model with quantized n-grams to achieve over 170K token context on a single 96GB GPU card, with performance up to 110 tokens per second using INT4 quantization and memory-mapped disk access.
Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
The article benchmarks llama.cpp and ik_llama.cpp for running the Qwen3.8-Flash-Next model on dual RTX 3060 hardware, showing that optimizing tensor splitting and ubatch settings can dramatically improve prefill speed from 36 t/s to 400 t/s, with comparisons of RAM usage.