@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…

X AI KOLs Timeline Tools

Summary

A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.

Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger GPU, he built an inference engine. That’s how Strata happened. On the same machine: llama.cpp: ~15 tok/s Strata: up to 65.1 tok/s And the hardware isn’t some giant AI workstation. It’s: RTX 5070 12GB 64GB DDR5-5600 Ryzen 5 7600 Windows The interesting part is that Strata isn’t trying to be a generic inference engine. It was built specifically around Qwen3.8-Flash-Next and the kind of setup where the model has to work with a relatively small GPU alongside system RAM. At 128K context, the reported numbers look like this: Q2_0 65.1 tok/s decode 543 tok/s prompt processing IQ2_XS 52.0 tok/s decode 472 tok/s prompt processing IQ3_XXS 44.8 tok/s decode 414 tok/s prompt processing That’s a pretty dramatic difference from the original ~15 tok/s result. And this is exactly why I find local AI inference so interesting. People often compare models as if the model itself determines the experience. It doesn’t. The stack underneath matters. Model architecture. Quantization. KV cache. GPU memory. System RAM. CPU/GPU transfers. CUDA kernels. Speculative decoding. And eventually, the inference engine tying all of it together. A model that feels painfully slow in one runtime can become surprisingly usable after someone spends the time optimizing the actual workload. In this case, Strata was designed around Qwen3.8-Flash-Next, CUDA, system-RAM offloading and RCO-GSQ quantized builds. And the best part? It’s open source. This is the kind of local AI project I love seeing. Someone hits a performance wall, decides the runtime is the problem, builds a better solution for their exact hardware, and then gives the result back to everyone else. You don’t always need a bigger GPU. Sometimes you need a better engine.
Original Article
View Cached Full Text

Cached at: 09/26/26, 07:04 PM

Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable.

So instead of buying a bigger GPU, he built an inference engine.

That’s how Strata happened.

On the same machine:

llama.cpp: ~15 tok/s Strata: up to 65.1 tok/s

And the hardware isn’t some giant AI workstation.

It’s:

RTX 5070 12GB 64GB DDR5-5600 Ryzen 5 7600 Windows

The interesting part is that Strata isn’t trying to be a generic inference engine.

It was built specifically around Qwen3.8-Flash-Next and the kind of setup where the model has to work with a relatively small GPU alongside system RAM.

At 128K context, the reported numbers look like this:

Q2_0 65.1 tok/s decode 543 tok/s prompt processing

IQ2_XS 52.0 tok/s decode 472 tok/s prompt processing

IQ3_XXS 44.8 tok/s decode 414 tok/s prompt processing

That’s a pretty dramatic difference from the original ~15 tok/s result.

And this is exactly why I find local AI inference so interesting.

People often compare models as if the model itself determines the experience.

It doesn’t.

The stack underneath matters.

Model architecture.

Quantization.

KV cache.

GPU memory.

System RAM.

CPU/GPU transfers.

CUDA kernels.

Speculative decoding.

And eventually, the inference engine tying all of it together.

A model that feels painfully slow in one runtime can become surprisingly usable after someone spends the time optimizing the actual workload.

In this case, Strata was designed around Qwen3.8-Flash-Next, CUDA, system-RAM offloading and RCO-GSQ quantized builds.

And the best part?

It’s open source.

This is the kind of local AI project I love seeing.

Someone hits a performance wall, decides the runtime is the problem, builds a better solution for their exact hardware, and then gives the result back to everyone else.

You don’t always need a bigger GPU.

Sometimes you need a better engine.

FHILY👑 (@Oluwaphilemon1): Qwen3.8 Flash is starting to look very interesting on a pair of RTX 3090s.

After testing different quantization levels, 3.5bpw looks like a pretty sweet spot between model quality and inference speed.

The setup:

• Qwen3.8 Flash • 2× RTX 3090 • 3.5bpw • ExLlamaV3 backend

Similar Articles

2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s

Reddit r/LocalLLaMA

The user describes a hardware configuration with dual RTX 3090 GPUs and an AMD EPYC CPU for running the Qwen3-Flash-Next model using llama.cpp, achieving 38 tokens per second in single-stream inference, and seeks advice on whether to add a third GPU or upgrade the CPU to improve performance, especially for running multiple parallel agents.