@ProTekkFZS: Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can't lie. Using the llama.cpp fork from @no_stp_on_snek …

X AI KOLs Following News

Summary

User reports successfully running a 35B-parameter mixture-of-experts model at 768K context length using Q4_K_M quantization and YaRN on an RTX 3090 via a llama.cpp fork, offloading only 8 experts to CPU while maintaining acceptable performance.

Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can't lie. Using the llama.cpp fork from @no_stp_on_snek for turboquant, only offloading 8 experts to cpu and still getting acceptable performance. 10k prompt with good recall.
Original Article
View Cached Full Text

Cached at: 04/21/26, 08:09 AM

Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can’t lie. Using the llama.cpp fork from @no_stp_on_snek for turboquant, only offloading 8 experts to cpu and still getting acceptable performance. 10k prompt with good recall.

Similar Articles

Qwen3.6:35b UD Q4_K_M 80 tok/s on Nvidia P40

Reddit r/LocalLLaMA

A user shares achieving 80 tok/s on a Qwen3.6 35B model with Q4_K_M quantization and 100k context on a single Nvidia P40 using TheTom's TurboQuant fork of llama.cpp, highlighting various optimizations.