@ProTekkFZS: Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can't lie. Using the llama.cpp fork from @no_stp_on_snek …
Summary
User reports successfully running a 35B-parameter mixture-of-experts model at 768K context length using Q4_K_M quantization and YaRN on an RTX 3090 via a llama.cpp fork, offloading only 8 experts to CPU while maintaining acceptable performance.
View Cached Full Text
Cached at: 04/21/26, 08:09 AM
Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can’t lie. Using the llama.cpp fork from @no_stp_on_snek for turboquant, only offloading 8 experts to cpu and still getting acceptable performance. 10k prompt with good recall.
Similar Articles
Qwen3.6:35b UD Q4_K_M 80 tok/s on Nvidia P40
A user shares achieving 80 tok/s on a Qwen3.6 35B model with Q4_K_M quantization and 100k context on a single Nvidia P40 using TheTom's TurboQuant fork of llama.cpp, highlighting various optimizations.
First local-LLM tuning attempt: Qwen3.8-27B true Q4_K_M at 13.2 tok/s near 50-61K context on RTX 5080 16GB
A user details their first attempt at tuning the Qwen3.8-27B model with Q4_K_M quantization on an RTX 5080 16GB, achieving about 13.2 tokens per second at 50-61K context by selectively offloading FFN tensors to CPU to improve performance.
7900XTX 24GB vram, can finally fit Q6K+MTP with Qwen 3.6 27B at 131k context
A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.
Qwen3.6-35B-A3B APEX on a Single RTX 3090 - Getting the Most Out of It
A detailed guide on running the Qwen3.6-35B-A3B APEX model on an RTX 3090, comparing two llama.cpp forks and quantization methods for optimal speed and quality.
EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s
An experiment running the Qwen3.8-2.4T-A95B MoE model locally on dual consumer GPUs (RTX 5090 + 5060 Ti) with llama.cpp, achieving ~0.8 tok/s with MTP speculative decoding enabled.