If you have a 3090, or other 30xx for local LLMs, I have something for you
Summary
The article presents a custom fork of llama.cpp optimized for NVIDIA Ampere architecture GPUs like the 3090, achieving over 90 tokens per second for 27B parameter models with large context windows, making local inference faster than API speeds.
Similar Articles
GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP) - Machine Learning, LLMs, & AI
A fork of llama.cpp has been created and optimized for AMD GFX906 GPUs, improving performance on Mi50, Mi60, Radeon VII, and GCN HIP for machine learning and LLM applications.
@TheAhmadOsman: Why do I focus on Inference Engines/Software Stacks for your hardware? - 2x RTX 3090s: ~14.5 tok/s → ~64 tok/s moving t…
Comparison of inference engine performance on different hardware: moving from baseline to vLLM with TP=2 on 2x RTX 3090s improves from ~14.5 tok/s to ~64 tok/s, and on RTX PRO 6000 moving to Sglang improves from ~32 tok/s to ~110 tok/s. Recommends vLLM/Sglang for CUDA/multi-GPU and llama.cpp for edge devices.
@ItsmeAjayKV: Update on 3090: Now with Qwen 3.6-35b-a3b moe (q6_k_xl). Crossed 90 t/s for the very first time, no MTP yet, prefill sp…
A user reports achieving over 90 tokens per second inference speed with Qwen 3.6-35b-a3b MoE model on an RTX 3090 using llama.cpp, with prefill speeds exceeding 1000 t/s, indicating practical local deployment of large language models on consumer hardware.
I forked Ninfer 3090 and converted it to run on the CMP170HX - doubled my Qwen3.6-35B from llama.cpp
A user forked the Ninfer inference engine to run on CMP170HX GPUs, doubling performance for the Qwen3.6-35B model with llama.cpp and sharing the configuration details.
dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode
A GitHub fork of llama.cpp optimized for dual AMD 7900 XTX GPUs, significantly improving decode speed for the Qwen 3.8 Q8 model to 82 tokens per second.