A practical guide covering common inference bottlenecks — including runtime selection (Ollama, llama.cpp, vLLM, SGLang), hardware-specific compilation flags, quantization choices (IQ vs Q_K vs NL), and model selection by task — with best practices for optimizing LLM inference within hardware constraints.
Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing runtime are far fewer than the people that are trying to build apps or offer AI solutions. So I've been lurking around this community, and I've noticed a lot of people who seem to have massively suboptimal setups for their hardware, and I've grouped the biggest errors into several buckets. The purpose of this guide is to expose common inference bottlenecks and provide best practices for avoiding them within your hardware constraints. RUNTIME: I. Choosing the Right Runtime This is the biggest mistake I see. Choosing the correct runtime for your architecture and model is the most important decision to make. In general, here are some rules to help you determine what runtime to use. Firstly, Ollama is never optimal. Its just the simplest. If you want quick and easy and have extra RAM, it's a good place to start. It's very user friendly and requires less setup. But it just won't offer best inference speeds. If your model requires cpu offload, then llama.cpp will be your best choice. If not and you're solely in gpu, VLLM will likely provide the best results. It's really as simple as that for 90% of cases. SGLang may be worth it if your workload involves Langgraph, as it is highly optimized for the tooling. Otherwise, stick to the above. Mainline branches are best, with community forks offering only highly niche performance boosts (i.e., for specific models/configurations, but are generally under optimized and not well maintained). II. Optimizing and Maintaining Runtime The other big mistake people make with runtime is failing to compile it with hardware specific flags. Not going to go through all of them here, Google can help you out. Just search "optimal runtime compilation flags for [runtime] using [GPU, CPU, RAM type]." The most missed/missed flags tend to be for architecture specific optimizations. Those are crucial. Runtime should be recompiled (with optimal flags) any time *any* of the following occur: - System updates - Kernel/driver updates - Running a model released or modified later than your last compile - You haven't recompiled in over a month (recent updates often contain kernel or path optimizations) MODEL CHOICE: I. Quantization: Quantization. Such a big word. Such little meaning. All you need to know is that it makes a model smaller. There are a million Q_K_X_&$&$&$ quant sizes, so I'm not going to go over them individually. Rather, I will provide basic principles. - IQ quants are generally the best for the size. If you're choosing between IQ4_XS or Q4_K_M, IQ4_XS is both a smaller VRAM footprint and higher complexity. - Nonlinear (NL) quants are only ever going to be better if you have CPU offload. Even then, IQ quants often offer extra context space vs NL quants and thus are preferable. - If it's a quant you've never encountered, read the docs. It more than likely is highly optimized for that specific model and is indeed the one you should choose. Searching it or asking chatgpt *will give you the wrong answer every single time for custom quants.* You will only encounter these with custom tuned models. - Standard Q_K_M quants are best if you have absolutely no hardware constraints for the model you're running, as path optimizations are the best. If you have no hardware constraints, though, you could be running a better model. This is only useful for running simple models for simple tasks. II. Task Certain models excel at certain tasks. This is subjective and preference based but this is my list: Coding: for API, anthropic. Hands down the best models. Opus and Sonnet 5.5 both excel in performance and their low token usage per task makes them more affordable than previous iterations. Deepseek models are the best budget choice. Qwen models are the clear winner for local inference on all fronts. Writing: Opus/sonnet for technical writing, Gemini for creative writing. Gemma for local creative writing. Video: Wan 2.2 for local in most cases will get it done, chatgpt and copilot both have surprisingly robust free image/video gen, Veo is the best paid. HARDWARE: Buy a budget box, or build your own from parts. I've managed to squeeze better performance out of an RTX 3060 and 128 gb RAM than a DGX spark across all categories for multiple models. The spark has an edge for dense models, but I was able to run higher complexity models overall on the other setup for 1/5 the price. AMD and Intel lag significantly on speed per price, but I've heard Intel has had some major gains recently. Have not confirmed myself though. MODEL OPTIMIZATIONS: I. Spec Decode (MTP) - If you have CPU offload, spec decode will *always* slow you down. The extra overhead compute isn't worth it if you don't have at least several hundred Mb/s bandwidth, which your CPU won't. - MTP is sometimes a baked in feature, and sometimes requires a special secondary model. Ensure you know how it works for your model and what flags to run. - Each model will be optimized for exactly 0-1 type of spec decode. Figure out which one it is (or isnt) rather than wasting your time testing methods. II. Model Tuning Just to show the kind of command optimization you can get, here is my sample command for running Qwen 3.8 Flash-Next, a 156b parameter model, on 12 gb VRAM (and 128 gb RAM) at 10 token/s decode and 150 token/s profile at 200k context: ~/llama.cpp/build/bin/llama-server --flash-attn on --batch-size 1024 --ubatch-size 1024 --no-warmup --cache-reuse 256 --jinja --host 0.0.0.0 --port 8090 --presence-penalty 0.0 --repeat-penalty 1.0 -m ~/llama.cpp/LLM/Qwen3.8-Flash-IQ4_XS/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf --temp 0.95 --top-k 20 --top-p 0.97 --min-p 0.05 -np 1 --chat-template-kwargs '{"enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "xhigh"}' --threads-batch 16 --threads 8 --gpu-layers 150 --n-cpu-moe 48 -c 200000 --override-tensor per_layer_token_embd.weight=CPU -ctv q8_0 -ctk q8_0 --cache-ram 8192 --checkpoint-min-step 512 --ctx-checkpoints 4 --kv-unified --reasoning-preserve --load-mode mmap+mlock That's a lot, right? It's every possible optimization you could apply. I'll go through them individually. This is llama.cpp specific, but you'll find the same flags with slightly different syntax apply to other runtimes. -flash-attn (-fa) on: forces flash attention optimizations and paths. Explicitly set to on to override any fallback. Auto can be optimal if the model is recent and is not yet optimized. -batch/-ubatch: batch is the decode chunks, ubatch is the prefill chunks. They must be multiples of one another, otherwise you're adding compute. Equal to one another is ideal for CPU offload, and a 2x-4x higher batch is optimal for full GPU loads. You'll need to play with these values to optimize. Batch/ubatch should be a power of 2 to optimize architecture. Intervals of 256 is typically good enough for testing. --no-warmup: prevents initial model poll to load weights. Removes unnecessary latency -- cache-reuse x: instructs the model to reuse cache values and scan for similarity at x token intervals -jinja: highly underutilized and important flag. Utilizes native chat template kwargs to ensure output consistency. --presence-penalty: flat penalty rate to words that appear in text. Used mainly for creative writing to prevent repetitive prose. --repeat-penalty: reduces liklihood of already used tokens being reused. Best used for preventing loops in thinking agents. -temp: model temperature– how creative the model is. 0 is completely deterministic, 1 is creative freedom. Top-k: hard cutoff that keeps only the k most likely words. Each model will have recommended k values for thinking/instruct setups. Low key reduces hallucinations at the cost of repetition and loss of creativity Top-p: includes P percentage of possible tokens. It reduces liklihood of hallucination dynamically. Min-p: dynamic cutoff based on highest probability token. If the biggest probability token is 50% and min-p is 0.05, then the bottom 0.025 (2.5%) liklihood tokens will be excluded. Reduces noise without hampering creativity terribly. Chat template kwargs: explicit chat template activations; newer runtime compilations should have flags for these. Controls model reasoning, reasoning effort, and internal chain of thought storage. -threads (-t): number of cores used for decode. Set equal to physical cores (hyperthreading will thrash cores and degrade results) -threads-batch(-tb): number of cores used for prefill. Set to double the number of physical cores, as hyperthreading helps here. --gpu-layers (-ngl) : total layers on GPU. Fit as many as you can without OOM. --n-cpu-moe: number of MoE layers offloaded to cpu. For MoE models, you should always offload these first and keep all layers on GPU if possible. Offload as few as possible to CPU. --override-tensor...: tells the runtime to offload the n-gram table if needed. For qwen 3.8 flash specifically. -ctk/-ctv: k and v cache quantization. K cache should *never* be below q8 unless youre running on less than 80k context. V cache can be q4 up to 150k context without issues, for the most part. Generally, q8 for both will be best for speed and is my recommendation to start with. --cache-ram: sets RAM aside for cache allocation to ensure it doesn't go to swap --context-checkpoints: the amount of checkpoints captured. Generally, you don't need as high as the defaults do. Leave the default if you have extra RAM, otherwise you may want to lower it. --kv-unified: tells all instances to run on the same kv cache pool rather than allocating individual cache. --reasoning-preserve: tells the runtime to retain CoT traces for evaluation. Prevents the model from getting stuck or looping as much when thinking. --load-mode: tells the runtime how to load in the model. No mmap generally loads slower but is more stable. Mmap+mlock (or just mlock) is a balance of both with fast loading and page faults initially but it stabilizes as you run it, mmap alone is fast but will cause constant page faults and slows down inference, especially on models with CPU offloading. I hope this guide helps! I'd be willing to answer any specific questions or make any additions if there are additional areas the community agrees are major uncovered inference bottlenecks. Some claims are based on my personal experience and I am open to data based claim revisions or anecdotal counterclaims, so feel free to provide. Happy tuning!
A comprehensive guide to optimizing local LLM inference on consumer hardware, covering tools like llama.cpp, vLLM, and LM Studio, with practical advice on memory hierarchy, layer placement, and common failure modes.
This guide explains the discipline of AI inference engineering, covering the split between prefill and decoding phases, the shift from closed to open models, and optimization techniques for latency, throughput, and cost.
A tweet lists key projects to build in inference engineering for understanding production LLM systems, including inference servers, paged KV cache, speculative decoding, quantization libraries, and guardrails.
A blog post argues that hyper-specialized, hardware- and model-specific inference runtimes (Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo) are now outperforming general-purpose engines like llama.cpp, vLLM, and Ollama by 2.5-4x on fixed consumer hardware, with Strata reaching 53-90 tok/s decoding Qwen3.8-Flash-Next on a single 12GB RTX 4070.