i finally switched from windows to linux and got a 30-50% boost in speed.
Summary
The user switched from Windows to Linux and replaced llamacpp with vllm, achieving a 30-50% speed improvement in AI inference.
Similar Articles
Speed difference between Windows 11 and Linux with llama.cpp: a myth when using medium and large MoE models
User benchmarks show no significant speed difference between Windows 11 and Linux when running large MoE models with llama.cpp, debunking a common myth. Tests on a multi-GPU setup with models like Qwen 3.5 122B, 397B, and MiniMax 2.7 yield nearly identical prompt processing and token generation speeds.
Llama.cpp PR 8% speed boost
A llama.cpp PR moves sampling from CPU to GPU, yielding 8% faster tokens on an RTX 5090 and ~4% on a Tesla P40 for Qwen3.6-35B inference.
I’m switching to Linux, is Ubuntu the most compatible with local AI?
A user asks about Ubuntu's compatibility for local AI tools like vLLM, llama.cpp, and ComfyUI when switching to Linux.
AMD Users: Have you tried the llamma.cpp AMD-Ecosystem branch? Up to 2x PP Speed
The article discusses an AMD-specific branch of llama.cpp that significantly boosts prompt processing speed for AMD users, with up to 2x faster performance on dense models using ROCm/Hip, though with some trade-offs in other metrics.
@pupposandro: 2.5x faster than llama.cpp on Strix Halo. We just shipped DFlash + PFlash for the AMD Ryzen AI MAX+ 395 iGPU (gfx1151, …
A new toolset (DFlash + PFlash) achieves 2.5x faster inference than llama.cpp on AMD Ryzen AI MAX+ 395 iGPU, demonstrating significant speedups for Qwen3.6-27B with 128 GiB unified memory.