FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

Reddit r/LocalLLaMA Tools

Summary

A hands-on benchmark on a single RTX 3090 shows llama.cpp is 2-3x faster and much quicker to first token when the MoE model fits in VRAM, while FreeToken wins 7x on time-to-first-token with 32 concurrent users on gpt-oss-120b (63 GB) that overflows VRAM.

No content available
Original Article
View Cached Full Text

Cached at: 10/01/26, 02:36 PM

# Which local inference engine you should use Source: [https://comendeiro.substack.com/p/which-local-inference-engine-you](https://comendeiro.substack.com/p/which-local-inference-engine-you) On a model that fits comfortably in a 24 GB GPU,[FreeToken](https://github.com/FlashML-org/FreeToken), an engine built specifically to run big Mixture\-of\-Experts models on consumer hardware, was**2–3× slower**than plain llama\.cpp, and kept me waiting**5–6× longer**for the first token\. Then I loaded a model 2\.6× bigger than my vRAM\. Now FreeToken delivered the first token**7\.2× faster**at 32 concurrent users\. [![](https://substackcdn.com/image/fetch/$s_!vXj2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5896afb3-937a-4cc0-9ff4-f4de7746eaea_1776x1024.png)](https://substackcdn.com/image/fetch/$s_!vXj2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5896afb3-937a-4cc0-9ff4-f4de7746eaea_1776x1024.png) *A log\-scale chart of how many times faster FreeToken is to first token\. On a model that fits in VRAM it sits at about 0\.18×, meaning five to six times slower, and crashes at 8 users\. On a model bigger than VRAM it rises from 1\.8× to 7\.2× faster as load grows\. Each line is the ratio of llama\.cpp’s time\-to\-first\-token to FreeToken’s\. Above the dashed line, FreeToken wins\.* This does not mean “FreeToken is good” or “FreeToken is bad\.” This post is about where FreeToken has and advantage, why it’s there, and the handful of traps that could have faked the answer\. **In short:** - **Model fits in VRAM → use llama\.cpp\.**2\.2–3\.2× the throughput and a 5–6× faster first token\. FreeToken*cannot*keep 4\-bit experts on the GPU at all \(I’ll show you the line of code\), and it ran out of memory at 8 concurrent requests where llama\.cpp served 32 in less VRAM\. - **Model doesn’t fit → use FreeToken\.**FreeToken’s time\-to\-first\-token stayed around 9 seconds for anywhere from 2 to 8 users\. llama\.cpp’s climbed to 139 seconds at 32\. - **Neither engine escapes the cliff\.**Spilling to system RAM costs about 10× in generation speed: 10–17 tokens/s, versus 113–163 when llama\.cpp could hold the model in memory\. - **FreeToken’s PCIe link was pinned at its ceiling\.**A faster bus*should*help a lot\. A desktop with an RTX 3090 \(24 GB vRAM\), a Ryzen 7 3700X, 125 GB of DDR4 and a PCIe 3\.0 x16 slot\. Engines: llama\.cpp \(build b10482\) and FreeToken 0\.1\.2, both serving an OpenAI\-compatible API\. NVIDIA’s AIPerf as load generator\. I measured the PCIe link at**13\.1 GB/s**\. The 3090’s GDDR bandwidth is**936 GB/s**\. VRAM is about 71× faster than the bus feeding it\. Mixture\-of\-Experts models make that gap tempting to cheat\. A 26B MoE might activate only ~4B parameters per token, so why keep all 26B in VRAM? Park the experts in system RAM, cache the popular ones on the GPU, and stream the rest\. That’s FreeToken’s bet\. llama\.cpp bets the other way: quantise hard, keep everything resident, and if something doesn’t fit, compute the overflow experts on the CPU\. Which bet wins depends on whether the model fits\. Comparing two inference engines usually means comparing two*files*\. llama\.cpp reads GGUF; FreeToken reads safetensors\. A GGUF Q4 and an NVFP4 checkpoint are different quantisations, so any speed gap mixes “engine” with “encoding\.” There’s one exception\. FreeToken ships a GGUF loader for Gemma\-4\. For that model, both engines can read the**identical bytes**\. It took three tries\. Two public Gemma\-4 GGUFs died inside FreeToken with a bare`AssertionError`and no message\. To find out which tensor, I injected a small reporter at startup without touching FreeToken’s source \(a`sitecustomize\.py`on the Python path, which the spawned workers inherit\): ``` [MISMATCH] model.embed_tokens.qweight ``` `model expects shape \(262144, 2310\)` `file has shape \(262144, 2992\)` At this model’s hidden size of 2816, a Q6\_K row packs into 2310 bytes and a Q8\_0 row into 2992\. FreeToken’s Gemma\-4 loader hard\-codes the Q4\_0 weights with a Q6\_K embedding table\. However the public file had Q8\_0 embeddings\. The fix was to build the file to spec with llama\.cpp’s own quantiser: one 14\.4 GB file, read by both engines\. \(caveat: the embedding went Q8\_0 → Q6\_K, a genuine double quantisation\.\) Three more controls that can alter the outcome: - **llama\.cpp has two prompt caches\.**`\-\-no\-cache\-prompt`turns off one\. A separate 8 GiB host\-RAM cache \(`\-\-cache\-ram`\) stays on by default, and my server log showed it evicting 281 MiB entries mid\-benchmark\. I only caught it by reading the log\. - **Reasoning models hide tokens\.**Gemma\-4 and gpt\-oss “think\.” If the server splits those tokens into a separate field, a client counting normal output undercounts them, differently on each engine\. I turned thought\-parsing off on both\. - **Pin the output length\.**Every request was forced to generate 256 tokens\. At concurrency one, llama\.cpp generated**114 tokens/s**; FreeToken**51**\. First token in 0\.38 s versus 2\.0 s\. Adding users didn’t help FreeToken: it sat at 51–53 tok/s from 1 to 4 concurrent requests, while llama\.cpp peaked at 163\. This is by design: ``` if (is_moe and expert_quant not in ("none", "fp8_block") ``` `and not is\_offload\_moe\_backend\(config\.moe\_backend\)\):` `raise ValueError\(f"\{expert\_quant\} experts require \-\-moe\-backend offload or cpu"\)` Only unquantised BF16 or FP8 experts may live on the GPU\. Every 4\-bit format has to be offloaded\. At 26B, BF16 is 52 GB and FP8 is 26 GB, so on a 24 GB card FreeToken \(at version 0\.1\.2\) has no resident mode for this model\. It streams experts from system RAM that llama\.cpp simply keeps in VRAM\. On identical weights, that’s most of Round 1’s story\. It also hit a wall: FreeToken crashed with a CUDA out\-of\-memory error at 8 concurrent requests, and tripling its activation headroom didn’t move the ceiling\. llama\.cpp served 32 on the same file using 2\.6 GB*less*VRAM\. \(Hold that thought — the ceiling turns out to be more interesting than it looks\.\) One oddity worth flagging\. In a quick early check with the server sized for 4 users, FreeToken decoded a hair*faster*than llama\.cpp: 6\.9 ms per token against 7\.3\. Sized for 32 users, the same engine on the same file took 11\.7 ms\. The reason seems to be that its sliding\-window memory pool is reserved per user slot*before*the expert cache is sized, so provisioning for a crowd shrinks the cache that makes single\-user decoding fast \(80% of expert slots resident at 4\-way, 39% at 32\-way\)\. I also tried Qwen3\.6\-35B\-A3B, hoping it would force a spill\. It didn’t: llama\.cpp held it fully resident at 32 users in 23\.5 GB, and it ran just as fast as the 26B model \(116 vs 114 tok/s\)\. That’s the MoE speed promise:**total parameters are nearly free; active parameters are what you pay for\.**To force a real spill, I needed something bigger\. gpt\-oss\-120b is 63 GB of weights against 24 GB of VRAM\. Now both engines*must*spill and do the same job\. They do it differently\. llama\.cpp computes the experts that don’t fit on the**CPU**\. \(I tuned`\-\-n\-cpu\-moe 26`so ten layers sit on the GPU and the card is actually used; the default left 20 GB of VRAM idle, which would have made llama\.cpp an unfair opponent to beat\.\) FreeToken**streams**the experts each step needs across PCIe and computes them on the GPU\. Move the computation to the data, or the data to the computation\. [![](https://substackcdn.com/image/fetch/$s_!XHW-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e758609-d88f-4587-a7a9-f98838b6eb89_1770x978.png)](https://substackcdn.com/image/fetch/$s_!XHW-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e758609-d88f-4587-a7a9-f98838b6eb89_1770x978.png) *Time to first token on gpt\-oss\-120b versus concurrent requests\. llama\.cpp climbs from 8 seconds to 139 seconds\. FreeToken stays near 9 seconds through 8 users and reaches 19 seconds at 32\.* At one user, FreeToken returned the first token in 4\.7 s against 8\.4 s\. The gap grows with load because**FreeToken’s first\-token latency barely moves**: about 9\.4 s at every load from 2 to 8 users, and 19 s at 32\. Meanwhile, llama\.cpp’s keeps climbing to 8 s, 17 s, 30 s, 58 s, 95 s, 139 s\. My theory is: prefill touches every expert, so streaming the set once and sharing it across a batch amortises the cost, while per\-request expert compute on 8 CPU cores serialises\. It doesn’t win everything\. Per\-token latency was actually*worse*on FreeToken up to 8 users \(696 ms vs 409 ms at 8\-way\)\. Throughput was a near\-tie at low load and, at 16–32 users, FreeToken pulled ahead by about 18–27%\. And the out\-of\-memory wall from Round 1? FreeToken completed all six load levels here, including 32 users\. The difference is attention geometry\. Gemma\-4’s sliding window is 1,024 tokens; gpt\-oss’s is 128\. The per\-slot window pool that ate 13 GiB on Gemma is tiny here\. The “ceiling” was an interaction with large sliding windows, not a general limit of the engine\. [![](https://substackcdn.com/image/fetch/$s_!l1jW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2065922c-0ada-4ac9-bee9-6d1f94d075c2_1764x982.png)](https://substackcdn.com/image/fetch/$s_!l1jW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2065922c-0ada-4ac9-bee9-6d1f94d075c2_1764x982.png) *Generation speed for one user\. With a model that fits in 24 GB, llama\.cpp reaches 114 tokens per second and FreeToken 51\. With gpt\-oss\-120b, which is 2\.6 times the VRAM, llama\.cpp reaches 10\.6 and FreeToken 11\.3\.* The drop between the two groups is about 10×, which dwarfs any difference between engines\. \(Different models with similar active\-parameter counts, roughly 4B and 5B, so it’s a rough comparison, but the message is stable\.\)**For MoE on a consumer card, where the weights live is more important than the engine\.** [![](https://substackcdn.com/image/fetch/$s_!y002!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31e4954b-4667-4ff1-b4c3-c533bb05f21c_1770x750.png)](https://substackcdn.com/image/fetch/$s_!y002!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31e4954b-4667-4ff1-b4c3-c533bb05f21c_1770x750.png) *Two panels across 1 to 32 concurrent requests\. Left: output tokens per second for both engines, between about 10 and 17\. Right: PCIe traffic into the GPU\. FreeToken sits at the 13\.1 GB/s measured ceiling throughout; llama\.cpp uses about 3 to 4 GB/s until 32 users\.* FreeToken on gpt\-oss held the PCIe link at 12\.5–13\.4 GB/s at*every*load level, right at the ceiling I’d measured earlier\. Meanwhile its throughput rose from 11 to 17 tok/s\. With the bus already full, that can only mean fewer bytes per token: about 1\.1 GB per token at one user, 0\.77 GB at 32, because concurrent requests share expert fetches\. llama\.cpp, by contrast, used a quarter of the bus\. It was limited by CPU compute and the GPU↔CPU round trips\. Then a surprise\. I shrank FreeToken’s expert cache from 10% of all expert slots to 5\.6% and then 2\.8%, expecting more bus traffic and slower tokens\. Throughput didn’t move: 18\.7, 18\.2, 18\.1 tok/s\. Bytes per token stayed flat too\. gpt\-oss looks up 4 of 128 experts in each of 36 layers, 144 lookups per token across 4,608 distinct experts, so a cache covering ~10% doesn’t catch enough of them to matter\. \(I couldn’t get a larger cache to load in the configurations I tried, so I can’t say where it would start to help\.\) At this scale, cache tuning isn’t a substitute for bandwidth\. [![](https://substackcdn.com/image/fetch/$s_!MITw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa721ddf9-625f-42fa-a970-139c430d6618_1776x1024.jpeg)](https://substackcdn.com/image/fetch/$s_!MITw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa721ddf9-625f-42fa-a970-139c430d6618_1776x1024.jpeg) *Projection for FreeToken on gpt\-oss\-120b at 32 users\. PCIe 3\.0 measured at 17 tokens per second; PCIe 4\.0 projected at 34; PCIe 5\.0 projected at 68 if only the bus mattered, but a rough compute ceiling of 40 to 55 and a host\-RAM limit would likely cap it lower\.* If the bus is the limit and bytes\-per\-token stays the same, PCIe 4\.0 should roughly**double**FreeToken’s throughput, to about 34 tok/s\. PCIe 5\.0 would run into other limits first\. My rough guess at a compute\-side ceiling is 40–55 tok/s, extrapolated from a different model, so take it lightly\. There’s a second issue: this box’s DDR4 reads about 28 GB/s, while a Gen 5 link would need ~52\. On this kind of CPU, Gen 5 would be starved before it started\. llama\.cpp, limited by the CPU instead of the bus, should gain far less, which would widen FreeToken’s lead on newer hardware\. All of that hangs on one assumption I couldn’t test directly: that bytes\-per\-token doesn’t change when the bus gets faster\. The real test is to downclock the link to Gen 2 and see whether throughput halves\. I didn’t want to touch link speed on a machine that has other jobs to do\. If you have a Gen 4 board and a 3090, I’d love to see your numbers\. - **My own theory, wrong, and how I caught it\.**At 32\-way, FreeToken refused to start with a cache budget of*negative 42\.6 GB*\. The number lined up so neatly with one flag’s units that I was sure I knew why\. I changed the flag and got the error again, byte for byte\. A cause that produces the identical error after you remove it isn’t the cause\. The real culprit: with prefix caching off, FreeToken reserves a fixed sliding\-window pool per request slot roughly 58 GiB for 33 slots on Gemma\-4’s 262k\-token context\. The fix was switching to radix caching \(13 GiB\) for that phase, which means FreeToken got prefix caching that llama\.cpp didn’t\. I expect the effect is small, since my prompts were unique, but I didn’t measure it\. - **The docs’ driver requirement is stricter than the reality\.**FreeToken asks for driver r580\+ and CUDA 13\. My box runs 570 and CUDA 12\.8, and FreeToken works there when built from source\. PyPI’s torch 2\.11\.0 is the CUDA 13 build, so the published wheel ships binaries linked against`libcudart\.so\.13`\.`pip install`looked perfect… and failed on first use\. - **A dead engine looks like a slow one\.**When FreeToken’s backend ran out of memory mid\-sweep, its web frontend kept accepting requests and my load generator sat there for an hour\. - **Live cache resize only shrinks\.**FreeToken’s resize command can shrink the expert cache but won’t grow it back, even to its own starting size, while`status`keeps advertising a range it will refuse\. These are normal rough edges for a 0\.1\.x project, and several are reportable upstream\. One more fairness note: Ampere has no FP4 tensor cores, so FreeToken’s FP4 formats get dequantised on the fly here\. A Blackwell card might tell a different story\. - **The model fits in your VRAM:**llama\.cpp\. It isn’t close\. - **It doesn’t fit, one user:**roughly a wash on speed \(10\.6 vs 11\.3 tok/s\), though FreeToken starts answering sooner\. - **It doesn’t fit, many users \(concurrent agents\), first\-token latency matters, power isn’t your constraint, and the model doesn’t use a big sliding window:**FreeToken is worth trying, and its advantage grows with load\. - **One run per cell\.**The only repeat differed by 9\.4%, so gaps under ~10% aren’t established\. - **Multiple workloads:**I always run 1,024 tokens in, 256 out\. - **Tail latency\.**My load generator was on Wi\-Fi with ±25 ms of jitter, comparable to per\-token latency, so I report means only\. - **Identical files only in Round 1\.**gpt\-oss was natively MXFP4 on both sides but not byte\-identical, and I didn’t sweep llama\.cpp’s`\-\-n\-cpu\-moe`beyond the value that filled VRAM\. - **Output quality\.**I didn’t run the check that both engines answer comparably\. And I only used one type of GPU and a 2020 PC build\. If you run this on a different GPU or PCIe generation, or you think I’ve got something wrong, tell me, especially about that bytes\-per\-token assumption\. #### Discussion about this post ### Ready for more?

Similar Articles

Llama.cpp PR 8% speed boost

Reddit r/LocalLLaMA

A llama.cpp PR moves sampling from CPU to GPU, yielding 8% faster tokens on an RTX 5090 and ~4% on a Tesla P40 for Qwen3.6-35B inference.

MTP+GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 - llama.cpp

Reddit r/LocalLLaMA

A user benchmarks token generation speed on llama.cpp with the GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 flag, comparing performance with and without MTP (Multi-Token Prediction). Results show a significant speedup from 49 tok/s to 64 tok/s when MTP is enabled on an RTX5090 with a Qwen3.6-27B model.