The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.
Summary
MegaCapybara 是一个针对 NVIDIA RTX 5090 优化的开源 LLM 推理引擎,支持 Qwen3.8-27B,单智能体可达 500+ tokens/s,多智能体并发最高 2000+ tokens/s,具备投机解码、共享上下文池和智能 VRAM/RAM/磁盘缓存管理。
View Cached Full Text
Cached at: 10/03/26, 02:56 AM
perkel666/MegaCapybara
Source: https://github.com/perkel666/MegaCapybara
MegaCapybara
The fastest inference engine for Qwen3.8-27B on the NVIDIA GeForce RTX 5090.
Up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, on Windows 11 and Linux.
A launcher sets it up and shows, before you load, what every setting costs in accuracy, speed and memory.
Download for Windows · Download for Linux · Models on Hugging Face · Usage guide
Twelve coding agents working at once on one RTX 5090 with the Tiny model, at about 1,500 tokens/s.
Features
Fast: built for the RTX 5090
State-of-the-art inference for Qwen3.8-27B on the RTX 5090, for one agent and for many at once:
- Kernels tuned to this card. The weights are stored in MXFP4 and MXFP6, 4- and 6-bit formats that the RTX 5090’s tensor cores compute natively. Each step takes barely longer than reading the weights from memory once, which is the card’s hard limit.
- Speculative decoding. A small DFlash2 drafter guesses the next few tokens and the model checks them all in one step: the same answers, 3-4 times faster on code.
- Confidence scheduling, from DSpark. Every round, each conversation drafts only as many tokens as are likely to be accepted.
| Tokens per second | Tiny | Small | Medium | Large | XXL |
|---|---|---|---|---|---|
| One agent coding | 540 | 495 | 440 | 377 | 313 |
| 8 agents coding, in total | 1,890 | 1,754 | 1,638 | 1,427 | 1,254 |
Greedy decoding with thinking off, on an RTX 5090 with its memory overclocked by 20%. A stock card is somewhat slower.
Built for a team of agents
Agents never work in step: one reads a 100K-token codebase while another writes a two-line fix. So instead of a fixed slice each, MegaCapybara gives them one shared pool of context:
- any agent can use the model’s full context when its task needs it;
- no memory sits reserved for agents that are idle;
- long tasks are never cut short or compacted early.
Fixed slots are one switch away if you prefer them.
Before every GPU step, the scheduler decides what runs:
- Continuous batching. All agents that are writing get their next tokens in the same step, up to 12 at once. A new request joins at the next step, and its prompt is read a piece at a time between the steps, so even a 100K-token prompt never freezes the agents that are writing.
- Batched speculative decoding. Each conversation drafts its own tokens, and one pass checks the drafts of all of them together.
- Paged memory. The pool is split into pages of 512 tokens. A conversation takes pages as it grows and gives them back when it is done.
- First come, first served. The request at the front of the queue starts as soon as there is room for its prompt, so small requests can never keep a big one waiting forever.
- Prefix reuse. A returning agent goes back to where its conversation is cached, and only the new part of its prompt is read.
- Queue-aware eviction. When memory runs short, idle conversations move from VRAM to RAM, then to disk. The ones no waiting request needs go first, then the ones the queue needs last: the idea behind Belady’s optimal policy. In our multi-agent test this served 20% more prompt tokens than plain LRU eviction.
- Preemption by swapping. If every agent that is writing needs more room at the same moment, the youngest one moves to RAM whole. It comes back before any new request starts and continues exactly where it stopped: nothing is lost and nothing is computed twice.
Pick up where you left off
When an agent sends its next request, its conversation is usually still cached: in VRAM, in RAM (back in about 3 ms) or on disk (back in seconds). The reply starts right away instead of reading a long prompt again from the start.
See what each setting costs before you load
Every model file carries its own accuracy scores, measured against Qwen’s original weights, and every setting has a measured cost. The launcher adds them up for exactly the settings you pick:
- how close the answers stay to the original;
- how fast it runs;
- how much VRAM, RAM and disk it needs.
Switch to a smaller model or a 4-bit KV cache and you see right away what you trade for the extra speed or context.

How close to the original
Top-1 agreement is how often the model picks the same next token as Qwen’s original weights: higher is closer. Unsloth’s GGUF quants are widely considered the state of the art for running models locally, so here is how our model files compare with them, size for size.
Unsloth’s stay a little closer at the same size. Ours use the 4- and 6-bit formats that the RTX 5090 computes natively, which is where MegaCapybara’s speed comes from.
Every setting explained
Hover any ? in the launcher for a plain explanation of its setting: an animated picture, the options compared on
your PC, and their pros and cons. The animations on this page are taken from those panels.

Models in one click
Press the arrow next to the model to list every model file on Hugging Face, then download one with a click:
- 8 connections at once;
- resumes after an interruption;
- checked against its SHA-256 hash;
- no account needed.

Works with your tools
MegaCapybara speaks both OpenAI’s and Anthropic’s APIs, including tool calls, thinking and images. Point your client at it:
| Client | Base URL |
|---|---|
| OpenCode, Cline, Continue, Open WebUI and other OpenAI-compatible clients | http://127.0.0.1:8080/v1 |
| Claude Code and Anthropic’s SDKs | http://127.0.0.1:8080 |
- On your network: set a password first, with the lock next to the address in the launcher.
- On Linux: the same launcher runs on X11, Wayland and WSL2, or the server runs alone over SSH.
Setup for each client: docs/USAGE.md.
Prefer the terminal?
The launcher is optional. The presets folder has ready-made start scripts, .bat on Windows and .sh on Linux:
the launcher’s three presets and six more, one for each model size and number of agents. Run one and the server
starts with a live dashboard showing the speed, prompt reading, each agent’s context, and what every conversation is
doing.
- Edit a preset: each script explains every setting right above it. Change the values, or copy the file to make your own.
- Or copy from the launcher: set everything up there and press Copy at its bottom right. You get the full command line for those settings, ready for a terminal or a script.


Every option is in docs/USAGE.md and in megacapybara --help.
Get started
- Download the archive for your system from Releases and unpack it anywhere.
- Run
MegaCapybaraLauncher. Windows may warn you once because the programs are not code-signed: click More info, then Run anyway. - Press the arrow next to the model and download one. Medium is a good start.
- Pick a preset and press Load server.
- Point your client at
http://127.0.0.1:8080/v1and set its context size to the number under Context size for your front end.
Requirements
- An NVIDIA GeForce RTX 5090 with driver R580 or newer.
- Windows 11 x64, or Linux x86-64 with glibc 2.38 or newer (Ubuntu 24.04, Debian 13, Fedora 39 or newer; WSL2 works too).
- 32 GB of RAM or more, and 15-30 GB of disk per model.
Nothing else to install: the programs are self-contained and the folder can live anywhere you can write to. On Linux, the launcher uses the desktop’s own X11, Cairo, Pango and libcurl, and offers to install any that are missing.
Models
Converted from Qwen/Qwen3.8-27B, in five sizes: perkel/Qwen3.8-27B-MC.
| Size | VRAM | KL divergence | Top-1 agreement | Good for |
|---|---|---|---|---|
| Tiny | 12.70 GiB | 0.0582 | 92.7% | the most speed and context |
| Small | 13.67 GiB | 0.0478 | 93.8% | speed, with a little more accuracy |
| Medium | 15.39 GiB | 0.0298 | 95.2% | the balance, and the presets’ choice |
| Large | 18.66 GiB | 0.0081 | 97.7% | answers close to the original |
| XXL | 23.58 GiB | 0.0051 | 98.3% | the closest to the original, less room for context |
Both measured against Qwen’s original BF16 weights on 81,880 held-out tokens. KL divergence: 0 is identical, lower is closer. Top-1 agreement: how often the most likely next token is the same.
The same five sizes also come uncensored, converted from an abliterated release (one with the refusals removed): perkel/Qwen3.8-27B-Uncensored-MC.
License
MegaCapybara is released as programs for Windows and Linux for now; the source code will be published later.
MegaCapybara is under the MIT License, (c) 2026 Perkel’s Software Corner. The model weights and the
libraries it includes keep their own licenses: see THIRD-PARTY-NOTICES.md. Each release
lists its changes in its notes and in the package’s CHANGELOG.txt.
Support
If MegaCapybara is useful to you, you can support it on Patreon. The money goes to the GPUs needed to support more cards and multi-GPU.
Similar Articles
@iotcoi: Qwen3.6-27B-FP8 + Dflash + DDTree, 256k context, 10 agents ~200 tokens/sec max decode 136t/s average on a single tiny G…
Quantized 27B Qwen3.6 model achieves 200 tok/s peak (136 avg) with 256k context and 10 agents on a single 49W GB10 GPU using Dflash+DDTree optimizations.
Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings
Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.
Qwen 3.6 27B Speculative Decoding Bench: Pushing ~100 TPS on a single RTX 3090
A detailed benchmark comparing speculative decoding engines for Qwen 3.6 27B on a single RTX 3090, showing ik_llama achieving ~100 tokens per second in code generation. Results include decode TPS, TTFT, VRAM usage, and context degradation across 5 engine variants.
Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
Nifer is a tool that achieves 700 tokens per second inference on Qwen 3.6 35B without thinking, specifically optimized for the RTX 5090, with support for full 250k context.
@servasyy_ai: https://x.com/servasyy_ai/status/2106563331976991217
This long post is a real-world test of running Qwen3.8-27B on two second-hand V100 32G GPUs. By rewriting the kernel of the inference engine NInfer, adding dual-GPU support and speculative decoding, generation speed on a 200K-token code task jumped from 36 tok/s to roughly 100 tok/s, beating the 80.9 tok/s of a single 4090 48G. The post also analyzes how VRAM bandwidth, quantization, and NVLink affect inference performance.