@sgl_project: We pushed some updates to the RTX 5090 / RTX Pro 6000 recipes in the Qwen3.8-27B cookbook http://docs.sglang.io/cookboo…

X AI KOLs Timeline Tools

Summary

SGLang has updated its deployment recipes for the Qwen3.8-27B model on RTX 5090 and RTX Pro 6000 hardware, adding variants for different configurations with tuning options.

We pushed some updates to the RTX 5090 / RTX Pro 6000 recipes in the Qwen3.8-27B cookbook http://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B… • Added recipe variants for non-spec, MTP, and DSpark • Included high throughput and low latency options All recipes are good starting points to tinker. Expect to tune flags and configs for the use case you actually care about. Thanks for all the feedback from the community!
Original Article
View Cached Full Text

Cached at: 08/18/26, 12:22 AM

We pushed some updates to the RTX 5090 / RTX Pro 6000 recipes in the Qwen3.8-27B cookbook http://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B…

• Added recipe variants for non-spec, MTP, and DSpark • Included high throughput and low latency options

All recipes are good starting points to tinker. Expect to tune flags and configs for the use case you actually care about.

Thanks for all the feedback from the community!


Qwen3.8-27B - SGLang Documentation

Source: https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B

Deployment

Install SGLang

For all methods and hardware platforms, see theofficial SGLang installation guide. The two paths below match thePython / Dockertoggle in the command panel.

  • Python (pip / uv)
  • Docker

Then run thePythonoutput of the command panel below in that environment.

Pick your card + checkpoint precision to generate the launch command. The model runs single-GPU on every supported card — H200, RTX PRO 6000, RTX 5090 and DGX Spark — and ships one operating point.

Mamba ratio calculator

How --mamba-full-memory-ratio is calculated

Hybrid GDN models split post-weight memory into a worst-case-reservedGDN state pool(sets the concurrency ceiling) and a pagedattention KV pool, divided by\-\-mamba\-full\-memory\-ratio. Every parameter below exceptLand the target concurrency is read live from the Deploy panel and Playground selection; the balanced value is the per-request cost ratio:

  • S— state slots per running request:extra\_buffer=5(default),extra\_buffer\_lazy=4,no\_buffer=3, disabled radix cache=1. For the twoextra\_bufferstrategies,SGLANG\_OPT\_MAMBA\_SKIP\_DECODE\_LOCK=1frees one slot, andextra\_bufferfrees one more with the overlap scheduler off; the calculator reads both knobs.
  • D— verify intermediate states under speculative decoding:\-\-speculative\-num\-draft\-tokensfor EAGLE/MTP (4 at the recommended 3/1/4);\-\-speculative\-dspark\-block\-size \+ 1for DSPARK, where the block size falls back to the draft checkpoint’sblock\_sizewhen the flag is omitted (7 forRadixArk/Qwen3\.8\-27B\-DSpark, soD = 8); 0 with speculation off or with\-\-enable\-linear\-replayssm\-spec, which keeps the verify intermediates on a fixed ring instead of per-request slots.
  • state\_bytes— one state slot, from the fixed geometry (48 GDN layers x 48 heads x 128 x 128 at\-\-mamba\-ssm\-dtype, plus bf16 conv state): 153.9 MB at fp32, 78.4 MB at bf16.
  • kv\_bytes\_per\_token— 16 attention layers x GQA 4 x 256 x K+V: 32.8 KB at fp8, 65.5 KB at bf16.
  • L— average total request length in tokens: input + output.

\-\-max\-mamba\-cache\-size = target\_concurrency x Sis the equivalent explicit pin and overrides the ratio; the calculator emits it alongside.Dis not a term here: the engine divides the state pool bySalone and sizes the speculative verify buffer separately, so foldingDinto the pin would over-provision the pool. After boot, verify with themax\_running\_requestsline in the server log — it should not be capped below your target concurrency.

Playground

The Playground is where you experiment withSGLang features beyond the recipes above. The Deploy panel emits this model’s documented launch recipes; the Playground lets you turn on additional knobs on top of whichever cell the Deploy panel is currently showing.

1. Model Introduction

Qwen3.8-27Bis a dense hybrid Gated Delta Networks (GDN)vision-languagemodel: a 27B causal language model paired with a vision encoder, with native image and video understanding alongside text. SGLang serves it through the Qwen3-VL path, so the vision tower is live on the recipes below.The language model is 64 layers, laid out as 16 repeats of3 × (Gated DeltaNet → FFN)followed by1 × (Gated Attention → FFN)— 48 linear-attention layers to 16 full-attention ones. Gated DeltaNet runs 48 value heads and 16 QK heads at head_dim 128; Gated Attention is GQA 24/4 at head_dim 256 with a 64-dim rotary slice. Hidden size is 5120 over a 17,408-dim FFN, and the checkpoint ships an MTP head trained with multiple steps. Context is 262,144 tokens natively, extensible to 1,000,000. The serving-relevant architecture is identical to Qwen3.6-27B.Thinking mode is on by default and can be disabled per request; reasoning depth is tunable withreasoning\_effort, andpreserve\_thinkingretains reasoning context from earlier messages.

ModelQuantizationWeightsQwen3.8-27BBF16Qwen/Qwen3.8-27BQwen3.8-27B-FP8FP8 (blockwise)Qwen/Qwen3.8-27B-FP8Qwen3.8-27B-NVFP4NVFP4 W4A4 + FP8 projectionsRadixArk/Qwen3.8-27B-NVFP4The NVFP4 checkpoint declareskv\_cache\_quant\_algo: FP8; SGLang’s default\-\-kv\-cache\-dtype autohonors it, so the KV pool runs infp8\_e4m3with the checkpoint’s calibration scales automatically.

2. Configuration Tips

  • SM120/SM121 (RTX PRO 6000 Blackwell, RTX 5090, DGX Spark): use\-\-attention\-backend flashinfer;trtllm\_mhais SM100-only. MTP with the FlashInfer backend requires a FlashInfer build whose prefillplanacceptsuniform\_q\_len(newer than 0.6.15.post1); otherwise run spec with\-\-attention\-backend triton. On DGX Spark the 128GB is unified memory shared with the host CPU, so all three checkpoints fit, and its cells reuse the RTX PRO 6000 recipe verbatim rather than a separate operating point.Validated on SM121 / aarch64: all 36 configurations (3 checkpoints x Speculative Decoding x Serving Strategy x Mamba SSM Dtype) booted and served on GB10 underlmsysorg/sglang:qwen38\-27bat ISL 8192 / OSL 1024, concurrency 1. That is boot-and-serve coverage only — no throughput or acceptance-length numbers — and it includes the FlashInferplan/uniform\_q\_lenpath above, which raised no arity error on that image. Two host quirks when reproducing on GB10: docker GPU access is CDI-only (\-\-device nvidia\.com/gpu=all, as nonvidiaruntime is registered), andnvidia\-smireportsNot Supportedfor memory because it is unified with the CPU — gate a relaunch onMemAvailablein/proc/meminfoinstead.
  • H200 (SM90): BF16 and FP8 only — the card has no FP4 tensor cores, so the NVFP4 checkpoint’s MLP would fall back to the Marlin W4A16 weight-only path and its cell is greyed out. The H200 recipes use 32768-token prefill chunks (SM90 prefill is fast enough that a big chunk barely stalls decode, unlike the SM120 guidance below), and the FlashInfer GDN prefill backend engages by default under them.\-\-attention\-backend fa3is a valid alternative, measured slightly faster at bs=1.
  • MTP:\-\-speculative\-algorithm EAGLE \-\-speculative\-num\-steps 3 \-\-speculative\-eagle\-topk 1 \-\-speculative\-num\-draft\-tokens 4uses the in-checkpoint MTP head. (This recipe was originally documented withNEXTN, an alias ofEAGLE— same algorithm.)
  • DSpark: the trained draft model is a separate checkpoint — add\-\-speculative\-algorithm DSPARK \-\-speculative\-draft\-model\-path RadixArk/Qwen3\.8\-27B\-DSpark(the Playground’s Speculative Decoding card emits this pair). DSpark doesnottake\-\-speculative\-num\-draft\-tokens: its verify window is\-\-speculative\-dspark\-block\-size(gamma)+ 1, and gamma is auto-inferred from the draft checkpoint when the flag is omitted (7 for this checkpoint, so D = 8). ThatDis a term in the balanced ratio —r = \(S \+ D\) x token\_equiv / L, wheretoken\_equivis the state slot expressed in KV tokens,state\_bytes / kv\_bytes\_per\_token(4698 at fp32 state / 2394 at bf16, over fp8 KV) — so DSpark needs a materially higher\-\-mamba\-full\-memory\-ratiothan no-spec at the sameS, and pinning a different gamma changes the ratio with it. MTP is the opposite case: with\-\-enable\-linear\-replayssm\-specits draft intermediates move onto a fixed ring, soD = 0and the ratio returns to the no-spec value. Thecalculatorapplies both rules.
  • Hardware fit: FP8 weights ~28.5GB (not serviceable beyond bs≤2 on 32GB cards); NVFP4 weights ~16.5GB (recommended for RTX 5090-class GPUs).
  • \-\-mamba\-radix\-cache\-strategy extra\_buffer\_lazylowers the state cost per request from 5 slots to 4 at no accuracy cost. On small-VRAM cards (RTX 5090 32GB) the state pool bounds concurrency long before KV does — prefer loweringS(lazy strategy, or\-\-disable\-radix\-cachefor S=1); thecalculatorre-derives the ratio for the newS. The balanced ratio itself is VRAM-independent.
  • \-\-mamba\-ssm\-dtype: the GDN state slot is153.9 MB atfloat32(the checkpoint’s declared precision) and78.4 MB atbfloat16, so bf16 roughly halves the state pool and hands the difference to KV — measured on an RTX 5090 with no speculation, 97,280 KV tokens at bf16 against 68,588 at fp32. On 32GB cards it also decides whether a config fits at all: EAGLE needs\-\-mem\-fraction\-static 0\.94at fp32 but 0.92 at bf16. Speed isnota one-way trade — with speculative decoding fp32 sometimes wins (NVFP4 + EAGLE: 152.9 vs 144.5 tok/s/user) and sometimes loses (FP8 + EAGLE: 106.3 vs 116.1); measure both for your quantization. Treatbfloat16as an accuracy gate and validate it for your workload. On SM120 both precisions run the Triton linear-attn prefill path — the FlashInfer GDN prefill fast path gates on SM100, where its validated domain is in fact a bf16 state pool — so no dtype forces an extra flag here. One interaction to know:\-\-enable\-linear\-replayssm\-specauto-selects fp32 state when\-\-mamba\-ssm\-dtypeis unset, and an explicit non-fp32 value logs a state-drift warning at boot. The SSM dtype row always emits the flag explicitly, so the bf16 + EAGLE cells run with that warning — accounted for in their validation.
  • \-\-chunked\-prefill\-size 2048: decode steps stall behind each prefill chunk on hybrid GDN models, and 8192-token chunks stall them ~600ms at a time. 2048 keeps decode inter-token latency smooth under mixed load and also improves single-wave TTFT.

3. Agent Harnesses

Agent harnesses drive the model through the OpenAI-compatible endpoint — or, for Claude Code, through SGLang’s Anthropic-compatible one — so any of them works once three things line up.The parsers ship in the command.Every recipe above carries\-\-reasoning\-parser qwen3 \-\-tool\-call\-parser qwen3\_coder, because without them a harness receives tool calls as raw text instead of structuredtool\_calls. TheParserscard in thePlaygroundis therefore an opt-out — both chips start on, and turning one off strips its flag.qwen3\_coderis the right tool-call parser for this checkpoint: its chat template instructs the model to reply with an inner<function=…\>/<parameter=…\>block nested in<tool\_call\></tool\_call\>, which is exactly what that parser decodes. The Hermes parser (\-\-tool\-call\-parser hermes) reads adifferentpayload — bare JSON inside<tool\_call\>— so pointing a Hermes-format harness at this model without switching the flag yields tool calls that never parse.\-\-reasoning\-parser qwen3matches the template’senable\_thinkingtoggle, which defaults to on.**Endpoint and model id.**The base URL ishttp://<host\>:30000/v1. Themodelstring a harness sends must equal the server’s\-\-model\-path— the OpenAI/v1/modelsname defaults to it — unless you override it with\-\-served\-model\-name, which is usually worth doing to keep harness configs short.SGLang also serves an Anthropic-compatible/v1/messages, which is what§3.3uses. It converts each request to the OpenAI shape, hands it to the same chat-serving path, and converts the response back — so the parser flags above apply there identically.Auth.\-\-api\-keyis unset by default, so the server accepts unauthenticated requests. Harnesses that insist on a key can send any placeholder; set\-\-api\-keyon the server if the endpoint is reachable beyond localhost.

3.1 OpenCode

OpenCodereaches a self-hosted endpoint through a provider entry inopencode\.json.

Register SGLang as an OpenCode provider

Store the credential first — pickOther, give the provider an id, and enter any placeholder when the server has no\-\-api\-key:

Then declare the provider inopencode\.json:

npmselects the transport —@ai\-sdk/openai\-compatibleis the one for a plain OpenAI-shaped endpoint.apiKeyis optional and takes a"\{env:VAR\_NAME\}"reference rather than a literal. Themodelskeys are the ids sent on the wire, so they must match the served model name. Confirm with/models.

3.2 Pi

Pi(@earendil\-works/pi\-coding\-agent) registers providers from an extension rather than a config file.

Register SGLang as a Pi provider

api: "openai\-completions"is what selects the OpenAI-compatible transport, andapiKeytakes a$ENV\_VARreference rather than a literal.contextWindowis the checkpoint’s native 262,144; setmaxTokensto whatever output cap you want per turn. Confirm registration withpi \-\-list\-models.

3.3 Claude Code

Claude Code speaks the Anthropic API, so it points at SGLang’s/v1/messagesrather than the OpenAI endpoint.

Point Claude Code at SGLang

ANTHROPIC\_BASE\_URLis the server origin — Claude Code appends/v1/messagesitself, so leave the/v1suffix off:

The two credential variables travel in different headers:ANTHROPIC\_AUTH\_TOKENgoes out asAuthorization: Bearer,ANTHROPIC\_API\_KEYasx\-api\-key. Either satisfies a server started without\-\-api\-key; with\-\-api\-keyset, pick the variable matching the header your server reads. A credential variable also takes precedence over a saved claude.ai login for that session.The same pair can live in a settings file instead, which persists across shells and wins over a shell export:

Run/statusin Claude Code to confirm which base URL and credential source the session picked up.

3.4 Hermes Agent

Hermes Agent(Nous Research, MIT) selects a self-hosted endpoint through its setup wizard or its config file.

Point Hermes Agent at SGLang

Equivalently, in~/\.hermes/config\.yaml:

For several endpoints at once, declare them underproviders:and switch with/model custom:<name\>mid-session:

SGLang (@sgl_project): The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:

  • 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark
  • 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model

Similar Articles

Here's your recipes for models for the most popular hardware

Reddit r/LocalLLaMA

A curated collection of LLM inference recipes and performance benchmarks for popular hardware configurations, including NVIDIA RTX Pro 6000 Blackwell, H100, AMD Strix Halo, and RTX 5090, with detailed batch size, quantization, and throughput data.