@PyTorch: While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering …
Summary
SGLang provided Day-0 support for DeepSeek-V4, and collaboration between LMSys and NVIDIA engineering teams achieved up to 5x throughput increase in production, with improvements shown on the SemiAnalysis InferenceX dashboard.
View Cached Full Text
Cached at: 06/24/26, 03:57 AM
While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering teams has taken its production performance to the next level.
According to the public SemiAnalysis InferenceX dashboard, the GB300 disaggregated lane (DeepSeek-V4 Pro, FP4, 8K/1K) saw a 5x throughput increase—surging from ~2,200 to ~11,200 tok/s/GPU at identical interactivity levels. These updates sustain high throughput much deeper into target interactivity ranges most deployments target, while also driving a 2.9x lift on the Blackwell Ultra aggregated lane.
Find the full technical breakdown in the comments below:
Similar Articles
@sgl_project: We just added recipes for DeepSeek-V4-Flash-Vision & DeepSeek-V4-Flash-0731 on 2x DGX Spark. https://docs.sglang.io/coo…
SGLang has added deployment recipes for DeepSeek-V4-Flash-Vision and DeepSeek-V4-Flash-0731 models on 2x DGX Spark hardware, with support for various configurations and optimizations.
@TheAhmadOsman: DeepSeek V4 Flash is ~70% smaller in size than GLM 5.2 It also beats GLM 5.2 which was the SoTA model just about a mont…
DeepSeek V4 Flash is about 70% smaller than GLM 5.2 yet outperforms it, implying state-of-the-art-level AI could run on consumer hardware like an RTX 5090 much sooner than expected.
@Ex0byt: Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a sing…
Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.
Optimised DSv4-Flash for 2x GH200: 10,000 tok/s PP, >300 tok/s TG on SGLang
A detailed benchmark of DeepSeek V4 Flash on a dual GH200 workstation, comparing SGLang and vLLM at 1M context with DSpark speculative decoding, finding SGLang faster at ~317 vs ~276 tok/s.
@PyTorch: LightSeek (@lightseekorg) TokenSpeed brings Day-0 optimized FP4 inference support for Thinking Machines Lab’s Inkling a…
LightSeek's TokenSpeed provides Day-0 optimized FP4 inference support for Thinking Machines Lab's Inkling model on NVIDIA and AMD accelerators, built on PyTorch and collaborating with vLLM.