@PyTorch: While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering …
Summary
SGLang provided Day-0 support for DeepSeek-V4, and collaboration between LMSys and NVIDIA engineering teams achieved up to 5x throughput increase in production, with improvements shown on the SemiAnalysis InferenceX dashboard.
View Cached Full Text
Cached at: 06/24/26, 03:57 AM
While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering teams has taken its production performance to the next level.
According to the public SemiAnalysis InferenceX dashboard, the GB300 disaggregated lane (DeepSeek-V4 Pro, FP4, 8K/1K) saw a 5x throughput increase—surging from ~2,200 to ~11,200 tok/s/GPU at identical interactivity levels. These updates sustain high throughput much deeper into target interactivity ranges most deployments target, while also driving a 2.9x lift on the Blackwell Ultra aggregated lane.
Find the full technical breakdown in the comments below:
Similar Articles
@TheAhmadOsman: DeepSeek V4 Flash is ~70% smaller in size than GLM 5.2 It also beats GLM 5.2 which was the SoTA model just about a mont…
DeepSeek V4 Flash is about 70% smaller than GLM 5.2 yet outperforms it, implying state-of-the-art-level AI could run on consumer hardware like an RTX 5090 much sooner than expected.
@Ex0byt: Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a sing…
Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.
Optimised DSv4-Flash for 2x GH200: 10,000 tok/s PP, >300 tok/s TG on SGLang
A detailed benchmark of DeepSeek V4 Flash on a dual GH200 workstation, comparing SGLang and vLLM at 1M context with DSpark speculative decoding, finding SGLang faster at ~317 vs ~276 tok/s.
@PyTorch: LightSeek (@lightseekorg) TokenSpeed brings Day-0 optimized FP4 inference support for Thinking Machines Lab’s Inkling a…
LightSeek's TokenSpeed provides Day-0 optimized FP4 inference support for Thinking Machines Lab's Inkling model on NVIDIA and AMD accelerators, built on PyTorch and collaborating with vLLM.
DeepSeek open sources DSpark, a new framework to speed up LLM inference by up to 85% (18 minute read)
DeepSeek open-sourced DSpark, an MIT-licensed framework using speculative decoding to accelerate LLM inference by up to 85%, with support for multiple model families including its own DeepSeek-V4, Alibaba's Qwen, and Google's Gemma.