A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing
Summary
An observation of Ling-3.0-flash (124B) on one DGX Spark generating 15,128 tokens in a single response with stable decode throughput around 35.6 tok/s, highlighting long-context decoding performance.
Similar Articles
@LotusDecoder: DeepSeek-V4.1-Flash-0910 decode 400 token/s 😋 Would deploying this to my home DGX spark also achieve this speed?
A user on X/Twitter asks if deploying the DeepSeek-V4.1-Flash-0910 model on a home DGX Spark could achieve a decode speed of 400 tokens per second.
Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box
An independent benchmark by sudoingX shows the Ling-3.0-flash model runs at 38.7 tok/s on a single DGX Spark with official INT4 quantization, 2.4x faster than DeepSeek V4 Flash on the same hardware, after a correction clarifying the quants do work.
GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context
The user benchmarks the GLM-5.3-Flash AI model on a DGX Station, achieving ~206 tokens per second in single-stream inference with a 1 million context window, and shares a Docker command for setup.
Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks
Achieved 181 tokens per second aggregate throughput on the Qwen3.8-Flash-Next model using a 2x DGX Spark cluster with optimizations like NVMe mapping and speculative decoding.
@onusoz: 16x parallel Gemma-4-26B-A4B-NVFP4 runs 18 output tokens/s, aggregate 300 tok/s 1 DGX Spark with 128 GB unified memo…
@onusoz demonstrates running 16 parallel instances of NVIDIA's quantized Gemma-4-26B-A4B-NVFP4 model on a single DGX Spark with 128GB unified memory, achieving 300 tok/s aggregate, showcasing high concurrency without flashinfer.