@Ex0byt: Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a sing…

X AI KOLs Timeline Models

Summary

Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.

Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a single DGX Spark. sglang inference + custom mega-kernel. Pure beauty. https://t.co/vRpHIFHqOO
Original Article
View Cached Full Text

Cached at: 06/24/26, 12:23 PM

Update: the road to GLM-5.2: we’re getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a single DGX Spark. sglang inference + custom mega-kernel. Pure beauty. https://t.co/vRpHIFHqOO

Similar Articles

Deepseek V4 flash performance on DGX Spark

Reddit r/LocalLLaMA

A Reddit user shares their experience running DeepSeek V4 Flash on a dual-ASUS GX10 DGX Spark setup, detailing performance metrics, configuration, and power consumption, with throughput benchmarks across various context lengths.