@TechMDAI: GLM-5.3-Flash EXL3-3.0bpw by @0xSero @BrandonMusicKy @LottoLabs @localmaxxing 193.8 tk/s
Summary
The article highlights the GLM-5.3-Flash EXL3-3.0bpw AI model with an inference speed of 193.8 tokens per second, attributed to multiple contributors.
View Cached Full Text
Cached at: 08/30/26, 10:13 PM
GLM-5.3-Flash EXL3-3.0bpw by @0xSero @BrandonMusicKy @LottoLabs @localmaxxing
193.8 tk/s https://t.co/aJIBGyb0W1
Similar Articles
GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context
The user benchmarks the GLM-5.3-Flash AI model on a DGX Station, achieving ~206 tokens per second in single-stream inference with a 1 million context window, and shares a Docker command for setup.
GLM-5.3-Flash
Release of GLM-5.3-Flash, an AI language model optimized for fast inference and performance updates.
@Yuchenj_UW: GLM-5.3 at 310 tok/s! Databricks inference is #1 in both speed and latency, again. On our internal Databricks coding be…
GLM-5.3 achieves 310 tokens per second on Databricks inference, leading in both speed and latency, and is the strongest open-source model for coding, competitive with Fable 5 and Opus 4.8.
@Ex0byt: Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a sing…
Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.
@philipkiely: https://x.com/philipkiely/status/2069212319746506968
Baseten announces the world's fastest API for the GLM-5.2 open model, achieving over 280 tokens per second via NVFP4 quantization, disaggregated inference, and other optimizations.