bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
Summary
Tim Dettmers, creator of bitsandbytes, is teasing a new quantization method that reportedly runs GLM 5.3 on a single DGX Spark at 7 tokens/s, with cautious optimism from the community.
Similar Articles
4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
Running a 4-bit quantized version of GLM-5.2 (753B MoE) on 4 DGX Spark machines achieves 70.8% on Terminal-Bench 2.1, compared to 81.0% from the full model.
@Ex0byt: Update: the road to GLM-5.2: we're getting there, folks! non-quantized, non-pruned DeepSeek-v4-Flash. 11tok/s on a sing…
Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.
GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode
Describes deployment and benchmarking of the quantized GLM-5.2-Int4-Int8Mix model on an 8-node DGX Spark (GB10) cluster using a custom vLLM fork, achieving ~1,200 t/s prefill and ~35 t/s decode with MTP tool calling.
@TheAhmadOsman: Luke Alonso has uploaded an NVFP4 of GLM 5.2 467GB, would fit on 4x DGX Sparks (~$20k)
Luke Alonso uploaded an NVFP4 quantized version of GLM 5.2 (467GB) that can fit on 4x DGX Sparks hardware, costing approximately $20k.
GLM 5.2 on 4x Sparks reasonable?
A user asks about the feasibility of running GLM-5.2 at 4-bit quantization on four Ascend GX10s or DGX Sparks, wondering about speed and memory for 100k context.