bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s

Reddit r/LocalLLaMA News

Summary

Tim Dettmers, creator of bitsandbytes, is teasing a new quantization method that reportedly runs GLM 5.3 on a single DGX Spark at 7 tokens/s, with cautious optimism from the community.

Don't get too hyped + take with a grain of salt as there have been an endless amount of quantization schemes with big promises that never really became a thing. Tim Dettmers is a pretty well known researcher though, so maybe something will come of this. Time shall tell. Another tweet about the method, DS4 Pro on a single B300 (288 GB VRAM): https://xcancel.com/Tim_Dettmers/status/2087624491362820364
Original Article

Similar Articles

GLM 5.2 on 4x Sparks reasonable?

Reddit r/LocalLLaMA

A user asks about the feasibility of running GLM-5.2 at 4-bit quantization on four Ascend GX10s or DGX Sparks, wondering about speed and memory for 100k context.