GLM 5.2 on 4x Sparks reasonable?

Reddit r/LocalLLaMA News

Summary

A user asks about the feasibility of running GLM-5.2 at 4-bit quantization on four Ascend GX10s or DGX Sparks, wondering about speed and memory for 100k context.

So GLM-5.2 is obviously a very good model, and I'm wondering how fast it would run on four Ascend GX10s / DGX Sparks. I can't find any data online at all. Wouldn't it be possible to run a 4bit quant on 4*128=512GB unified memory? What would the prompt processing and output token/sec be for e.g. 100k context?
Original Article

Similar Articles

Thinking about grabbing 4x Ascend GX10s

Reddit r/LocalLLaMA

A user considers buying four Ascend GX10s to run GLM5.2, citing performance numbers like 400-500 tok/s prompt processing and ~15 tok/s output at 128k context, and plans for future open-source models.