How long can I expect to wait until the local ~30B A3B frontier catches up to GLM 5.3 Flash quality?

Reddit r/LocalLLaMA News

Summary

The user queries whether local AI models around 30B parameters can achieve GLM 5.3 Flash quality within a year, given current hardware constraints like 16GB RAM and 8GB VRAM.

The jump from Qwen3 Coder 30B A3B to current-day Qwen 3.6 35B A3B is crazy, especially with all the fine-tunes, and that was around 6 months (I didn't care for local AI back then, or AI at all, apart from as a toy so I don't know). Is around a year until I will never need cloud without buying ridiculously expensive hardware (or any extra hardware at all, just what I have; 16 GB RAM + 8 GB VRAM) a reasonable estimate? Or can I daydream about it happening even faster?
Original Article

Similar Articles

GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra

Reddit r/LocalLLaMA

Optimizations for GLM 5.3 Flash on Apple M3 Ultra achieve up to 550 t/s prefill and 38 t/s inference speed through kernel fusion and efficient memory use, without quality loss.