CPU-only GLM 5.2: Epyc and 512GB RAM

Reddit r/LocalLLaMA Models

Summary

GLM 5.2 is optimized for CPU-only inference on AMD Epyc processors with 512GB RAM.

No content available
Original Article

Similar Articles

Show HN: Getting GLM 5.2 running on my slow computer

Hacker News Top

Colibrì is a pure C inference engine that runs the 744B GLM-5.2 MoE model on consumer hardware with ~25GB RAM by streaming experts from disk, achieving ~2.2-2.8 tokens/second with speculative decoding.

Cheapest way to run GLM 5.x locally that's not a unified memory system?

Reddit r/LocalLLaMA

A discussion on the cheapest local hardware setups for running GLM 5.x and similarly sized models at 4-bit quantization, including CPU-only and multi-GPU options, with a user sharing their experience running Minimax 2.7 and Qwen 3.6 on a 5900X + 128GB DDR4 + 7900XT setup.