Tag
This post details running GLM 5.2 on a 4xGB10 setup with a 100G switch, achieving ~25 tok/s decode and ~650 tok/s prefill at 330k context. It includes hardware costs, performance benchmarks with Depth Prefill, and notes on model pruning for longer context.
Theo praises GLM 5.2 as an incredible model but criticizes the notion that it is self-hostable and comparable to Fable, sparking debate about model accessibility.
GLM 5.2 from Z.ai emerges as a strong open-weights competitor to frontier models like Opus and GPT, but the real story is the impending margin collapse in AI inference as costs decrease and competition intensifies.
Ahmad Osman highlights improvements in Tencent Hy3 over its preview version and compares it to GLM 5.2, which is twice its size, while sharing a prediction about running similar intelligence on an RTX 5090.
Testing GLM 5.2 with FP8 quantization and FP8 KV cache on H200 yields a score of 79.8% on Terminal-Bench 2.1, with one timeout not rerun.
China's open-source model GLM 5.2 achieves performance comparable to top US AI models at low cost, challenging the high-price monopoly of US AI companies and risking the bursting of the Silicon Valley AI bubble.
Engineers successfully serve GLM 5.2 on AMD MI355X at 2626 tok/s per node and 213 tok/s single stream, achieving ~80% of B200 throughput at over 2x lower cost than Blackwell.
Ahmad Osman predicts that within 18 months, a GPU like the RTX 5090 will be able to host intelligence equivalent to GLM 5.2.
GLM 5.2 demonstrates impressive capabilities in connecting scriptural themes and references when used with RAG for Bible study, outperforming other models in providing deeper insights.
GLM-5.2 is now selectable in Claude Code via Hugging Face Inference Providers and hf-claude, making it easier to integrate open models into developer workflows.
Z.ai launches ZCode, an agentic development environment for its GLM-5.2 model, challenging existing AI coding tools with deep integration, multi-device support, and competitive pricing.
GLM 5.2, an open frontier-scale AI model, is now available on Microsoft Foundry with AMD MI300X hardware, enabling efficient Codex goal execution.
A user shares positive impressions of the GLM-5.2 model, noting its impressive performance and cost-efficiency compared to Deepseek, and reflects on the educational value and limitations of AI agents for coding and learning.
GLM 5.2 demonstrates its ability to build a working CV (curriculum vitae) application end-to-end, showcasing AI-assisted development.
This article evaluates whether GLM 5.2 is suitable for production use by testing it on a complex multi-file computer vision implementation task.
According to rumors, a certain new model from zAI is at least as strong as Fable 5 in cybersecurity capabilities, but information is limited, only citing a Wall Street Journal article discussing GLM 5.2.
An open-weight model, GLM 5.2 from Zhipu AI, beats Claude Code in IDOR detection benchmarks at a fraction of the cost, though it still trails Semgrep's purpose-built multimodal pipeline. The article explores how much of vulnerability-detection performance comes from the model versus the harness around it.
NVIDIA released an NVFP4 quantized checkpoint of GLM-5.2, a 744B MoE model (40B active) optimized for reasoning and coding, with day-0 support in SGLang.
A tweet discussing how GLM 5.2 reveals enterprise trends toward local compute and post-trained models, with opposing views on the future of open-source AI.
Fireworks is offering a managed service for reinforcement learning training on GLM 5.2 that ensures numerical identity between training and inference via batch invariance and zero-KLD alignment, previously only available to top frontier labs. This allows anyone to customize and surpass frontier quality.