Tag
Announcement of model release week featuring Meta Spark 1.1, Grok 4.5, and GLM 5.2, with AI Gateway now supporting Muse Spark 1.1 for agentic tasks.
Colibrì is a pure C inference engine that runs the 744B GLM-5.2 MoE model on consumer hardware with ~25GB RAM by streaming experts from disk, achieving ~2.2-2.8 tokens/second with speculative decoding.
Jun Song announces the final development stage of a new MLX engine, achieving 41.8 tok/s on a MacBook with a 256k context window and only ~4% quality loss, representing a significant performance improvement.
AlpinDale reports running GLM-5.2-FP8 on 4 nodes of RTX 4090 (48GB) achieving ~28 tok/s decode over 10Gbit ethernet, with plans to optimize using DSpark.
NVIDIA officially offers a free GLM 5.2 model service, with an RPM of about 50. Registration using a Chinese phone number is sufficient for verification.
The user compared planet images generated by three models (GLM-5.2, Fugu Ultra, and Fable 5) using the same prompt. The results each have their own characteristics, with the user favoring Fable 5.
Chamath argues that companies should build their own AI intelligence using models like GLM to avoid leaking competitive advantage, emphasizing cost efficiency and control.
Describes a method using OMP and an advisor to address knowledge loss in REAP models, steering the 'lobotomy' out of GLM.
Based on NousResearch's Hermes Studio and MoA multi-model coordination mechanism, the combination of DeepSeek and GLM-5.2 can generate high-quality dynamic web pages, despite the longer generation cycle and redundant output.
A report on running the GLM5.2 language model across 5 AMD Radeon Pro 6000 GPUs and an NVIDIA RTX 5090, detailing the high cost and technical challenges.
Maka's Harness project improved the self-check mechanism, enabling DeepSeek Flash V4 to achieve evaluation results close to GLM-5.2 on the terminal-bench sample set, completing 10 programming agent tasks with only 4 RMB and a 97.5% cache hit rate.
RedHatAI releases a preview DSpark speculator for GLM-5.2-FP8, the first DSpark draft model for a non-DeepSeek frontier model, achieving ~1.5× faster decode on 4×B300 via vLLM nightly. The checkpoint is a work-in-progress, with training details and acceptance metrics provided.
ZCode is an AI programming development tool launched by the GLM team. It is deeply integrated with the GLM-5.2 model, offering features such as long-range task management and remote Bot control, and supports multiple subscription plans.
This article details the latest progress of GLM-5.2 in Agentic RL, including the introduction of slime infrastructure, shifting from GRPO to PPO for handling long trajectories, and an online anti-cheat mechanism; it also explores Qwen's research on verifier quality, proposing three dimensions of scalability, faithfulness, and robustness, and designs multiple verification strategies for different tasks to improve the reliability of reward signals.
DeepSeek releases version 4 of its GLM model, version 5.2 PRO.
GLM 5.2, a new version of the GLM language model, has been released, demonstrating improved performance.
Moonshot AI's Kimi and Zhipu AI's GLM have achieved notable results on frontier code benchmarks.
GLM 5.2 is optimized for CPU-only inference on AMD Epyc processors with 512GB RAM.
A hobbyist compares a heavily quantized GLM 5.2 (Q1_S) against a high-quant Qwen 27B (Q8) on a code generation task, finding that the lower-quant larger model significantly outperforms the higher-quant smaller model in quality and completeness.
Zhipu AI launches a free tier for GLM 5.2 with usage limits.