@browser_use: Open-weights models have officially caught up We tried GLM 5.2 in BrowserCode > Near Opus-level score > Cheapest model …
Summary
Open-weights models have caught up with proprietary ones, with GLM 5.2 achieving near Opus-level scores in browser agent tasks at low cost. Other models like Minimax M3 and Kimi k2.7 also show notable improvements.
View Cached Full Text
Cached at: 06/20/26, 02:36 PM
Open-weights models have officially caught up
We tried GLM 5.2 in BrowserCode > Near Opus-level score > Cheapest model to date
This task took $0.18 Try it now ↓ https://t.co/aR24ZDU1pj
Alexander Yue (@Alezander907): GLM 5.2 is a huge improvement for browser agents, offering near opus level score, beating GPT 5.5
Minimax M3 is a sonnet level score at just $0.30 input, my new best value model (cheaper than deepseek v4 pro)
Kimi k2.7 is a +9% improvement from k2.6 but is outclassed by M3
Similar Articles
GLM-5.2 is the first open-weights model to cross 80% on Terminal-Bench and beats every other open model available
GLM-5.2 is the first open-weights model to exceed 80% on Terminal-Bench, surpassing all other open models and even Gemini, making it a frontier-level model at a fraction of the cost.
GLM 5.2 vs. Opus
GLM 5.2 is a new open-weights model from Z.ai, compared against Claude Opus in a 3D game coding task. Opus performed faster and cleaner, but GLM 5.2 offers compelling cost and accessibility advantages.
GLM-5.2 just dropped open weights and it already looks weirdly strong for coding
GLM-5.2 has been released with open weights under MIT license, featuring a 1M context window and two reasoning effort modes. Early benchmarks show it performing strongly in coding tasks, making it worth testing beyond benchmark screenshots.
GLM-5.2 is the new leading open weights model on Artificial Analysis
Z ai's GLM-5.2 has become the new leading open weights model on the Artificial Analysis Intelligence Index, scoring 51 and outperforming competitors like MiniMax-M3 and DeepSeek V4 Pro. The model features 744B total parameters, 40B active, MIT license, and 1M context window.
@browser_use: Opus 4.7 GLM 5.2 We're benchmarking models on frontend design. We run each model on Browser Use v4 > One prompt from th…
Opus 4.7 and GLM 5.2 are being benchmarked on frontend design using Browser Use v4; results are shared via a link.