Agent Arena Code - Very good result (preliminary) for GLM and Qwen!
Summary
Preliminary results from Agent Arena Code show good performance for GLM and Qwen models, indicating advancements in open-weight AI models.
Similar Articles
GLM-5.2 is the new leading open weights model on Artificial Analysis
Z ai's GLM-5.2 has become the new leading open weights model on the Artificial Analysis Intelligence Index, scoring 51 and outperforming competitors like MiniMax-M3 and DeepSeek V4 Pro. The model features 744B total parameters, 40B active, MIT license, and 1M context window.
GLM-5.2 is a step change for open agents
Z.ai released GLM-5.2, an open-weight AI model that represents a step change for open agents, with strong benchmark performance and community hype, positioning it as the only open model competing with top closed models from OpenAI and Anthropic.
Agent Arena
Agent Arena is the first public arena for AI agents, allowing users to test and compare AI agents in a competitive environment.
GLM-5.2 is probably the most powerful text-only open weights LLM
Chinese AI lab Z.ai released GLM-5.2, a 753B parameter open weights LLM with a 1M token context window under MIT license, achieving top scores on the Artificial Analysis Intelligence Index and ranking second on the Code Arena WebDev leaderboard.
@rohanpaul_ai: Arena just released a real-world agent leaderboard that ranks AI models by how well they complete actual user jobs, not…
Agent Arena is a new leaderboard that evaluates AI models on real-world agentic tasks such as coding, research, and file analysis, using signals like task success, steerability, and recovery, with GPT-5.5 High leading.