Released a model tuned for agent testing work that other models refuse. AgentDojo 97.5% utility.
Summary
A fine-tuned model based on GLM-5.2, abliterated and specialized for agent testing and red teaming, achieving 97.5% benign utility on AgentDojo and strong coding benchmarks.
Similar Articles
We released an abliterated + fine-tuned GLM-5.2. High scores on adversarial benchmarks while keeping coding performance.
Released an abliterated and fine-tuned version of GLM-5.2 (abliterated-model-large) that achieves high scores on adversarial and agent benchmarks while maintaining coding performance. The model is available via API with zero data retention and no built-in policy.
Prime Agent - a new coding harness surpassing Codex/CC/PI
Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.
Is it agentic enough? Benchmarking open models on your own tooling
This blog post introduces a benchmark methodology for evaluating how well open models perform on agentic coding tasks, focusing not just on accuracy but on the efficiency of the agent's process. It provides a customizable tooling harness using the pi coding agent and tests across models and library revisions.
Agent Arena Code - Very good result (preliminary) for GLM and Qwen!
Preliminary results from Agent Arena Code show good performance for GLM and Qwen models, indicating advancements in open-weight AI models.
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
This paper proposes AgentDoG 1.5, a lightweight and scalable alignment framework for AI agent safety, using taxonomy-guided training with minimal samples to achieve performance comparable to leading closed-source models.