@sumeetrm: LongCoT is adding two new leaderboards! Due to the interest in agents (particularly RLMs), we’re adding a “Restricted H…
Summary
LongCoT introduces two new agent leaderboards (Restricted & Open Harness), with GPT 5.2 RLM topping the Open Harness at 25.12%.
View Cached Full Text
Cached at: 04/21/26, 10:51 AM
LongCoT is adding two new leaderboards! Due to the interest in agents (particularly RLMs), we’re adding a “Restricted Harness” and an “Open Harness” leaderboard. GPT 5.2 RLM from our paper is SOTA on “Open Harness” at 25.12%. We expect tool-use SOTA to exceed this very soon! On
Similar Articles
@billxbf: Excited to release Polar, our Agent RL rollout infra for real-world harnesses. Be it Codex, Claude Code, OpenClaw, Herm…
Polar is an agent RL rollout infrastructure that allows using real-world harnesses as training environments without code changes, supporting models like Codex, Claude Code, OpenClaw, and Hermes.
Prime Agent: A Self-Improving RLM Harness
Prime Agent is an open-source harness that uses recursive subagents and persistent computation to extend language models' long-horizon capabilities across coding and reasoning tasks, significantly improving performance on benchmarks like ARC-AGI-3.
@omarsar0: Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great …
This tweet highlights the new DataSpace benchmark for data agents, showing that harness choice significantly impacts accuracy across multimodal models and agent harnesses, with the best accuracy reaching 66.34% and the benchmark remaining unsaturated.
OpenRouter Cloud Agents leaderboard - 9 July 2026
OpenRouter's Cloud Agents leaderboard highlights top agents by token consumption, with Gitlawb leading at 8.34B tokens, followed by Ito and Roo Code, reflecting rapid growth in cloud agent platforms.
Prime Agent - a new coding harness surpassing Codex/CC/PI
Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.