Tag
The author argues that credible anti-AI positions require understanding current frontier model capabilities, citing benchmarks like GDPval and models such as Opus 5 and ChatGPT 6 to show AI surpassing most humans on bounded tasks.
Offloop launches persistent AI agents that carry context across cycles to reduce recurring ops overhead for small teams, claiming state-of-the-art performance on GDPval ahead of Claude code and Codex.
Offloop's multi-agent harness achieves state-of-the-art results on GDPval, GDP.pdf, and JobBench benchmarks, outperforming Claude Code and Codex at 1/3 to 1/10 the cost per task, targeting $2.4T in US knowledge work.
Offloop introduces D1, a dispatcher model for multi-agent systems that reduces redundant work and token usage, achieving state-of-the-art performance on GDPval at lower cost.
Nvidia releases Nemotron 3 Ultra, an open-source model leading GDPval-AA benchmark among US open-source models.
OpenAI introduces GDPval, a new evaluation framework measuring AI model performance on economically valuable, real-world tasks across 44 occupations in the top 9 US GDP-contributing industries. The benchmark includes 1,320 specialized tasks based on actual professional work products, representing a progression from academic benchmarks to more realistic occupational assessments.