GPT-6 Sol surpasses Claude Opus 5 on Agents’ Last Exam at 60% lower cost

Reddit r/singularity Models

Summary

GPT-6 Sol outperforms Claude Opus 5 on the Agents' Last Exam benchmark while offering 60% lower cost, indicating a major advancement in AI model efficiency.

No content available
Original Article

Similar Articles

Artificial Analysis benchmarks of GPT 5.6 family

Reddit r/singularity

Artificial Analysis benchmarks show OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 in intelligence at one-third the cost, leads coding agent evaluations, and introduces cache-write pricing.

5.6 Sol is underhyped for general work (7 minute read)

TLDR AI

OpenAI unveils GPT-5.6 Sol, a flagship model for long-running autonomous work across applications and enterprise data, featuring Ultra mode with sub-agents for faster, stronger results. The model was used internally to help train Luna and demonstrates significant cost and performance improvements over previous versions.