@PrajwalTomar_: A SMALLER TECH COMPANY just surpassed Claude and GPT 73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-…
Summary
Simular's computer-use agent sai_borg achieved 73% accuracy on the OSWorld 2.0 benchmark, surpassing Claude and GPT models at about two-thirds the cost per task, with all trajectories made public.
View Cached Full Text
Cached at: 08/27/26, 07:32 PM
A SMALLER TECH COMPANY just surpassed Claude and GPT
73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-thirds the cost per task.
The difference is @sai_borg doesn’t re-think a job it already knows. It saves the successful run as code and replays it, so the second run costs a fraction of the first.
It works on its own computer in the cloud, clicks through real websites like a person, and checks its own work before calling it done.
Every trajectory is public. OpenAI and Anthropic submitted nothing.
I’m running it on my weekly admin this week. Will report back
Simular (@SimularAI): Our computer-use agent @sai_borg just beat Opus 5 and GPT-5.6 Sol on OSWorld 2.0.
Sai scored 73% on the CUA benchmark where each complex task takes a skilled human over an hour to execute.
And Sai did it at about 2/3 the cost per task of Opus and GPT 💸
The unlock is our
Similar Articles
@sashimikun_void: GPT-5.5 outperformed Claude Opus 4.8 on the DEEPSWE benchmark. Opus 4.8 takes twice as long, generates three times the …
GPT-5.5 outperforms Claude Opus 4.8 on the DEEPSWE benchmark, achieving higher scores with lower cost and less token bloat.
@VraserX: GPT-5.5 is still the king. GPT-5.5 destroys Claude Opus 4.8 at almost half the cost and about double the speed. OpenAI …
A tweet claims that OpenAI's GPT-5.5 outperforms Claude Opus 4.8 at nearly half the cost and double the speed, asserting OpenAI's continued dominance in AI.
Open-source models are closing the coding gap with GPT/Claude/Gemini ~1.5x faster than the frontier is advancing, and on decontaminated benchmarks a 27B model already beats Claude Opus 4.8 [live dashboard + analysis]
A live dashboard and statistical analysis shows open-source coding models are closing the gap with closed models at 1.5x the rate, with a 27B model already surpassing Claude Opus on decontaminated benchmarks. Tool-call reliability remains the main bottleneck.
@PrajwalTomar_: https://x.com/PrajwalTomar_/status/2075532429641809935
GPT 5.6 is a family of three tiers (Sol, Terra, Luna) priced significantly lower than competing models like Claude's Fable 5, achieving top scores on coding benchmarks but falling short on ambiguous, high-complexity tasks where Fable excels, suggesting a role-based division where Fable serves as a manager and Sol as a senior worker.
@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.