@PrajwalTomar_: A SMALLER TECH COMPANY just surpassed Claude and GPT 73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-…

X AI KOLs Following Models

Summary

Simular's computer-use agent sai_borg achieved 73% accuracy on the OSWorld 2.0 benchmark, surpassing Claude and GPT models at about two-thirds the cost per task, with all trajectories made public.

A SMALLER TECH COMPANY just surpassed Claude and GPT 73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-thirds the cost per task. The difference is @sai_borg doesn't re-think a job it already knows. It saves the successful run as code and replays it, so the second run costs a fraction of the first. It works on its own computer in the cloud, clicks through real websites like a person, and checks its own work before calling it done. Every trajectory is public. OpenAI and Anthropic submitted nothing. I'm running it on my weekly admin this week. Will report back
Original Article
View Cached Full Text

Cached at: 08/27/26, 07:32 PM

A SMALLER TECH COMPANY just surpassed Claude and GPT

73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-thirds the cost per task.

The difference is @sai_borg doesn’t re-think a job it already knows. It saves the successful run as code and replays it, so the second run costs a fraction of the first.

It works on its own computer in the cloud, clicks through real websites like a person, and checks its own work before calling it done.

Every trajectory is public. OpenAI and Anthropic submitted nothing.

I’m running it on my weekly admin this week. Will report back

Simular (@SimularAI): Our computer-use agent @sai_borg just beat Opus 5 and GPT-5.6 Sol on OSWorld 2.0.

Sai scored 73% on the CUA benchmark where each complex task takes a skilled human over an hour to execute.

And Sai did it at about 2/3 the cost per task of Opus and GPT 💸

The unlock is our

Similar Articles

@PrajwalTomar_: https://x.com/PrajwalTomar_/status/2075532429641809935

X AI KOLs Timeline

GPT 5.6 is a family of three tiers (Sol, Terra, Luna) priced significantly lower than competing models like Claude's Fable 5, achieving top scores on coding benchmarks but falling short on ambiguous, high-complexity tasks where Fable excels, suggesting a role-based division where Fable serves as a manager and Sol as a senior worker.