@browser_use: Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5…

X AI KOLs Following News

Summary

The Browser Use Bench v2 benchmark has updated its Pareto frontier, showcasing GPT-6 models from OpenAI outperforming and being more cost-effective than Claude Opus 5.5 from Anthropic.

Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5x cheaper) > GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus) OpenAI is in its own league 🔥 All models available to try on our cloud. [x-axis is log scale] https://t.co/ko7ANkW0va
Original Article
View Cached Full Text

Cached at: 09/23/26, 02:08 PM

Browser Use Bench v2 Pareto frontier got completely redrawn today

> Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5x cheaper) > GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus)

OpenAI is in its own league 🔥

All models available to try on our cloud.

[x-axis is log scale] https://t.co/ko7ANkW0va

Similar Articles

@browser_use: Astra is a monster at browser use

X AI KOLs Timeline

Astra achieved 77.3% on the Browser Use Benchmark v2, far surpassing Opus 5 (50.5%) and GPT-5.6 Sol xhigh (49.1%), with 22 of 60 tasks earning full marks compared to zero for Opus 5.