@browser_use: Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5…
Summary
The Browser Use Bench v2 benchmark has updated its Pareto frontier, showcasing GPT-6 models from OpenAI outperforming and being more cost-effective than Claude Opus 5.5 from Anthropic.
View Cached Full Text
Cached at: 09/23/26, 02:08 PM
Browser Use Bench v2 Pareto frontier got completely redrawn today
> Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5x cheaper) > GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus)
OpenAI is in its own league 🔥
All models available to try on our cloud.
[x-axis is log scale] https://t.co/ko7ANkW0va
Similar Articles
@browser_use: Opus 5 and GPT-5.6 Sol are neck-and-neck on this!
Alexander Yue introduces a new browser-use benchmark where Opus 5 and GPT-5.6 Sol show similar performance, emphasizing the benchmark's robust design with verified rubrics for LLM judges.
@omarsar0: The efficiency frontier! Where do you think GPT-5.6 will land?
Discussion of recent benchmark results for Claude Opus 4.8 and GPT-5.5 on DeepSWE Bench, with speculation about future GPT-5.6 performance and efficiency trends.
@browser_use: Astra is a monster at browser use
Astra achieved 77.3% on the Browser Use Benchmark v2, far surpassing Opus 5 (50.5%) and GPT-5.6 Sol xhigh (49.1%), with 22 of 60 tasks earning full marks compared to zero for Opus 5.
@sashimikun_void: GPT-5.5 outperformed Claude Opus 4.8 on the DEEPSWE benchmark. Opus 4.8 takes twice as long, generates three times the …
GPT-5.5 outperforms Claude Opus 4.8 on the DEEPSWE benchmark, achieving higher scores with lower cost and less token bloat.
@VraserX: GPT-5.5 is still the king. GPT-5.5 destroys Claude Opus 4.8 at almost half the cost and about double the speed. OpenAI …
A tweet claims that OpenAI's GPT-5.5 outperforms Claude Opus 4.8 at nearly half the cost and double the speed, asserting OpenAI's continued dominance in AI.