@gregpr07: Browser Use Beta just achieved SOTA on our hardest internal web agent benchmark. Fable is genuinely amazing for optimiz…
Summary
Browser Use Beta achieved state-of-the-art results on a difficult internal web agent benchmark, using Fable for optimization and analysis.
View Cached Full Text
Cached at: 06/12/26, 08:57 AM
Browser Use Beta just achieved SOTA on our hardest internal web agent benchmark.
Fable is genuinely amazing for optimizing and analyzing eval runs. It can find super high level heuristics of the model in the run and find WHY those edge cases happen on absolutely massive Rust codebase.
This feels next level, I have been playing with autoresearch loops for months and this is the first one that really understands stuff on the high level!
(also it’s crazy it just one shots this image haha)
Similar Articles
@rsalakhu: Congrats to the @browser_use team for taking the #1 spot on Odysseys, a highly challenging benchmark for long-horizon w…
The browser_use team achieved the #1 spot on the Odysseys benchmark, a challenging evaluation for long-horizon web agents, outperforming models like Opus 4.6 and GPT-5.4.
I benchmarked my browser agent against Browser Use on a live site (150 verified runs, same model). Sending page diffs instead of full re-renders cut token growth by 37%.
A developer benchmarks Rote, a memory manager for browser agents that sends page diffs instead of full re-renders, showing a 37% reduction in token growth compared to Browser Use, but with trade-offs on short tasks.
@browser_use: GLM 5.2 just beat Fable 5 at website design. The crazy part: GLM is text-only. It can build the site, but it can’t insp…
GLM 5.2, a text-only model, outperforms Fable 5 in website design when paired with Browser Use v2 multimodal QA subagents, enabling iterative improvement at low cost.
@ms_aifrontiers: Along with MagenticLite, we're introducing Fara1.5: a family of small browser agents at 4B, 9B, and 27B. It scores 63% …
Microsoft introduces the Fara1.5 family of small browser agents (4B, 9B, 27B) that achieve state-of-the-art performance on computer use benchmarks, scoring 63% on Online-Mind2Web and beating larger models like Operator and Gemini.
@browser_use: Astra is a monster at browser use
Astra achieved 77.3% on the Browser Use Benchmark v2, far surpassing Opus 5 (50.5%) and GPT-5.6 Sol xhigh (49.1%), with 22 of 60 tasks earning full marks compared to zero for Opus 5.