@Miles_Brundage: BREAKING: massively improved SOTA score on Clear AVERI Pronunciation Guide Bench, via my colleague Carly
Summary
Miles Brundage announces a state-of-the-art (SOTA) score improvement on the Clear AVERI Pronunciation Guide Bench achieved by colleague Carly.
View Cached Full Text
Cached at: 06/04/26, 01:58 AM
BREAKING: massively improved SOTA score on Clear AVERI Pronunciation Guide Bench, via my colleague Carly https://t.co/6ucOvO5wX5
Similar Articles
We have a new SimpleBench king
A new top score has been achieved on the SimpleBench benchmark, nearly matching the human baseline.
@nick_kango: One more task to add to my twitter benchmark collection:) Btw, Opus 4.8 and all the SOTA models passed when i tried tha…
Nick Kang adds a new task to his Twitter benchmark collection; Claude Opus 4.8 and other SOTA models pass, while Sonnet 4.6 and Grok 4.3 fail. Alfin remarks on Opus 4.8's dangerous capabilities.
New SOTA every week
A remark on the constant stream of new state-of-the-art results being achieved weekly in AI.
@cline: Claude Opus 5 takes #1 on SWE-Bench at 97%, and claims Fable 5 level intelligence at half the price. Incredible that in…
Anthropic released Claude Opus 5, achieving 97% on SWE-Bench and claiming Fable 5-level intelligence at half the price, marking rapid SOTA improvement.
@patpcj: Fable-5/Mythos dropped this morning, so we tested it on agentic search - and it’s the new SOTA. The performance gap is …
Fable-5/Mythos achieves new SOTA on agentic search but is expensive for self-hosting, while open-weight Harness-1 offers a cost-effective alternative with fewer query restrictions.