Tag
OpenAI's model experiences a significant regression on the SimpleBench benchmark, indicating a drop in performance.
Claude Fable 5 achieves 81.9% on the Simplebench leaderboard, taking the top position.
A new top score has been achieved on the SimpleBench benchmark, nearly matching the human baseline.