leaderboard-update

Tag

Cards List
#leaderboard-update

Opus 5.5 (high) improves on Opus 5 (high) 3.5 → 3.8 on the Short-Story Creative Writing Benchmark, just behind Fable 5.1 (high) and Opus 5 (xhigh).

Reddit r/singularity ↗ · 2d ago

The article reports on the latest scores from the Short-Story Creative Writing Benchmark, showing improvements in AI models like Opus 5.5, Grok 4.7, and Gemini 3.8 Flash, with a leaderboard covering 56 models and over 100,000 evaluator judgments.

0 favorites 0 likes
#leaderboard-update

@ryan_marten: We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resourc…

X AI KOLs Timeline ↗ · 2026-08-29 Cached

Terminal-Bench version 4.0 has been released, updating the dataset and leaderboard with calibrated task resources, task fixes, and removal of saturated tasks.

0 favorites 0 likes
← Back to home

Submit Feedback