Tag
The article reports on the latest scores from the Short-Story Creative Writing Benchmark, showing improvements in AI models like Opus 5.5, Grok 4.7, and Gemini 3.8 Flash, with a leaderboard covering 56 models and over 100,000 evaluator judgments.
Terminal-Bench version 4.0 has been released, updating the dataset and leaderboard with calibrated task resources, task fixes, and removal of saturated tasks.