Opus 5.5 (high) improves on Opus 5 (high) 3.5 → 3.8 on the Short-Story Creative Writing Benchmark, just behind Fable 5.1 (high) and Opus 5 (xhigh).

Reddit r/singularity News

Summary

The article reports on the latest scores from the Short-Story Creative Writing Benchmark, showing improvements in AI models like Opus 5.5, Grok 4.7, and Gemini 3.8 Flash, with a leaderboard covering 56 models and over 100,000 evaluator judgments.

https://github.com/lechmazur/writing/ Grok 4.7 (high) makes a substantial jump over Grok 4.6 (high): −2.8 → 0.4. MiMo V2.6 Pro (thinking) improves sharply over V2.5 Pro: −0.7 → 1.6. Gemini 3.8 Flash (high) advances over Gemini 3.7 Flash (high): −0.7 → 0.2. DeepSeek V4.1 Flash (high) enters at −0.5. The judging panel has been updated. New comparisons draw from nine model families, including Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash, and Grok 4.7. The Creative Writing Benchmark tests how well models turn constrained briefs into complete 600–800-word stories. Each brief requires 10 elements, including a character, object, setting, motivation, and tone, that must meaningfully shape the story. Judges assess prose, originality, coherence, characterization, and how effectively those ingredients work together. Models write to the same prompts. The latest comparisons use three judges from different model families, excluding the writers’ own families. Each story pair is shown in both orders to reduce position bias. The leaderboard combines earlier and newer judging panels and now covers 56 models and 102,592 evaluator judgments.
Original Article

Similar Articles