Opus 5.5 (high) improves on Opus 5 (high) 3.5 → 3.8 on the Short-Story Creative Writing Benchmark, just behind Fable 5.1 (high) and Opus 5 (xhigh).
Summary
The article reports on the latest scores from the Short-Story Creative Writing Benchmark, showing improvements in AI models like Opus 5.5, Grok 4.7, and Gemini 3.8 Flash, with a leaderboard covering 56 models and over 100,000 evaluator judgments.
Similar Articles
Opus 5.5 just surpassed the professional human baseline on a benchmark testing whether AI can write complete scripts with the right voice, substance, pacing, hooks and minimal AI slop
The Opus 5.5 AI model has surpassed professional human writers on a benchmark evaluating script quality, including voice, substance, pacing, and minimal AI-generated errors.
Opus 5.5 dominates on all three performance metrics from ArtificialAnalysis.ai
Opus 5.5 dominates all three performance metrics from ArtificialAnalysis.ai and offers a cost-effective alternative to Fable 5.1.
Opus 5 benchmarks (30.2% on ARC-AGI3!!!)
Opus 5 achieves 30.2% on the ARC-AGI3 benchmark, marking a notable performance improvement.
Gemini 3.5 Flash improves over Gemini 3.1 Pro on the Short Story Creative Writing Benchmark: -2.3 → -1.8.
Gemini 3.5 Flash outperforms Gemini 3.1 Pro on a short story creative writing benchmark, improving from -2.3 to -1.8 in head-to-head comparisons.
Anthropic releases Opus 5.5 with lower prices and Fable-level performance
Anthropic has released Opus 5.5, a new AI model with lower prices and performance matching or exceeding larger models like Fable, featuring improved communication and alignment with safety pacing efforts.