The May 2025 Sonnet still beats Sonnet 5 on livebench's coding score. On agentic coding it lose to it by 27 points. So what's the difference?
Summary
The May 2025 Sonnet beats Sonnet 5 on LiveBench's general coding score but loses by 27 points on agentic coding, highlighting differences in benchmark performance.
Similar Articles
@reach_vb: GPT-5.6 Luna Max scores 13.3 points above Sonnet 5 Max on DeepSWE v1.1 while Sonnet costs 44x as much. DeepSWE tests co…
GPT-5.6 Luna Max outperforms Sonnet 5 Max on the DeepSWE v1.1 coding benchmark at a much lower cost, and also shows strong performance compared to Gemini 3.7 Flash Medium.
Benchmark of sonnet 5 (good improvement)
Sonnet 5 demonstrates good improvement in benchmarks.
Claude Sonnet 5 Artificial Analysis Results & Comparison
Provides analysis and comparison of Claude Sonnet 5's performance across benchmarks.
Claude Sonnet 5 Benchmarks
Anthropic's Claude Sonnet 5 model benchmarks are released, showing performance improvements.
Sonnet 5 - its updated tokenizer maps the same text to more tokens (roughly 1.0–1.35× depending on content), so cost per task can be higher.
Anthropic released Claude Sonnet 5 with improved reasoning, tool use, and coding, but its updated tokenizer maps text to more tokens (up to 1.35×), increasing effective cost per task despite the same listed price; introductory pricing applies until August 31, 2026.