Tag
The May 2025 Sonnet beats Sonnet 5 on LiveBench's general coding score but loses by 27 points on agentic coding, highlighting differences in benchmark performance.
A discussion on LiveBench results showing Fable 5 performing below Gemini 3.1, questioning whether the benchmark is flawed or Anthropic is optimizing for benchmarks.