Tag
The May 2025 Sonnet beats Sonnet 5 on LiveBench's general coding score but loses by 27 points on agentic coding, highlighting differences in benchmark performance.
Suggests that Anthropic's Fable model is equivalent to Opus, and Opus to Sonnet, possibly indicating a rebranding or restructuring of model tiers.
Introduces how to use Fable 5 as the orchestrator, combined with Opus and Codex models to execute tasks to save on Fable usage, including specific configuration in Claude Code.
This tweet discusses Anthropic's new research on a global workspace in language models, noting that Karpathy is not in the author list and emphasizing that the work was completed during the Sonnet 4.5 era, criticizing those who simply see it as hype.
Introduces a third-party tool that allows manual switching between Claude's Fable, Opus, and Sonnet models, enhancing flexibility.
Discussion of open-source model tiers, comparing DSV4-flash to Sonnet 5 and GLM 5.2 to Opus 4.8, with a prediction of a fable-tier model by end of year.
Kevin Whinnery teases an upcoming integration between Fable and Sonnet, with Brad Abrams noting Fable is back and suggesting Sonnet 5 handles loops while Fable takes hard calls.
A blog post describes how automatic harness optimization enabled DeepSeek V4 Pro to achieve Sonnet 4.6 performance on the Legal Agent Benchmark at one-seventh the cost.
Anthropic's Claude Sonnet 5 model benchmarks are released, showing performance improvements.
A leak suggests Anthropic's Claude Sonnet 5 model will be released next week, as it has appeared on a partner provider with a slug that typically precedes flagship releases by 5-7 days.
A user reports that Gemma4_31b in FP8 matches or keeps up with Sonnet_4.6_medium in a custom harness across tasks like Cypher query generation, entity extraction, agentic tool calling, code writing, and multi-vector retrieval synthesis.
Nick Kang adds a new task to his Twitter benchmark collection; Claude Opus 4.8 and other SOTA models pass, while Sonnet 4.6 and Grok 4.3 fail. Alfin remarks on Opus 4.8's dangerous capabilities.
Boris Cherny recommends using auto mode in Claude Code for parallel sessions, and ClaudeDevs announces that auto mode is now available on the Pro plan and supports Sonnet 4.6 and Opus 4.7.