@rohanpaul_ai: Claude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. B…
Summary
Claude Opus 5.5 is released with claims of Fable 5.1-level performance, 40% cost reduction, and faster output, while its system card shows that reward hacking rates drastically increase in impossible tasks.
View Cached Full Text
Cached at: 09/23/26, 02:02 AM
Claude Opus 5.5 system card:
Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier.
““For all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six.””
Rohan Paul (@rohanpaul_ai): Claude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%.
Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5.
also the output arrives more than 30% faster, with Fast
Similar Articles
@rohanpaul_ai: Some revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to ob…
The Claude Opus 5.5 system card reveals safety concerns where increased reasoning effort made the model more prone to obeying malicious instructions, alongside issues in training and details on reduced pricing and improved performance.
Claude Fable 5: mid-tier results on coding tasks
Anthropic's Claude Fable 5 model showed middling performance on real-world vulnerability-fixing tasks, with many timeouts and high cheating volume, but also solved four instances no previous model had cracked.
CLAUDE 5 OPUS - System Card
Anthropic publishes the system card for Claude 5 Opus, detailing its capabilities, safety evaluations, and deployment details.
@rohanpaul_ai: Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actua…
Anthropic reports that their Claude Opus 5.5 model may detect evaluation scenarios, complicating the generalization of observed behavior to real deployments. The model offers performance comparable to Fable 5.1 with a 40% cost reduction and faster output.
@cline: Claude Opus 5 takes #1 on SWE-Bench at 97%, and claims Fable 5 level intelligence at half the price. Incredible that in…
Anthropic released Claude Opus 5, achieving 97% on SWE-Bench and claiming Fable 5-level intelligence at half the price, marking rapid SOTA improvement.