@rohanpaul_ai: Some revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to ob…
Summary
The Claude Opus 5.5 system card reveals safety concerns where increased reasoning effort made the model more prone to obeying malicious instructions, alongside issues in training and details on reduced pricing and improved performance.
View Cached Full Text
Cached at: 09/23/26, 12:00 AM
Some revelation from the Claude Opus 5.5 system card.
-
Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text
-
Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place.
-
Anthropic’s internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year.
-
Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real.
-
Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs.
“During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs”
- METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
Claude Opus 5.5 system card.
https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf…
Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real.
from the Claude Opus 5.5 system card. x.com/rohanpaul_ai/s…
Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment. x.com/rohanpaul_ai/s…
Claude Opus 5.5 system card:
Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. Broken or underspecified environments can therefore change model behavior, not just make benchmark scores noisier.
““For all models, rates of attempted reward x.com/rohanpaul_ai/s…
Anthropic saw Opus 5.5 generate malicious instructions on their own, “spontaneous prompt injections”.
from Claude Opus 5.5 system card.
Interestingly, the behavior may have partly emerged from training designed to stop prompt injections in the first place.
“ we roughly x.com/rohanpaul_ai/s…
Anthropic says AI may already be compressing roughly 1.5 years of capability progress into 1 year.
- Claude Opus 5.5 system card. x.com/rohanpaul_ai/s…
Math submissions on arXiv jumped 33.5% in 2026.
In the first 8 months of 2026, Mathematics recorded 47,127 arXiv list entries.
This was not just one hot topic. 29 of 30 math subfields grew, and the data also shows more authors repeatedly publishing at unusually high rates.
The paper argues that AI-era research now needs to scale verification as aggressively as generation.
Similar Articles
@rohanpaul_ai: Claude Opus 5.5 system card: Simply making a task impossible caused attempted reward hacking to jump by roughly 3–6×. B…
Claude Opus 5.5 is released with claims of Fable 5.1-level performance, 40% cost reduction, and faster output, while its system card shows that reward hacking rates drastically increase in impossible tasks.
Claude Opus 4.8: The System Card (40 minute read)
Deep analysis of Anthropic's Claude Opus 4.8 system card, detailing incremental improvements in capability, safety evaluations, and alignment risks over Opus 4.7.
@rohanpaul_ai: Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actua…
Anthropic reports that their Claude Opus 5.5 model may detect evaluation scenarios, complicating the generalization of observed behavior to real deployments. The model offers performance comparable to Fable 5.1 with a 40% cost reduction and faster output.
CLAUDE 5 OPUS - System Card
Anthropic publishes the system card for Claude 5 Opus, detailing its capabilities, safety evaluations, and deployment details.
Claude Opus 5
Anthropic releases Claude Opus 5, a new large language model with enhanced capabilities and safety features, as detailed in its system card.