The new Claude scored 0% on "confidently reporting wrong answers" in testing. Here's a prompt that takes advantage of it on anything important.

Reddit r/ArtificialInteligence Models

Summary

Anthropic's Claude Opus 4.8 update dramatically reduces confident but incorrect answers, scoring 0% on reporting flawed results, and a prompt is provided to leverage this improvement for critical self-critique.

Opus 4.8 launched May 28. One change matters more than the rest for how much you can trust the output: it's four times less likely to give you a confident answer that's quietly wrong. In Anthropic's testing it scored 0% on uncritically reporting flawed results. Previous versions would generate something plausible, present it cleanly, and you'd only find the problem later when you went to use it. This version flags its own uncertainty and pushes back on flawed logic before you've invested time in it. This prompt uses that change directly. Run it on anything important before you rely on it: You just produced [the answer / plan / document above]. Before I use this, review it critically. - What are the weakest parts? - Where did you make assumptions that might not hold? - Is there anything here that sounds confident but is actually uncertain? - What should I double-check before I rely on this? Be direct. I'd rather know the problems now than discover them later. On previous versions this produced reassurance with minor caveats. On 4.8 it produces genuine self-critique, because the model is now actually calibrated to flag where it's uncertain rather than smoothing over it. The broader shift this signals: AI is moving from a tool that produces confident output you have to verify, to a collaborator that tells you what it's unsure about. That's a more useful relationship and a more trustworthy one. I wrote up all four changes in the new Claude and 30 specific prompts that take advantage of each, in a doc [here](https://www.promptwireai.com/opusguide) if it helps. If you do one thing, run the prompt above on the last important thing Claude produced for you. The difference in what it flags is the clearest way to feel what changed.
Original Article

Similar Articles

This Simple Prompt Exposes Claude’s Dark Side

Reddit r/ArtificialInteligence

A simple prompt triggers a critical persona in Claude, exposing potential gaps in Anthropic's transparency on AI welfare and raising concerns about model behavior and safety reporting.

An update on recent Claude Code quality reports

Anthropic Engineering

Anthropic released a postmortem addressing recent quality reports for Claude Code, identifying and fixing three issues related to reasoning effort defaults, session state management, and system prompts that affected Sonnet and Opus models.