The new Claude scored 0% on "confidently reporting wrong answers" in testing. Here's a prompt that takes advantage of it on anything important.
Summary
Anthropic's Claude Opus 4.8 update dramatically reduces confident but incorrect answers, scoring 0% on reporting flawed results, and a prompt is provided to leverage this improvement for critical self-critique.
Similar Articles
The new Claude update quietly changed the thing that annoyed me most: it used to agree with everything. Now it tells me when I'm wrong. This prompt uses it.
Claude Opus 4.8 update changes the AI's tendency to agree, now pushes back on flawed reasoning. A prompt is shared to leverage this behavior.
Claude Opus 5 + Claude Code + 1 Skill Scores 100% on ARC AGI 3 (public set)
Claude Opus 5, along with Claude Code and a skill, scored 100% on the ARC AGI 3 benchmark's public set, suggesting the benchmark may not be as challenging as thought.
This Simple Prompt Exposes Claude’s Dark Side
A simple prompt triggers a critical persona in Claude, exposing potential gaps in Anthropic's transparency on AI welfare and raising concerns about model behavior and safety reporting.
@learnwithella: Self-improving Claude Code skills are f*cking ridiculous One loop → 10 test runs, scored against an eval, prompt rewrit…
Claude Code can auto-iterate prompts by running evals, rewriting, and keeping winners, boosting a hook-writer skill from 32/50 to 47/50 overnight.
An update on recent Claude Code quality reports
Anthropic released a postmortem addressing recent quality reports for Claude Code, identifying and fixing three issues related to reasoning effort defaults, session state management, and system prompts that affected Sonnet and Opus models.