@levie: Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities. At Box, we've been testi…
Summary
Claude Opus 5 has been released, showing significant improvements over Opus 4.8 in enterprise benchmarks across due diligence, life sciences, legal, technology, and healthcare tasks, with gains of 12-30%.
View Cached Full Text
Cached at: 07/24/26, 07:14 PM
Claude Opus 5 is out now, and it’s a huge jump over Opus 4.8 on a wide number of capabilities.
At Box, we’ve been testing Claude Opus 5 with the Box AI Agent on Box’s Complex Work Eval, our agentic benchmark that puts models through real enterprise document work end-to-end across a variety of industries. We saw meaningful gains in performance across some of the most complex enterprise tasks on unstructured data.
Here are a few examples of some of the wins we saw in testing:
-
Due diligence (+17%): On a transaction due-diligence review, Opus 5 worked through the full set of required findings and flagged the ones Opus 4.8 missed. It stays thorough as the checklist grows instead of catching the obvious items and stopping.
-
Life Sciences (+30%): On a target-identification task, it intersected several ranked datasets under a strict matching rule to find the targets common to all of them. Opus 4.8 over-included partial matches and dropped genuine ones.
-
Legal (+12%): On a clause-by-clause contract review against a policy, Opus 5 scored every item and correctly cleared the clauses that were acceptable under an exception where Opus 4.8 both missed items and mis-scored the exception.
-
Technology (+19%) and Healthcare (+13%) showed the same pattern: complete, precise, multi-step analysis over messy source material.
Overall, the reasoning, analytical, and data processing skills of Opus 5 outshine Opus 4.8 meaningfully. Going to be very powerful for enterprise agentic use-cases. You’ll be able to build agents with Opus 5 in the Box AI Studio shortly.
Claude (@claudeai): Introducing Claude Opus 5.
It’s a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.
Similar Articles
Claude Sonnet 5 is out and the gap with Opus 4.8 is smaller than I expected
Anthropic released Claude Sonnet 5, which achieves benchmark scores very close to Opus 4.8 at a significantly lower price, making it a compelling option for agentic tasks despite potential real-world gaps.
@FinanceYF5: 官方发布:
Anthropic releases Claude Opus 4.8, building on Opus 4.7 with sharper judgment and longer independent work capability, available at the same price.
@mfpiccolo: Opus 4.8 is out. Here is the the verdict from @iiidevs lead engineer: did a stress test it’s just another llm can’t rea…
Anthropic released Claude Opus 4.8, an incremental update over Opus 4.7 with sharper judgment and longer autonomous work capability, though some engineers remain skeptical about its code generation without extensive guidance.
Introducing Claude Opus 4.6
Anthropic announces Claude Opus 4.6, an upgraded version of their smartest model designed for better planning, longer task retention, and increased autonomy.
@danshipper: BREAKING: Claude Opus 5 is OUT NOW! And…it’s a hard model to love. We’ve spent the last week @every testing it across c…
An early vibe check on Claude Opus 5 finds it breaks backward compatibility with existing workflows but shows brilliance when starting from scratch, occupying an awkward middle ground between genius and generalist models.