@levie: Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities. At Box, we've been testi…

X AI KOLs Following Models

Summary

Claude Opus 5 has been released, showing significant improvements over Opus 4.8 in enterprise benchmarks across due diligence, life sciences, legal, technology, and healthcare tasks, with gains of 12-30%.

Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities. At Box, we've been testing Claude Opus 5 with the Box AI Agent on Box's Complex Work Eval, our agentic benchmark that puts models through real enterprise document work end-to-end across a variety of industries. We saw meaningful gains in performance across some of the most complex enterprise tasks on unstructured data. Here are a few examples of some of the wins we saw in testing: * Due diligence (+17%): On a transaction due-diligence review, Opus 5 worked through the full set of required findings and flagged the ones Opus 4.8 missed. It stays thorough as the checklist grows instead of catching the obvious items and stopping. * Life Sciences (+30%): On a target-identification task, it intersected several ranked datasets under a strict matching rule to find the targets common to all of them. Opus 4.8 over-included partial matches and dropped genuine ones. * Legal (+12%): On a clause-by-clause contract review against a policy, Opus 5 scored every item and correctly cleared the clauses that were acceptable under an exception where Opus 4.8 both missed items and mis-scored the exception. * Technology (+19%) and Healthcare (+13%) showed the same pattern: complete, precise, multi-step analysis over messy source material. Overall, the reasoning, analytical, and data processing skills of Opus 5 outshine Opus 4.8 meaningfully. Going to be very powerful for enterprise agentic use-cases. You'll be able to build agents with Opus 5 in the Box AI Studio shortly.
Original Article
View Cached Full Text

Cached at: 07/24/26, 07:14 PM

Claude Opus 5 is out now, and it’s a huge jump over Opus 4.8 on a wide number of capabilities.

At Box, we’ve been testing Claude Opus 5 with the Box AI Agent on Box’s Complex Work Eval, our agentic benchmark that puts models through real enterprise document work end-to-end across a variety of industries. We saw meaningful gains in performance across some of the most complex enterprise tasks on unstructured data.

Here are a few examples of some of the wins we saw in testing:

  • Due diligence (+17%): On a transaction due-diligence review, Opus 5 worked through the full set of required findings and flagged the ones Opus 4.8 missed. It stays thorough as the checklist grows instead of catching the obvious items and stopping.

  • Life Sciences (+30%): On a target-identification task, it intersected several ranked datasets under a strict matching rule to find the targets common to all of them. Opus 4.8 over-included partial matches and dropped genuine ones.

  • Legal (+12%): On a clause-by-clause contract review against a policy, Opus 5 scored every item and correctly cleared the clauses that were acceptable under an exception where Opus 4.8 both missed items and mis-scored the exception.

  • Technology (+19%) and Healthcare (+13%) showed the same pattern: complete, precise, multi-step analysis over messy source material.

Overall, the reasoning, analytical, and data processing skills of Opus 5 outshine Opus 4.8 meaningfully. Going to be very powerful for enterprise agentic use-cases. You’ll be able to build agents with Opus 5 in the Box AI Studio shortly.

Claude (@claudeai): Introducing Claude Opus 5.

It’s a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

Similar Articles

@FinanceYF5: 官方发布:

X AI KOLs Timeline

Anthropic releases Claude Opus 4.8, building on Opus 4.7 with sharper judgment and longer independent work capability, available at the same price.

Introducing Claude Opus 4.6

YouTube AI Channels

Anthropic announces Claude Opus 4.6, an upgraded version of their smartest model designed for better planning, longer task retention, and increased autonomy.