@levie: At Box, we've been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is an…

X AI KOLs Timeline Models

Summary

Box tested Claude Sonnet 5.5 in early access and reported significant performance gains in complex work evaluations across financial services, legal, life sciences, and public sector, with faster processing and reduced token usage.

At Box, we've been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is another strong jump from Sonnet 5 on knowledge work with enterprise content. Across the board we saw a 4 pt overall improvement on our hardest tests, but it was roughly 2.4× faster to a finished deliverable on 12% fewer tokens. We also saw major wins in key verticals, like a +18 percentage point improvement in Financial Services, +7 pt in Legal, +8 pt in Life Sciences, and +7 pt in the Public Sector. Here are some specific wins on certain tasks that show the power of Sonnet 5.5: * Financial services (+48 points): a due-diligence review of a trading-tool acquisition, where the deal book's arithmetic doesn't hold up. Sonnet 5.5 caught the miscalculated interest totals and the mispriced options and flagged them, and it also finished this work about 61% faster. * Legal (+14 points): a commercial lease review. Sonnet 5.5 refused to invent a "standard" market benchmark for renewal terms and notice periods, and instead flagged what the lease was genuinely missing. * Life sciences (+22 points): writing up a field trial from the raw measurement data. Sonnet 5.5 got the sample standard deviations right across both treatment groups. Sonnet 5 missed nearly all of them. And Sonnet 5.5 did it about 41% faster. * Public sector (+11 points): a report on a math intervention program. Sonnet 5.5 pulled the right student progress-monitoring figures out of the underlying data, about 71% faster. Customers will be able to build AI agents with Sonnet 5.5 in the Box AI Studio shortly.
Original Article
View Cached Full Text

Cached at: 09/28/26, 07:37 PM

At Box, we’ve been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is another strong jump from Sonnet 5 on knowledge work with enterprise content.

Across the board we saw a 4 pt overall improvement on our hardest tests, but it was roughly 2.4× faster to a finished deliverable on 12% fewer tokens. We also saw major wins in key verticals, like a +18 percentage point improvement in Financial Services, +7 pt in Legal, +8 pt in Life Sciences, and +7 pt in the Public Sector.

Here are some specific wins on certain tasks that show the power of Sonnet 5.5:

  • Financial services (+48 points): a due-diligence review of a trading-tool acquisition, where the deal book’s arithmetic doesn’t hold up. Sonnet 5.5 caught the miscalculated interest totals and the mispriced options and flagged them, and it also finished this work about 61% faster.

  • Legal (+14 points): a commercial lease review. Sonnet 5.5 refused to invent a “standard” market benchmark for renewal terms and notice periods, and instead flagged what the lease was genuinely missing.

  • Life sciences (+22 points): writing up a field trial from the raw measurement data. Sonnet 5.5 got the sample standard deviations right across both treatment groups. Sonnet 5 missed nearly all of them. And Sonnet 5.5 did it about 41% faster.

  • Public sector (+11 points): a report on a math intervention program. Sonnet 5.5 pulled the right student progress-monitoring figures out of the underlying data, about 71% faster.

Customers will be able to build AI agents with Sonnet 5.5 in the Box AI Studio shortly.

Claude (@claudeai): Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.

It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

Similar Articles