@levie: At Box, we've been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is an…
Summary
Box tested Claude Sonnet 5.5 in early access and reported significant performance gains in complex work evaluations across financial services, legal, life sciences, and public sector, with faster processing and reduced token usage.
View Cached Full Text
Cached at: 09/28/26, 07:37 PM
At Box, we’ve been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is another strong jump from Sonnet 5 on knowledge work with enterprise content.
Across the board we saw a 4 pt overall improvement on our hardest tests, but it was roughly 2.4× faster to a finished deliverable on 12% fewer tokens. We also saw major wins in key verticals, like a +18 percentage point improvement in Financial Services, +7 pt in Legal, +8 pt in Life Sciences, and +7 pt in the Public Sector.
Here are some specific wins on certain tasks that show the power of Sonnet 5.5:
-
Financial services (+48 points): a due-diligence review of a trading-tool acquisition, where the deal book’s arithmetic doesn’t hold up. Sonnet 5.5 caught the miscalculated interest totals and the mispriced options and flagged them, and it also finished this work about 61% faster.
-
Legal (+14 points): a commercial lease review. Sonnet 5.5 refused to invent a “standard” market benchmark for renewal terms and notice periods, and instead flagged what the lease was genuinely missing.
-
Life sciences (+22 points): writing up a field trial from the raw measurement data. Sonnet 5.5 got the sample standard deviations right across both treatment groups. Sonnet 5 missed nearly all of them. And Sonnet 5.5 did it about 41% faster.
-
Public sector (+11 points): a report on a math intervention program. Sonnet 5.5 pulled the right student progress-monitoring figures out of the underlying data, about 71% faster.
Customers will be able to build AI agents with Sonnet 5.5 in the Box AI Studio shortly.
Claude (@claudeai): Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
Similar Articles
@levie: We've been running Anthropic's Claude Sonnet 5 through the Box AI Complex Work Eval, our agentic benchmark that puts mo…
Box ran Claude Sonnet 5 through its agentic benchmark, finding it surpasses Sonnet 4.6 in complex enterprise tasks like due diligence and cost analysis. Sonnet 5 will soon be available in Box AI Studio.
@levie: At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured d…
Box tested Claude Opus 5.5 and found it delivers significant performance improvements over Opus 5 for complex enterprise knowledge tasks, with major gains in token efficiency, speed, and cost.
@levie: At Box, we’ve been testing Fable 5.1 in early release against our complex enterprise work eval. Fable 5.1 delivers a hu…
Box tested Fable 5.1, which delivered a 7 percentage point improvement over Fable 5 in complex enterprise tasks, with notable gains in financial services, technology, and public sector, and it will be available in Box AI Studio alongside new models Claude Fable 5.1 and Claude Mythos 5.1.
@rohanpaul_ai: Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices. Over…
Claude Sonnet 5.5 is released by Anthropic, offering major improvements in benchmark scores, cost efficiency, and speed compared to Sonnet 5, with a 70.6% score on Terminal-Bench 4.0.
@cline: The new Sonnet 5 achieves Opus 4.8 level performance on Terminal-Bench for less than half the cost. Importantly for --y…
New Sonnet 5 model achieves Opus 4.8 level performance on Terminal-Bench at less than half the cost, with improved refusal of prompt injection attacks, now available in Cline.