Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Summary
This paper introduces BusinessCaseBench, a benchmark of business case questions from 18 disciplines with expert grading rubrics. It finds that frontier AI models already score highly and show rapid improvement, with implications for business education and professional work.
View Cached Full Text
Cached at: 07/20/26, 09:36 AM
# Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning Source: [https://arxiv.org/html/2607.16057](https://arxiv.org/html/2607.16057) \\journaltitle\\DOI\\vol\\access\\appnotes Working Paper \\corresp \[∗\\ast\]To whom correspondence should be addressed:[me@ajayp\.app](https://arxiv.org/html/2607.16057v1/email:[email protected]) 0Year 0Year 0Year Kartik Hosanagar\\ORCID0000\-0002\-6442\-9434Ramayya Krishnan\\ORCID0000\-0001\-9935\-2468Chris Callison\-Burch\\ORCID0000\-0001\-8196\-1943Karim Lakhani\\ORCID0000\-0002\-5535\-8304Mitch Weiss\\orgdivOperations, Information and Decisions Department,\\orgnameThe Wharton School, University of Pennsylvania,\\orgaddress\\streetPhiladelphia,\\postcode19104,\\statePA,\\countryUSA\\orgdivHeinz College of Information Systems and Public Policy,\\orgnameCarnegie Mellon University,\\orgaddress\\streetPittsburgh,\\postcode15213,\\statePA,\\countryUSA\\orgdivDepartment of Computer and Information Science,\\orgnameUniversity of Pennsylvania,\\orgaddress\\streetPhiladelphia,\\postcode19104,\\statePA,\\countryUSA\\orgdivTechnology and Operations Management Unit,\\orgnameHarvard Business School, Harvard University,\\orgaddress\\streetBoston,\\postcode02163,\\stateMA,\\countryUSA\\orgdivEntrepreneurial Management Unit,\\orgnameHarvard Business School, Harvard University,\\orgaddress\\streetBoston,\\postcode02163,\\stateMA,\\countryUSA \(Date\) ###### Abstract Large language models \(LLMs\) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem\-solving, and coding and agentic tool\-use\. What remains poorly measured is AI progress on the analytical knowledge work white\-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi\-stakeholder settings, weighing trade\-offs, and producing defensible, structured analyses\. This gap is even more pronounced for subjective components of such work, where success can be challenging to define\. The “case method” form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert\-written instructor case solution\. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years\. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry\-level professional roles, where such skills have historically anchored early\-career work\. ###### keywords: knowledge work, analytical reasoning, business, education, benchmark, AI, large language models \\otherabstract \[Significance statement\]Business school trains students in decision\-making and analytical reasoning: reading financial statements, evaluating market and competitive landscapes, designing operations, assessing strategy—shaping analysts, managers, and leaders in business and white\-collar roles\. We show frontier AI models already perform these skills at a high level across eighteen business disciplines, graded against expert\-written standards\. A longitudinal comparison within one model family documents roughly a 23\-percentage\-point gain over two years\. We contribute BusinessCaseBench, a validated, discipline\-spanning benchmark for AI in business reasoning and knowledge work that shifts the empirical question: not whether AI can do the kind of work business education trains for, but how business education and professional roles may change as analytic capability redistributes between humans and machines\. Figure 1:The evaluation pipeline used to construct and score BusinessCaseBench\. Case narratives and open\-ended questions are paired with expert\-written reference solutions from the instructor case solution\. The reference solutions are transformed into equally\-weighted checklist rubrics\. A frontier AI model receives the case and question, produces an attempted solution, and an LLM\-as\-judge model scores the solution against each rubric criterion compared to the reference solution\. Scores aggregate to Standard scoring \(partial credit\) and Complete Answer scoring \(all criteria satisfied\) metrics reported throughout the paper\. We use human annotators to validate the automatic grading\.## 1Introduction Large language model \(LLM\) benchmark scores have risen rapidly\(gpt3\_few\_shot\_learners;singh\_openai\_2026\), yet the benchmarks driving those scores largely assess a narrow set of well\-defined competencies: factual recall, narrow question answering, mathematical problem\-solving, and coding and agentic tool\-use\(joshi\_triviaqa\_2017;zellers\_hellaswag\_2019;hendrycks\_measuring\_2021;chen\_evaluating\_2021;hendrycks\_measuring\_2021\-1;jimenez\_swe\-bench\_2024;merrill\_terminal\-bench\_2026\)\. That design is well suited when correctness is unambiguous and answers can be checked against known facts, computed values, or enumerated choices\. This pattern fails to capture, however, the capabilities required in the kind of analytical knowledge work white\-collar professionals perform every day\. That work calls for synthesizing complex information, exercising judgment under uncertainty and incomplete information, strategic and adversarial reasoning in multi\-stakeholder settings, weighing trade\-offs, and producing defensible, structured analyses—especially where success depends on subjective judgment and reasoning through multiple options rather than a single verifiable answer\. This is the kind of work business schools train students to perform and is economically valuable to organizations and firms yet is scarcely measured in existing benchmarks\(patwardhan\_gdpval\_2025;wang\_how\_2026\)\. A close analogue in medicine illustrates the same measurement gap and a constructive precedent\. High LLM scores on the United States Medical Licensing Examination \(USMLE\) led many to infer that LLMs had matched clinical reasoning, yet USMLE questions are multiple\-choice proxies optimized for administrative scale rather than for the open\-ended synthesis high\-performing clinicians perform on real cases\.nori\_sequential\_2025instead evaluated models onNew England Journal of Medicinecase challenges—clinically demanding and expert\-written diagnostic puzzle narratives with eventual solutions—using each case and its resolution to test reasoning under conditions closer to authentic practice with all of the ambiguity and uncertainty of real clinical work\. Frontier models performed strongly on this design, even though such case challenges are meant to be rare and difficult to solve, leading to new insights about the strength of medical reasoning in frontier AI\. The formulation validated expert\-written professional cases as a more challenging and realistic alternative to narrow exam\-style benchmarks like USMLE\. Business school case studies occupy a similar role in professional education that NEJM medical case challenges occupy in clinical training and education\. They are curated, high\-stakes, expert\-written narratives that place learners in realistic decision contexts that require synthesis of complex information, judgment under uncertainty and incomplete information, strategic reasoning across multiple stakeholders, weighing trade\-offs, and defensible structured analysis rather than narrow memorization or factual recall, as illustrated in Table[1](https://arxiv.org/html/2607.16057#S1.T1)\. Publisher licensing of these cases and their use in the educational setting also limits pre\-training exposure and contamination risk, making these cases a strong candidate for evaluation\(Xu2024BenchmarkDC\)\. These cases are often created by business school professors, with experience in conducting field studies, to describe a historical situation at a real organization or firm or a hypothetical situation at a fictional firm to train students for a particular business discipline\(rebeiz\_insider\_2011;Nohria2021CaseMethod\)\. Prior efforts to benchmark LLMs on business\-relevant tasks have remained narrowly scoped, with some targeting numerical reasoning within finance\(chen\_finqa\_2022;koncel\-kedziorski\_bizbench\_2024;wu\_bloomberggpt\_2023;xie\_finben\_2024\), others evaluating domain\-adjacent language tasks such as format\-following or ad\-copy generation\(xia\_fofo\_2024;liu\_llms\_2025;wang\_enterprise\_2025\), and a growing line of work assessing LLM agents operating in enterprise software systems or domain\-specific interactive environments\(drouin\_workarena\_2024;boisvert\_workarena\_2025;huang\_crmarena\_2025;huang\_crmarena\-pro\_2025;li\_investorbench\_2024;xu\_theagentcompany\_2025;patwardhan\_gdpval\_2025\)\. A few recent works have measured AI performance with business cases or simulations\(ai\_biz\_impact\_3;dellacqua2026jagged;allen2026strategy\), but are limited in their breadth of coverage across business disciplines and kinds of knowledge work\. These prior benchmarks and studies provide useful signal on specific capabilities relevant to business and knowledge work, but none measures open\-ended analytical work across the full breadth of business disciplines under conditions that mirror authentic professional judgment\. We extend case\-grounded evaluation to business disciplines through a design pattern and automated pipeline for benchmarking complex analytical knowledge work\. Using business school case studies paired with reference solutions derived from instructor case solutions, we generate structured, equally\-weighted checklist rubrics\. Model responses are then evaluated against those rubric criteria for scoring, following the pipeline illustrated in Fig\.[1](https://arxiv.org/html/2607.16057#S0.F1)\. These evaluations yield BusinessCaseBench, a discipline\-spanning capability map of frontier LLMs across 615 questions drawn from 238 cases spanning eighteen business disciplines\. Each question is additionally mapped to the O\*NET occupational taxonomy\(onet\_online\_2026\), so results can be read by professional work activity and occupation and used to characterize where implied AI impact concentrates\. Finally, a four\-model within\-family generational trajectory documents how performance has evolved over two years\. Several findings follow for how business education and labor markets should adapt as frontier AI capability on this class of work rises\. Top frontier models achieve rubric\-graded accuracy above 87% under partial credit scoring, suggesting that aggregate performance on these tasks, many of which involve reasoning and judgment over uncertain and ambiguous information, is now high\. Improvement across model generations is broad\-based, with approximately a 23\-percentage\-point gain within one model family over two years, and is not confined to quantitative subtasks where older models previously had the most room to improve\(hendrycks\_measuring\_2021\-1\)\. Discipline is associated with far wider differences in model difficulty than question type, with gaps between questions from fictional and real cases, numerical and non\-numerical questions, and subjective and objective questions each small relative to the variation across disciplines\. Finally, genuine capability ceilings are rare\. Fewer than 7% of questions defeat every frontier model, indicating that the required knowledge is distributed across the frontier rather than absent from it and that complete and thorough responses to highly open\-ended analytical tasks remain the main challenge and opportunity for human\-AI collaboration\. Table 1:Representative examples from BusinessCaseBench of case questions\. ## 2Overview BusinessCaseBench evaluates open\-ended analytical knowledge work using 238 business school cases paired with instructor solutions, yielding 615 questions across eighteen disciplines from “Strategy” and “Finance” to “Leadership & Organizational Behavior,” with the full compositional breakdown reported in SI Appendix, Section[A1 Final dataset schema](https://arxiv.org/html/2607.16057#Sx3.SSx1)\. The benchmark spans subjective and open\-ended analytical questions as well as objective, numerical reasoning questions\. We classify each question as derived from a case about a fictional firm or a real firm, as well as whether the question is numerical or non\-numerical, and whether it is subjective or objective\. The benchmark includes 245 questions on fictional firms and 370 on real firms, 184 numerical and 431 non\-numerical questions, and 339 subjective and 276 objective questions\. Each question is mapped to the O\*NET occupational taxonomy developed by the U\.S\. Department of Labor\(onet\_online\_2026\)at the Work Activity \(WA\), Intermediate Work Activity \(IWA\), and Detailed Work Activity \(DWA\) levels\. Together these labels cover 24 WAs, 55 IWAs, and 108 DWAs, so results can be read against professional work classifications as well as academic disciplines; see SI Appendix, Section[A1 Final dataset schema](https://arxiv.org/html/2607.16057#Sx3.SSx1)for tabular summaries of these mappings\. For every question, the instructor case solutions supply a reference solution that we extract and use to generate a checklist\-style rubric, as illustrated in Fig\.[1](https://arxiv.org/html/2607.16057#S0.F1)\. A model receives the full case narrative and exam\-style question prompt and produces an open\-ended attempted solution\. The attempted solution is then evaluated with an LLM\-as\-judge protocol\(zheng\_judging\_2023\)against rubric criteria using the reference solution\. We report two complementary scores throughout\. Standard scoring awards partial credit as the rubric\-weighted fraction of criteria satisfied\. It captures how much of what the instructor case solution expected the model actually produced, even when not every element is present\. Complete Answer scoring, shown in Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2), is stricter and all\-or\-nothing\. It marks whether a model satisfies every criterion on a question’s rubric\. Standard scoring summarizes how much of the instructor rubric a response covers and, especially on open\-ended and subjective questions, often provides the more faithful gauge of answer quality when full checklist satisfaction is neither required nor always attainable\. Complete Answer scoring complements it by recording how often responses nonetheless satisfy every criterion—a conservative completeness check when partial credit scores are already high\. Precise formal definitions and the scoring protocol appear in the Methods section and SI Appendix, Section[6](https://arxiv.org/html/2607.16057#Sx3.T6)\. For a subset of questions, three trained annotators independently authored a rubric and graded each assigned response before seeing any automated rubric or score\. This blinded protocol checks whether the evaluation, automated rubrics, and automated grading accurately measure an attempted solution’s success or failure in a way that directionally correlates with human judgement; see the Methods section and SI Appendix, Section[B3 LLM\-as\-judge grading prompt](https://arxiv.org/html/2607.16057#Sx4.SSx3)for further details\. Primary comparisons use three frontier AI models, OpenAI GPT\-5\.4\(singh\_openai\_2026\), Anthropic Claude Sonnet 4\.6\(anthropic\_system\_2025\), and Google Gemini 3 Flash Preview\(deepmind\_gemini\_2025\), evaluated on all 615 questions with pinned model identifiers and sampling settings documented in SI Appendix, Section[6](https://arxiv.org/html/2607.16057#Sx3.T6)\. A within\-family generational analysis evaluates four successive OpenAI models spanning roughly two years \(GPT\-4 Turbo to GPT\-5\.4\) on the same fixed benchmark\(openai\_gpt\-4\_2024\)\. This comparison tracks how AI performance on this class of work has evolved over roughly two years of model development\. Figure 2:Frontier model performance under Standard scoring and Complete Answer scoring\.A\)Mean scores for GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview on all 615 BusinessCaseBench questions; bars show Standard scoring \(full height\) with Scores under Complete Answer scoring overlaid \(lighter segment\)\. Error bars are bootstrapped 95% confidence intervals\.B\)The same two metrics by business discipline \(only showing disciplines withn≥5n\\geq 5questions\); disciplines are ordered by ascending mean Standard scoring\. Scores under Complete Answer scoring are uniformly lower than those under Standard scoring and widen the spread across models and disciplines, indicating that high partial credit often coexists with incomplete multi\-criterion satisfaction\. ## 3AI performance across the full business canon Across the full 615\-question benchmark, the organizing pattern is broad competence under partial credit, much lower full rubric satisfaction at the same scale, and a discipline ordering that all three models share\. Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)summarizes this structure for GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview under Standard and Complete Answer scoring\. ### 3\.1Aggregate performance Under Standard scoring, all three frontier models perform at a uniformly high level on the full benchmark\. Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)A places Claude Sonnet 4\.6 at 88\.4% \(95% CI \[87\.1, 89\.7\]\), GPT\-5\.4 at 87\.2% \(95% CI \[85\.8, 88\.7\]\), and Gemini 3 Flash Preview at 81\.6% \(95% CI \[80\.0, 83\.4\]\)\. The top\-to\-bottom spread is only 6\.8 percentage points, and confidence intervals for the two leading models overlap throughout\. These scores indicate that, when rubric criteria are credited independently, frontier models already cover most of what instructor case solutions expect on open\-ended business case work spanning the full discipline canon\. ### 3\.2AI outputs as drafts, not verdicts High partial credit scores hide a stricter picture\. Complete Answer scoring requires every rubric criterion on a question to be satisfied, and full satisfaction is far rarer\. In Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)A, the lighter overlaid bars show that Claude Sonnet 4\.6 reaches it on 49\.6% of questions \(95% CI \[45\.9, 53\.8\]\), GPT\-5\.4 on 47\.6% \(95% CI \[43\.9, 51\.4\]\), and Gemini 3 Flash Preview on 32\.0% \(95% CI \[28\.6, 35\.6\]\)\. Even the leading model leaves more than half of questions without a fully complete answer by instructor standards\. The top\-to\-bottom spread widens to 17\.6 percentage points, roughly 2\.6 times the spread under Standard scoring, because a response can earn high partial credit while still missing one or more required elements\. Complete Answer scoring is intentionally a conservative lower\-bound measure of performance in this sense\. Any missed criterion fails the question, even when the satisfied portion of the rubric would already support a strong partial credit grade that many instructors might treat as analytically sufficient on subjective, open\-ended questions where full rubric satisfaction might be an unreasonably strict expectation\. High Standard scores therefore describe analytically strong drafts that capture much of the expected analysis, not outputs that meet the full scope of the instructor rubric without review\. ### 3\.3Discipline\-level variation Discipline\-level performance is structured rather than noisy\. Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)B orders the sixteen disciplines with at least five questions by ascending mean Standard score\. SI Appendix, Table S[11](https://arxiv.org/html/2607.16057#Sx6.T11)reports the underlying cross\-model means: Standard scores range from 80\.1% in Marketing & Sales to 95\.0% in Business & Government Relations, a spread that exceeds the gap among the three models themselves, and under Complete Answer scoring the same disciplines separate more sharply, from 25\.8% in Operations & Service Management to 82\.5% in Decision Analysis\. All three frontier models rank disciplines in nearly the same order, with high rank concordance \(Kendall’sWW= 0\.79 under Standard scoring and 0\.89 under Complete Answer scoring\(kendall\_problem\_1939\)\), indicating that difficulty reflects discipline\-specific task structure more than any single provider’s training idiosyncrasies\. Figure 3:Frontier AI performance by O\*NET Intermediate Work Activities \(IWAs\)\. Each cell is an IWA with at least three BusinessCaseBench questions \(n≥3n\\geq 3\); color encodes the mean score under Standard scoring averaged across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview on those questions \(red tones indicate higher implied AI impact to those work activities\)\. IWAs are sorted by implied AI impact\. Open\-ended advisory and opportunity\-identification activities \(e\.g\., identifying organizational opportunities, advising on financial matters\) rank among the hardest; structured analytical and explanatory activities approach ceiling performance\. ## 4Characterizing model capabilities and strengths Aggregate scores hide how case framing, question metadata, and cross\-model patterns organize the remaining gap\. We therefore characterize what predicts residual difficulty and how benchmark questions map onto professional work activities\. ### 4\.1Fictional cases are slightly easier than real ones Questions from real cases score slightly higher than questions from fictional cases on the full benchmark \(86\.5% vs\. 84\.6% under Standard scoring, a 1\.9 percentage\-point gap\), a pattern attributable largely to a discipline\-specific outlier in Finance rather than a uniform fiction–real effect\. In six of the nine disciplines with sufficient questions in both arms, questions from fictional cases score at or above questions from real cases, consistent with real narratives introducing longer contexts, more ambiguity, and extraneous detail that complicate synthesis and make for a more challenging case question\. Finance and Accounting reverse the pattern: fictional finance questions score 15\.0 percentage points below real finance questions \(70\.9% vs\. 85\.9%\), and Accounting shows a smaller 2\.7\-point gap in the same direction, with Finance standing out as the lone large negative outlier in SI Appendix, Fig\. S[D5 Fictional\-vs\-real performance delta by discipline](https://arxiv.org/html/2607.16057#Sx6.SSx5)\. We hypothesize that questions from real cases in these disciplines often anchor on named companies whose financial profiles may appear in public filings or financial press, giving models a pre\-training edge that questions from fictional cases, which strip away that exposure, do not provide\. The Finance outlier therefore likely reflects how pre\-training interacts with case framing rather than a general advantage of fictional material\. ### 4\.2Difficulty is largely case\- and question\-specific Characterizing what makes a question hard largely resists coarse categorical metadata like question type or discipline, and most of the variation in difficulty lies beyond these labels in demands specific to a given case or question\. SI Appendix, Fig\. S[D4 Scoring by question\-type strata](https://arxiv.org/html/2607.16057#Sx6.SSx4)places mean performance across the six question\-type strata in a narrow 82\.3–87\.2% band under Standard scoring\. Numerical questions score lowest \(82\.3%\) and non\-numerical highest \(87\.2%\), subjective and objective questions differ by only 0\.8 percentage points, and questions from fictional versus real cases differ by 1\.9 points when pooled across all questions, although we note earlier that for some disciplines the gap is larger\. Complete Answer scoring widens these gaps by question type modestly but leaves them small relative to the spread across disciplines that Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)shows\. A regression of mean frontier Standard score on question\-type strata and discipline categories, using effect coding with disciplines of fewer than five questions pooled into a single Other category and with standard errors clustered by case, confirms that these coarse labels explain little of the variation in scores\. Question\-type tags alone account for 2\.6% of question\-level score variation \(adjustedR2R^\{2\}\), discipline alone for 3\.6%, and both together for 5\.1%, and under cross\-validation the two blocks predict held\-out cases comparably\. Each block still carries a detectable association with difficulty\. Adding discipline after question type is jointly significant \(p<0\.001p<0\.001\), as is adding question type after discipline \(p=0\.007p=0\.007\)\. Once discipline and the other strata are accounted for, scores on fictional and real cases no longer differ in a meaningful way, and numerical questions remain modestly harder \(−\-5\.9 points\)\. Most of the difference between harder and easier questions therefore lies in case\- and question\-specific demands that coarse question\-type and discipline metadata do not capture, such as longer ambiguous narratives, multi\-part rubric structure, and more open\-ended advisory or evaluative reasoning\. A variance decomposition across the238238source cases supports that reading\. Questions from the same case have moderately correlated difficulty \(intraclass correlation0\.220\.22\), knowing which case a question comes from predicts its score far better than the coarse metadata labels do \(adjustedR2=0\.22R^\{2\}=0\.22vs\.0\.050\.05\), and nearly half of the variation in scores occurs among questions within the same case\. Full nested\-model statistics, adjustedR2R^\{2\}, cross\-validated comparisons, and the case\-level decomposition appear in SI Appendix, Section[D5 Fictional\-vs\-real performance delta by discipline](https://arxiv.org/html/2607.16057#Sx6.SSx5)\. ### 4\.3Difficulty is rarely absolute Few benchmark questions defeat every evaluated model\. SI Appendix, Table S[15](https://arxiv.org/html/2607.16057#Sx6.T15)reports that only 43 of 615 questions \(7\.0%\) have a maximum Standard score at or below 70% across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview, and that only 521 of 8,369 rubric criterion instances \(6\.2%\) never receive credit from any of the three\. A cross\-model oracle that takes, per question, the best score among the three frontier models reaches 92\.8%, 4\.5 percentage points above the leading single model\. What one model misses on a given question is therefore often captured by another, so remaining error reflects incomplete satisfaction of multi\-part rubrics more often than absolute capability ceilings; SI Appendix, Section[D6 Score variation by question type, discipline, and case](https://arxiv.org/html/2607.16057#Sx6.SSx6)documents the underlying complementarity pattern\. ### 4\.4From cases to an occupational map Fig\.[3](https://arxiv.org/html/2607.16057#S3.F3)maps benchmark questions to O\*NET Intermediate Work Activities, bridging academic disciplines to occupational work activities and occupations to assess the implied AI impact that BusinessCaseBench would predict\. Across 29 Intermediate Work Activities \(IWAs\) with at least three benchmark questions, mean Standard scores range from 70\.9% on “Identify business or organizational opportunities” to ceiling performance on bounded explanatory and mathematical analysis activities\. The hardest activities combine open\-ended advisory framing with case\-specific quantitative or evaluative reasoning, whereas well\-scoped data interpretation and performance evaluation tasks approach saturation\. Because O\*NET’s work activity taxonomy involves overlapping and partially redundant categories at different levels of granularity\(onet\_limitations\), we treat these mappings as indicative rather than definitive\. Implied AI impact on occupations, which SI Appendix, Fig\. S[Implied AI impact on O\*NET occupations](https://arxiv.org/html/2607.16057#Sx6.SSx8.SSSx1)provides, should be read alongside the coverage bands in that figure that show how thoroughly each job’s relevant work activities are exercised in the benchmark\. Figure 4:Within\-family performance trajectory for four OpenAI models evaluated on the fixed 615\-question BusinessCaseBench \(GPT\-4 Turbo, GPT\-4\.1, GPT\-5, and GPT\-5\.4;≈\\approxtwo years\)\.A\)Standard scoring rubric\-weighted scores with 95% bootstrap confidence bands; aggregate gain from earliest to latest model is annotated\.B\)Scores under Complete Answer scoring on the same models\.C\)Standard scoring decomposed by case type and question type \(fictional vs\. real cases; numerical vs\. non\-numerical; subjective vs\. objective\)\.D\)Scores under Complete Answer scoring for the same six strata\. All strata improve from GPT\-4 Turbo to GPT\-5\.4\. ## 5A generational comparison within one model family How quickly has AI capability on this class of work improved? Holding the benchmark fixed, we evaluate four successive OpenAI releases spanning roughly two years, from GPT\-4 Turbo through GPT\-4\.1, GPT\-5, and GPT\-5\.4, to trace the generational trajectory of performance on open\-ended analytical knowledge work\. We choose the OpenAI model family as we are able to reliably access all four models for retrospective evaluation\. ##### Four\-generation performance trajectory Fig\.[4](https://arxiv.org/html/2607.16057#S4.F4)A and Fig\.[4](https://arxiv.org/html/2607.16057#S4.F4)B summarize the trajectory on all 615 questions under Standard scoring and Complete Answer scoring\. Mean Standard score rises from 63\.9% for GPT\-4 Turbo \(95% CI \[62\.0, 65\.9\]\) to 87\.2% for GPT\-5\.4 \(95% CI \[85\.8, 88\.7\]\), a gain of 23\.3 percentage points across the four models\. The stricter Complete Answer score moves in parallel, from 13\.2% \(95% CI \[10\.7, 15\.8\]\) to 47\.6% \(95% CI \[43\.9, 51\.4\]\), a 34\.4\-point gain that shows improvement on full and complete answers, not only on partial credit\. Intermediate releases step up the curve in sequence, with GPT\-4\.1 reaching 80\.7% under Standard scoring and GPT\-5 crossing 84\.9% before the latest model approaches the high partial credit levels reported in Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)\. ##### Categorizing the distribution of improvement The gains are broad\-based rather than confined to easy slices of the benchmark\. Fig\.[4](https://arxiv.org/html/2607.16057#S4.F4)C and Fig\.[4](https://arxiv.org/html/2607.16057#S4.F4)D decompose the same four\-model trajectory by fictional versus real cases and by numerical, non\-numerical, subjective, and objective questions\. Under Standard scoring, fictional and real cases improve by essentially the same margin \(23\.0 and 23\.4 percentage points\), numerical questions by 29\.0 points, and non\-numerical questions by 20\.8 points\. Under Complete Answer scoring every stratum gains at least 30 points, and non\-numerical and subjective questions improve as much as or more than their numerical and objective counterparts \(for example, \+36\.2 points on non\-numerical questions versus \+30\.4 on numerical ones\)\. That pattern runs counter to a hypothesis in which progress is concentrated on quantitative reasoning alone, the kind of capability early models struggled with and that popular benchmarks measure—and that emphasis has led to improvements in numerical precision and mathematical reasoning\(hendrycks\_measuring\_2021\-1\)\. SI Appendix, Fig\. S[D3 OpenAI performance trajectory by discipline](https://arxiv.org/html/2607.16057#Sx6.SSx3)extends the same pattern to disciplines in panels A and B: all twelve disciplines with at least fifteen questions post double\-digit Standard gains, including large moves in Economics \(\+31\.7 points\), Accounting \(\+28\.8 points\), and Finance \(\+28\.7 points\), where GPT\-4 Turbo began well below today’s levels\. Even disciplines that were already relatively strong in 2024, such as Business Ethics and Strategy, continue to gain under Complete Answer scoring\. The most recent step from GPT\-5 to GPT\-5\.4 is smaller \(\+2\.3 points on Standard scoring and \+1\.1 points on Complete Answer scoring\), indicating slowing improvement, which is expected as scores approach the ceiling\. In summary, the two\-year arc documents a substantial gain in capability on the analytical work this benchmark measures\. ## 6Extended results on the latest frontier of ultra\-large models After the primary experiments reported above were completed, Anthropic released Claude Fable 5, a Mythos\-class model substantially larger than previously released public systems and marketed with particular emphasis on software engineering capability\. Deployment was temporarily interrupted under U\.S\. government export controls related to cybersecurity concerns\(anthropic\_redeploying\_fable\_2026\)\. As this release fell near the end of our evaluation period for this research, we report Claude Fable 5 as an after\-the\-fact Extended Result on BusinessCaseBench\. The purpose of this extension is to document how a contemporaneous ultra\-large frontier model performs on the open\-ended knowledge\-work tasks measured here\. Aggregate performance for Claude Fable 5 is nearly indistinguishable from Claude Sonnet 4\.6\. Under Standard scoring, Fable 5 attains 88\.0% \(95% CI \[86\.4, 89\.5\]\), compared with 88\.4% for Sonnet 4\.6 \(SI Appendix, Table S[9](https://arxiv.org/html/2607.16057#Sx6.T9)\)\. Under Complete Answer scoring, Fable 5 attains 50\.9% \(95% CI \[47\.3, 54\.8\]\), a 1\.3\-percentage\-point difference relative to Sonnet 4\.6 at 49\.6% \(SI Appendix, Table S[10](https://arxiv.org/html/2607.16057#Sx6.T10)\)\. Confidence intervals for the two models overlap on both metrics\. Such similarity is consistent with substantial overlap in training data between successive models within a provider family\. SI Appendix, Table S[11](https://arxiv.org/html/2607.16057#Sx6.T11)displays the discipline\-level breakdown for Fable 5; across disciplines, Fable 5 remains close to Sonnet 4\.6 and does not rearrange the harder and easier domains previously identified\. On BusinessCaseBench, Claude Fable 5’s similar performance profile to Claude Sonnet 4\.6 may reflect this release class’s greater emphasis on agentic software engineering capabilities like cybersecurity\. The Extended Result therefore leaves the study’s qualitative conclusions unchanged\. Frontier models already achieve high Standard scores, Complete Answer satisfaction remains substantially harder, and the latest ultra\-large Anthropic release does not alter the capability picture this benchmark establishes for business knowledge work\. One provisional conclusion that may be drawn, however, is that optimizing for software engineering workflows does not necessarily improve the capabilities of these models on this class of knowledge work\. ## 7Discussion Frontier AI models already perform well on open\-ended business case questions graded against expert\-written instructor case solutions, yet this economically central work remains weakly represented in the benchmarks that track LLM progress\. On this benchmark, aggregate performance under Standard scoring is high, and capability gains within one model family over roughly two years are large, as the generational trajectory in Fig\.[4](https://arxiv.org/html/2607.16057#S4.F4)shows\. Residual difficulty is characterized more by gaps on certain business disciplines and occupational activity than by simple categorization \(numerical vs\. non\-numerical or subjective vs\. objective questions\)\. Few questions systematically defeateveryevaluated frontier AI model, and a cross\-model oracle performs well above the best single model\. These patterns suggest that the remaining gap reflects incomplete satisfaction of multi\-part rubrics more often than truly missing domain knowledge, a distinction Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)makes especially visible under Complete Answer scoring, since even leading systems satisfy every rubric criterion on fewer than half of questions\. Gaps in attempted solutions are frequently model\-specific rather than absolute, and much of what one model misses may be captured by another model on the same question\. The empirical question has shifted from whether frontier models can perform the analytic knowledge work business cases demand to how completely they do so, where they remain weakest, and how business schools and firms should respond as capability improves\. Our evaluation rests on the same curated artifacts professional schools use to train judgment, reasoning, and analytical skills\. Treating expert\-written cases as evaluation instruments is increasingly attractive across specialized domains because they approximate the real task more closely than convenient proxies like standardized exams do, because they carry realistic context, simulate information gaps and ambiguity, require judgment under uncertainty, and their long\-form narratives and analytical solutions align more closely with realistic deployed LLM use than short multiple\-choice/free\-response questions found in standardized exams\.nori\_sequential\_2025demonstrated a challenging medical diagnosis and reasoning benchmark by using complex NEJM medical cases as a more realistic alternative than measuring medical reasoning with standardized exams like USMLE\. Our study and benchmark extend that design pattern to business school cases to test analytical reasoning skills required in realistic knowledge work\. In both works, license\-restricted cases impose an openness cost, but they also strengthen validity because verbatim pretraining exposure is less plausible and high scores are harder to dismiss as memorization\(Xu2024BenchmarkDC\)\. In this evaluation regime, models earn credit for satisfying articulated expectations rather than matching a single, short gold string, which better mirrors a quality response, at the price of requiring rubrics that are well specified and stably scored\. The case format also lowers barriers to benchmark construction for specialized domains\. Domain experts without technical training can describe authentic situations in the narrative form—the context at hand, what was known and unknown, the options weighed, trade\-offs, and the reasoning behind a decision, captured in a reference solution\. Once cases and solutions are written in that structure, automated extraction, rubric synthesis, and checklist grading can scale capability measurement across models and time, asnori\_sequential\_2025and BusinessCaseBench both illustrate\. More challenging benchmarks, in turn, help push model development further\. This approach still requires validation that the automated rubrics and grading match human judgements of quality and skill—a step we take in our study through human annotation\. Several limitations bound these results and point to further evaluation work in this domain\. While we do not find any evidence of verbatim contamination of case materials in common large pre\-training corpora, frontier models are trained on proprietary datasets that may or may not have exposure to some case materials that would be difficult to verify without open access\. As we discuss earlier, some cases based on real firms may hinder unbiased evaluation as similar data may be published under public market regulatory filings or in the financial press, although we find the effect to largely systematically affect a single discipline \(Finance\)\. As is commonly known, LLM\-as\-judge scoring retains residual risks, brittleness on long answers, and numerical lapses, even after human validation\(zheng\_judging\_2023\)\. In our grading, high partial credit under Standard scoring can coexist with failed completeness under Complete Answer scoring for cases where full\-credit is unreasonably strict and a strong partial credit answer may be a sufficient baseline of expert human performance\. Rubric\-guided scoring generally assesses only whether each stated criterion is accurately satisfied, but it does not audit the response for extraneous, unsubstantiated, or factually incorrect content outside the checklist\. The benchmark is in English, and other languages, institutional settings, and pedagogical norms may yield different profiles of skills and capability\. The occupational mapping in Fig\.[3](https://arxiv.org/html/2607.16057#S3.F3)is informative but incomplete both because O\*NET work activities are broad and may only be loosely connected to some of the skills tested in a case and because many work activities appear rarely in the case corpus\. Results should therefore be read alongside the occupation coverage bands shown in SI Appendix, Fig\. S[Implied AI impact on O\*NET occupations](https://arxiv.org/html/2607.16057#Sx6.SSx8.SSSx1)\. The evaluation is also single\-turn\. Iterative clarification, information gathering, active negotiation, and formal accountability in firms lie outside what we measure and such areas may be opportunities for human\-AI teaming\(ai\_biz\_impact\_1\)\. Progress on AI for knowledge work will require a portfolio of benchmarks that separate discrete skills from integrative judgment and tie models to economically meaningful tasks\(patwardhan\_gdpval\_2025;wang\_how\_2026\)\. We contribute BusinessCaseBench, a discipline\-spanning case benchmark with rubric\-graded scoring that sits between narrow quantitative/analytical suites\(chen\_finqa\_2022;koncel\-kedziorski\_bizbench\_2024;xie\_finben\_2024\)and agent benchmarks focused on enterprise software use\(drouin\_workarena\_2024;xu\_theagentcompany\_2025\)\. Performance against instructor case solution standards is already high and rising within at least one model family\. As Fig\.[3](https://arxiv.org/html/2607.16057#S3.F3)summarizes, models are comparatively strong on structured analytic tasks and more exposed on open\-ended advisory and analysis tasks\. They also often fall short of full rubric satisfaction under Complete Answer scoring even when partial credit under Standard scoring is substantial, a gap evident in Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)\. These findings from BusinessCaseBench have immediate implications for business education and white\-collar labor\. Case pedagogy trains the judgment, synthesis, and defensible reasoning that MBA programs and undergraduate curricula were designed to build, skills that have historically anchored early\-career analytical work\. With frontier models already scoring high under Standard scoring and improving quickly within at least one model family, producing a strong draft answer is increasingly inexpensive\. The design problem for business schools is therefore less whether students can generate plausible case analyses and more how curricula cultivate verification, thoroughness, and the ability to recognize what a fully complete answer requires, especially on tasks where Standard scores are already high but Complete Answer satisfaction remains elusive\(ai\_biz\_impact\_2\)\. For labor markets, the occupational mapping in Fig\.[3](https://arxiv.org/html/2607.16057#S3.F3)and the occupation coverage bands in SI Appendix, Fig\. S[Implied AI impact on O\*NET occupations](https://arxiv.org/html/2607.16057#Sx6.SSx8.SSSx1)help identify where entry\-level analytical work appears most exposed to AI\. They also suggest where human advantage may remain, particularly in work that depends on integrating stakeholder perspectives, exercising accountability, and navigating organizational context\. Such capabilities lie beyond the scope of single\-turn benchmarks and are therefore undermeasured in evaluations such as ours\.\. Those remaining gaps are natural sites for human–AI teaming rather than wholesale substitution\(dellacqua2026jagged;krakowski2025human\)\. The within\-family trajectory we document is a lower bound on how quickly capability can improve\. It does not forecast every model provider’s path, but it argues against treating today’s Standard scores as a ceiling on what organizations and firms will soon see in practice\. BusinessCaseBench also has implications for AI evaluation and model development\. As many widely used AI benchmarks approach saturation, measuring further progress will increasingly require evaluations grounded in economically meaningful knowledge work rather than narrow factual recall, mathematical reasoning, or coding tasks alone\. BusinessCaseBench identifies a subset of open\-ended analytical reasoning that remains challenging for frontier models despite their strong overall performance\. Models often produce analytically strong drafts while omitting important considerations, failing to integrate competing constraints, or stopping short of a complete recommendation\. These harder cases therefore provide a natural target for future model development, including post\-training, reinforcement learning, and synthetic data generation that reward synthesis, completeness, trade\-off analysis, and judgment under uncertainty\. Subsequent work can pair BusinessCaseBench with interactive protocols, organizational deployment studies, and direct human baselines to connect measured capability to institutional outcomes\. ## 8Methods ### 8\.1Benchmark construction BusinessCaseBench was constructed from licensed business school case PDFs and their paired instructor case solution PDFs, distributed by publishers under restricted use agreements; see SI Appendix, Section[E2 Scope of coverage](https://arxiv.org/html/2607.16057#Sx7.SSx2)for the list of released and withheld artifacts\. All LLM usage on these materials was conducted under Zero Data Retention \(ZDR\) policies\. Each evaluated instance associates a case narrative with an exam\-style question prompt, a reference solution drawn directly from the instructor case solution, an equally\-weighted checklist rubric, question\-type metadata, and O\*NET work\-activity tags\. The dataset schema is described in SI Appendix, Section[A Benchmark construction](https://arxiv.org/html/2607.16057#Sx3), and the compositional breakdown of BusinessCaseBench questions is reported in SI Appendix, Section[A1 Final dataset schema](https://arxiv.org/html/2607.16057#Sx3.SSx1)\. All PDFs were converted to linearized markdown using OCR from the olmOCR toolkit\(poznanski\_olmocr\_2025\), with theallenai/olmOCR\-2\-7B\-1025\-FP8vision\-language model served via vLLM\(kwon\_efficient\_2023\)applied in batch over the full corpus and followed by an LLM\-assisted cleaning pass on the resulting text\. We audited verbatim exposure of these materials in large web\-crawled pre\-training corpora using InfiniGram\(Liu2024InfiniGram\)and found no evidence of contamination in commonly used datasets\. SI Appendix, Section[E Validity, limitations, and reproducibility](https://arxiv.org/html/2607.16057#Sx7)reports the audit protocol and results for C4\(raffel2020exploring\), The Pile\(pile\), RedPajama\(together2023redpajama\), Dolma\(soldaini\-etal\-2024\-dolma\), and DCLM\(li2025datacomplmsearchgenerationtraining\)\. Prior to question construction, each case was independently labeled as depicting a fictional organization or a named real focal organization\. Questions and reference solutions were extracted solely from material appearing explicitly in the instructor case solution, not inferred from the case narrative, and any instance without an explicit reference solution in the instructor case solution was excluded\. Retained instances were standardized through LLM rewrite passes into self\-contained exam\-style question prompts, and an automated quality gate discarded any extraction that failed basic fidelity checks on the question–reference\-solution triple\. We further annotated each instance for numerical versus non\-numerical content, subjective versus objective framing, business discipline, and O\*NET taxonomy\. Although the construction pipeline drew on LLM assistance from question extraction through metadata annotation \(google/gemini\-2\.5\-proandgoogle/gemini\-2\.5\-flash,\(comanici\_gemini\_2025\)\), every retained question, metadata, and reference solution underwent manual review and inspection for quality before inclusion\. For each instance, an equally\-weighted checklist rubric was constructed from the exam\-style question prompt and the reference solution in the instructor case solution, following established best practices for rubric\-guided automated evaluation\(min\_factscore\_2023;lin\_wildbench\_2024;wei\_rocketeval\_2025\), with binary credit awarded per criterion and every criterion carrying equal weight for assigning partial credit\. Every question was additionally mapped to O\*NET Work Activity \(WA\), Intermediate Work Activity \(IWA\), and Detailed Work Activity \(DWA\) labels, with counts by level reported in SI Appendix, Section[A1 Final dataset schema](https://arxiv.org/html/2607.16057#Sx3.SSx1)\. Construction and evaluation workflows were orchestrated with DataDreamer\(patel\_datadreamer\_2024\)for reproducible LLM pipelines\. The evaluation pipeline code is released alongside the paper \(see Data availability\)\. ### 8\.2Frontier AI models and their evaluation The primary cross\-model analysis compares three frontier AI models, OpenAI GPT\-5\.4\(singh\_openai\_2026\), Anthropic Claude Sonnet 4\.6\(anthropic\_system\_2025\), and Google Gemini 3 Flash Preview\(deepmind\_gemini\_2025\)\. SI Appendix, Section[6](https://arxiv.org/html/2607.16057#Sx3.T6)lists the API identifiers and sampling settings for each\. To trace capability development within a single model family, we additionally evaluated four successive OpenAI releases spanning roughly two years, from GPT\-4 Turbo through GPT\-4\.1, GPT\-5, and GPT\-5\.4\(openai\_gpt\-4\_2024;singh\_openai\_2026\), on the same fixed question set\. Each solver model received the complete case narrative and exam\-style question prompt in a single turn, with no tool access, retrieval augmentation, or multi\-turn clarification; the inference prompt template, provider settings, and token limits appear in SI Appendix, Section[6](https://arxiv.org/html/2607.16057#Sx3.T6)\. Grading was conducted independently of generation by a fixed LLM\-as\-judge\(zheng\_judging\_2023\)held constant across all solver models\. We use the smaller, faster, and efficient Gemini 2\.5 Flash \(google/gemini\-2\.5\-flash\) model\(comanici\_gemini\_2025\)for the judging task as is typically standard in LLM\-as\-judge setups\. The judge evaluates each attempted solution criterion by criterion, awarding binary credit \(0 or 1 point\) against the checklist rubric and reference solution, following the protocol in SI Appendix, Section[B2 Inference prompt template](https://arxiv.org/html/2607.16057#Sx4.SSx2)\. All seven models were evaluated on the same 615 questions, with only the solver varying across runs\. To assess whether the choice of an LLM judge model reshapes comparative conclusions, we regraded answers from Claude Sonnet 4\.6, GPT\-5\.4, and Gemini 3 Flash Preview with three family\-matched small judge models \(Gemini 2\.5 Flash, Claude Haiku 4\.5, and GPT\-5\-mini\) on alln=615n=615questions\. We measured agreement in two complementary ways\. Rank\-order concordance used Kendall’s coefficient of concordance \(WW\) and pairwise Kendall’sτ\\tauon mean scores across the three solvers\. Relative\-score concordance used pairwise Pearson correlations of each judge’s scores over those same solvers, capturing whether judges agree on how far models sit apart while remaining invariant to absolute score level\. All three judges produced the same overall ranking \(W=1\.0W=1\.0; all pairwiseτ=1\.0\\tau=1\.0\), and relative score differences likewise agreed almost exactly \(pairwiser∈\[0\.999,1\.000\]r\\in\[0\.999,1\.000\]\)\. At the question level, agreement was only moderate \(mean pairwiseτ≈0\.56\\tau\\approx 0\.56; mean pairwiser≈0\.61r\\approx 0\.61; identical ranking on 35% of questions\), as expected from item\-level noise, yet these disagreements left the overall rank ordering unchanged\. We therefore conclude that the choice of LLM judge model does not significantly reshape comparative conclusions on BusinessCaseBench\. ### 8\.3Human annotation interface and blinded validation protocol To validate automated rubrics and grading scores against human judgments of answer quality, three annotators with business school grading experience \(Annotators A, B, and C\) completed a blinded evaluation protocol on a stratified random sample\. Recruitment and training procedures are described in SI Appendix, Section[C1 Interface and protocol](https://arxiv.org/html/2607.16057#Sx5.SSx1)\. The workflow was implemented in a custom web interface, shown in SI Appendix, Fig\. S[C1 Interface and protocol](https://arxiv.org/html/2607.16057#Sx5.SSx1), with staged unblinding that ensured human rubrics and scores were not shaped by automated outputs; see SI Appendix, Section[C Human annotation and blinded validation](https://arxiv.org/html/2607.16057#Sx5)for the full protocol\. For each assigned question, annotators first judged whether the exam\-style question prompt and reference solution were usable and well\-formed to further independently validate the quality of the extracted questions and solutions from the case PDFs\. Then, annotators independently authored their own checklist rubric and graded the model\-generated attempted solution before viewing the automated rubric or grade\. Subsequent stages asked annotators to evaluate the automated rubric and grading and to compare both layers directly to their own\. Grading open\-ended case responses is necessarily somewhat subjective as different instructors and graders may disagree on acceptance criteria and partial credit thresholds\. In this blinded protocol, annotators wrote independent rubrics and often assigned different partial credit scores to the same model response\. Given human annotators do not agree upon a single score, we do not expect automated scores to match any single human score exactly\. Validation statistics are reported on ten assignments in which all three annotators judged the same case, question, and model response\. The results support that the automatically generated rubrics and the LLM\-as\-judge procedure are directionally consistent with expert judgment and suitable for measuring relative progress across models, disciplines, and time: automated Standard scores correlated with human partial credit scores at Spearmanρ=0\.54\\rho=0\.54; annotators rated 100% of the automatically generated rubrics acceptable or mainly acceptable and 96% of the LLM\-as\-judge Standard scores acceptable or mainly acceptable; and in 74% of the assignments they preferred the automated score or judged it as equivalent to their own\. Further details and agreement statistics appear in SI Appendix, Section[C2 Annotator recruitment and training](https://arxiv.org/html/2607.16057#Sx5.SSx2)\. ### 8\.4Reporting of evaluation metrics and statistics We report two primary outcome metrics, formalized as follows\. Let modelmmanswer questionjj, which haskjk\_\{j\}equally\-weighted rubric criteria, and letcmji∈\{0,1\}c\_\{mji\}\\in\\\{0,1\\\}denote whether criterioniiis satisfied under LLM\-as\-judge grading\. The per\-question*Standard score*is smj=1kj∑i=1kjcmji,s\_\{mj\}=\\frac\{1\}\{k\_\{j\}\}\\sum\_\{i=1\}^\{k\_\{j\}\}c\_\{mji\},\(1\)a normalized value in\[0,1\]\[0,1\]\(reported as a score percentage in figures\)\. The model\-level Standard score is the unweighted mean over theN=615N=615benchmark questions, s¯m=1N∑j=1Nsmj\.\\bar\{s\}\_\{m\}=\\frac\{1\}\{N\}\\sum\_\{j=1\}^\{N\}s\_\{mj\}\.\(2\)*Complete Answer scoring*marks whether every criterion is satisfied on a question\. Define the indicator amj=𝟙\[smj=1\],a\_\{mj\}=\\mathbbm\{1\}\[s\_\{mj\}=1\],\(3\)and the*Complete Answer score*as its mean over questions, a¯m=1N∑j=1Namj\.\\bar\{a\}\_\{m\}=\\frac\{1\}\{N\}\\sum\_\{j=1\}^\{N\}a\_\{mj\}\.\(4\)The same definitions apply when reporting means over a subset of questions within a discipline, question\-type strata, or per\-IWA\. We report full leaderboards and decompositions of results in SI Appendix, Section[C3 Rubric and grading validation and agreement results](https://arxiv.org/html/2607.16057#Sx5.SSx3)\. Uncertainty, when provided, is quantified through nonparametric bootstrap resampling over questions, drawingB=500B=500resamples with replacement for each metric and reporting percentile 95% confidence intervals\. ## 9Acknowledgments The authors thank Daniel Rock and Mark Yatskar for an early discussion during the conceptualization of this work\. The authors thank the anonymous reviewers for their valuable suggestions\. ## 10Competing interest The authors declare that they have no competing interests\. ## 11Supplementary material Supplementary material is available below\. ## 12Funding This work is supported in part by funds from the U\.S\. Department of War\. ## 13Author contributions Conceptualization: A\.P\., K\.H\., R\.K\., and C\.C\.B\.; Data curation: A\.P\. and K\.H\.; Methodology, experiments, and analysis: A\.P\.; Manuscript writing: A\.P\.; Manuscript review and editing: A\.P\., K\.H\., R\.K\., C\.C\.B\., K\.L\., and M\.W\.; Resources: K\.H\., R\.K\., K\.L\., and M\.W\. ## 14Data availability All data and code \(where license terms allow us to share the data\) to reproduce the results in this article have been deposited to Zenodo at[https://doi\.org/10\.5281/zenodo\.20211020](https://doi.org/10.5281/zenodo.20211020)\. ## References ## Supplementary material Supplementary Information for “Frontier AI performance across the business disciplines: a case\-grounded benchmark of knowledge work and analytical reasoning” providing information about additional results, reproducibility details, and experimental details\. ## Table of contents ## A Benchmark construction ### A1 Final dataset schema Each evaluated item is one JSON record: case metadata, full case and instructor case solution text, a standalone exam\-stylequestion, asolutionderived from instructor case solution, a checklistgrading\_rubric, question\-type flags, and O\*NET work\-activity tags\. The illustrative record below is a representative example of the format; proprietary text fields are shown as ellipses\. ``` { "case_name": "ILLUSTRATIVE-001", "case_title": "Acme Corp: European Expansion", "case_summary": "In late 2023, Acme Corp---a $480M U.S. manufacturer---weighed entering the European market...", "case_clean_text": "...", "instructor_case_solution_clean_text": "...", "question": "Provide an analysis of Acme Corp’s strategic position and make a recommendation on...", "solution": "An analysis should size the European TAM at roughly 11B EUR and note that three incumbents...", "task_description": "Evaluate a business’s market entry and channel development strategy.", "numerical": false, "subjective": true, "grading_rubric": [ "Answer identifies European market size (TAM) correctly.", "Answer cites competitive threat from incumbent concentration.", "Answer computes or discusses break-even economics for entry.", "Answer notes regulatory or CE-mark compliance risk.", "Answer states a clear enter-or-defer recommendation for 2024.", "..." ], "discipline": "Strategy", "work_activity": "Making Decisions and Solving Problems", "work_activity_id": "4.A.2.b.1", "intermediate_work_activity": "Advise others on business or operational matters.", "intermediate_work_activity_id": "4.A.4.b.6.I05", "detailed_work_activity": "Measure effectiveness of business strategies or practices.", "detailed_work_activity_id": "4.A.2.a.1.I02.D03" } ``` ### A2 Composition tables The following tables summarize the composition of the benchmark used in our analysis\. They report the overall coverage and break down the questions in the dataset by strata, discipline, and by the O\*NET taxonomy of work activities, providing a compact view of the distribution of questions in the benchmark\. Table 2:Benchmark composition: overall question\-level metadata\.Table 3:Benchmark composition by business discipline\. Non\-num\. = non\-numerical\.Table 4:Benchmark composition by O\*NET Work Activity \(WA\)\. Rows are ordered by question count descending\.Work ActivityWA IDnnAnalyzing Data or Information4\.A\.2\.a\.4245Making Decisions and Solving Problems4\.A\.2\.b\.1108Developing Objectives and Strategies4\.A\.2\.b\.489Judging the Qualities of Things, Services, or People4\.A\.2\.a\.128Provide Consultation and Advice to Others4\.A\.4\.b\.627Interpreting the Meaning of Information for Others4\.A\.4\.a\.124Processing Information4\.A\.2\.a\.219Selling or Influencing Others4\.A\.4\.a\.615Estimating the Quantifiable Characteristics of Products, Events, or Information4\.A\.1\.b\.313Evaluating Information to Determine Compliance with Standards4\.A\.2\.a\.39Thinking Creatively4\.A\.2\.b\.28Getting Information4\.A\.1\.a\.16Organizing, Planning, and Prioritizing Work4\.A\.2\.b\.65Identifying Objects, Actions, and Events4\.A\.1\.b\.14Coaching and Developing Others4\.A\.4\.b\.54Documenting/Recording Information4\.A\.3\.b\.62Training and Teaching Others4\.A\.4\.b\.32Monitor Processes, Materials, or Surroundings4\.A\.1\.a\.21Interacting With Computers4\.A\.3\.b\.11Resolving Conflicts and Negotiating with Others4\.A\.4\.a\.71Staffing Organizational Units4\.A\.4\.c\.21Updating and Using Relevant Knowledge4\.A\.2\.b\.31Performing Administrative Activities4\.A\.4\.c\.11Communicating with Persons Outside Organization4\.A\.4\.a\.31Table 5:Benchmark composition by O\*NET Intermediate Work Activity \(IWA\)\. Rows ordered by question count descending\.Intermediate Work ActivityIWA IDnnAnalyze business or financial data\.4\.A\.2\.a\.4\.I11117Advise others on business or operational matters\.4\.A\.4\.b\.6\.I0588Evaluate programs, practices, or processes\.4\.A\.2\.a\.1\.I0248Analyze business or financial risks\.4\.A\.2\.a\.4\.I0344Research organizational behavior, processes, or performance\.4\.A\.1\.a\.1\.I0934Develop financial or business plans\.4\.A\.2\.b\.2\.I0933Investigate organizational or operational problems\.4\.A\.1\.a\.1\.I2231Develop organizational policies, systems, or processes\.4\.A\.2\.b\.4\.I0127Analyze market or industry conditions\.4\.A\.2\.a\.4\.I0226Develop business or marketing plans\.4\.A\.2\.b\.2\.I0320Assess characteristics or impacts of regulations or policies\.4\.A\.2\.a\.4\.I0913Analyze data to improve operations\.4\.A\.2\.a\.4\.I0713Develop organizational or program goals or objectives\.4\.A\.2\.b\.2\.I2713Prepare financial documents, reports, or budgets\.4\.A\.3\.b\.6\.I019Calculate financial data\.4\.A\.1\.b\.3\.I038Advise others on financial matters\.4\.A\.4\.b\.6\.I116Evaluate the characteristics, usefulness, or performance of products or technologies\.4\.A\.2\.a\.1\.I076Analyze scientific or applied data using mathematical principles\.4\.A\.2\.a\.4\.I046Examine financial activities, operations, or systems\.4\.A\.2\.a\.3\.I035Resolve personnel or operational problems\.4\.A\.4\.a\.7\.I035Develop contingency or emergency response plans\.4\.A\.2\.b\.2\.I195Identify business or organizational opportunities\.4\.A\.1\.b\.1\.I024Evaluate personnel capabilities or performance\.4\.A\.2\.a\.1\.I044Develop sustainable organizational or business policies or practices\.4\.A\.2\.b\.2\.I204Develop models of systems, processes, or products\.4\.A\.2\.b\.2\.I214Explain financial information\.4\.A\.4\.a\.1\.I043Evaluate condition of financial assets, property, or other resources\.4\.A\.2\.a\.1\.I093Reconcile financial data\.4\.A\.2\.a\.2\.I043Evaluate project feasibility\.4\.A\.2\.a\.1\.I103Determine operational methods or procedures\.4\.A\.2\.b\.1\.I042Maintain sales or financial records\.4\.A\.3\.b\.6\.I102Develop research plans or methodologies\.4\.A\.2\.b\.2\.I232Plan work activities\.4\.A\.2\.b\.6\.I022Research laws, precedents, or other legal data\.4\.A\.2\.a\.4\.I081Develop systems or practices to mitigate or resolve environmental problems\.4\.A\.2\.b\.2\.I161Present research or technical information\.4\.A\.3\.b\.6\.I031Examine materials or documentation for accuracy or compliance\.4\.A\.2\.a\.3\.I011Research historical or social issues\.4\.A\.1\.a\.1\.I181Prepare reports of operational or procedural activities\.4\.A\.3\.b\.6\.I151Monitor external affairs, trends, or events\.4\.A\.1\.a\.2\.I081Investigate criminal or legal matters\.4\.A\.1\.a\.1\.I031Develop professional relationships or networks\.4\.A\.4\.a\.4\.I011Analyze environmental or geospatial data\.4\.A\.2\.a\.4\.I011Evaluate production inputs or outputs\.4\.A\.2\.a\.1\.I051Estimate project development or operational costs\.4\.A\.1\.b\.3\.I021Collect data about consumer needs or opinions\.4\.A\.1\.a\.1\.I141Develop marketing or promotional materials\.4\.A\.2\.b\.2\.I121Monitor operations to ensure adequate performance\.4\.A\.1\.a\.2\.I021Explain regulations, policies, or procedures\.4\.A\.4\.a\.1\.I021Evaluate the quality or accuracy of data\.4\.A\.2\.a\.2\.I011Process digital or online data\.4\.A\.3\.b\.1\.I061Prepare legal or regulatory documents\.4\.A\.3\.b\.6\.I141Monitor individual behavior or performance\.4\.A\.1\.a\.2\.I061Gather data about operational or development activities\.4\.A\.1\.a\.1\.I061Prepare proposals or grant applications\.4\.A\.3\.b\.6\.I071Table 6:Benchmark composition by O\*NET Detailed Work Activity \(DWA\)\. Rows ordered by question count descending\.Detailed Work ActivityDWA IDnnAnalyze business or financial data\.4\.A\.2\.a\.4\.I11\.D0457Develop operating strategies, plans, or procedures\.4\.A\.2\.b\.4\.I01\.D0237Advise others on business or operational matters\.4\.A\.4\.b\.6\.I05\.D1033Analyze financial information\.4\.A\.2\.a\.4\.I11\.D0229Measure effectiveness of business strategies or practices\.4\.A\.2\.a\.1\.I02\.D0327Analyze operational data to evaluate operations, processes or products\.4\.A\.2\.a\.4\.I07\.D0327Develop business or market strategies\.4\.A\.2\.b\.2\.I03\.D0126Analyze data to inform operational decisions or activities\.4\.A\.2\.a\.4\.I07\.D1223Analyze data to assess operational or project effectiveness\.4\.A\.2\.a\.4\.I07\.D0919Analyze market conditions or trends\.4\.A\.2\.a\.4\.I02\.D0418Assess risks to business operations\.4\.A\.2\.a\.4\.I03\.D0115Analyze risks to minimize losses or damages\.4\.A\.2\.a\.4\.I03\.D0314Analyze financial records or reports to determine state of operations\.4\.A\.2\.a\.4\.I11\.D0613Develop organizational policies or programs\.4\.A\.2\.b\.4\.I01\.D0113Analyze data to identify or resolve operational problems\.4\.A\.2\.a\.4\.I07\.D1412Develop financial or business plans\.4\.A\.2\.b\.2\.I09\.D0212Analyze costs and benefits of proposed designs or projects\.4\.A\.2\.a\.4\.I05\.D0810Apply mathematical models of financial or business conditions\.4\.A\.2\.a\.4\.I11\.D0310Develop organizational goals or objectives\.4\.A\.2\.b\.2\.I27\.D0210Analyze budgetary or accounting data\.4\.A\.2\.a\.4\.I11\.D019Apply mathematical principles or statistical approaches to solve problems in scientific or applied fields\.4\.A\.2\.a\.4\.I04\.D019Evaluate effectiveness of personnel policies or practices\.4\.A\.2\.a\.1\.I02\.D048Calculate financial data\.4\.A\.1\.b\.3\.I03\.D017Establish business management methods\.4\.A\.2\.b\.4\.I01\.D057Develop marketing plans or strategies\.4\.A\.2\.b\.2\.I03\.D037Develop sustainable business strategies or practices\.4\.A\.2\.b\.2\.I20\.D027Recommend organizational process or policy changes\.4\.A\.4\.b\.6\.I05\.D076Analyze industry trends\.4\.A\.2\.a\.4\.I02\.D025Develop organizational methods or procedures\.4\.A\.2\.b\.4\.I01\.D045Examine financial records to ensure compliance with policies or regulations\.4\.A\.2\.a\.3\.I03\.D055Analyze operational or research data\.4\.A\.2\.a\.4\.I07\.D025Analyze data to identify trends or relationships among variables\.4\.A\.2\.a\.4\.I04\.D025Analyze financial records to improve budgeting or planning\.4\.A\.2\.a\.4\.I11\.D054Analyze forecasting data to improve business decisions\.4\.A\.2\.a\.4\.I07\.D084Evaluate civic projects or public policies\.4\.A\.2\.a\.1\.I02\.D074Determine causes of operational problems or failures\.4\.A\.2\.b\.1\.I02\.D014Analyze market or customer related data\.4\.A\.2\.a\.4\.I02\.D064Advise others on ways to improve processes or products\.4\.A\.4\.b\.6\.I05\.D114Prepare financial documents, reports, or budgets\.4\.A\.3\.b\.6\.I01\.D024Evaluate applicable laws and regulations to determine impact on organizational activities\.4\.A\.2\.a\.4\.I09\.D044Calculate data to inform organizational operations\.4\.A\.2\.a\.4\.I07\.D013Develop sustainable organizational policies or practices\.4\.A\.2\.b\.2\.I20\.D033Interpret research or operational data\.4\.A\.2\.a\.4\.I07\.D053Develop plans for programs or services\.4\.A\.2\.b\.2\.I27\.D033Evaluate potential of products, technologies, or resources\.4\.A\.2\.a\.1\.I07\.D033Evaluate program effectiveness\.4\.A\.2\.a\.1\.I02\.D013Develop contingency plans to deal with organizational emergencies\.4\.A\.2\.b\.2\.I19\.D033Maintain financial or account records\.4\.A\.3\.b\.6\.I10\.D032Prepare financial documents\.4\.A\.3\.b\.6\.I01\.D012Maintain knowledge of business operations\.4\.A\.2\.b\.3\.I01\.D172Conduct quantitative failure analyses of operational data\.4\.A\.2\.a\.4\.I12\.D012Conduct research on social issues\.4\.A\.1\.a\.1\.I18\.D032Evaluate project designs to determine adequacy or feasibility\.4\.A\.2\.a\.1\.I10\.D012Negotiate contracts with clients or service providers\.4\.A\.4\.a\.7\.I02\.D102Devise research or testing protocols\.4\.A\.2\.b\.2\.I23\.D012Recommend investments to clients\.4\.A\.4\.b\.6\.I11\.D022Analyze impact of legal or regulatory changes\.4\.A\.2\.a\.4\.I09\.D032Develop emergency response plans or procedures\.4\.A\.2\.b\.2\.I19\.D042Determine pricing or monetary policies\.4\.A\.2\.b\.4\.I01\.D062Monitor financial indicators\.4\.A\.1\.a\.2\.I03\.D051Evaluate employee performance\.4\.A\.2\.a\.1\.I04\.D041Resolve operational performance problems\.4\.A\.4\.a\.7\.I03\.D021Analyze market research data\.4\.A\.2\.a\.4\.I02\.D081Investigate legal issues\.4\.A\.1\.a\.1\.I03\.D041Make decisions in legal cases\.4\.A\.2\.b\.1\.I05\.D011Develop marketing plans or strategies for environmental initiatives\.4\.A\.2\.b\.2\.I03\.D021Develop environmental sustainability plans or projects\.4\.A\.2\.b\.2\.I08\.D031Develop scientific or mathematical models\.4\.A\.2\.b\.2\.I26\.D031Advise others on human resources topics\.4\.A\.4\.b\.6\.I05\.D081Assess product or process usefulness\.4\.A\.2\.a\.1\.I07\.D051Evaluate reports or designs to determine work needs\.4\.A\.1\.a\.1\.I02\.D121Identify investment opportunities or strategies\.4\.A\.1\.b\.1\.I02\.D031Advise others on financial matters\.4\.A\.4\.b\.6\.I11\.D031Assess financial status of clients\.4\.A\.2\.a\.1\.I09\.D021Evaluate the effectiveness of counseling or educational programs\.4\.A\.2\.a\.1\.I02\.D021Establish interpersonal business relationships to facilitate work activities\.4\.A\.4\.a\.4\.I01\.D041Determine the value of goods or services\.4\.A\.2\.b\.1\.I01\.D021Advise others on analytical techniques\.4\.A\.4\.b\.6\.I05\.D051Develop program goals or plans\.4\.A\.2\.b\.2\.I27\.D011Design research studies to obtain scientific information\.4\.A\.2\.b\.2\.I23\.D041Analyze risks related to investments in green technology\.4\.A\.2\.a\.4\.I03\.D021Implement advertising or marketing initiatives\.4\.A\.2\.b\.1\.I09\.D041Prepare analytical reports\.4\.A\.3\.b\.6\.I03\.D051Resolve personnel problems\.4\.A\.4\.a\.7\.I03\.D031Measure environmental characteristics\.4\.A\.1\.a\.2\.I09\.D031Recommend changes or corrective procedures\.4\.A\.4\.b\.6\.I05\.D121Develop detailed project plans\.4\.A\.2\.b\.6\.I02\.D081Analyze data to inform personnel decisions\.4\.A\.2\.a\.4\.I07\.D111Monitor external factors impacting operations\.4\.A\.1\.a\.2\.I08\.D041Develop proposals for current or prospective customers\.4\.A\.3\.b\.6\.I07\.D011Determine operational procedures\.4\.A\.2\.b\.1\.I04\.D041Document organizational or operational procedures\.4\.A\.3\.b\.6\.I08\.D301Identify strategic business investment opportunities\.4\.A\.1\.b\.1\.I02\.D041Communicate with clients about products, procedures, and policies\.4\.A\.4\.a\.1\.I02\.D061Identify opportunities to improve operational efficiency\.4\.A\.1\.b\.1\.I02\.D061Assess the cost effectiveness of products, projects, or services\.4\.A\.2\.a\.1\.I07\.D011Estimate costs of products, services, or materials\.4\.A\.2\.b\.1\.I01\.D011Develop database parameters or specifications\.4\.A\.2\.b\.2\.I06\.D011Prepare data for analysis\.4\.A\.3\.b\.1\.I06\.D071Explain technical product or service information to customers\.4\.A\.4\.a\.1\.I01\.D051Develop operating strategies, plans, or procedures for green or sustainable operations\.4\.A\.2\.b\.2\.I20\.D011Conduct scientific research of organizational behavior or processes\.4\.A\.1\.a\.1\.I09\.D031Research topics in area of expertise\.4\.A\.2\.b\.3\.I01\.D101Analyze jobs using observation, survey, or interview techniques\.4\.A\.2\.a\.4\.I07\.D101Evaluate environmental or sustainability projects\.4\.A\.2\.a\.4\.I05\.D071Recruit personnel\.4\.A\.4\.c\.2\.I01\.D051Forecast economic, political, or social trends\.4\.A\.2\.a\.4\.I02\.D031 ## B Model evaluation protocol ### B1 Evaluated models The following table lists every evaluated model: provider, API identifier, approximate release date, and inference role\. For LLM requests, we use temperature 0; where the API supports it, we enable provider reasoning \(approximatelymediumeffort, excluded from the graded response\)\. Each solver completion is capped at 10,000 tokens for reasoning and final answer combined\. Table 7:Models evaluated in this study\. All evaluations are single\-turn with temperature set to 0\. The frontier trio \(GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview\) are used for primary cross\-model comparisons; the four OpenAI models constitute the within\-family generational trajectory\. Claude Fable 5 is used for an after\-the\-fact Extended Result as a new model release that occurred during the course of this research; it is not part of the frontier trio\. ### B2 Inference prompt template Each solver model receives the full case text and the benchmark question in a single\-turn prompt: ``` You are given a business case and a question about the case. You must output an answer that thoroughly addresses the question and instructions, with full reasoning and justification for your answer. The Question: ’’’ {question} ’’’ The Case: ’’’ {case_clean_text} ’’’ ``` ### B3 LLM\-as\-judge grading prompt A fixed judge model \(google/gemini\-2\.5\-flash\) grades each attempted answer at temperature 0 \(provider reasoning disabled\), with the same 10,000\-token completion cap\. The judge assigns 0 or 1 point per rubric criterion and returns the total score between<<and\>\>brackets: ``` You are an experienced judge of LLM-generated answers. You are given an LLM-generated answer to a question about a business case and a grading rubric for the model’s answer to that question. Remember that business cases sometimes involve subjective questions meant to test critical-thinking and reasoning skills that have no objective correct answers. Other times, there are objective questions that have a single correct answer. You will be provided a rubric of criteria items to grade the LLM-generated answer against along with an example of a gold standard answer or description of an answer to the question taken from the instructor case solution for this business case. **Your task:** Reason through each rubric criterion and grade the LLM-generated answer against it. Each criterion is worth 1 point; assign 0 or 1 point per criterion. **Output:** Write out your reasoning for each criterion, then return your final total score for the LLM-generated answer between << and >> brackets. Question: ’’’ {question} ’’’ The LLM-Generated Answer: ’’’ {model_answer} ’’’ The Grading Rubric: ’’’ {grading_rubric_list} ’’’ Gold Standard Answer: ’’’ {solution} ’’’ Case Summary: ’’’ {case_summary} ’’’ ``` ## C Human annotation and blinded validation ### C1 Interface and protocol Annotators complete a seven\-step*Rubric & Score Validation*workflow in a web interface \(SI Appendix, Fig\. S[C1 Interface and protocol](https://arxiv.org/html/2607.16057#Sx5.SSx1)\)\. Staged unblinding keeps human\-made rubrics and grades from being biased or shaped by automatic model outputs\. Throughout, annotators can read the case PDF, instructor case solution PDF, question, and expert\-written reference solution; the attempted answer and automated scores are revealed only when required and after the annotator has created their own rubric and graded the attempted solution against it\. 1. 1\.Evaluate question and solution\.Judge whether the benchmark question and instructor reference solution are usable and well\-formed \(*no*attempted answer;*no*automated rubric or score\)\.*Options:*yes / no on the question; yes / no on the solution; optional brief justification \(early exit if either is no\)\. 2. 2\.Write an independent rubric\.Author a checklist of grading criteria \(*still no*attempted answer;*no*automated rubric or score\)\.*Options:*free\-form criterion lines \(minimum three\)\. 3. 3\.Grade the attempted answer\.The model’s response is shown for the first time; mark which human\-rubric criteria it satisfies \(*still no*automated rubric or score\)\.*Options:*met / not met per criterion \(human grading score computed\)\. 4. 4\.Evaluate the automated rubric\.The pipeline’s checklist rubric is shown\.*Options:*acceptable / mainly acceptable / unacceptable; optional justification\. 5. 5\.Evaluate the automated grading\.The LLM\-as\-judge score and reasoning are shown\.*Options:*acceptable / mainly acceptable / unacceptable; optional justification\. 6. 6\.Compare rubrics\.Side\-by\-side view of human and automated checklists\.*Options:*A strong alignment / B divergent but valid / C divergent and invalid; optional justification\. 7. 7\.Compare scores\.Side\-by\-side view of human and automated grading scores\.*Options:*A prefer automated / B equivalent / C prefer human; optional justification\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_annotation_interface.png) Human annotation interface used for blinded validation\. ### C2 Annotator recruitment and training Three annotators with training in business school case grading completed the validation layer \(denoted Annotators A, B, and C\)\. Before scoring live instances, each walked through the interface instructions and a practice assignment covering all seven steps\. Task instances were drawn from a stratified random sample of the benchmark\. The agreement analysis below focuses on ten annotation assignments in which all three annotators judged the same cases, benchmark questions, and model\-generated responses\. ### C3 Rubric and grading validation and agreement results SI Appendix, Table S[8](https://arxiv.org/html/2607.16057#Sx5.T8)summarizes validation metrics across the seven protocol steps\. The central question is whether an automatically generated rubric, scored consistently by a fixed LLM judge, is directionally consistent with expert judgment and receives acceptability endorsements from human graders who have independently produced their own rubrics and scores\. Table 8:Summary validation metrics from the blinded human\-annotation protocol on the ten annotation assignments \(same case, question, and model response\) judged independently by Annotators A, B, and C\.In Step 1, all three annotators rated each question and reference solution as usable and well\-formed \(100% yes on both dimensions; 100% agreement\)\. In Step 2, annotators wrote independent rubrics on every question \(mean 10\.7 criteria items; mean pairwise difference 3\.6 criteria items\)\. Variation in rubric lengths across human annotators is expected because grading open\-ended cases is somewhat subjective, even when graders work from the same case and reference solution materials\. In Step 3, human partial credit scores correlate with automated Standard scores at Spearmanρ=0\.54\\rho=0\.54\. For 40% of annotation assignments, the human and automated graded scores differed by no more than ten percentage points, compared with 46\.7% of inter\-annotator score pairs that fell within the same ten\-percentage\-point band\. Thus, automated–human score agreement is close to the level of agreement observed among independent human graders\. The automated judge is on average 6\.4 percentage points more lenient\. Agreement is often close on straightforward items and wider when annotators wrote rubrics of different granularity or applied different partial credit thresholds to the same response\. After viewing the automated outputs, annotators rated the automatically generated rubric acceptable or mainly acceptable on 100% of completed Step 4 judgments and rated the LLM\-as\-judge Standard score acceptable or mainly acceptable on 96% of Step 5 judgments\. The one unacceptable grading judgment occurred on a case where human scores ranged from 0% to 100% across annotators with different rubrics indicating high disagreement among the annotators\. In Step 6, 96% of rubric comparisons were classified as strongly aligned or divergent but still valid \(A or B\), and one comparison was rated invalid \(C\)\. In Step 7, 74% of score comparisons preferred the automated Standard score or judged the two scores equivalent \(A or B\)\. These results do not establish exact agreement between automated and human grading scores, and they should not be read that way given the subjectivity visible in Steps 2–3 and 7\. They do support a narrower claim that is sufficient for benchmark use: retained questions and instructor\-derived reference solutions passed human quality screening; expert graders endorsed the automatically generated rubric and LLM\-as\-judge Standard scores after working independently; and automated Standard scores correlate with human partial credit scores in the same direction \(Spearmanρ=0\.54\\rho=0\.54\)\. The validation evidence indicates that the automatic grading procedure is directionally aligned with expert judgment and appropriate for tracking relative progress across models, disciplines, and model generations\. ## D Extended results ### D1 Full leaderboards under both scoring regimes The tables below report aggregate performance across all 615 questions under Standard scoring and Complete Answer scoring, with bootstrapped 95% confidence intervals \(B=500B=500resamples\)\. Table 9:Full aggregate leaderboard under Standard scoring on all 615 benchmark questions\. 95% confidence intervals are bootstrapped \(B=500B=500resamples\)\. Models are sorted by mean score descending\. The four OpenAI models are evaluated on the same fixed question set, enabling direct within\-family comparison; GPT\-5\.4 also serves as the OpenAI representative in primary cross\-model comparisons\. Claude Fable 5 is reported below the rule as an after\-the\-fact Extended Result from a new model release during the course of this research; it is not part of the primary ranking\.Table 10:Full aggregate leaderboard under Complete Answer scoring on all 615 benchmark questions\. 95% confidence intervals are bootstrapped \(B=500B=500resamples\)\. Models are sorted by Complete Answer score descending\. See SI Appendix, Table S[9](https://arxiv.org/html/2607.16057#Sx6.T9)for Standard scoring\. Claude Fable 5 is reported below the rule as an after\-the\-fact Extended Result from a new model release during the course of this research; it is not part of the primary ranking\. ### D2 Per\-discipline by per\-model score matrix This matrix reports mean Standard scores for every evaluated model within each of the eighteen business disciplines, highlighting where capability is uniform and where models diverge\. Table 11:Per\-discipline performance matrix for the frontier trio under Standard scoring and Complete Answer scoring\. Mean is the unweighted average of the three frontier models from the primary experiments only\. Disciplines are ordered by ascending mean Standard scoring\. Bold values in the frontier\-trio columns indicate the highest score within each scoring block and discipline row among those three models\. Claude Fable 5 appears to the right of each Mean \(dotted rule\) as an after\-the\-fact Extended Result from a new model release during the course of this research; it is excluded from the Mean\. Fable 5 scores are bolded when strictly higher than all three frontier models in that scoring block\. L&OB = Leadership & Organizational Behavior; HRM = Human Resource Management; E&I = Entrepreneurship & Innovation; IT = Information Technology; B&GR = Business & Government Relations\. ### D3 OpenAI performance trajectory by discipline The figure tracks the four OpenAI models on the twelve disciplines with at least fifteen questions, under both scoring regimes, to show how generational gains vary by discipline\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_openai_trajectory_by_discipline.png) Discipline\-level OpenAI model family trajectories \(twelve disciplines withn≥15n\\geq 15questions\)\.A\)Standard scoring by model; lines highlight the three largest and three smallest discipline\-level gains from GPT\-4 Turbo to GPT\-5\.4\.B\)Scores under Complete Answer scoring by model for the same disciplines\. Every tracked discipline improves under both metrics, but gain magnitudes and terminal levels differ\. ### D4 Scoring by question\-type strata We report mean frontier performance \(under Standard and Complete Answer scoring\) by various question\-type strata, including whether questions are derived from fictional versus real cases, numerical versus non\-numerical questions, and subjective versus objective questions\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_question_type_strata_standard_vs_complete_scoring.png) Performance by case type and question type under Standard scoring and Complete Answer scoring\. Grouped bars report mean scores with 95% bootstrap confidence intervals for fictional vs\. real cases, numerical vs\. non\-numerical questions, and subjective vs\. objective questions, with each question averaged across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview\. Under Standard scoring all six strata fall in a narrow band; scores under Complete Answer scoring widen these gaps modestly but remain small relative to between\-discipline variation documented in Fig\.[2](https://arxiv.org/html/2607.16057#S2.F2)\. ### D5 Fictional\-vs\-real performance delta by discipline The figure plots, for each discipline with sufficient questions in both arms, the mean frontier score on fictional case questions minus the mean on real case questions, surfacing the Finance discipline as a major outlier as discussed in the main text\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_fiction_vs_real_performance_delta_by_discipline.png) Within\-discipline gap between fictional and real case questions under Standard scoring\. For each discipline with sufficient questions in both arms \(nfictional≥5n\_\{\\mathrm\{fictional\}\}\\geq 5andnreal≥5n\_\{\\mathrm\{real\}\}\\geq 5\), bars show the mean Standard score on fictional case questions minus the mean Standard score on real case questions, averaged across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview\. Positive values indicate higher scores on fiction; Finance is a pronounced negative outlier \(fiction scores far below real\), whereas most other disciplines show small fictional case advantages or near parity\. ### D6 Score variation by question type, discipline, and case To complement the descriptive strata and discipline comparisons in the main text, we fit descriptive ordinary least squares models predicting each question’s mean frontier Standard score \(average of GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview\)\. Scores lie on the 0–1 scale, so coefficients can be read as percentage\-point differences in average score\. The independent variables are binary indicators for numerical, subjective, and fictional\-case questions, together with discipline fixed effects coded using sum\-to\-zero contrasts\. Under this coding, discipline coefficients represent deviations from the adjusted average discipline effect\. Disciplines represented by fewer than five questions \(Social Enterprise,n=1n=1; Management Communications,n=4n=4\) are pooled into a single Other category before fitting: cells this sparse produce high\-leverage observations whose robust standard errors are unreliable, so the pooled Other term is included only as a control and its coefficient is not interpreted\. Because multiple questions from the same case may share unobserved difficulty, inference uses standard errors clustered by case \(238 clusters;N=615N=615\)\. SI Appendix, Table S[12](https://arxiv.org/html/2607.16057#Sx6.T12)reports nested models\. Taken at face value, question\-type tags alone explain 3\.1% of question\-level score variation, discipline alone explains 6\.1%, and both together explain 8\.0%\. On an adjustedR2R^\{2\}basis, the question\-type tags explain 2\.6%, discipline 3\.6%, and the two together 5\.1%\. Each adds information beyond the other\. Adding discipline after the question\-type tags raises adjustedR2R^\{2\}by 0\.025 \(clustered WaldF\(16,237\)=3\.11F\(16,237\)=3\.11,p<0\.001p<0\.001\), and adding the question\-type tags after discipline raises it by 0\.015 \(F\(3,237\)=4\.16F\(3,237\)=4\.16,p=0\.007p=0\.007\)\. Discipline keeps a modest edge in adjustedR2R^\{2\}\(3\.6% vs\. 2\.6%\), but both sets of labels leave most differences between harder and easier questions to case\- and question\-specific demands that coarse metadata do not capture\. As a stronger check, we asked how well each set of labels predicts scores it was not fitted on, using cross\-validation \(1010\-fold,5050repeats\)\. When entire cases are held out, so that the model must predict scores for cases it has never seen, the question\-type tags predict slightly better than discipline \(mean out\-of\-sampleR2R^\{2\}of about0\.0100\.010vs\.−0\.004\-0\.004, where the negative value means the predictions are no better than guessing the overall average score\)\. Discipline is essentially a property of the case, spread across many small categories, so what it explains within this sample does not carry over to new cases\. We therefore read discipline and question type as each carrying a detectable but small association with difficulty, rather than as evidence that discipline is the stronger predictor in general\. Table 12:Nested ordinary least squares models of the mean frontier Standard score \(average of GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview from the primary experiments;0–11scale\) on question\-type strata and discipline \(N=615N=615questions\)\. Discipline uses effect \(sum\-to\-zero\) coding, with disciplines of fewer than five questions pooled into a single Other category\. AdjustedR2R^\{2\}is reported alongside rawR2R^\{2\}because the discipline block carries many more parameters than the three question\-type indicators\.Δ\\DeltaAdj\.R2R^\{2\}is the gain in adjustedR2R^\{2\}from adding a block of predictors, andppis from a joint WaldFFtest with standard errors clustered by case \(238 clusters\)\.SI Appendix, Table S[13](https://arxiv.org/html/2607.16057#Sx6.T13)reports coefficients from the joint model\. In that specification, numerical questions score lower than non\-numerical ones, while subjective vs\. objective and fictional vs\. real framing are not distinguishable from zero once the other predictors are included\. Relative to the overall discipline mean, Business & Government Relations and Strategy score higher and Operations & Service Management scores lower atp<0\.05p<0\.05; Decision Analysis is marginally higher \(p=0\.051p=0\.051\), whereas Finance is not distinguishable from zero \(p=0\.117p=0\.117\)\. These descriptive decompositions show that coarse metadata leave most question\-level score variation unexplained\. The case\-level analysis below asks how much of that residual is shared within cases\. Table 13:Coefficients from the joint model of SI Appendix, Table S[12](https://arxiv.org/html/2607.16057#Sx6.T12), with the same outcome, sample, and case\-clustered standard errors\. Effect coding compares each discipline to the overall mean across disciplines\. The pooled Other category comprises the disciplines with fewer than five questions \(Social Enterprise,n=1n=1; Management Communications,n=4n=4\), which avoids unstable estimates from near\-empty cells\. The Other coefficient is included as a control and is not interpreted\. Coefficients are in score points on the0–11scale \(e\.g\.,−0\.05\-0\.05≈\\approx5 percentage points\)\.To test whether the difficulty left unexplained by coarse metadata is tied to the underlying case, we compare how much scores vary between the238238source cases versus among questions within the same case\. Questions drawn from the same case have moderately correlated difficulty \(intraclass correlation0\.220\.22\), and average scores differ across cases by more than chance alone would produce \(F\(237,377\)=1\.72F\(237,377\)=1\.72,p<0\.001p<0\.001\)\. Knowing which case a question comes from also predicts its score far better than the coarse metadata labels do \(adjustedR2=0\.22R^\{2\}=0\.22vs\.0\.050\.05for question\-type tags and discipline combined, SI Appendix, Table S[14](https://arxiv.org/html/2607.16057#Sx6.T14)\)\. Accounting for case identity on top of the metadata model raises adjustedR2R^\{2\}from0\.050\.05to0\.240\.24, and the case effects remain jointly significant after the metadata block \(F\(236,359\)=1\.64F\(236,359\)=1\.64,p<0\.001p<0\.001\), so the case\-level signal is not simply the metadata labels restated\. Even so, nearly half of the variation in scores occurs among questions within the same case, and the question\-type tags explain only about1%1\\%of that remaining variation\. Difficulty is therefore partly a property of the case, which motivates the case\-clustered standard errors above, but it remains substantially a property of the individual question\. Table 14:How much question\-level score variation is explained by case versus by coarse metadata\. The outcome, sample, and metadata model \(question type plus pooled discipline\) are as in SI Appendix, Table S[12](https://arxiv.org/html/2607.16057#Sx6.T12)\. The case intraclass correlation \(ICC\) is from a one\-way ANOVA with unequal cluster sizes across the238238cases \(60 contribute a single question and the rest contribute between two and ten\)\. Case fixed effects use one indicator per case\. The last row reports the adjustedR2R^\{2\}gain from adding case identity to the metadata model, and the nestedFFtests whether that gain is jointly zero\. Its numerator carries 236 rather than 237 degrees of freedom because fictional vs\. real status is constant within a case and is absorbed by the case indicators\. A sensitivity analysis restricting to cases with at leastkkquestions is reported in the text below\.Because cases contribute unequal numbers of questions \(60 of 238 cases contribute only one question, while the remainder contribute between two and ten\), we recompute the case ICC and case fixed\-effects fit after successively dropping cases below each size threshold\. Restricting to the 178 cases with at least two questions yields ICC=0\.211=0\.211and adjustedR2=0\.211R^\{2\}=0\.211\(N=555N=555\)\. Raising the threshold to three, four, or five questions per case leaves the ICC in a narrow 0\.20–0\.23 band and the adjustedR2R^\{2\}in a 0\.20–0\.22 band\. The case\-level association is therefore not an artifact of singleton cases or of the unequal cluster\-size distribution\. ### D7 Cross\-model complementarity The table summarizes how much performance rises when one takes, per question, the best score among the frontier trio \(oracle ceiling\) and how often all three models jointly fail the same questions or rubric criteria to gauge “universal hardness”\. Table 15:Cross\-model oracle ceiling and universal hardness incidence for the frontier trio \(GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview\)\. Universal hardness threshold is Standard scoring≤0\.70\\leq 0\.70on a question by all three frontier models\. ### D8 Extended O\*NET taxonomy results These supplementary analyses extend the main\-text IWA results \(Fig\.[3](https://arxiv.org/html/2607.16057#S3.F3)\) to both O\*NET occupations and the finer Detailed Work Activity \(DWA\) taxonomy\. #### Implied AI impact on O\*NET occupations The occupation chart \(SI Appendix, Fig\. S[Implied AI impact on O\*NET occupations](https://arxiv.org/html/2607.16057#Sx6.SSx8.SSSx1)\) groups O\*NET jobs by how densely the benchmark covers their relevant IWAs and reports mean frontier scores within each coverage band, demonstrating the implied AI impact on O\*NET occupations according to the benchmark\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_ai_impact_on_onet_occupations.png) Frontier AI performance on O\*NET occupations represented in the benchmark, grouped by breadth of relevant Intermediate Work Activity \(IWA\) coverage\. Occupations are partitioned into high, medium, and low coverage bands according to how many IWAs relevant to the occupation are represented sufficiently in the benchmark \(n≥3n\\geq 3\); within each band, color bars encode the mean score under Standard scoring averaged across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview on benchmark questions linked to relevant IWAs for that occupation, with the total number of contributing questions annotated\. Occupations are sorted by implied AI impact within each coverage band\. The figure summarizes which occupations the benchmark covers most densely and where implied AI impact concentrates at the job level\. #### Implied AI impact by detailed work activity The DWA heatmap \(SI Appendix, Fig\. S[Implied AI impact by detailed work activity](https://arxiv.org/html/2607.16057#Sx6.SSx8.SSSx2)\) extends the main\-text IWA view to the finest O\*NET work\-activity level\. \{Figure\}![[Uncaptioned image]](https://arxiv.org/html/2607.16057v1/figures/sfig_ai_impact_by_detailed_work_activity.png) Frontier AI performance by O\*NET Detailed Work Activities \(DWAs\)\. Each cell is a DWA with at least three benchmark questions \(n≥3n\\geq 3\); color encodes the mean score under Standard scoring averaged across GPT\-5\.4, Claude Sonnet 4\.6, and Gemini 3 Flash Preview on those questions, withnnand score annotated\. DWAs are sorted by implied AI impact\. The finer DWA taxonomy reveals additional granularity in which specific professional subtasks resist full rubric satisfaction relative to aggregate IWA\-level patterns\. ## E Validity, limitations, and reproducibility ### E1 Contamination evidence We audit verbatim exposure of evaluation materials using InfiniGram\(Liu2024InfiniGram\), which supports efficient substring search over large, commonly used web\-crawled pre\-training corpora \(C4\(raffel2020exploring\), The Pile\(pile\), RedPajama\(together2023redpajama\), Dolma\(soldaini\-etal\-2024\-dolma\), and DCLM\(li2025datacomplmsearchgenerationtraining\)\)\. For each case and its instructor case solution, we draw three random substrings of 5–10 tokens and query whether each appears in those indexes\. We find no hits for the cases used in the evaluation pipeline, so the indexed corpora provide no evidence of verbatim contamination on this pre\-training contamination probe\. ### E2 Scope of coverage The benchmark is single\-turn and each instance supplies full case context in one prompt with no clarification or iterative evidence gathering \(contrast with interactive diagnostic or agentic work benchmarks\)\. This is intentional by design as this is how cases are generally designed to be used and helps isolate improvement on knowledge work and analytical reasoning without confounding on tool use capabilities or conversational dialogue capabilities\. It is English\-only and grounded in the business school case\-method tradition \(long narrative cases with instructor\-derived reference standards\) so it measures analytic and advisory synthesis under that particular pedagogical format\. Results should be read as capability under this controlled protocol for evaluating business knowledge and reasoning\. ### E3 Released vs\. withheld artifacts and access pathway ##### Released\. Evaluation harness code, pinned model identifiers, hyperparameters, and prompts \(including those in SI Appendix, Section[6](https://arxiv.org/html/2607.16057#Sx3.T6)\), structure of benchmark metadata and rubrics, and aggregated model outputs and scores\. ##### Withheld\. Full case and instructor case solution PDFs and text, which publisher and clearinghouse licenses prohibit redistributing openly\. ##### Access pathway\. Qualified researchers can reconstruct the evaluation set by obtaining cases through standard business school case clearinghouses and instructor\-access channels, then repeating the procedure on locally held materials\. License restrictions reduce open reusability but also limit verbatim inclusion of proprietary cases in pre\-training corpora—the same constraint that allows such cases to be a valid measure of progress on knowledge work with low contamination risk\.
Similar Articles
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work
This paper presents QuestBench, a benchmark built by students to evaluate deep research systems across humanities and social science domains. Results show that even advanced systems like GPT-5.5 pass only 57.58% of questions, highlighting failures in trustworthiness.
Frontier and Center: Who evaluates the evaluations? (12 minute read)
Google Data Cloud's frontier AI team discusses a new approach to evaluating AI agents using information theory to create a meta-benchmark called Discovery Bench that measures how vague a query can be before an agent fails, providing a more nuanced map of agent capabilities than simple pass/fail exams.
Design and Report Benchmarks for Knowledge Work
This paper proposes a three-step framework for designing and reporting benchmarks for knowledge work AI, emphasizing alignment between benchmark tasks and real-world work activities. It derives 18 work activities from the O*NET database and analyzes three existing benchmarks (GDPval, OfficeQA Pro, APEX-SWE) to demonstrate gaps between benchmark scores and actual work capability.
Open-World Evaluations for Measuring Frontier AI Capabilities
This paper argues that traditional benchmarks both overestimate and underestimate frontier AI capabilities, and proposes 'open-world evaluations'—long-horizon, real-world tasks assessed qualitatively—as a complementary approach. The CRUX project is introduced, with a demonstration where an AI agent successfully published an iOS app to the App Store with minimal intervention.
JobBench: Aligning Agent Work With Human Will
JobBench is a benchmark built from worker surveys to evaluate AI agents on tasks that workers most want automated, covering 130 tasks across 35 professions with detailed rubrics.