How a new benchmark uses business schools’ signature teaching tool to measure AI against the analytical work careers are built on.
Listen to this article:
Most of us are already comfortable handing AI short, one-off tasks like document summaries, email drafts, and data analysis, but we are less comfortable trusting its results on more complex work. Can AI stay with a complicated business problem long enough to interpret messy evidence, weigh competing priorities, exercise judgment under uncertainty, and arrive at a recommendation that can be defended? Understanding AI’s ability to complete that level of work has largely been elusive because conventional benchmarks tend to reward quick factual recall, narrow question answering, mathematics, coding, or tool use. But in “Frontier AI Performance Across the Business Disciplines: a Case-Grounded Benchmark of Knowledge Work and Analytical Reasoning,” co-authored by chair and co-founder of the HBS AI Institute Karim R. Lakhani, a team of researchers tackled this challenge with BusinessCaseBench, a tool that repurposes the business school case method to test frontier AI models on genuine analytical knowledge work. What they found reframes the AI performance debate: today’s models are remarkably capable, but an important gap remains.
Key Insight: Can AI Get an A?
“A model receives the full case narrative and exam-style question prompt and produces an open-ended attempted solution.” [1]
BusinessCaseBench contains 615 questions drawn from 238 business cases across 18 disciplines, including strategy, finance, accounting, operations, business ethics, economics, and leadership. Models might be asked to write a judicial opinion on a pharmaceutical patent settlement, or value an automaker under changing currency assumptions. The benchmark also maps questions to O*NET work activities, connecting academic disciplines to recognizable forms of professional work. Each case was paired with an instructor-written solution. The researchers converted that solution into an equally weighted checklist rubric, gave a frontier model the full case and question, and then used a fixed LLM judge to score the response criterion by criterion. Models answered in one turn, without tools, retrieval, or clarification. That isolates analytical reasoning, but it also means the benchmark does not measure interactive investigation, negotiation, implementation, or accountability inside real organizations. The benchmark reports two scores: Standard scoring, which gives partial credit for each rubric item satisfied, and a stricter Complete Answer scoring, which only counts a response as a success if every single criterion is met.
Key Insight: Strong Drafts, Rarely Finished Work
“Fewer than 7% of questions defeat every frontier model.” [2]
On the surface, AI models look like star students. Under Standard scoring, Claude Sonnet 4.6 hit 88.4%, GPT-5.4 reached 87.2%, and Gemini 3 Flash Preview landed at 81.6%. Those scores show that the models generally included a large share of the elements specified in the instructor-derived rubrics. Under the stricter Complete Answer measure, however, the same models scored 49.6%, 47.6%, and 32.0% respectively. Even the best model failed to satisfy every rubric criterion on more than half the questions. The researchers’ practical interpretation is that current models are best treated as first-pass analyses whose coverage must be checked before they can be considered complete. But it’s also improving fast: tracing four OpenAI generations across roughly two years, Standard scores climbed from 63.9% to 87.2%, while Complete Answer scores more than tripled, from 13.2% to 47.6%.
Key Insight: Uniquely Difficult
“The hardest activities combine open-ended advisory framing with case-specific quantitative or evaluative reasoning.” [3]
Simple categories do not explain very much about how well LLMs perform. Whether a question was subjective or objective, numerical or not, or drawn from a real company or a fictional one accounted for only a small portion of the performance differences. The business discipline had more explanatory power. Standard scores ranged from about 80% in Marketing & Sales to 95% in Business & Government Relations, a wider spread than the gap between the models themselves. Structured, well-defined activities like crunching financial data or evaluating a program against clear criteria sit near ceiling performance. More ambiguous advisory work, like identifying a business opportunity, ranks among the hardest tasks by a wide margin. However, the underlying business case actually proved more informative than discipline or question type, suggesting that difficulty depends heavily on the particular narrative and question specifics. It seems that difficulty is often baked into a particular narrative with its ambiguity, its distractors, and its multi-part demands.
Why This Matters
The researchers also devised a “cross-model oracle” that selected the highest score achieved by any of the three frontier models on each question. It reached 93%, well above any single model alone, and reveals an important pattern: what one model misses, another often captures. That finding has real implications for how organizations should actually deploy their AI. If different frontier models stumble on different questions, then the smartest move for many businesses may not be to go ‘all in’ on a single provider. For business leaders and executives, the next level of AI transformation might be to broaden your AI portfolio if you haven’t already done so. Leverage cross-model complementarity to route specific tasks, use additional models to audit omissions or conflicting answers, and preserve human review for unresolved disagreement and consequential judgments.
References
[1] Patel, Ajay, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, and Karim R. Lakhani, “Frontier AI Performance Across the Business Disciplines: a Case-Grounded Benchmark of Knowledge Work and Analytical Reasoning,” arXiv preprint arXiv:2607.16057v2 (July 2026): 3.
[2] Patel et al., “Frontier AI Performance Across the Business Disciplines,” 3.
[3] Patel et al., “Frontier AI Performance Across the Business Disciplines,” 6.
Meet the Authors

Ajay Patel is CEO & Co-Founder of Plasticity.

Kartik Hosanagar is the John C. Hower Professor of Technology and Digital Business and a Professor of Marketing at The Wharton School of the University of Pennsylvania.

Ramayya Krishnan is William W. and Ruth F. Cooper Professor of Management Science and Information Systems at Heinz College at Carnegie Mellon University.

Chris Callison-Burch is a Professor of Computer and Information Science at the University of Pennsylvania.

Karim R. Lakhani is the Dorothy & Michael Hintze Professor of Business Administration at Harvard Business School. He specializes in technology management, innovation, digital transformation, and artificial intelligence. He is also the Co-Founder and Faculty Chair of the HBS AI Institute and the Founder and Co-Director of the Laboratory for Innovation Science at Harvard (LISH).
Watch a video version of the Insight Article here.