How a new benchmark uses business schools’ signature teaching tool to measure AI against the analytical work careers are built on.
Most of us are already comfortable handing AI short, one-off tasks like document summaries, email drafts, and data analysis, but we are less comfortable trusting its results on more complex work. Can AI stay with a complicated business problem long enough to interpret messy evidence, weigh competing priorities, exercise judgment under uncertainty, and arrive at a recommendation that can be defended? Understanding AI’s ability to complete that level of work has largely been elusive because conventional benchmarks tend to reward quick factual recall, narrow question answering, mathematics, coding, or tool use. But in “Frontier AI Performance Across the Business Disciplines: a Case-Grounded Benchmark of Knowledge Work and Analytical Reasoning,” co-authored by chair and co-founder of the HBS AI Institute Karim R. Lakhani, a team of researchers tackled this challenge with BusinessCaseBench, a tool that repurposes the business school case method to test frontier AI models on genuine analytical knowledge work. What they found reframes the AI performance debate: today’s models are remarkably capable, but an important gap remains.
Why This Matters
The researchers also devised a “cross-model oracle” that selected the highest score achieved by any of the three frontier models on each question. It reached 93%, well above any single model alone, and reveals an important pattern: what one model misses, another often captures. That finding has real implications for how organizations should actually deploy their AI. If different frontier models stumble on different questions, then the smartest move for many businesses may not be to go ‘all in’ on a single provider. For business leaders and executives, the next level of AI transformation might be to broaden your AI portfolio if you haven’t already done so. Leverage cross-model complementarity to route specific tasks, use additional models to audit omissions or conflicting answers, and preserve human review for unresolved disagreement and consequential judgments.
Link to the HBS AI Institute Insight Article
Link to the Research Paper
Sign up for our newsletter to stay up to date with HBS AI Institute news and research