A new framework argues that AI evaluation should start where AI actually gets used.
Real-world AI use is messy. Uneven user expertise, shifting workflows, institutional dynamics, and a range of other factors all have an effect on what we get out of AI, but we also want to have a clear picture of whether AI investments are driving worker productivity and firm value. The new article “Toward Causal Field Evaluations of AI Systems,” co-written by HBS AI Institute PI Iavor Bojinov, makes the case that the messy real world is exactly the place to evaluate the efficacy of AI. It presents a new framework for thinking more rigorously about how AI systems can be evaluated through causal field evaluations: randomized experiments run in the actual setting where a new system will be used.
Why This Matters
One of the most popular AI integration plans in business today is barely a plan at all: hand everyone an AI chatbot, encourage them to experiment, and hope for the best. For business leaders and executives, this research should make you think twice about that plan, not because experimentation is wrong, but because unstructured experimentation might not be answering the questions you truly care about. When you simply let adoption happen, the people who lean in differ from those who don’t, and the feedback data you collect could be almost entirely preference data. A stronger playbook starts by envisioning what a credible evaluation will look like in your company, and which guardrails and outcomes will tell you whether the tool is actually improving work and value. In a business environment where AI models change quickly, workflows are fragile, and implementation costs are real, that discipline can mean the difference between adopting AI that feels good and implementing AI that actually works.
Link to the HBS AI Institute Insight Article
Link to the Research Paper
Sign up for our newsletter to stay up to date with HBS AI Institute news and research