A new framework argues that AI evaluation should start where AI actually gets used.
Listen to this article:
Real-world AI use is messy. Uneven user expertise, shifting workflows, institutional dynamics, and a range of other factors all have an effect on what we get out of AI, but we also want to have a clear picture of whether AI investments are driving worker productivity and firm value. The new article “Toward Causal Field Evaluations of AI Systems,” co-written by HBS AI Institute PI Iavor Bojinov, makes the case that the messy real world is exactly the place to evaluate the efficacy of AI. It presents a new framework for thinking more rigorously about how AI systems can be evaluated through causal field evaluations: randomized experiments run in the actual setting where a new system will be used.
Key Insight: The Evaluation Landscape and Its Limits
“A lab evaluation supports credible causal claims but may not generalize to the deployment population.” [1]
The authors theorize that AI evaluations vary along two major dimensions: whether they occur in the lab or in the field, and whether they rely on preference data or outcome data. Lab evaluations, such as benchmarks and adversarial tests, are useful for controlled comparison and model development, but they often strip away the real-life factors that shape performance in practice.
The authors note that results that look decisive in controlled settings may not travel: a model that leads in one evaluation can fall behind in real deployment scenarios, and the size of measured gains often shrinks outside the lab. Field-based approaches are closer to real deployment, but tools like monitoring dashboards and user feedback may not actually establish whether the AI system was the factor responsible for a particular outcome.
The authors also emphasize that preference data can be misleading: people may prefer outputs that look subjectively better or feel easier to use even if they go on to produce worse downstream results. They cite field research at Procter & Gamble showing that generative AI can reshape teamwork and knowledge work itself. That kind of effect cannot be captured by asking whether one isolated response looks better than another. Therefore, for deployment decisions, organizations need to move beyond what users like toward what actually happens through outcome data, while understanding that trade offs will be inevitable. For example, AI speed gains could come with quality costs, while tighter safeguards could reduce flexibility or efficacy for advanced users.
Key Insight: The Case for Causal Field Evaluations
“The most direct way to answer a causal question is to define the experiment you would run if you could, and then assess how far each alternative falls from that ideal.” [2]
Against this backdrop of disconnect between controlled AI evaluations and actual outcomes in the field, the authors present causal field evaluations as the strongest target for assessing AI systems. These evaluations use randomized experiments in the actual deployment environment and measure meaningful outcomes such as learning, productivity, code quality, safety, revenue, or retention. Their strength is that they address two core weaknesses of other methods at once: they take place in the real-world setting where the AI will be used, and randomization helps support credible claims about cause and effect. Even when a full randomized field trial is impractical, the authors argue that defining the ideal experiment clarifies what an organization wants to learn, what assumptions alternatives require, and what evidence is still missing.
Why This Matters
One of the most popular AI integration plans in business today is barely a plan at all: hand everyone an AI chatbot, encourage them to experiment, and hope for the best. For business leaders and executives, this research should make you think twice about that plan, not because experimentation is wrong, but because unstructured experimentation might not be answering the questions you truly care about. When you simply let adoption happen, the people who lean in differ from those who don’t, and the feedback data you collect could be almost entirely preference data. A stronger playbook starts by envisioning what a credible evaluation will look like in your company, and which guardrails and outcomes will tell you whether the tool is actually improving work and value. In a business environment where AI models change quickly, workflows are fragile, and implementation costs are real, that discipline can mean the difference between adopting AI that feels good and implementing AI that actually works.
Bonus
This article highlights the importance of using AI to learn faster, but pairing that speed with disciplined evidence about what actually works. For a related, marketing-focused look at how AI can expand the scope of business experimentation, check out Larger, Faster, Cheaper: The Future of Market Research with AI.
References
[1] Arbour, David, Iavor Bojinov, Avi Feller, and Tu Ni, “Toward Causal Field Evaluations of AI Systems,” Harvard Data Science Review, 8(2) (May 11, 2026): 5.
[2] Arbour et al., “Toward Causal Field Evaluations of AI Systems,” 9.
Meet the Authors

David Arbour is Senior Research Scientist at Adobe.

Iavor Bojinov is James Dinan and Elizabeth Miller Associate Professor of Business Administration at Harvard Business School and co-PI of the HBS AI Institute AI and Data Science Operations Lab hosted within the Laboratory for Innovation Science.

Avi Feller is an associate professor in the Goldman School of Public Policy and the Department of Statistics at UC Berkeley.

is a postdoctoral fellow at the HBS AI Institute.
Watch a video version of the Insight Article here.