This series introduces Harvard Business School AI Institute Associates Program projects which aim to answer important questions at the intersection of artificial intelligence and digital technologies in business and society.
This article shares insights from Hannah. K. Galvin, Assistant Professor of Pediatrics, Harvard Medical School, Chief Health Information Officer, Cambridge Health Alliance who is pursuing research on the topics of artificial intelligence and organizations.
1. What drew you to this area of research and how did you first become involved in this work?
As both a practicing physician and Chief Health Information Officer at a Harvard-affiliated safety net health system, I noticed that clinicians were rapidly adopting large language models to answer clinical questions in real time, yet health systems had limited evidence about whether these tools actually improve clinical decision-making or introduce new patient safety risks.
Traditionally, new clinical technologies undergo extensive evaluation before incorporation into patient care. Generative AI has largely skipped that process. Physicians are already using these tools at scale, often outside formal organizational oversight, while the evidence base has lagged far behind.
That realization motivated our team to design what is now one of the first prospective studies examining how clinicians use AI during actual patient care rather than on board-style examination questions or simulated clinical vignettes. More broadly, my research focuses on helping health systems move beyond enthusiasm or skepticism toward evidence-based adoption of AI that improves care while protecting patients.
2. What are some common misconceptions or barriers around the problem you’re working to solve?
Real world medicine is fundamentally different. Clinicians manage uncertainty, incomplete information, competing diagnoses, social determinants of health, and continuously evolving evidence. Those complexities are difficult to reproduce in standardized benchmarks.
Another misconception is that the question we should be focused on is whether AI should replace physicians. In reality, I think the more important question at this time is how AI changes physician decision-making. Even highly-accurate systems can influence clinical reasoning in subtle ways that can affect patient care and outcomes.
A significant barrier is the lack of reproducible evaluation methods. Every health system is currently being asked to decide whether to deploy LLM’s for clinical decision making, yet few have a standardized framework for assessing and governing these types of AI tools specifically. We hope this study helps to start address this gap.
3. What research is being done on this topic and how is your approach or perspective unique?
Although several recent studies have become evaluating LLM’s in clinical workflows, much of the evidence in this space still comes from benchmarks, curated cases, and controlled experiments. Those studies have been invaluable for measuring technical capability, but relatively little is known about how clinicians independently use these tools during routine patient care and how that use affects clinical reasoning and decision-making.
Our study instead follows residents during routine clinical practice across internal medicine, family medicine, and psychiatry. Every AI query is compared against established reference resources such as UpToDate, society guidelines, or PubMed before influencing patient care, and specialty-matched attending physicians independently evaluate the clinical appropriateness of the resulting decisions. We then compare identical real-world clinical prompts across multiple publicly available large language models, including OpenEvidence, ChatGPT, Claude, Gemini, and Microsoft Copilot.
We are also developing a reproducible evaluation framework that other health systems can use when assessing emerging AI tools. Rather than evaluating a single product, we hope to establish a methodology for responsible AI adoption across healthcare.
4. What excites you most about this work and its potential impact?
Healthcare is at an inflection point. AI has the potential to improve quality, reduce cognitive burden, expand access to expertise, and accelerate evidence-based medicine. But those benefits will only be realized if health systems understand how these tools perform, where they add value, and under what circumstances they can be used safely.
What excites me most is helping move the conversation from opinions to evidence. Instead of asking whether AI is “good” or “bad,” we can begin asking much more useful questions: Which systems perform best? For what clinical tasks? Under what conditions? Where do safety concerns emerge? How should organizations monitor these tools after deployment?
If we can answer those questions rigorously, we can accelerate responsible adoption while improving patient safety.
5. How do you hope working with the HBS AI Institute will amplify the impact of your work?
The challenges surrounding clinical AI extend far beyond medicine. They involve organizational decision-making, governance, implementation, trust, incentives, regulation, and responsible innovation.
The HBS AI Institute brings together scholars across these disciplines. Collaborating with researchers studying strategy, operations, organizational behavior, and technology adoption creates opportunities to translate findings from healthcare into broader principles for evaluating and governing AI in high-risk industries.
I hope the Institute will help connect rigorous clinical research with other organizational and business leaders responsible for implementing AI at scale. Generating evidence is only the first step. Ensuring that evidence informs organizational decisions is equally important.
6. What changes do you hope to see in your field as a result of the work being done in this area?
I hope the field moves beyond relying primarily on benchmark performance and controlled evaluations as we develop a much deeper understanding of how large language models influence care in real clinical environments. The critical question is no longer simply whether an AI system can produce the right answer, but how its recommendations interact with clinician judgment, workflow, patient context, and the realities of clinical practice.
As the evidence base grows, I hope health systems will be able to make more informed, task-specific decisions about where AI meaningfully improves care, where additional safeguards are needed, and where current systems are not yet ready for use. I also hope studies like ours help establish reproducible methods for evaluating AI within local clinical environments, so organizations can generate their own evidence rather than relying solely on vendor claims or generic benchmark performance.
Ultimately, I would like clinical AI evaluation to become a routine part of health-system quality and safety infrastructure, not a one-time exercise performed before deployment.
7. What’s an essential area in which AI and digital technologies will reshape the way business or society operate in the long run that we may not be considering?
Much of the conversation focuses on improving AI models themselves. I think the larger long-term challenge will be learning how to govern people working alongside increasingly capable AI systems.
Organizations will need to understand not only whether an AI system is accurate, but how it changes human judgment, confidence, decision-making, and behavior. Two systems with similar technical performance can produce very different outcomes depending on how people interpret, trust, challenge, or defer to their recommendations.
That means competitive advantage will increasingly come from building evidence-based governance and workflow systems that continuously evaluate how humans and AI perform together, rather than evaluating either in isolation. Understanding and managing that interaction may ultimately matter more than incremental gains in model performance.
The Harvard Business School AI Institute Associates Program supports and accelerates faculty research into the ways AI and digital technologies are reshaping companies, organizations, society, and practice.