The Digital Data Design Institute at Harvard is now the Harvard Business School AI Institute.

Think You’re a Good Leader? Try Leading an AI Team

Think You’re a Good Leader? Try Leading an AI Team

Researchers found that an AI-based assessment closely tracked leaders’ impact on human teams.

Judging a leader with a successful team is like judging the conductor of a great orchestra: the outcome reflects both the person directing the group and the abilities of everyone around them. Measuring leadership therefore requires a way to separate the individual’s influence from the circumstances and people surrounding them. Artificial intelligence may offer a new way to approach that problem. The paper “Measuring Human Leadership Skills with AI Agents,” co-written by 2026 HBS AI Institute Associate David Deming, investigates whether performance with AI agents can provide a practical measure of a person’s ability to lead human teams in a comparable experimental task. 

Key Insight: One Test, Two Kinds of Teammates

“While many studies show the value of good leadership, we know relatively little about how to measure individual differences in leadership skills” [1]

The researchers designed a pre-registered laboratory experiment and recruited 249 people to act as leaders. The leaders were given two separate versions of the test. In one version, they led three human followers; in the other, they led three AI agents built on GPT-4o. Roughly half the leaders did the human version first, while the other half started with AI. The leaders worked with six different human teams to separate their own contribution from the quality of any particular group (and they similarly completed six different puzzles with AI teammates). The task itself was a modified ‘Hidden Profile’ problem, where the information required to solve each puzzle was divided among the participants. Some clues eliminated possible answers, while others were irrelevant distractors. Leaders had to question followers, uncover private information, manage limited time, and combine the evidence into a final judgment. Every puzzle was written from scratch in order to match difficulty across the two versions of the test and to reduce the chance that the model had encountered the answers in its training data. Before leading teams, participants also completed assessments of typing speed, task-specific puzzle skill, fluid intelligence, emotional perceptiveness, and economic decision-making. These measures allowed the researchers to distinguish leadership-related capabilities from more technical advantages.

Key Insight: The Leader, Not the Luck of the Draw

“More than half the variation in group performance on both tests can be explained solely by the identity of the leader.” [2]

Leaders who performed well with AI followers generally also performed well with human followers. The directly observed correlation between overall performance in the two assessments was 0.67. After statistically correcting for the fact that scores based on a limited number of puzzles contain measurement noise, the researchers estimated the relationship at 0.81. The size of the leadership effect itself was also striking. Swapping an average leader for one who’s a standard deviation above average (a ‘good’ leader) boosted team performance by about 0.65 standard deviations. In other words: a good leader correctly solved 53% of the problems, while an under-performing one solved only 10%. And critically, the same underlying leadership skills predicted success in both the AI and human versions of the test.

Key Insight: Successful Leaders Ask, Listen, and Reassess

“[G]ood leaders ask more questions and their teams engage in more conversational turn-taking.” [3]

If you assume the best leaders are the ones who talk the most, the data politely disagrees. What separated strong leaders was how they communicated: they asked more questions, generated more back-and-forth turn-taking, and they leaned on plural pronouns (‘we’ and ‘us’ rather than ‘I’), echoing earlier findings on inclusive leadership. The low-performing leaders looked very different, with final answers similar to uninformed guesses, suggesting that they had learned little from conversations with their teammates. Emotion mattered more with human followers: a leader’s use of positive, warm language reliably helped human teams but had little effect on the AI ones. Even so, emotional perceptiveness still predicted performance on the AI test after controlling for intelligence. The authors interpret this as evidence that the assessment captures interpersonal capabilities in addition to cognitive skill.

Why This Matters

For business leaders and executives, a promising implication of this research is that AI could become a practice environment for leadership itself. Managers can rehearse questioning, delegation, information gathering, and decision-making with AI teams, further enhanced by employing a range of personas. Instead of treating leadership development as an occasional workshop, leaders and organizations can build short, repeatable simulations around familiar challenges. The resulting practice could help leaders see whether they ask for the right information, create room for different perspectives, and turn a fragmented discussion into a clear course of action. AI simulations like these may be particularly useful for practicing the mechanics of coordination, but leaders should also ensure that their development addresses the trust, morale, and emotional judgment that shape real human teams. 

Bonus

Practicing leadership with AI is one step, being prepared to lead an AI-first organization is another. To learn how leaders can strengthen their organizations’ ability to sense change, rewire resources, and embed AI lessons into everyday operations, check out Mastering Change Resilience: The Key to AI-Driven Success

References

[1] Weidmann, Ben, Yixian Xu, and David J. Deming, “Measuring Human Leadership Skills with AI Agents,” NBER Working Paper 33662 (2025): 2. https://doi.org/10.3386/w33662 

[2] Weidmann, Xu, and Deming, “Measuring Human Leadership Skills with AI Agents,” 5.

[3] Weidmann, Xu, and Deming, “Measuring Human Leadership Skills with AI Agents,” 8.

Meet the Authors

Ben Weidmann is Director of Research at the Harvard Skills Lab.

Yixian Yu is a Research Fellow at the Harvard Skills Lab. 

David Deming

David J. Deming is the Danoff Dean of Harvard College, the William Henry Bloomberg Professor of Economics, the Isabelle and Scott Black Professor of Public Policy at Harvard Kennedy School, and 2026 Associate at the HBS AI Institute.

Engage With Us

Join Our Community

Ready to dive deeper with the HBS AI Institute? Subscribe to our newsletter, contribute to the conversation and begin to invent the future for yourself, your business and society as a whole.