The Digital Data Design Institute at Harvard is now the Harvard Business School AI Institute.

Out of the Loop?

Flat vector illustration of a multitasking AI robot with multiple arms managing various digital tasks simultaneously. Artificial intelligence for productivity, automation, and workflow management.

The case for rethinking when humans should review AI output

We tend to treat human oversight in AI adoption as the responsible choice. But as organizations increase their reliance on AI, the feasibility of that plan falters: if AI is deployed at scale, there will simply be too many touchpoints to keep a human in the loop on them all. The new paper “Optimal Human AI Coordination in Decision Workflows: Collaboration Paradox and Automation Cliffs,” co-written by HBS AI Institute Associate Michael Lingzhi Li, tackles this challenge head on. For a given workflow, should decisions be made by AI alone, humans alone, or humans with AI assistance? According to this study, the surprising answer is not to collaborate by default.

Key Insight: Workflow Coordination

“The model highlights that coordination decisions must account not only for the comparative accuracy of humans and AI, but also for how human engagement responds to workload and incentives.” [1]

In this theoretical paper, the authors consider a stream of tasks arriving over time and apply game theory and mathematics to model the results. Before those tasks are handled, a coordinator chooses a routing policy: what share of tasks will be AI-only, what share will be human-only, and what share will receive AI-assisted review, where a human may verify, edit, or override AI’s recommendations. The authors assume that assisted review is at least as good as either human-only or AI-only performance, and that assisted review doesn’t take more time than doing a task completely manually. The coordinator wants strong performance, but faces a hard constraint on human involvement. This creates what the authors call a generalized Nash equilibrium. In plain English, that means each side’s best choice depends on the other side’s choice. The coordinator’s routing plan determines the human’s workload, while the human’s review effort (how much time and energy to put into reviewing AI recommendations) determines whether the routing plan can meet a required performance standard. In the paper this performance standard is theoretical, but in the real world it could be a required accuracy rate, quality threshold, safety level, service-level standard, or any other minimum acceptable outcome metric. An equilibrium is the stable outcome where both choices fit together: the coordinator has no better feasible routing policy, and the reviewer has no reason to exert more effort.

Key Insight: Attention is a Strategic Constraint

“Human involvement is capacity constrained and costly.” [2]

In many models, human performance is assumed to be fixed: assign a task to a human, get a human-quality answer. This paper makes a more realistic assumption: people adjust their effort depending on workload, incentives, and the minimum standard they must meet. The authors identify a “collaboration paradox”: the performance of collaboration depends on human engagement, but if AI’s baseline performance meets the required standard, the human may rationally engage minimally, lowering the supposed value-add of collaboration. The second finding is “automation cliffs.” As AI performance improves, organizations might expect a smooth transition: slightly better AI, slightly more automation. When AI is below human performance, many arrangements can be optimal, including human-only work, but once AI crosses key performance thresholds, the equilibrium shifts abruptly toward AI-only operation. The result is that small technical improvements can trigger large reallocations of work. The authors mention that future work should include “empirical validation using data from real-world human-AI workflows.” [3] 

Why This Matters

For business leaders and executives, adding more human-in-the-loop steps may feel like a safer AI strategy, but this research shows why that can be misleading. If you’re stretching your people across too many AI-assisted decisions and workflows, the expected gains and ROI may not materialize. The right strategy might be full automation, or even no AI at all. Collaboration doesn’t need to be the default answer: it’s a design choice, and one that will decide whether your organization will rise above the competition.

Bonus

To move from theory to practice, and to see the complexity of how AI can help with some tasks, but actually hurt with others, check out Back to the Beginnings of AI at Work.

References

[1] Gu, Wei, Michael Li, and Shiziang Zhu, “Optimal Human AI Coordination in Decision Workflows: Collaboration Paradox and Automation Cliffs,” working paper (March 14, 2026): 2-3. https://dx.doi.org/10.2139/ssrn.6417798. [now renamed  “Should Humans Be In the Loop? Human-AI Collaboration Paradox and Automation Cliffs”]

[2] Gu et al., “Optimal Human AI Coordination in Decision Workflows,” 1.

[3] Gu et al., “Optimal Human AI Coordination in Decision Workflows,” 10.

Meet the Authors

Wei Gu is a Postdoctoral Researcher at the Heinz College of Information Systems and Public Policy at Carnegie Mellon University.

Michael Lingzhi Li

Michael Lingzhi Li is Assistant Professor of Business Administration and Mary Ellen Jay and Jeffrey Jay Fellow at Harvard Business School. He is also an Associate at the HBS AI Institute.

Shixiang Zhu

Shixiang Zhu is Assistant Professor of Data Analytics at the Heinz College of Information Systems and Public Policy at Carnegie Mellon University.

Watch a video version of the Insight Article here.

Engage With Us

Join Our Community

Ready to dive deeper with the HBS AI Institute? Subscribe to our newsletter, contribute to the conversation and begin to invent the future for yourself, your business and society as a whole.