The promise of AI creativity depends on knowing where human judgment still breaks down.
Listen to this article:
People love using AI as a brainstorming partner. The experience feels collaborative, fast, and often confidence-building. It’s tempting to assume that once AI enters the room, great ideas are near at hand. But the new paper, “Innovating with Generative AI: A Human Bottleneck Framework,” co-written by HBS AI Institute 2025 and 2026 Associate Julian De Freitas and HBS AI Institute Customer Intelligence Lab co-founder Ayelet Israeli, argues that this assumption overlooks a real constraint. AI can accelerate parts of the innovation process, but it also remains shaped and limited by distinct human bottlenecks: the psychological and behavioral mechanisms that limit how we imagine, notice, articulate, evaluate, and act on new ideas. Tracing this logic across the four traditional stages of innovation—ideation, screening and testing, preference measurement and consumer insight, and diffusion and market learning—the authors explore whether AI interventions at each stage alleviate these bottlenecks, shift them elsewhere, or make them worse.
Key Insight: The Anchoring Trap
“Once a line of thinking is established in a human mind, explicit instructions to diversity are largely ineffective.” [1]
The first stage of innovation is ideation, generating new concepts that are both novel and useful. One of the human bottlenecks at this stage is cognitive fixation. Ideation depends on escaping familiar mental models, but people often reach for ideas that are already mentally available. Expertise can actually intensify the problem: deep knowledge helps people solve known problems efficiently, but it can also route attention to well-worn solutions. LLMs risk making this worse: left unguided, they tend to generate typical ideas clustered around familiar patterns, and once humans see those outputs, their own thinking can become even further entrenched. The authors’ intervention, then, is to structure the AI system through Chain-of-Thought (CoT) Prompting: first asking it to sketch a batch of initial ideas, then explicitly instructing it to revise or extend those into bolder, more distinct ones before finalizing an answer. Because the model is optimized to follow that kind of meta-instruction, prior research cited by the authors suggests that CoT prompting measurably increases the idea diversity of AI outputs.
Key Insight: More Ideas Can Mean Worse Judgment
“Screening is therefore an active force shaping which innovations reach the market, and its failures are, by construction, invisible” [2]
The second stage is screening and testing: deciding which ideas to advance, refine, fund, or abandon. Because AI can dramatically increase idea volume, the authors focus on the bottleneck of cognitive load: human evaluators have limited attention, and high load makes people more likely to rely on shortcuts that lead to poorer decisions. Under high load, for instance, evaluators lean on surface cues like fluency and polish as a stand-in for quality. AI-generated ideas often arrive in smooth, well-structured language, which can make them feel higher quality than they actually are. As an intervention, the authors recommend combining LLM ratings with historical human expert ratings to decide which ideas need careful review. To illustrate the potential of this approach, the authors cite a separate study of 153 ideation contests and more than 74,000 submissions. In that setting, a screening architecture that combined LLM-generated ratings with historical expert ratings reduced human evaluation effort while preserving alignment with sponsor choices. The system was not simply replacing human judgment with AI judgment; it learned which patterns in prior human evaluations best predicted sponsor preferences, then used that signal to reduce the number of submissions requiring deeper human review.
Key Insight: Some Preferences Have No Words Yet
“[R]adical preference research requires identifying preferences that consumers themselves do not yet have language for.” [3]
The third stage is preference measurement and consumer insight: after ideas become more concrete concepts, firms need to understand what customers want, value, and might adopt. One bottleneck here is the articulability gap: consumers cannot always articulate preferences for products or experiences they have never encountered. Before people experienced touchscreen swipe-based navigation, in the authors’ example, they were unlikely to request it in a survey or focus group. Because AI language models learn from text that people have written, they face a structural limitation: they can only surface preferences that someone has already put into words somewhere. A want nobody has ever expressed simply isn’t in the data for the model to find. Even though AI can’t seem to solve this problem, the authors emphasize that it can help boost creativity in human workers and translate confusing, pre-existing customer feedback into actionable insights.
Key Insight: AI Solves Volume, Not Priority
“The cost and latency of market listening have collapsed.” [4]
The fourth stage is diffusion and market learning. Once a product launches, firms must understand how it spreads through a market and what post-launch signals reveal about the next round of innovation. The authors include the human bottleneck of attention and aggregation because modern products generate more feedback than human teams can absorb. The authors draw on prior work showing that this flood of text can contain rich market insights, while also noting that it can be biased: reviews skew toward extreme opinions, support tickets skew toward failures, and social media skews toward the most engaged and vocal customers. As an intervention, AI solves half of this problem: it can now extract and summarize at a level of quality that once required professional analysts and at a fraction of the time and cost. But the authors are careful to note that the hard question becomes deciding which of those signals is actually worth acting on next.
Why This Matters
Used well, AI can accelerate innovation. Used carelessly, it can create a false sense of progress. The implementation lesson is to build awareness of these tradeoffs into every AI use case. For business leaders and executives, getting this diagnosis right is where competitive advantage will actually accumulate. In their full paper, De Freitas, Israeli, and their co-authors explore 11 different human bottlenecks and AI interventions targeted towards them. For anyone involved in shaping AI strategy and deployment, every one is worth a deep read.
References
[1] De Freitas, Julian, Ayelet Israeli, Gideon Nave, Artem Timoshenko, and Olivier Toubia, “Innovating with Generative AI: A Human Bottleneck Framework,” Harvard Business School Working Paper, No. 26-094 (June 2026): 9
[2] De Freitas et al., “Innovating with Generative AI,” 14-15.
[3] De Freitas et al., “Innovating with Generative AI,” 25.
[4] De Freitas et al., “Innovating with Generative AI,” 30.
Meet the Authors

Julian De Freitas is an Assistant Professor of Business Administration in the Marketing Unit and Director of the Ethical Intelligence Lab at Harvard Business School, and Associate HBS AI Institute. His work sits at the nexus of AI, consumer psychology, and ethics.

Ayelet Israeli is a Principal Economist in the Store Economics and Science Team at Amazon. She was previously Marvin Bower Associate Professor of Business Administration at Harvard Business School, and co-founder of the Customer Intelligence Lab at the HBS AI Institute.

Gideon Nave is Carlos and Rosa de la Cruz Associate Professor at The Wharton School of the University of Pennsylvania.

Artem Timoshenko is Associate Professor of Marketing at the Northwestern University Kellogg School of Management.

Olivier Toubia is Claubinger Professor of Business and Senior VIce Dean for Curriculum and Instruction at Columbia Business School.