Review status
This summary is based on the full article text.
The DOI and source URL have been verified through the official INFORMS article page.
The article is a preregistered field experiment published in Organization Science. It studies generative AI in real product innovation work at Procter & Gamble. The study uses random assignment, expert evaluation of submitted solutions, survey measures of emotional responses, prompt and interaction records, and regression analysis. The findings should be interpreted in the context of one company, one broad industry setting, one-day virtual product development workshops, and the AI models available during the study period.
Research question
How does generative AI affect the core functions traditionally associated with teamwork in knowledge work?
More specifically, the article asks whether generative AI can reproduce or complement three pillars of human collaboration:
- performance enhancement;
- expertise integration across functional boundaries;
- social and emotional engagement during collaborative work.
The article also asks where AI contributes most in the innovation process: idea generation, idea selection, or final solution development.
Hypotheses
Not specified as formal numbered hypotheses.
The study is preregistered and experimentally tests the effects of AI and team configuration on performance quality and expertise integration. Emotional responses and top-tail innovation outcomes are also analyzed as important outcomes, with emotional responses described as a preregistered variable of interest but with unclear expected direction.
Method
The study uses a preregistered 2 x 2 field experiment conducted at Procter & Gamble between May and July 2024.
The setting is early-stage new product development. Participants worked on real product innovation challenges from their own business units rather than hypothetical cases. The tasks involved developing viable ideas for products, packaging, communication, retail execution, or related innovation opportunities.
The experiment involved 826 workshop participants. The main analyses focus on 791 participants who completed the workshop. Of these, 776 provided complete data including post-task surveys. The solution-level performance analyses use 550 observations, because two-person teams submitted one joint solution.
Participants were assigned to four conditions:
- individual without AI;
- two-person human team without AI;
- individual with AI;
- two-person human team with AI.
The human teams paired one commercial professional with one R&D professional. This was designed to mirror P&G’s cross-functional product innovation routines. Participants were randomly assigned within eight randomization clusters defined by four business units and two geographies. The four business units were Baby Care, Feminine Care, Grooming, and Oral Care. The two geographies were Europe and the Americas.
The GenAI tool was built on GPT-4 and accessed through Microsoft Azure. Participants in AI-enabled conditions received a one-hour training session on prompting and using the tool for consumer product goods-related tasks. Participants in the July workshop had access to GPT-4o. The article notes that only 38% of AI-enabled participants used the suggested prompts, and those who did not use the suggested prompts still outperformed control conditions.
The main dependent variable is solution quality, measured on a 1-10 scale by independent expert evaluators with business and technology backgrounds. Evaluators were blind to experimental condition and participant profile. Scores were standardized relative to the individual-without-AI control group. On average, each solution received more than three independent evaluations.
Additional outcomes include:
- novelty, feasibility, impact, and business potential of solutions;
- solution length;
- whether participants expected their solution to rank in the top 10%;
- technical versus commercial orientation of the solution;
- changes in positive emotions, based on enthusiasm, energy, and excitement;
- changes in negative emotions, based on anxiety, frustration, and distress;
- whether submitted solutions ranked in the top 10% of quality scores;
- retention of AI-generated content in final submissions.
The empirical strategy uses regression analysis. The baseline condition is individuals working without AI. The authors estimate models with treatment indicators, business-unit and date fixed effects, and controls including band level, company experience, gender, and prior AI usage at work and for personal purposes. Robust standard errors are used, with additional robustness checks using clustered standard errors and wild cluster bootstrap.
Results / key findings
The study finds that AI improves performance in early-stage innovation work.
In the baseline comparison, two-person teams without AI improved solution quality by 0.245 standard deviations compared with individuals without AI. This supports the traditional view that cross-functional teamwork improves performance. With fixed effects and controls, the estimate is 0.307 standard deviations.
AI produced larger gains. Individuals with AI improved solution quality by 0.373 standard deviations in the baseline model and 0.370 standard deviations in the model with fixed effects and controls. Teams with AI improved solution quality by 0.392 standard deviations in the baseline model and 0.463 standard deviations with fixed effects and controls.
A central result is that individuals using AI performed at a level comparable to teams without AI. The difference between teams with AI and teams without AI was not statistically significant in the main quality models. This suggests that AI can substitute for some performance benefits normally provided by human collaboration, at least in this early-stage product innovation task.
AI also led to much longer and more developed submissions. In the controlled model, individuals with AI produced solutions that were about 504 words longer than the control group, and teams with AI produced solutions that were about 552 words longer. Teams without AI produced only modestly longer solutions, about 57 additional words in the controlled model, with weaker statistical support.
The study finds that AI particularly helped participants less familiar with new product development work. For non-core-job employees, teams without AI did not show a reliable quality improvement in the controlled model. Individuals with AI improved by 0.360 standard deviations. For core-job employees, all three treatment groups improved performance: teams without AI improved by 0.377 standard deviations, individuals with AI by 0.457 standard deviations, and teams with AI by 0.455 standard deviations in the controlled model.
The study also finds that AI helped bridge functional silos. Without AI, commercial professionals tended to produce less technical, more commercially oriented ideas, while R&D professionals tended to produce more technical ideas. With AI, this difference largely disappeared. Both groups produced a more balanced mix of technical and commercial ideas. The article reports that quality did not significantly vary based on whether a solution was more technical or more commercial, suggesting that this balancing effect did not come at the cost of solution quality.
The emotional results show that AI use was associated with more positive self-reported emotional responses. In the baseline model, teams without AI increased positive emotions by 0.269 standard deviations, individuals with AI by 0.457 standard deviations, and teams with AI by 0.635 standard deviations. In the controlled model, the estimates were 0.257, 0.485, and 0.666 standard deviations respectively.
AI use was also associated with lower negative emotions, although the strength of the result varies by specification. In the baseline model, individuals with AI reduced negative emotions by 0.233 standard deviations, and teams with AI reduced negative emotions by 0.235 standard deviations. In the controlled model, the individual-with-AI effect remained statistically significant at -0.263 standard deviations, while the team-with-AI estimate weakened to -0.157 and was not statistically significant.
The study further finds a relationship between emotional experience and expected future AI use. Among AI users, larger increases in expected future AI use were associated with stronger positive emotional changes and lower negative emotional changes. In the controlled models, the coefficient was 0.638 for positive emotions and -0.663 for negative emotions. This evidence is correlational, not causal, but it suggests that positive experiences with AI may support future adoption.
The process analysis shows that AI mainly improved idea generation rather than idea selection. Participants first generated five ideas, selected one, and then developed the selected idea into a detailed solution. AI increased the average quality of generated ideas. However, teams without AI were better at selecting their best idea, choosing the highest-quality idea about 50% of the time, compared with roughly 37% for AI-enabled conditions. This suggests that AI acted more as a quality amplifier than as a decision enhancer.
The article also finds that AI did not simply compress the quality distribution. The gap between the highest- and lowest-quality ideas remained similar across conditions. This means AI raised the overall quality of ideas while preserving variance, which matters because innovation often depends on rare high-quality ideas rather than only average performance.
Top-tail outcomes show a strong result for AI-enabled teams. Teams with AI were 9.2 percentage points more likely to produce solutions in the top 10% of all submissions compared with individuals without AI. The control-group top-10% rate was 5.8%, so this corresponds to roughly three times the chance of producing a top-decile solution. Individuals with AI showed a smaller positive top-10% effect that was not statistically significant.
The study also finds a confidence gap. Although AI improved objective performance, AI-enabled participants were 9.2 percentage points less likely to expect their solution to rank in the top 10% compared with the control group. This suggests that participants using AI may have underestimated the quality of their AI-supported work.
In team collaboration, AI reduced dominance patterns between technical and commercial orientations. Teams without AI showed a more bimodal distribution of solution technicality, with a bimodality coefficient of 0.564. AI-enabled teams showed a more uniform distribution, with a bimodality coefficient of 0.482, while maintaining a similar range of technical content. This suggests that AI helped teams integrate technical and commercial perspectives more evenly.
Finally, the AI-use analysis shows substantial variation in how participants used the tool. Many AI-enabled participants retained more than 75% of AI-generated content in their final submissions, but a nontrivial share retained none. This indicates two different use patterns: some participants relied heavily on AI-generated text, while others used AI mainly for ideation, brainstorming, validation, or refinement.
Practical implications
For managers, the article suggests that generative AI can change how innovation teams are designed.
The most direct implication is that AI can help individuals perform at a level similar to small cross-functional teams in some early-stage innovation tasks. In the study, individuals with AI matched the quality of two-person human teams without AI. This does not mean teams are no longer useful. It means that AI can substitute for some collaborative functions, especially idea generation, broadening perspective, and developing more complete proposals.
The article also suggests that AI can help reduce functional silos. Commercial and R&D professionals normally bring different strengths, but they can also overemphasize their own functional perspective. AI helped both groups produce more balanced solutions. For organizations, this means AI may be useful not only for productivity, but also for knowledge integration across functions.
The findings matter for new product development. P&G’s early-stage innovation process depends on high-quality “seeds” that can later move through the innovation pipeline. AI improved the quality of those early ideas, and teams with AI were more likely to produce top-decile solutions. For firms where a small number of exceptional ideas can create large value, the top-tail effect may be strategically important.
The study also warns managers not to treat AI as a perfect evaluator. AI improved generated idea quality, but AI-enabled participants were worse at selecting their best idea. Teams without AI selected their highest-quality idea more accurately. This suggests that managers should separate AI-supported ideation from human-led evaluation. AI can help generate better raw material, but human judgment remains important for selection.
The emotional findings are also relevant for implementation. Participants using AI reported more enthusiasm, energy, and excitement, and in several specifications reported lower anxiety, frustration, and distress. This matters because technology adoption often fails when workers experience tools as threatening, confusing, or demotivating. In this study, the immediate emotional response to AI was mostly positive.
For managers, useful diagnostic questions include:
- Which tasks currently require teams mainly because individuals lack access to complementary expertise?
- Where can AI help employees cross functional boundaries without removing the need for expert review?
- Which work stages should use AI for idea generation, and which should preserve human judgment for selection?
- Are employees becoming overconfident in AI output, or underconfident in their own AI-supported work?
- Does AI reduce productive disagreement too much, especially in creative work where constructive tension may be useful?
- Should AI be used to improve average productivity, increase the chance of breakthrough ideas, or both?
Theoretical implications
The article contributes to research on teams by showing that AI can reproduce some benefits traditionally associated with human collaboration. In the experiment, individuals using AI achieved quality levels comparable to human teams without AI. This challenges a simple distinction between individual work and team work because AI can provide some functions usually supplied by teammates.
The article also contributes to knowledge management and expertise integration. Teamwork is often justified because different people hold different specialized knowledge. The study shows that AI can help workers reason across functional boundaries, allowing commercial professionals and R&D professionals to produce more balanced solutions. This suggests that AI may become a boundary-spanning mechanism inside organizations.
The article adds to research on AI in organizations by treating AI not only as a tool but as a cybernetic teammate. The term refers to a feedback-based system that can participate in collaborative processes. In this view, AI affects not only task output but also how expertise, motivation, emotional experience, and team boundaries are organized.
The results also contribute to innovation research by decomposing AI’s effect across the innovation process. AI’s main contribution was not better selection accuracy. Instead, AI raised the average quality of generated ideas while preserving variance. This supports the view that AI is especially powerful as a generator and amplifier of creative material, while human judgment remains valuable in evaluative selection.
The article further contributes to debates about exploration and exploitation in AI-supported work. AI-enabled teams were more likely to produce top-decile solutions, suggesting that AI may increase the chance of exceptional outcomes. At the same time, the article notes that AI-assisted solutions may become more semantically similar, raising questions about long-term diversity in organizational innovation.
Limitations
The study was conducted inside one company, Procter & Gamble. The findings may not generalize fully to other firms, industries, organizational cultures, or types of knowledge work.
The industry context is consumer packaged goods and early-stage product innovation. Results may differ in software development, scientific research, consulting, public administration, healthcare, or manufacturing operations.
The experiment used one-day virtual workshops. This setting captures an important part of early-stage innovation work but does not capture long-term team dynamics, repeated collaboration, extended conflict, trust development, rework cycles, or implementation after idea generation.
The teams were small cross-functional pairs, typically one commercial and one R&D professional. Larger teams, established teams, teams with similar expertise, or teams with stronger prior relationships may respond differently to AI.
The study reflects the capabilities of GPT-4 and GPT-4o at the time of the experiment. The effects may change as AI models, interfaces, and organizational AI systems improve.
The emotional outcomes are self-reported and measured around a short task. They capture immediate emotional responses, not long-term psychological effects, worker identity changes, or sustained motivation.
The article shows that AI improved idea generation but weakened selection accuracy in this setting. The study does not fully establish whether this happened because of AI sycophancy, overreliance, lower human effort, weaker internalization of ideas, or other mechanisms.
The study does not observe downstream commercialization outcomes. The evaluated solutions were relevant to P&G’s innovation pipeline, but the article does not track whether AI-supported ideas became successful products in the market.
Future research
Future research could examine how the benefits of AI collaboration evolve as employees become more experienced with prompting and AI-supported workflows.
Researchers could test whether AI has similar effects in larger teams, established teams, face-to-face teams, or teams with different functional compositions.
Future studies could investigate how organizations should divide labor between AI-supported idea generation and human-led idea selection.
Another important research direction is whether AI increases breakthrough innovation across different industries and task types, or whether the top-tail effect depends on product innovation settings.
Researchers could study whether AI-enabled boundary spanning leads to real knowledge development over time or mainly provides temporary access to expertise.
Future research could examine whether positive early emotional experiences with AI create self-reinforcing adoption cycles, and whether those cycles improve or weaken long-term learning.
A further research direction is whether AI-supported work increases semantic similarity across ideas over time, potentially improving average quality while reducing diversity in organizational search.
Finally, future studies could run longer organizational field experiments that track AI-supported ideas beyond early ideation into evaluation, resource allocation, implementation, and market outcomes.