The Hidden Cost of AI Assistance
Diverse human groups produce diverse ideas. These ideas serve as the raw material for innovation. Researchers found that while diverse groups of people come up with more varied and creative ideas than AI can mimic, the way we use AI might accidentally destroy that advantage. While using AI to generate ideas can make everyone's work sound too similar, using AI only to polish human-made ideas helps keep that creativity alive. Even when scientists tried to make AI act like diverse people, the AI still could not match the richness of real human thought without becoming nonsensical.
Protecting the Engine of Innovation
The core tension in this research lies in the distinction between human diversity and collective diversity. Human diversity refers to the varied backgrounds, values, and cognitive repertoires of individuals. Collective diversity is the resulting variety of ideas produced by a group. Much like a biological ecosystem relies on many species to maintain resilience, creative industries rely on many perspectives. This ensures a broad "solution space"—the total range of possible ideas available for a given problem.
The authors argue that generative AI poses a dual threat. First, everyday AI assistance might act as a homogenizer. It pulls different people's unique expressions toward a single, common statistical mean. Second, there is the risk of "silicon sampling." This is the attempt to replace diverse human contributors with simulated personas. Developers often assume that AI can replicate the variance of a real population.
The Mechanics of Creative Erosion
To investigate these threats, the authors conducted a creative metaphor experiment involving 644 writers. The study crossed two linguistic backgrounds—native English (L1) and non-native English (L2)—with three AI workflows. There was a human-only control. There was also "AI ideation" (where the AI suggests the initial ideas). Finally, there was "AI refinement" (where the human generates the idea and the AI merely polishes the explanation).
The point at which AI enters the workflow is decisive. In the AI ideation condition, collective diversity was compressed for everyone. Crucially, this process erased the natural advantage held by L2 writers. These writers typically contribute more diverse ideas than L1 writers [Figure 2A]. However, when using the AI refinement workflow, both the diversity of the ideas and the L2 advantage were largely preserved [Figure 2A].
The study then tested if AI could simulate this diversity through "persona engineering." This involves creating digital identities based on real participants' backgrounds. The researchers deployed three model families (GPT-4o mini, Claude Sonnet 4.6, and Qwen3.5-27b). They tried to boost variety using high "sampling temperatures." Temperature is a parameter that increases randomness in model output. Despite these efforts, every simulated pool fell below every human pool in diversity [Figure 3A]. Furthermore, the authors found a "diversity-fluency frontier." Attempts to force models to be more diverse through high temperature resulted in "degenerate text." This means the models produced gibberish or lost grammatical coherence [Figure 3D].
Identifying the Incentive Gap
One striking implication is the misalignment between individual rewards and collective health. The authors found that while AI ideation harms the diversity of the overall pool, it actually boosts individual creativity ratings [Figure 4A, 4B]. This creates a dangerous incentive. A writer is personally rewarded for using a tool that contributes to "groupthink" and shrinks the total pool of ideas.
The researchers report a "floor-raising" effect. This improves the average quality while narrowing the range of ideas. This effect was evident in how AI ideation affected the distribution of creative scores. It helped lower-performing writers reach a baseline of competence. However, it failed to lift the "ceiling" of the most original thinkers. Conversely, the AI refinement workflow offered a way to align these interests. This was especially true for L2 writers who used their native language to ideate before translating to English. This "translate-late" approach allowed them to retain unique conceptual insights [Figure 4C, 4D].
Limits of the Framework
While the findings are robust, the authors note several boundaries. The study focused exclusively on the production of short metaphors. It is unclear if these patterns hold for long-form scientific or visual innovation. Additionally, the researchers operationalized diversity specifically through linguistic backgrounds. Other forms of diversity, such as gender or ethnicity, may interact with AI differently. Finally, the study looked at independent production rather than team dynamics. It does not capture how humans might inspire or conform to one another in a live setting.
Designing for Survival
The research suggests that organizations must change how they integrate these tools. The goal should be to keep humans "upstream." This means humans should be responsible for the initial spark of ideation. AI should be relegated to "downstream" tasks like stylistic refinement. By protecting the moment of original thought, we ensure that the unique cognitive variety of a diverse workforce is not traded away for mere efficiency.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 17 / 17
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 99,428
Wall-time: 213.4s
Tokens/s: 465.8