Feed 0% source
AI/ML AI-generated

Human diversity fuels collective creativity that large language models cannot simulate or sustain

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

The Hidden Cost of AI Assistance

Diverse human groups produce diverse ideas. These ideas serve as the raw material for innovation. Researchers found that while diverse groups of people come up with more varied and creative ideas than AI can mimic, the way we use AI might accidentally destroy that advantage. While using AI to generate ideas can make everyone's work sound too similar, using AI only to polish human-made ideas helps keep that creativity alive. Even when scientists tried to make AI act like diverse people, the AI still could not match the richness of real human thought without becoming nonsensical.

Protecting the Engine of Innovation

The core tension in this research lies in the distinction between human diversity and collective diversity. Human diversity refers to the varied backgrounds, values, and cognitive repertoires of individuals. Collective diversity is the resulting variety of ideas produced by a group. Much like a biological ecosystem relies on many species to maintain resilience, creative industries rely on many perspectives. This ensures a broad "solution space"—the total range of possible ideas available for a given problem.

The authors argue that generative AI poses a dual threat. First, everyday AI assistance might act as a homogenizer. It pulls different people's unique expressions toward a single, common statistical mean. Second, there is the risk of "silicon sampling." This is the attempt to replace diverse human contributors with simulated personas. Developers often assume that AI can replicate the variance of a real population.

The Mechanics of Creative Erosion

To investigate these threats, the authors conducted a creative metaphor experiment involving 644 writers. The study crossed two linguistic backgrounds—native English (L1) and non-native English (L2)—with three AI workflows. There was a human-only control. There was also "AI ideation" (where the AI suggests the initial ideas). Finally, there was "AI refinement" (where the human generates the idea and the AI merely polishes the explanation).

The point at which AI enters the workflow is decisive. In the AI ideation condition, collective diversity was compressed for everyone. Crucially, this process erased the natural advantage held by L2 writers. These writers typically contribute more diverse ideas than L1 writers [Figure 2A]. However, when using the AI refinement workflow, both the diversity of the ideas and the L2 advantage were largely preserved [Figure 2A].

The study then tested if AI could simulate this diversity through "persona engineering." This involves creating digital identities based on real participants' backgrounds. The researchers deployed three model families (GPT-4o mini, Claude Sonnet 4.6, and Qwen3.5-27b). They tried to boost variety using high "sampling temperatures." Temperature is a parameter that increases randomness in model output. Despite these efforts, every simulated pool fell below every human pool in diversity [Figure 3A]. Furthermore, the authors found a "diversity-fluency frontier." Attempts to force models to be more diverse through high temperature resulted in "degenerate text." This means the models produced gibberish or lost grammatical coherence [Figure 3D].

Identifying the Incentive Gap

One striking implication is the misalignment between individual rewards and collective health. The authors found that while AI ideation harms the diversity of the overall pool, it actually boosts individual creativity ratings [Figure 4A, 4B]. This creates a dangerous incentive. A writer is personally rewarded for using a tool that contributes to "groupthink" and shrinks the total pool of ideas.

The researchers report a "floor-raising" effect. This improves the average quality while narrowing the range of ideas. This effect was evident in how AI ideation affected the distribution of creative scores. It helped lower-performing writers reach a baseline of competence. However, it failed to lift the "ceiling" of the most original thinkers. Conversely, the AI refinement workflow offered a way to align these interests. This was especially true for L2 writers who used their native language to ideate before translating to English. This "translate-late" approach allowed them to retain unique conceptual insights [Figure 4C, 4D].

Limits of the Framework

While the findings are robust, the authors note several boundaries. The study focused exclusively on the production of short metaphors. It is unclear if these patterns hold for long-form scientific or visual innovation. Additionally, the researchers operationalized diversity specifically through linguistic backgrounds. Other forms of diversity, such as gender or ethnicity, may interact with AI differently. Finally, the study looked at independent production rather than team dynamics. It does not capture how humans might inspire or conform to one another in a live setting.

Designing for Survival

The research suggests that organizations must change how they integrate these tools. The goal should be to keep humans "upstream." This means humans should be responsible for the initial spark of ideation. AI should be relegated to "downstream" tasks like stylistic refinement. By protecting the moment of original thought, we ensure that the unique cognitive variety of a diverse workforce is not traded away for mere efficiency.

Figures from the paper

Figure 1
Fig. 1. Study design. (A) Native (L1) and non-native (L2) English writers completed four creative metaphor tasks either without AI, with AI-generated ideas (AI ideation), or with AI polishing their own ideas (AI refinement); L2 writers chose their preferred language of AI interaction. (B) The same tasks were simulated by LLM personas built from the real participants' backgrounds, with simulated diversity strengthened through model choice, native-language prompting, and higher sampling temperature. The resulting human and simulated metaphor pools were compared on collective diversity.
Figure 2
Fig. 2. Collective diversity of the metaphor embeddings. For each metaphor, we computed its cosine similarity to the centroid of all other metaphors in the same condition and scenario; higher mean similarity indicates less collective diversity. (A) Mean centroid similarity (± SEM) across the three conditions (human-only, AI ideation, AI refinement) for native (L1) and non-native (L2) English writers. (B) The same, with L2 split by language of AI interaction, L2 (EN) vs. L2 (Non-EN), the latter present only in the AI conditions. (C, D) Kernel density of the per-metaphor centroid similarity (further right = more similar = less diverse), by condition, for (C) L1 and (D) L2 writers.
Figure 3
Figure 3 — from the original paper
Figure 4
Fig 4. Individual creativity and its trade-off with collective diversity. An independent panel of evaluators (N = 351) rated each metaphor on four items, blind to condition; novelty and surprise form the divergent subscale, and aptness and explanation quality form the convergent subscale. (A, B) Mean divergent (A) and convergent (B) rating (± SEM) for native (L1) and non-native (L2) writers. The asterisks indicate the significance level of paired contrasts with the human-only condition within each language group (* p < .05, ** p < .01, *** p < .001). (C, D) Incentive landscape for L2 writers: individual divergent (C) and convergent (D) rating against the collective diversity of the corresponding pool (x-axis, mean
Figure 5
Figure 5 — from the original paper
Figure 6
Fig. S3. Subgroup decomposition of the collective diversity in Fig. 3B by persona background.
Novelty
0.0/10
Impact
0.0/10
Overall
0.0/10
#ai#nlp#creativity#human-computer interaction#diversity
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 17 / 17

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 99,428
Wall-time: 213.4s
Tokens/s: 465.8

Next up

AI Guardrails Shift Healthcare Decisions: Field Experiment Reveals Directiona...

8.3/10· 6 min