Feed 0% source
Social science AI-generated

The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

The Algorithmic Flattening of Sound

Researchers studied how AI music tools like Suno and Lyria 3 change the variety of music they create. They found that one system makes music within a genre sound too similar. Another system makes different genres sound too much alike. This could potentially narrow the diversity of what we hear.

Two Paths to Acoustic Sameness

The core question is whether generative music systems—AI capable of creating complete audio tracks from text—introduce a systematic tendency toward "homogenization." In this context, homogenization refers to a reduction in acoustic variation. Think of it like a spice factory. If the machines produce every dish with the exact same level of salt, the unique flavor of individual cuisines disappears.

The authors of this study, Zoe Slendebroek and Dana´e Metaxa, investigate whether these systems produce music that is flatter than human-produced music. They define this flattening through three lenses: rhythm and timing, timbre (the characteristic color of a sound), and dynamics (the variation in loudness). The goal is to see if AI is fundamentally reshaping the structural diversity of musical genres.

The Background You Need

To audit these systems, the researchers rely on Music Information Retrieval (MIR). MIR is a field that uses computational methods to extract descriptive features from audio. These include tempo, spectral shape, or harmonic content. Instead of using subjective human labels, MIR allows for a mathematical comparison of sound.

The study situates this problem within a history of musical standardization. Historically, homogenization occurred "downstream." This means radio consolidation or streaming algorithms selected a narrow range of popular music. However, the authors argue that generative AI moves this process "upstream." Rather than just filtering existing music, these models participate in the actual production of new material.

Because these models are trained to minimize "loss" (a mathematical penalty for deviating from training data), they gravitate toward frequent patterns. The authors note a critical tension here. Many training datasets are heavily skewed toward Western musical traditions. Some estimates suggest up to 94% of data is Western. Consequently, models may treat Western aesthetic norms as a universal default. This can render other musical traditions "illegible" (difficult for the system to recognize or represent accurately) to the algorithm.

How The Argument Works

The researchers conducted a black-box audit of two commercial systems, Suno and Lyria 3. They tested four genres: Afrobeats, K-pop, Dance Pop, and Heavy Metal. They used two experimental designs to separate system tendencies from user instructions. In Experiment 1, they used "MIR-steered prompting." They wrote prompts based on the actual acoustic features of human tracks. In Experiment 2, they used "null prompting." They provided only the genre name to see what the models produced by default.

The study first examines "prompt fidelity"—how well the AI follows instructions. The authors report that Suno shows near-zero fidelity. This means it largely ignores specific acoustic requests in a prompt .

Figure 1
Figure 1: Prompt fidelity: Pearson r between encoded target feature value and AI output value (Experiment 1). Suno tracks prompts only weakly ( | r | < 0 . 26 across all features and genres). Lyria shows moderate tempo tracking ( r = 0 . 30 -0 . 57 ) but limited fidelity elsewhere.

Lyria 3 shows moderate success in tracking tempo. However, it fails to follow instructions regarding harmony or texture . This finding is crucial. It suggests that the observed patterns of sameness are "learned priors"—the model's built-in biases.

The audit reveals two distinct ways these systems flatten music. The authors report that Lyria 3 reduces within-genre diversity. Its outputs cluster more tightly around a genre's center than human tracks do .

Figure 2
Figure 2: Within-genre acoustic diversity by system and genre (RQ1). Bars show the AI/Human ratio of mean pairwise distances within each genre (values > 1 = AI more diverse than human; < 1 = AI more homogeneous). Lyria consistently falls at or below the human baseline (blue dashed line); Suno exceeds it in three of four genres, with Afrobeats showing the largest deviation ( 1 . 58 × ).

Conversely, Suno does not compress within-genre variety. Instead, it "collapses" the boundaries between genres. This makes different musical styles sound acoustically indistinguishable .

Figure 3
Figure 3: Genre separation ratio by system (RQ2). Higher values indicate more acoustically distinct genre categories. Suno's separation ratio falls 36% below the human baseline, indicating that genre boundaries collapse in its output space. Lyria matches human genre separation ( +1% ).

Furthermore, the two systems do not converge on a single "AI sound." The authors find that Suno and Lyria 3 are actually quite different. They are more acoustically distant from each other than two random samples of human music would be .

Figure 4
Figure 4: System convergence ratio by genre (RQ3). Each bar shows the distance between the Suno and Lyria centroids divided by the expected distance between two random human splits. All ratios exceed 1 . 0 (dashed line), meaning the two AI systems are more acoustically different from each other than two random human subsets would be.

Finally, the researchers demonstrate that a standard machine learning classifier can distinguish AI tracks from human ones with near-perfect accuracy (AUC = 0.991). This score represents an extremely high level of reliability. The classifier succeeds primarily by detecting subtle irregularities in rhythmic timing and timbral dynamics .

Figure 5
Top Features: Al vs. Human Discriminability (AUC = 0.991 ± 0.003)

What This Lets Us See

This research shifts the conversation from whether AI music is "good" to how it affects musical ecosystems. It reveals that homogenization is not a single phenomenon. It can manifest in different architectural ways. A system might preserve the "idea" of a genre (like Lyria) while stripping away its nuance. Or, it might preserve internal variety while blurring the lines between genres (like Suno).

These patterns have implications for "platform legibility." Streaming services use automated classifiers to organize playlists. The authors suggest that AI-generated music may be "easier" for these platforms to sort and amplify. This happens because the music has predictable rhythmic and spectral features. This creates a feedback loop. If AI music is easier to recommend, it occupies more space on platforms. This could eventually shift the "reference distribution" (the standard expectation) of what a genre sounds like for human listeners.

Where The Edges Are

The authors acknowledge several limitations. Because this was a "black-box" audit, the study describes what the models do. It cannot explain the internal causal mechanisms driving these behaviors. Additionally, the human reference sets were limited to 30-second previews. This might underestimate the true diversity in full-length compositions.

The study also notes that MIR features are often rooted in Western production norms. This means the audit might miss highly nuanced musical elements. Examples include specific regional "grooves" or micro-timing variations. These are vital to practitioners but difficult to formalize mathematically. Finally, the researchers clarify that these results audit specific commercial behaviors. They do not necessarily predict the output of every generative model in existence.

Figures from the paper

Figure 6
Figure 6: Word cloud of Suno null-prompt titles (100 tracks per genre). Word size proportional to frequency; color indicates the genre in which the word is most frequent (Afrobeats = amber, K-pop = pink, Dance Pop = blue, Heavy Metal = grey). The vocabulary is narrow and several words (e.g., 'Anvil,' 'Pulse') appear prominently across multiple genres, illustrating cross-genre leakage in title generation.
Novelty
0.0/10
Overall
0.0/10
#generative AI#music information retrieval#algorithmic bias#cultural justice#homogenization
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 93% (passed)
Claims verified: 16 / 16

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 96,340
Wall-time: 205.3s
Tokens/s: 469.2