Feed 0% source
AI/ML AI-generated

Language Models Agree With Each Other, Not With Readers

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

Language Models Converge with Each Other Far More Than with Naturalistic Human Readers

Researchers have found that AI models tend to pick the same important sentences from documents much more often than real people do. While newer AI models are getting better at reading like humans, they are also getting much more similar to each other. This creates a widening gap. This observation challenges the idea that increasing model capability automatically brings AI closer to human judgment.

The Growing Gap in Machine Consensus

The central question is whether language models are becoming more like us, or simply more like each other. As Large Language Models (LLMs)—the engines behind tools like ChatGPT—become more sophisticated, a common claim is that they are beginning to "homogenize" (become increasingly similar). However, most studies measuring this effect use human comparisons where the humans are specifically instructed to follow the same rules as the AI.

The authors of this study argue that such designs are flawed. If you give two people the same instruction manual, they will naturally agree more than two people reading a book for pleasure. This creates an "instruction artifact" (a bias caused by the experimental setup). In this case, the measured similarity reflects the prompt rather than a genuine alignment between machine and human thought. Instead, the paper seeks to measure model convergence against a naturalistic human baseline. This baseline consists of people highlighting text on a social platform for their own personal reasons, without any external instructions or compensation.

The Background You Need

To understand this study, one must first understand the concept of "extractive salience"—the process of identifying and selecting the most important parts of a text. In traditional NLP (Natural Language Processing), this often involves picking a subset of sentences that represent the core meaning of a document.

Previous research has established that models do indeed tend to produce similar results. Some studies suggest that using AI for writing assistance reduces the diversity of ideas produced by humans. Others have noted that as models become more capable, their errors become more similar.

Crucially, the authors identify a flaw in how "agreement" is typically calculated. If two entities both pick sentences from the beginning of a document, they will appear to agree simply because of their position. To solve this, the paper employs a "position-controlled agreement estimator." This is a mathematical correction for a known bias in extractive tasks (tasks where a model selects existing text). It calculates "excess agreement"—the overlap between two sets of sentences minus the overlap you would expect to see by pure chance. This calculation accounts for the fact that some sentences are easier to pick because of their location or length.

How The Argument Works

The researchers constructed a massive, diverse panel to test this phenomenon. They analyzed 18 model "arms" (distinct versions of models) representing 11 different vendors. This panel spanned three generations of technology from 2024 to 2026. They compared these against a dataset of 2,523 "mark sets" created by real readers across 120 web documents.

The core logic follows a comparative scale. The authors first established a "human yardstick" by measuring how much two uninstructed readers agreed with one another. They found this agreement to be +0.040. This represents the baseline level of agreement between two spontaneous human readers. They then measured how much models agreed with each other. The study finds that the median agreement across 153 model pairs is +0.093. This is more than double the human benchmark.

The paper demonstrates that this convergence is a graded property. Small models, such as those in the 8B parameter range, show agreement levels roughly equivalent to humans (+0.041). However, as models scale up and become more recent, the gap widens significantly. The authors report that while agreement with human readers rises alongside capability, the agreement between models climbs much faster.

To ensure the results weren't just a byproduct of "machine style," the authors investigated whether models were choosing sentences based on surface features like vocabulary or sentence length. They found that when controlling for depth and length, models and readers choose sentences with the same structural properties. They simply choose different ones.

Implications for Model Evaluation

These findings have serious implications for how we use AI to evaluate other AI. A common practice is "LLM-as-a-judge," where one model evaluates the performance of another. However, this study shows that frontier models from rival labs reach remarkably high levels of agreement (+0.203).

This suggests that models may converge on a shared "machine" logic that differs from human preference. If models are more similar to each other than they are to humans, an evaluation pipeline based on multiple models might simply reinforce a machine-centric monoculture. Such a system might consistently reward certain types of outputs while failing to capture the diverse nuances of actual human readers.

What This Lets Us See

This research provides a sobering perspective on the trajectory of AI development. It suggests that "capability" does not necessarily buy "humanity." Instead, as models get smarter, they seem to be converging on a mathematical centroid that is distinct from human interest.

The findings imply that a population simulated from several highly capable models is not a reliable proxy for a human population. Because models agree with each other far more than they agree with the messy, diverse reality of human readers, any "synthetic population" generated by AI will likely lack human diversity. Furthermore, the study reveals that even rival frontier models—such as those from OpenAI and Anthropic—reach remarkably high levels of agreement. This suggests the industry is moving toward a functional monoculture in how machines interpret information.

Where The Edges Are

The authors are transparent about the limitations of their framework. Most notably, they acknowledge an inherent asymmetry. Models were given a specific task (ranking importance), whereas the human readers were acting naturally without a goal. This means the study cannot fully decouple the effect of "having an instruction" from the effect of "having more capability."

Additionally, the study could not test whether this convergence persists across entirely different training data. All models in the panel were trained on undisclosed datasets. Finally, the "human ceiling"—the maximum possible agreement between readers—was estimated from a subset of documents. The authors also note that the identification of recurring users on the highlighting platform remains an unverified assumption.

Novelty
0.0/10
Impact
0.0/10
Overall
0.0/10
#ai#nlp#language_models#convergence
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 1
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 83% (passed)
Claims verified: 14 / 15

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 125,732
Wall-time: 249.8s
Tokens/s: 503.4

Related
Next up

When Shippers Become Algorithms: LLM Agents Drive Market Concentration in Fre...

8.3/10· 6 min