Companies are increasingly using AI to act like human survey respondents to save time and money. This practice, often called creating "silicon samples," involves prompting large language models (LLMs) with specific demographics. These profiles include attributes like age, race, or political affiliation. The goal is to simulate how a particular group might respond to products, policies, or market trends.
While the appeal of replacing expensive human studies with instant API calls is obvious, the reliability of these synthetic users is highly contested. Some researchers suggest LLMs can reproduce aggregate population patterns. Others warn that they are fundamentally unsafe substitutes. Before this paper, the field lacked a unified way to answer a critical question. Is a synthetic user actually providing any more information than a simple lookup table of "what people with these demographics usually say"?
This paper provides a sobering answer. By benchmarking models across two massive, independent datasets—U.S. social attitudes and cross-cultural values—the authors demonstrate that LLMs fail to outperform basic statistical baselines at the individual level. Perhaps more dangerously, they systematically over-stereotype. They treat identity as far more predictive of opinion than it is in real life.
The illusion of aggregate alignment
Current research into "silicon sampling" often suffers from a measurement bias. It focuses on aggregate fidelity (how well the model mimics the overall distribution of answers in a population). Researchers frequently report that LLMs successfully mimic the "average" person in a group. This leads to a sense of "aggregate alignment." However, the authors argue that this is a dangerous metric to rely on in isolation. A model can mimic the average person while failing to predict how a specific individual within that group behaves.
Furthermore, existing benchmarks often fail to include non-LLM baselines. If a paper claims a model has 60% accuracy, that number is meaningless without context. We must know what a trivial predictor could achieve. Without a "yardstick"—a baseline representing the irreducible information contained in demographics alone—we cannot know if the LLM is performing "intelligence." It might simply be regurgitating a demographic lookup table. As the authors show, when you introduce these baselines, the perceived utility of the LLM often evaporates .
Anchoring simulation in baseline reality
To expose these gaps, the authors implement a rigorous evaluation framework. This approach is "baseline-anchored." Instead of measuring LLMs in a vacuum, they pit them against a hierarchy of non-LLM predictors. These predictors are trained on held-out human data. The most critical is the demographic lookup. This predictor simply returns the most common answer for a given set of demographics. If an LLM cannot beat this simple lookup, it adds zero marginal value to decision-making.
The experimental protocol is applied across four models. These range from the 8B-parameter Llama-3.1 to frontier-scale Claude models. The researchers use two distinct prompting styles: 1. Style A (Single-answer): The model predicts one discrete answer for a specific persona. 2. Style C (Distribution): The model returns a full probability distribution (a list of likelihoods for every possible answer) in a JSON format.
The authors test these styles across two vastly different domains. They use the General Social Survey (GSS) for U.S. attitudes and the World Values Survey (WVS) for global values. This ensures their findings are not just artifacts of a specific dataset or model family.
Evidence of individual failure and subgroup distortion
The results are strikingly consistent across all tested models and domains. At the individual level, LLMs fail to provide an advantage over the simplest predictors. On the GSS, no model even ties the demographic lookup baseline. On the WVS, the failure is even more pronounced. Every model falls 11 to 22 percentage points behind the baseline .
Even when using "distance-aware" metrics (which give partial credit for being close on an ordinal scale), the LLMs remain inferior.
More concerning is the discovery of "demographic over-determination," or stereotyping. The authors introduce a stereotyping index ($\Delta\eta^2$). This measures the difference in how much a demographic attribute explains the variance in answers for a model versus real humans. The paper finds that $\Delta\eta^2$ is almost universally positive .
Models do not just mirror human biases. They caricature them. For example, in the GSS, political views explain only about 1.5% of the variation in trust in banks among real Americans. However, the models behave as if politics explains roughly 67% of that variation.
This has massive consequences for decision-making. The authors perform a "decision-impact analysis" on a segment-targeting task. They find that models exaggerate the differences between groups. They inflate the perceived "gap" between demographic segments by two to fourfold .
This leads to a high "wrong-target rate." A company using synthetic users would be directed to the wrong demographic segment in 50% of U.S. cases. In cross-cultural contexts, this error happens even more frequently.
Limits of scaling and prompt stability
One might hope that using a larger, more "intelligent" model would solve these issues. However, the evidence suggests otherwise. Frontier models, such as Claude Sonnet, often stereotype as strongly as—or more strongly than—smaller 8B models . Scaling capability does not seem to translate to increased sociological nuance.
There are also significant practical hurdles regarding model reliability. Smaller models struggle with the "distribution prompt" (Style C). They produce high rates of invalid or unparseable JSON outputs. For instance, the Llama-8B model produced an 85% invalid rate on the WVS dataset. Additionally, the results are highly sensitive to "prompt-surface" changes. Merely reversing the order of answer options can flip many predictions. Changing how a persona is described also causes shifts. This suggests the "knowledge" being extracted is often a fragile byproduct of phrasing.
Verdict: Proceed with extreme caution
If you are building a decision-support system that relies on synthetic users, do not trust the aggregate metrics alone. The paper demonstrates that "looking right" at a population level is a poor proxy for being right at the subgroup level.
If your goal is to understand how specific segments differ, current LLMs will likely manufacture "spurious splits." These are differences that look statistically significant in the simulation but do not exist in the real world. Until synthetic user frameworks incorporate the rigorous baseline-anchored validation proposed here, they should be treated as high-risk tools. They systematically exaggerate social polarization.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: lesswrong_skeptic
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 97% (passed)
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 113,665
Wall-time: 189.4s
Tokens/s: 600.2