LLMs Can Mimic Aggregate Urban Attitudes but Fail to Represent Population Structure
Researchers have recently begun testing if large language models (LLMs) can act as low-cost proxies for real residents. They ask models about new housing developments near their homes. While some models get the "average" answer right, they often fail to capture critical group differences. These include differences between renters and homeowners or various political affiliations. This study investigates whether these digital surrogates can truly replicate the complex social and spatial structures that drive urban public opinion.
The trap of aggregate matching
In urban planning, understanding public sentiment is rarely about finding a single mean value. Decisions regarding affordable housing depend heavily on how attitudes shift based on proximity to a site. These shifts also vary across different demographics. For instance, a homeowner might view a new development through the lens of property value. Conversely, a renter might prioritize housing availability.
Current research suggests that LLMs can effectively predict the average results of social science survey experiments. However, this creates a dangerous blind spot. A model might predict a correct "average" support level for a project. Yet, it might fail to realize that Republicans and Democrats react in opposite ways. Such a failure means the model has failed to represent the actual population. As illustrated in the conceptual framework, a valid simulation must do more than achieve behavioral correspondence (matching observed actions).
It must also preserve the underlying population structure and measurement stability. Relying on aggregate means alone masks the risk of flattening essential diversity.
Measuring the urban persona
To test this, the authors developed a multi-level validation framework. They tested eight open-weight LLMs against a US-based affordable-housing survey experiment. This human benchmark involved 843 respondents. The core mechanism involves simulating "sessions." In each session, the model is given a specific persona. This persona is defined by political party and housing tenure (the type of housing arrangement, such as renting or owning). The model then evaluates housing proposals at two different distances: 1/8 mile and 2 miles from home.
The researchers implemented a tiered evaluation strategy: 1. Spatial-behavioral correspondence: Does the model show the same "proximity penalty" (the drop in support as a project gets closer) seen in humans? 2. Population-structural correspondence: Does the model maintain specific differences between subgroups, such as the owner–renter contrast? 3. Measurement stability: Do the results hold up when the prompt is changed? This includes altering question order or adding context.
The study also employed an exploratory diagnostic called Representational Similarity Analysis (RSA). This technique analyzes "hidden states" (the internal mathematical vectors used to process information). It seeks to see if the model's internal geometry tracks physical distance in a way that mirrors human cognition .
Success in the mean, failure in the cells
The results reveal a striking mismatch between global accuracy and local reliability. The authors report that Qwen 2.5 14B was the only model to meet the pre-specified equivalence criterion for the primary owner–renter proximity contrast ($\theta$). It yielded an estimate of -0.242 compared to the human benchmark of -0.285 [Table 1]. This means the model's estimate fell within the acceptable margin of error for the human result. On the surface, this looks like a success.
However, the paper finds that this aggregate match masks profound structural failures. When the authors looked at individual "cells" (specific combinations like "Republican mortgage holders"), the models faltered. For example, the authors report that Qwen 2.5 14B attenuated the Republican contrast and exaggerated the Independent one. Furthermore, the models tended to compress within-group variation. The authors measure a median model-to-human variance ratio of only 0.099 for Qwen [Table 2]. This low number means the models produce much more uniform, "robotic" responses than diverse human populations.
The instability of the measurement process was also evident. The authors find that randomized question order significantly shifted estimates for several models .
For Qwen, the shift in the $\theta$ estimate was +0.367. This discrepancy does not occur in the human benchmark.
Limits of the digital surrogate
The study highlights several critical reasons why these models cannot yet be used as reliable replacements for human constituents. First, the researchers observe that "rationale-first" prompting fundamentally changes the outcome. This involves asking the model to explain its reasoning before making a choice. The authors report that for Gemma 3 12B, agreement with direct choices dropped to just 64.7% when the explanation preceded the selection [Table 3]. This suggests that LLM explanations may be "post hoc rationalizations" (explanations created after the fact) rather than transparent windows into decision-making.
Second, the models exhibit selective nonresponse. The authors note that certain models, such as Mistral 7B, failed to produce valid outputs for specific identity groups. This leads to a "coverage" problem. If a model only successfully simulates Democratic profiles, its conclusions will be biased. It will fail to represent the intended population.
Finally, the authors caution that the study does not establish internal psychological mechanisms. They note that they can observe behavioral differences. However, they cannot prove whether the models "think" like humans or simply follow statistical patterns.
The verdict: not for policy deployment
If you want a tool for rapid, high-level scenario screening, LLMs are a promising starting point. They can help generate hypotheses for future studies. However, if your goal is to replace resident surveys for actual planning decisions, the answer is a firm not yet.
The evidence is clear. An LLM can pass a "behavioral Turing test" by hitting the right average numbers. Yet, it can simultaneously misrepresent the people it is meant to simulate. For practitioners, the takeaway is vital. Treat LLM-generated attitudes as speculative simulations rather than empirical evidence. Until models can preserve the messy, heterogeneous, and spatially-anchored reality of human social structures, they remain tools for researchers. They are not yet substitutes for the public.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 11 / 11
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 154,683
Wall-time: 287.8s
Tokens/s: 537.5