Researchers are increasingly using large language models (LLMs) to automate tasks like classifying political text, summarizing policy choices, or simulating human respondents. This expansion turns the political character of model outputs into a serious concern for anyone deploying AI in public-facing or sensitive domains. However, measuring a model's "geopolitical preference"—its inherent leanings on international issues—is notoriously difficult. Current audits typically rely on simple surveys or a handful of contested questions. These often fail to capture a model's consistent stance on complex, multi-layered diplomatic issues.
A new study from the University of Oxford attempts to solve this by treating LLMs as respondents in a massive, historical voting exercise. By analyzing how models react to the full text of thousands of United Nations General Assembly (UNGA) resolutions, the researcher found a startling mismatch. Even models developed by American companies frequently express positions that align more closely with Russia than with the United States.
Beyond simple surveys and polls
Existing efforts to audit the political bias of AI often fall short. They lack a rigorous way to recover latent preferences (hidden, underlying viewpoints). Most researchers use "elicitation designs," which are structured ways of asking a subject for their opinion. While effective for simple topics, these methods struggle with geopolitics. In this field, a stance isn't just a single "yes" or "no." It is a position in a multidimensional landscape of competing interests.
Current approaches often assume that a model's geopolitical orientation can be guessed from its developer's home country. This study challenges that presumption. The authors argue that looking at a model's reaction to a small set of questions is insufficient. Instead, they look at how models behave across a vast history of international conflict and cooperation. As seen in, models do not approach the task with a common threshold.
For instance, GPT-5 supports 97.3% of resolutions. Meanwhile, DeepSeek opposes 44.8%. Relying on a simple survey might miss these fundamental differences in how models weigh "support" versus "opposition."
Recovering latent positions from UN votes
To move beyond surface-level agreement, the paper adopts a "dynamic ordinal ideal-point approach" borrowed from international relations. Think of an "ideal point" as a coordinate on a map. Just as you can estimate a person's favorite climate by looking at which cities they visit, this method estimates a model's "geopolitical coordinate." It does this by looking at which UN resolutions it supports, abstains from, or opposes.
The mechanism works in several sophisticated stages:
- The Latent Utility Model: The authors treat each vote as an expression of a hidden "utility" (a mathematical representation of preference). This utility is determined by the interaction between the actor's position and the specific characteristics of the resolution.
- Handling Ordinal Data: Unlike standard regression, this model handles "ordinal" data. These are categories like Support, Abstain, and Oppose that have a natural order but no fixed numerical distance between them.
- Bridging Time with Recurring Resolutions: The researchers use "recurring resolutions"—items that appear in multiple UN sessions—to act as anchors. This allows the model to maintain a consistent scale over time. It effectively bridges the gap between 1946 and 2025.
- Smoothing Trajectories: The method uses a "dynamic prior" (a mathematical assumption about how variables change over time) to ensure an actor's position doesn't jump erratically. This creates a smooth trajectory through the latent space.
By applying this to 5,555 divisive resolutions, the authors place models on the same "map" as the Permanent Five (P5) members of the UN Security Council.
A surprising distance from the US
The results of this spatial mapping are striking. The paper reports that in the twenty-first century, the modern models GPT-5, Claude Sonnet, and Gemini are actually closest to Russia among the P5 members. Specifically, the authors find that GPT-5's mean distance from Russia is 0.248. Its distance from the United States is much larger at 2.411 .
The discrepancy is most visible when examining specific subsets of contentious issues. The authors isolated 2,104 resolutions where the United States voted "No," but China and Russia voted "Yes." In this specific arena, the models leaned heavily toward the China/Russia position. The paper reports that GPT-5 supported 96.1% of these resolutions. DeepSeek was the outlier, supporting only 36.1% [Table 1].
Interestingly, the study notes that closeness to certain states can depend on the metric used. While the "ideal-point" estimate places the Western models near Russia, a simpler "S-score" (a measure of raw categorical agreement) might rank them as closer to China .
This highlights a critical lesson for engineers. A model might agree with a country on the frequency of its votes without sharing that country's underlying geopolitical logic.
Limits of the geopolitical map
While the methodology is robust, the authors define what the study does not prove. First, the study recovers "expressed choices," not necessarily stable internal beliefs. A model's response is a function of the specific prompt provided. Changing the language or assigning a national role could fundamentally alter the results.
Second, the analysis uses a one-dimensional scale. This simplifies the complex web of global politics into a single axis. This may compress or obscure issue-specific coalitions. For example, a model might align with Russia on maritime law but diverge on nuclear non-proliferation. The authors acknowledge this compression is a trade-off of the dimensionality reduction.
Finally, the study is limited by its reliance on English-language prompts and the specific corpus of adopted resolutions. Because the resolutions themselves often use highly normative, consensus-seeking language, the models might simply be reflecting a "prosocial" training bias. This is a tendency to agree with broadly accepted international principles. This might happen rather than the models expressing a calculated foreign policy.
Verdict: Audit, don't assume
Is this research ready for deployment in policy workflows? The answer depends on your task, but proceed with extreme caution.
If you are using LLMs to summarize international law, these models seem capable of capturing the prevailing normative spirit of the UN. However, if you use them to simulate the strategic reasoning of a nation, the paper proves that "developer nationality" is a poor proxy for "model alignment."
The takeaway for practitioners is clear. Sovereignty cannot be inferred from ownership alone. A model built in the US can still produce outputs that are strategically distant from US interests. Any organization integrating LLMs into sensitive policy workflows must conduct task-specific audits. These should be measured against actual policy corpora rather than relying on generic safety benchmarks.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 96% (passed)
Claims verified: 17 / 17
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 72,352
Wall-time: 177.1s
Tokens/s: 408.6