Frontier VLMs show fragmented Theory of Mind profiles across visual and abstract tasks
Researchers have recently begun testing advanced AI models on various "mind-reading" tasks. These are psychological tests designed to see if an agent can understand the internal states, perspectives, or intentions of others. This capacity is known as Theory of Mind (ToM). ToM is a cornerstone of social intelligence. However, a critical question remains: does a model possess a coherent, unified ability to reason about minds? Or does its performance fragment depending on the specific task it is performing?
The search for a unified cognitive profile
The central question investigated by Zhang et al. is whether frontier vision-language models (VLMs)—AI systems capable of processing both visual and textual information—present a consistent Theory of Mind profile. Specifically, the authors ask if a single model aligns with the same human reference group across multiple tasks. For example, does a model behave like a typically developing (TD) adult in every scenario? Or does its "personality" shift from one paradigm to the next?
To understand why this is difficult, one must recognize that Theory of Mind is not a monolithic skill. It encompasses diverse sub-capacities. These include perspective-taking (understanding that someone else sees something different than you do) and intention attribution (inferring a goal from observed motion). If a model performs well on one but fails on the other, it suggests a lack of a centralized "social engine." Instead, the AI may rely on task-specific heuristics (mental shortcuts).
Cracks in the single-score paradigm
Until now, much of the research into VLM social intelligence has relied on single-task benchmarks. Previous studies often used naturalistic stimuli. These included videos of human faces, dialogue, or complex household scenes. While these are realistic, they introduce "social-cue shortcuts." A model might correctly predict a character's intention by recognizing a facial expression. It might do this without actually understanding the character's mind.
Furthermore, existing benchmarks tend to isolate a single sub-capacity at a time. This approach creates a blind spot. It cannot detect "cross-task dissociation." This is a phenomenon where a model's cognitive profile changes fundamentally between different facets of social reasoning. As shown in the conceptual design, evaluating a model on only one dimension provides an incomplete picture.
It may lead to a misleading assessment of true social competence.
Decoupling social cues from mental reasoning
To bypass the trap of social shortcuts, the researchers designed a dual-benchmark suite. They used abstract, non-human stimuli. The first component is an adaptation of the Keysar Director Task. In this setup, the model sees 3D tabletop scenes with colored blocks and an opaque partition. The partition creates a "privileged-information asymmetry." This means the model can see certain blocks that a "director" figure cannot. The model must decide whether to act based on its own view or the director's restricted view.
The second component is an adaptation of the Frith-Happé animated-triangles task. Here, the model watches silent, 2D animations of geometric shapes. Because there are no faces or voices, the model cannot rely on social patterns. It must attempt to attribute intention (e.g., "the triangle is mocking the other") purely from the kinematics (the mathematics of motion) of the shapes.
The researchers evaluated a panel of nine frontier models. This panel included versions of Claude, GPT, Gemini, and Qwen. They utilized an "LLM-as-rater" jury to score the models' free-text descriptions. This jury ensured the grading followed established psychological rubrics like the Castelli rubric.
Fragmented profiles and the egocentric error
The findings reveal a striking lack of coherence across the model panel. On the Director Task, the models frequently fell into the "egocentric error." This is the tendency to pick an object based on one's own visual perspective. This happens instead of adopting the perspective of the person being addressed. The authors report that, without chain-of-thought reasoning, the panel made this error on 77.8% of trials. This failure mode is characteristic of young children rather than adults .
While explicit reasoning helped some models recover, the "know-but-don't-use" trap remained a major hurdle.
Even more surprising was the dissociation found in the animated-triangles task. Instead of showing a consistent profile, the models' performance shifted significantly. On the triangles, the panel exhibited a profound under-attribution of intention. The researchers found that the models' profiles sat more than three times closer to the mean of high-functioning autistic (HF-ASD) adults than to the mean of typically developing (TD) adults .
When these two tasks are viewed together in a joint coordinate system, the fragmentation becomes undeniable.
For eight of the nine models, the "nearest human reference" changed depending on the task. A model that appeared relatively adult-like in its perspective-taking on the Director Task would suddenly align with the HF-ASD profile on the animated-triangles task. No single model in the panel maintained a consistent proximity to the typically developing adult profile across both domains.
Implications for the future of social AI
This dissociation suggests that frontier VLMs do not possess a unified Theory of Mind. Instead, their "social" capabilities appear to be a collection of fragmented, task-specific patterns. If this finding generalizes, it implies that current scaling laws may not be sufficient. Simply making models larger or giving them more data may not produce true social intelligence. We may be building models that mimic social patterns without representing the actual mental states of others.
The study highlights two immediate consequences for the field. First, for practitioners building social robots, a single "ToM score" is an insufficient metric. Models must be evaluated across a diverse battery of complementary tasks. This avoids being misled by accidental successes. Second, for researchers, the results emphasize the necessity of using abstract stimuli. Using geometric motion or blocks helps ensure we are measuring genuine mental-state inference rather than mere pattern matching.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: science_essayist
Template: narrative_discovery
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 93% (passed)
Claims verified: 10 / 10
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 126,484
Wall-time: 242.4s
Tokens/s: 521.8