Feed 0% source
Neuroscience AI-generated

Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

Frontier VLMs show fragmented Theory of Mind profiles across visual and abstract tasks

Researchers have recently begun testing advanced AI models on various "mind-reading" tasks. These are psychological tests designed to see if an agent can understand the internal states, perspectives, or intentions of others. This capacity is known as Theory of Mind (ToM). ToM is a cornerstone of social intelligence. However, a critical question remains: does a model possess a coherent, unified ability to reason about minds? Or does its performance fragment depending on the specific task it is performing?

The search for a unified cognitive profile

The central question investigated by Zhang et al. is whether frontier vision-language models (VLMs)—AI systems capable of processing both visual and textual information—present a consistent Theory of Mind profile. Specifically, the authors ask if a single model aligns with the same human reference group across multiple tasks. For example, does a model behave like a typically developing (TD) adult in every scenario? Or does its "personality" shift from one paradigm to the next?

To understand why this is difficult, one must recognize that Theory of Mind is not a monolithic skill. It encompasses diverse sub-capacities. These include perspective-taking (understanding that someone else sees something different than you do) and intention attribution (inferring a goal from observed motion). If a model performs well on one but fails on the other, it suggests a lack of a centralized "social engine." Instead, the AI may rely on task-specific heuristics (mental shortcuts).

Cracks in the single-score paradigm

Until now, much of the research into VLM social intelligence has relied on single-task benchmarks. Previous studies often used naturalistic stimuli. These included videos of human faces, dialogue, or complex household scenes. While these are realistic, they introduce "social-cue shortcuts." A model might correctly predict a character's intention by recognizing a facial expression. It might do this without actually understanding the character's mind.

Furthermore, existing benchmarks tend to isolate a single sub-capacity at a time. This approach creates a blind spot. It cannot detect "cross-task dissociation." This is a phenomenon where a model's cognitive profile changes fundamentally between different facets of social reasoning. As shown in the conceptual design, evaluating a model on only one dimension provides an incomplete picture.

Figure 1
Figure 1 — from the original paper

It may lead to a misleading assessment of true social competence.

Decoupling social cues from mental reasoning

To bypass the trap of social shortcuts, the researchers designed a dual-benchmark suite. They used abstract, non-human stimuli. The first component is an adaptation of the Keysar Director Task. In this setup, the model sees 3D tabletop scenes with colored blocks and an opaque partition. The partition creates a "privileged-information asymmetry." This means the model can see certain blocks that a "director" figure cannot. The model must decide whether to act based on its own view or the director's restricted view.

The second component is an adaptation of the Frith-Happé animated-triangles task. Here, the model watches silent, 2D animations of geometric shapes. Because there are no faces or voices, the model cannot rely on social patterns. It must attempt to attribute intention (e.g., "the triangle is mocking the other") purely from the kinematics (the mathematics of motion) of the shapes.

The researchers evaluated a panel of nine frontier models. This panel included versions of Claude, GPT, Gemini, and Qwen. They utilized an "LLM-as-rater" jury to score the models' free-text descriptions. This jury ensured the grading followed established psychological rubrics like the Castelli rubric.

Fragmented profiles and the egocentric error

The findings reveal a striking lack of coherence across the model panel. On the Director Task, the models frequently fell into the "egocentric error." This is the tendency to pick an object based on one's own visual perspective. This happens instead of adopting the perspective of the person being addressed. The authors report that, without chain-of-thought reasoning, the panel made this error on 77.8% of trials. This failure mode is characteristic of young children rather than adults .

Figure 2
Figure 1: Schematic of the two-benchmark cross-task design. The Director Task adapts Keysar et al. (2000) and probes visual perspective-taking through the explicit-implicit gap x = P ( sub-prompt a correct ) -P ( sub-prompt b correct ) . The animated-triangles task adapts Frith-Happé clips and probes abstract intention attribution through the ToM-condition (Intent, Approp) profile. The right panel previews the joint cross-task plane (Section 5).

While explicit reasoning helped some models recover, the "know-but-don't-use" trap remained a major hurdle.

Even more surprising was the dissociation found in the animated-triangles task. Instead of showing a consistent profile, the models' performance shifted significantly. On the triangles, the panel exhibited a profound under-attribution of intention. The researchers found that the models' profiles sat more than three times closer to the mean of high-functioning autistic (HF-ASD) adults than to the mean of typically developing (TD) adults .

Figure 3
Figure 2: Director-Task per-model accuracy on the four color-word sub-prompts (Sec. 3.1). The panel contains nine models, twelve canonical scenes, and three trials per (model, scene, sub-prompt) cell. The (a)-(b) gap is the explicit-implicit gap. Seven of nine models score zero on (b), claude-opus-4.7 scores 1 / 12 , and gemini-3.0-pro is the exception at 12 / 12 . Model colors match Figs. 3, 4, and 5. Per-model accuracies appear in Appendix Table 2.

When these two tasks are viewed together in a joint coordinate system, the fragmentation becomes undeniable.

Figure 4
Figure 3: Per-model (Intent, Approp) profile on the animated-triangles ToM condition, using the metric from Section 3.2 and the anchor-blinded rubric from Appendix J. Each filled circle is one frontier VLM, marked by monogram and colored by lab. Stars mark the TD-adult and HF-ASD-adult group means from Castelli et al. (2002). All nine models lie closer to HFASD-adult than to TD-adult on this condition.

For eight of the nine models, the "nearest human reference" changed depending on the task. A model that appeared relatively adult-like in its perspective-taking on the Director Task would suddenly align with the HF-ASD profile on the animated-triangles task. No single model in the panel maintained a consistent proximity to the typically developing adult profile across both domains.

Implications for the future of social AI

This dissociation suggests that frontier VLMs do not possess a unified Theory of Mind. Instead, their "social" capabilities appear to be a collection of fragmented, task-specific patterns. If this finding generalizes, it implies that current scaling laws may not be sufficient. Simply making models larger or giving them more data may not produce true social intelligence. We may be building models that mimic social patterns without representing the actual mental states of others.

The study highlights two immediate consequences for the field. First, for practitioners building social robots, a single "ToM score" is an insufficient metric. Models must be evaluated across a diverse battery of complementary tasks. This avoids being misled by accidental successes. Second, for researchers, the results emphasize the necessity of using abstract stimuli. Using geometric motion or blocks helps ensure we are measuring genuine mental-state inference rather than mere pattern matching.

Figures from the paper

Figure 5
Figure 4: Joint coordinates of the nine frontier VLMs (filled circles, two-letter monograms per Fig. 3) and the two adult reference groups from Castelli et al. (2002) (stars). X axis: Director-Task explicit-implicit gap P ( a ) -P ( b ) (TD-adult 0 . 435 , HF-ASD-adult 0 . 340 ; sources in Sec. 3.3). Y axis: animated-triangles ToMcondition distance to TD-adult, a visualization anchor only; the triangles nearest-reference rule uses the full 2D(Intent, Approp) space. For eight of nine models the nearest adult reference differs between the two tasks; the full cross-tabulation appears in Appendix Table 8.
Figure 6
Figure 5: Per-model (Intent, Approp) profile on the animated triangles broken down by condition. Each panel shows the per-model points, colored by lab. The canonical ToM-condition panel is also reported as Figure 3 in the main text.
Novelty
0.0/10
Overall
0.0/10
#neuroscience#artificial intelligence#theory of mind#vision-language models#cognitive psychology
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: science_essayist
Template: narrative_discovery
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 93% (passed)
Claims verified: 10 / 10

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 126,484
Wall-time: 242.4s
Tokens/s: 521.8

Next up

When Shippers Become Algorithms: LLM Agents Drive Market Concentration in Fre...

8.3/10· 6 min