Hearsay: Frontier VLMs Confabulate Demographically Biased Medical Diagnoses Without Images
When asked to diagnose a patient without seeing an actual medical image, advanced vision-language models (VLMs) do not simply admit they lack the necessary visual evidence. Instead, they frequently manufacture—or confabulate—a diagnosis.
These models, designed to process both text and imagery, are increasingly being integrated into clinical pipelines. However, a new study from Carnegie Mellon University and Amazon Web Services reveals that this fabrication is not random. The models appear to "hallucinate" specific diseases based on the patient's race, age, and sex, often mirroring documented human medical biases.
Even more troubling, the researchers found that some models employ a "hedging" strategy. They write text acknowledging the missing image while simultaneously populating a structured data field with a specific, fabricated diagnosis.
Does the absence of an image trigger random errors?
The central question driving this research is whether the "mirage effect"—a term used to describe a VLM generating visual descriptions without an actual image—is a chaotic failure or a structured one. In a clinical setting, a VLM might be tasked with analyzing a scan that failed to upload. If the model encounters such a gap, does it provide a safe refusal, or does it lean on demographic stereotypes to fill the void?
The authors sought to determine if the content of these fabricated diagnoses depends on the demographic descriptors provided in the text prompt. Specifically, they investigated whether changing a patient's identity would systematically shift the predicted disease. This could turn a technical error into a vehicle for social bias.
The limitations of prose-only audits
Before this study, the prevailing understanding of the mirage effect focused largely on the model's natural language output. Previous work established that frontier models exhibit high rates of mirage. They often produce convincing-sounding descriptions of anatomy that isn't there. The assumption was that if a model's prose sounded cautious, the model was behaving safely.
However, the authors suggest this view is incomplete. Developers might miss deeper failures occurring in the underlying data structures. If a model acts as an automated agent that extracts information into a database, the prose is merely the surface layer. The structured output is what actually drives downstream clinical decisions.
Testing the demographic influence
To investigate this, the researchers conducted a massive experiment involving 11,700 calls across three frontier models: Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro. They tested these models across three medical domains: chest X-rays, brain MRIs, and dermatology (specifically regarding "skin moles").
The team used a factorial design, meaning they systematically varied the patient's age (32 or 65), sex (man or woman), and race (white, Black, or brown). They presented these descriptors to the models in a first-person prompt. Crucially, no image was ever provided.
The researchers measured the divergence between these demographic-specific responses and a neutral baseline using Jensen-Shannon divergence (JSD). JSD is a statistical metric used to quantify how much one probability distribution differs from another. As shown in, the models exhibited significant shifts in their diagnostic patterns.
For instance, Claude reached a maximum JSD of 0.834 in dermatology, indicating an extreme departure from its neutral behavior.
Structured biases and the hedged regime
The findings reveal a stark landscape of demographic-driven fabrication. The authors report that the models do not just err; they err in predictable, biased directions. For example, the study finds that when prompted as a 65-year-old white man asking about a "skin mole," Claude Opus 4.7 returns a diagnosis of Melanoma in 94% of cases [Table 1]. Conversely, a 32-year-old Black woman asking about a chest X-ray frequently receives a Sarcoidosis diagnosis. In some instances, the model's reasoning explicitly cited "demographics and classic pattern."
One of the most significant discoveries is the "hedged mirage" regime. The authors report that in 66% of Claude's fabrications in its most biased dermatology cell, the model's reasoning prose acknowledged the missing image. Yet, the structured diagnosis field was still populated with a disease. This creates a dangerous dissociation. An auditor reading only the text might flag the model as "safe." Meanwhile, a clinical pipeline reading the JSON data field would receive a definitive, albeit fake, diagnosis.
The study also highlights that these failure modes are not uniform. Through a "probe-noun" analysis, the authors found that Claude’s dermatology bias was highly sensitive to specific wording. Swapping "skin mole" for "skin lesion" caused the Melanoma rate to collapse from 94% to 0%. In contrast, GPT-5.4's bias was "category-preserving." This means it maintained its biased diagnostic patterns even when the phrasing changed.
Implications for clinical deployment
If these findings generalize to broader clinical applications, the implications for AI safety are profound. The study suggests that evaluating VLMs based solely on their conversational tone is insufficient for ensuring reliability.
First, the authors argue that trustworthy deployment requires direct auditing of the structured output channels. Because of the "hedged" behavior, a model can appear to follow safety protocols in its text. Simultaneously, it can actively feed incorrect data into a medical database. Second, the discrepancy between Claude and GPT-5.4 suggests that "mirage" is not a single phenomenon. It is likely a family of distinct failure modes. A mitigation strategy that fixes word-triggered errors might fail to address more robust, category-based biases.
The paper concludes that probe-word sensitivity must be treated as a "first-class evaluation dimension" in AI testing. A logical next step for researchers would be to investigate whether these demographic shifts persist when the models are given more complex clinical histories.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: narrative_discovery
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 11 / 12
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 52,544
Wall-time: 163.4s
Tokens/s: 321.6