Feed 0% source
Medicine AI-generated

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

The Judgment-Consequence Gap: LLMs Agree on Moral Responsibility but Refuse to Act on It

As large language models (LLMs) transition from digital assistants to active advisors in high-stakes sectors like healthcare, the values they encode become a matter of direct consequence for patient outcomes. In clinical settings, decisions regarding the allocation of scarce resources—such as donated organs or intensive care beds—often hinge on judgments of responsibility. This is especially true when a patient's own actions have contributed to their illness.

Current research focuses on whether LLMs can mimic human moral norms. Most studies treat moral reasoning as a single-step evaluation. This leaves a critical question unanswered: where exactly does the reasoning chain break? Do models differ from humans in how they assess responsibility, or in how they translate that assessment into a consequential decision? A new study from Pennsylvania State University suggests the latter. The researchers identified a profound "judgment-consequence gap" where models accurately diagnose moral culpability but systematically refuse to act upon it.

The Disconnect Between Assessment and Action

In human moral psychology, attributing responsibility is often a two-stage process. First, an observer makes a causal judgment (e.g., "this behavior led to this illness"). Second, they make an evaluative judgment (e.g., "therefore, this person deserves less help"). Decades of research in attribution theory show that humans reliably connect these stages. Humans often favor the "less-culpable" patient when resources are limited.

Existing benchmarks for LLM morality show that these models possess a broad familiarity with human norms. However, the authors of this paper argue that existing evaluations suffer from a "value-action gap." This is a phenomenon where a model's stated preferences diverge from its actual choices. The researchers sought to determine if LLMs fail because they lack moral knowledge. Alternatively, they might operate under a fundamentally different set of rules when applying that knowledge to life-or-death decisions.

Tracing the Chain of Moral Reasoning

To isolate the point of divergence, the researchers designed a layered experimental framework. This framework traces the reasoning process through three successive stages of increasing moral weight. Using clinical vignettes (detailed scenarios) adapted from studies on kidney transplants, lung cancer, and hip replacements, the study asks participants to evaluate:

  1. Behavioral Responsibility: Is the patient responsible for the health-harming behavior itself (e.g., smoking or heavy drinking)?
  2. Disease Responsibility: Is the patient responsible for the resulting illness?
  3. Deprivation Responsibility: Is the patient responsible for being denied a scarce resource?

Finally, the models must make a consequential allocation decision. Should the resource go to Patient A (the healthy patient), Patient B (the responsible patient), or be decided randomly?

The researchers tested 12 different model families, including Claude, GPT, Gemini, and Llama. They utilized both "reasoning" (models using extended thinking processes) and "non-reasoning" configurations. This allowed them to observe whether the gap resulted from shallow processing or a stable, deeply embedded normative commitment (a fixed adherence to a specific set of moral rules).

A Systematic Refusal to Penalize

The results reveal a striking divergence. On the first level—behavioral responsibility—LLMs and humans are in near-complete agreement. The human mean is 4.42 on a 5-point scale. The aggregate mean across all 19 LLM configurations is 4.433 .

Figure 1
Figure 1: Mean responsibility scores (5-point scale) by question and model group, aggregated across both behavior alteration conditions (continued vs. stopped). Error bars show 95% confidence intervals. 4

Both groups recognize that patients bear responsibility for their own harmful actions.

However, the consensus evaporates at the decision-making stage. While a majority of human participants (67.6%) prefer to allocate the resource to the less-culpable Patient A, LLMs overwhelmingly default to randomization .

Figure 3
Figure 3 — from the original paper

This means that instead of choosing the patient with the better prognosis, the models essentially "flip a coin." Even when models are equipped with extended reasoning capabilities, they do not move closer to the human pattern. Instead, they lean further into randomization.

The authors report that this is not a failure of comprehension. Through an analysis of reasoning traces, they found that "thinking" models frequently acknowledge the patient's responsibility in their internal monologue. They then explicitly invoke principles of "fairness" or "equal treatment" to justify a random choice. In fact, LLMs rate behavior-based allocation as significantly more "unfair" (5.6 on a 7-point scale) than humans do (3.9) .

Figure 4
Figure 4 — from the original paper

This indicates that LLMs are not failing to see the responsibility. They are actively choosing to ignore it in favor of an egalitarian (equality-focused) or contractualist (rule-based) framework. This framework prioritizes procedural equality over "desert," or what a person deserves based on their actions.

Furthermore, the study found that LLMs are uniquely sensitive to the patient's epistemic state (their level of knowledge). When a patient lacks access to information about health risks, LLMs sharply reduce responsibility judgments. This pattern is much more pronounced than in humans .

Limitations of the Framework

While the findings are robust, the study is constrained by several factors. First, the scenarios are "stylized." They are simplified abstractions designed to isolate variables. Real-world medical allocation involves a chaotic web of logistics and regulatory constraints. These vignettes cannot capture that complexity.

Second, the human baseline data used for comparison are primarily drawn from Western, educated populations. Moral norms regarding responsibility and "fairness" vary significantly across cultures. The "gap" observed here might look very different if evaluated against a more globally diverse human cohort. Finally, the study relies on single-turn or multi-turn interactions. It does not involve a truly deliberative, iterative dialogue between a clinician and an AI. Such a dialogue might alter how the model weighs conflicting moral duties.

The Verdict: A Risk of Partial Alignment

Is the LLM ready for clinical decision support? Not yet.

The research demonstrates that we are currently facing a problem of "weak alignment." Current training methods, such as Reinforcement Learning from Human Feedback (RLHF), appear to succeed at teaching models to describe human values. However, they fail to teach them how to compose those values into coherent decision-making rules.

The danger for practitioners is not that the AI will give a "wrong" answer to a moral question. The danger is that it will provide a "right" first step that masks a dangerous second step. A clinician might trust an AI that correctly identifies a patient's culpability. They might assume the model's subsequent recommendation will follow logically. Instead, the model may pivot to a completely different moral logic. It may treat the very act of considering culpability as an injustice. Until alignment research addresses the compositional rules that connect judgment to action, LLMs will remain unpredictable actors in the moral landscape of medicine.

Figures from the paper

Figure 2
Figure 2: Effect of behavioral change (stopped vs. continued) on responsibility scores. Error bars show 95% confidence intervals.
Figure 5
Figure 4: Mean Likert scores (7-point scale) by information condition for humans, reasoning models, and non-reasoning models across all four questions in the knowledge levels experiment. The shaded regions represent 95% confidence intervals.
Figure 6
Figure 5: Mean LLM scores for questions in the knowledge level vignette across all six information conditions. Error bars show 95% confidence intervals.
Novelty
0.0/10
Impact
0.0/10
Overall
0.0/10
#medicine#clinical#AI ethics#large language models#resource allocation
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: science_essayist
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 13 / 14

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 155,036
Wall-time: 265.5s
Tokens/s: 584.0

Related
Next up

When Shippers Become Algorithms: LLM Agents Drive Market Concentration in Fre...

8.3/10· 6 min