AI Chatbots Make Life-or-Death Choices Differently Than Humans in Kidney Transplant Scenarios

Reviewed byNidhi Govil

5 Sources

Share

Penn State researchers found AI chatbots prioritize starkly different attributes than humans when making high-stakes medical decisions. In kidney transplant scenarios, AI models fixated on single factors like drinking habits while showing unwarranted confidence, raising questions about AI's role in decision support for ethically sensitive situations.

AI Decision-Making Diverges from Human Values in High-Stakes Medical Scenarios

A Penn State University study presented at the 2026 Association for Computing Machinery Fairness, Accountability and Transparency (FAccT) conference reveals that AI chatbots make life-or-death choices fundamentally differently than humans

1

2

. Researchers tested large language models using hypothetical kidney transplant scenarios where multiple patients needed organs but only one was available. The findings expose critical gaps in how AI handles ethical decisions and moral nuance compared to human doctors.

Source: New York Post

Source: New York Post

"Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI's role in them right isn't optional," said Hadi Hosseini, associate professor of informatics and intelligent systems at Penn State University who led the study

3

. While researchers don't encourage using AI as a substitute for professional judgment in high-stakes medical decisions, understanding AI behavior becomes essential as organizations increasingly rely on these systems for recommendations

1

.

AI Fixates on Single Factors in Organ Allocation Decisions

The research team created head-to-head comparisons using two hypothetical patients described by attributes including age, number of dependents, health status, and drinking habits. They tested 14 scenarios, with 9 isolating single traits and 5 forcing trade-offs among multiple factors

2

. The team compared AI responses with decisions from 289 human participants in earlier academic studies on kidney allocation

2

.

Source: News-Medical

Source: News-Medical

"AI chatbots often diverge from human values in how they weigh a patient's traits," Hosseini explained. "They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do"

5

. While human respondents placed greater importance on age, favoring younger patients, many AI models prioritized lower alcohol consumption instead

3

. Human decisions considered multiple factors and remained context-sensitive, whereas language models often focused on single attributes

4

.

In one striking example, both patients were 55 years old with identical drinking habits. Humans chose the patient with two dependents 93% of the time, while Claude-3.5-Haiku selected the patient with no dependents 65% of the time

2

. DeepSeek models acted rigidly, repeatedly favoring younger patients even when younger candidates had worse health or drinking habits

2

.

AI Shows Overconfidence While Humans Express Indecision

Researchers added a "flip a coin" option to measure indecision, a key factor in human moral judgment

1

. They also tested alternatives like stating both patients deserved the kidney or that more information was needed. Humans used indecision across many ethically sensitive situations, but AI models almost never did

2

.

"Humans frequently express indecision perhaps because they don't want to accept agency," Hosseini noted. "AI models almost never do this: Even when directly given the option to 'flip a coin,' they overwhelmingly commit to a confident, deterministic answer instead. That's a meaningful gap, since real moral dilemmas often don't have one clearly correct answer"

1

. This overconfidence represents a fundamental difference in how AI handles ethical trade-offs compared to humans who recognize ambiguity in organ donation decisions

5

.

John Dickerson, chief executive officer at Mozilla.ai who collaborated on the study, emphasized this critical distinction: "When we allocate something scarce, whether it's a kidney, a job or access to some other resource, there isn't always a single objectively correct answer. Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don't"

3

.

Fine-Tuning Improves AI Alignment but Gaps Remain

The team tested whether fine-tuning could improve AI's role in decision support. They trained four open-source models using decisions from 132 people, each judging 40 kidney-allocation cases, creating 3,960 training examples

2

. Qwen-3-14B improved from 58.03% to 76.67% accuracy in two-choice tasks, while accuracy when indecision was allowed rose from 45.23% to 65.83%

2

. Models became more willing to express uncertainty after training, though they often hesitated when humans remained confident

2

.

Despite improvements, the research raises fundamental questions about AI alignment with human values in moral decision support. Large language models increasingly integrate into healthcare, supporting clinical workflows, diagnosis, treatment planning, and resource allocation

4

. These applications demand not only accuracy but alignment with human values and moral judgment in ethically sensitive situations

4

.

Source: Earth.com

Source: Earth.com

"The ethical stakes are high, and AI's role in such life-altering decisions requires deep reflection," Hosseini emphasized

3

. The study, conducted by Hosseini with doctoral student Samarth Khanna and undergraduate Leona Pierce from Penn State's College of IST, demonstrates that asking whether AI can make ethical decisions or align with human values sits at the core of today's AI discourse

1

. As individuals and firms increasingly rely on AI for high-stakes medical decisions and recommendations, continued research and governance involving policymakers and regulators becomes critical

1

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved