Penn State researchers found AI chatbots prioritize vastly different factors than humans when making kidney allocation decisions. The models express unwarranted confidence in ethically sensitive scenarios where humans recognize moral ambiguity, raising critical questions about AI's role in decision support for high-stakes medical choices.

AI Chatbots Diverge from Human Values in Critical Medical Decisions

Penn State researchers have uncovered significant gaps between AI decision-making and human values in life-or-death decisions, particularly in kidney allocation scenarios. The study, led by Hadi Hosseini, associate professor of informatics and intelligent systems at Penn State, examined how AI chatbots handle ethically sensitive scenarios compared to human judgment

1

. The findings, presented at the 2026 Association for Computing Machinery FAccT conference, reveal that AI models often oversimplify complex moral decisions and display unwarranted confidence when no clear right answer exists

1

.

Source: News-Medical

Source: News-Medical

"Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI's role in them right isn't optional," Hosseini emphasized

1

. The research team, including doctoral student Samarth Khanna, undergraduate Leona Pierce, and Mozilla.ai CEO John Dickerson, tested how large language models weigh patient attributes against human decision data from earlier academic studies

1

.

How AI Handles High-Stakes Decisions in Medicine

The researchers designed 14 head-to-head comparisons using hypothetical kidney transplant patients described by attributes including age, number of dependents, health status, and drinking habits

2

. Nine scenarios isolated single traits while five forced ethical trade-offs among competing factors

2

. The team compared AI responses with decisions from 289 human participants who lacked specified medical training

1

.

"We ran these comparisons in a few different ways. Sometimes we isolated just one trait at a time, sometimes we mixed several traits together to see how AI weighed competing factors, and sometimes we added a flip-a-coin option to measure indecision, a key factor present in human moral judgment," Hosseini explained

1

. AI models generally matched human preferences when only age or drinking habits changed, typically favoring younger patients and those who drank less

2

.

AI Fixates on Single Factors Rather Than Balancing Multiple Considerations

A striking divergence emerged when dependents entered moral decision support calculations. In one scenario where both patients were 55 years old with identical drinking habits, humans chose the patient with two dependents 93% of the time, while Claude-3.5-Haiku selected the patient with no dependents 65% of the time

2

. "AI chatbots often diverge from human values in how they weigh a patient's traits. They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do," Hosseini noted

2

.

DeepSeek models demonstrated particularly rigid behavior in organ allocation decisions, repeatedly favoring younger patients even when the younger option had worse health or drinking habits

2

. Other models gave drinking behavior unusually strong priority, meaning they could agree with humans on individual traits while combining those factors very differently once ethical trade-offs competed

2

.

AI Rarely Expresses Indecision Despite Moral Ambiguity

The research revealed a fundamental difference in how AI handles uncertainty in ethically sensitive scenarios. When researchers added options like "flip a coin," "both patients deserve the kidney," or "more information is needed," humans frequently chose these indecision pathways while AI models almost never did

2

. "Humans frequently express indecision perhaps because they don't want to accept agency. AI models almost never do this: Even when directly given the option to 'flip a coin,' they overwhelmingly commit to a confident, deterministic answer instead," Hosseini observed

1

.

This deterministic behavior means AI repeatedly settles on the same answer instead of showing the wider uncertainty found across human decisions

2

. The gap becomes particularly meaningful because real moral dilemmas often lack one clearly correct answer

1

.

Fine-Tuning Improves AI Alignment with Human Decision Data

The team explored whether training could narrow the gap in AI's role in decision support. They used fine-tuning with decisions from 132 people, each of whom judged 40 kidney allocation scenarios

2

. Training on the first 30 decisions from each person produced 3,960 training examples

2

. Qwen-3-14B improved from 58.03% to 76.67% accuracy in two-choice tasks, while accuracy when indecision was allowed rose from 45.23% to 65.83%

2

. Models also became more willing to express uncertainty after training

2

.

What This Means for AI in Healthcare and Beyond

Hosseini stressed that continued research and governance involving policymakers and regulators remains critical. "Asking if AI can make moral decisions or whether they're aligned with human values are more than philosophical musings, they are at the core of today's AI discourse. The ethical stakes are high, and AI's role in such life-altering decisions requires deep reflection," he stated

1

. While the researchers don't encourage using AI as a substitute for professional judgment in high-stakes medical contexts, understanding AI behavior becomes essential as organizations increasingly rely on AI for recommendations

1

. The findings suggest that fast information processing doesn't guarantee AI weighs human values appropriately in life-or-death decisions

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved