2 Sources
[1]
3 Questions: How to help students recognize potential bias in their AI datasets
Caption: Courses on developing AI models for health care need to focus more on teaching how to identify and address bias. Every year, thousands of students take courses that teach them how to deploy artificial intelligence models that can help doctors diagnose disease and determine appropriate
[2]
Q&A: How to help students recognize potential bias in their AI datasets
Every year, thousands of students take courses that teach them how to deploy artificial intelligence models that can help doctors diagnose disease and determine appropriate treatments. However, many of these courses omit a key element: training students to detect flaws in the training data used to
Share
Copy Link
MIT researcher Leo Anthony Celi highlights the need for AI courses to focus on identifying and addressing bias in healthcare datasets, emphasizing the importance of critical thinking and diverse perspectives in AI education.
Leo Anthony Celi, a senior research scientist at MIT's Institute for Medical Engineering and Science, has raised concerns about the lack of focus on bias detection in AI healthcare courses. In a recent paper, Celi highlights the importance of teaching students to identify and address potential biases in the datasets used to develop AI models for healthcare applications
1
2
.
Source: MIT
Celi points out that many AI models in healthcare are trained primarily on data from white males, leading to poor performance when applied to other demographic groups. He cites an example of pulse oximeters overestimating oxygen levels in people of color due to insufficient diversity in clinical trials
1
2
.The researcher also notes that medical devices and equipment are typically optimized for healthy young males, potentially compromising their effectiveness for other patient groups, such as elderly women with heart conditions
1
2
.An analysis of 11 AI courses revealed that only five included sections on dataset bias, with just two offering significant discussions on the topic. Celi argues that many courses focus primarily on model building and data visualization, neglecting the critical aspect of data quality and bias
1
2
.Celi suggests that at least 50% of course content should be dedicated to understanding the data, emphasizing that modeling becomes straightforward once the data is properly understood. He recommends incorporating a checklist of questions for students to evaluate data sources, including:
For instance, when working with ICU databases, students should consider potential sampling biases, such as the underrepresentation of minority patients who may not have equal access to ICU care
1
2
.
Source: Medical Xpress
To mitigate the effects of missing data resulting from social determinants of health and provider implicit biases, Celi and his team are exploring the development of a transformer model for numeric electronic health record data. This approach aims to model the underlying relationships between laboratory tests, vital signs, and treatments
1
2
.Related Stories
Since 2014, the MIT Critical Data consortium has been organizing datathons worldwide, bringing together healthcare professionals and data scientists to examine health and disease in local contexts. Celi emphasizes that critical thinking skills are best developed by bringing together people with diverse backgrounds and experiences
1
2
.As AI continues to play an increasingly important role in healthcare, addressing bias in datasets and AI models is crucial. By improving AI education to focus on critical thinking, data quality, and diverse perspectives, the healthcare industry can work towards developing more equitable and effective AI solutions for all patient populations.
Summarized by
Navi
[2]
05 Sept 2025•Health

12 Dec 2024•Science and Research

19 Dec 2024•Health

1
Technology

2
Technology

3
Policy and Regulation
