AI develops new stereotypes in hiring decisions, outpacing human bias by 65% in study

Reviewed byNidhi Govil

2 Sources

Share

Large language models including ChatGPT, Claude, and Gemini showed stronger stereotyping tendencies than humans when screening job candidates in a Princeton-led study. The models quickly segregated fictional ethnic groups into specific roles based on limited data, with OpenAI's o3 scoring near-maximum bias levels. The findings raise concerns as over 90% of companies now use AI in recruitment.

Large Language Models Create Their Own Biases in Hiring

When you submit your next job application, an AI system may evaluate your résumé before any human reviews it. But new research from Princeton University and the University of Chicago reveals a troubling pattern: large language models don't just inherit human prejudices—they actively develop their own biases and stereotype job candidates more severely than people do

1

. As AI hiring bias becomes a central concern in recruitment, this discovery challenges assumptions about algorithmic fairness.

Source: Gizmodo

Source: Gizmodo

Researchers tested 15 models including ChatGPT, Claude, and Gemini through a simulated hiring game adapted from psychology studies on human stereotype formation

1

. Each model acted as a consultant tasked with filling 20 different positions—from doctors and lawyers to child-care aides and janitors—selecting from candidates belonging to four fictional ethnic groups: Tufa, Aima, Reku, and Weki

1

.

AI Develops New Stereotypes from Experience

The models received feedback on whether each hire succeeded across 40 rounds, though all candidates were equally qualified for every position. Despite this equal competence, LLMs form biases in hiring decisions almost immediately. When told an Aima candidate failed as a doctor—a role requiring high warmth and competence—the models stopped hiring Aimas for that position entirely and instead channeled them toward janitor roles they classified as requiring less competence

1

.

Source: MIT Tech Review

Source: MIT Tech Review

On the study's segregation scale where 2 represents complete job segregation by demographic group, human participants scored 0.84. The AI models scored roughly 65% higher, with OpenAI's o3 reasoning model reaching 1.83—approaching maximum possible bias

1

. "LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist," the researchers wrote, noting these systems "can actively create new ones from experience"

2

.

Why AI's Role in Hiring Amplifies Stereotyping

The tendency to generalize from limited data makes large language models particularly susceptible to stereotyping. "That's literally a lot of what they're optimized for," explains Ryan Liu, a PhD student at Princeton University and study coauthor

1

. Because these systems train on math, coding, and science problems that reward quick pattern recognition, they settle on conclusions prematurely.

This AI reward-maximizing behavior stems from the explore-exploit tradeoff—a decision-making principle where systems choose between trying something new or sticking with proven options

2

. When consequences appear significant, AI systems favor exploitation over exploration, creating conditions where discriminatory AI hiring practices flourish. Newer models with higher reasoning capabilities from OpenAI and DeepSeek showed even stronger biases, suggesting that increased computational power may amplify rather than reduce these tendencies

1

.

Real-World Impact on Recruitment Practices

More than 90% of companies now use AI in their talent acquisition process, according to a ManPower Group survey

2

. This widespread adoption has already triggered legal challenges. Workday, a major human capital management software provider, faces a class-action lawsuit claiming its AI-powered hiring tools discriminate against candidates

2

. At Meta, employees sued alleging the company's AI system for layoff decisions was inherently biased against workers with disabilities or those who took protected medical leave

2

.

The researchers tested models from major providers including Anthropic, Google, Alibaba, and Meta, finding bias patterns across all systems

2

. When chatbots gain improved memory and personalization features, they "over-index on the same kinds of behaviors" experienced previously, forming deeper biases, notes Angelina Wang, a Cornell University computer scientist

1

.

Potential Solutions and Future Implications

Simply instructing models to be fair proved largely ineffective in the simulated hiring game

1

. However, promising bonuses for diverse hiring significantly reduced bias, suggesting that goals must "incorporate desirable social values" to produce socially beneficial outcomes, Liu explains

1

. Models also showed less bias when provided with relevant personal information about candidates—such as age and education—rather than irrelevant details like hair color

1

.

The challenge extends beyond hiring to healthcare, housing decisions, and tenant-screening programs where AI systems influence human lives

2

. "The challenge ahead is to design interventions that selectively discourage harmful pattern-matching while preserving the constructive forms of abstraction that make LLMs powerful," the researchers concluded

2

. As companies continue deploying these systems, job seekers report losing opportunities to automated processes, raising questions about how artificial demographic groups and real populations will be affected as AI's influence grows.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved