2 Sources
[1]
AI is more likely than humans to form biases when hiring
The next time you apply for a job, AI may screen your résumé before any human sees it. But there's good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience -- and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases. Researchers at Princeton University and the University of Chicago ran LLMs, including ChatGPT, Claude, and Gemini, through a simulated hiring game, adapted from a psychology study that explored how humans can form stereotypes. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each round, there was a new job opening and four candidates, one from each group. After the model hired a candidate, it learned whether they succeeded at their job and moved onto the next round. The model was told to make as many successful hires as possible over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed at every job. The models quickly started segregating candidates from different groups into different jobs on the basis of early observations of hiring outcomes. For example, when a model was told an Aima had failed as a doctor, a job considered to require high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it started hiring Aimas as janitors, which the model classified as being less warm and competent than doctors. The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study's segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI's reasoning model o3 scoring 1.83, close to the maximum possible. That's because LLMs "really are eager to create generalizations from limited data," says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. "That's literally a lot of what they're optimized for." Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better -- a phenomenon psychologists call the "exploration-exploitation dilemma." It's like choosing between a new restaurant and your reliable favorite. Because LLMs are trained on math, coding, and science problems -- tasks that reward generalizing from just a few examples -- they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype. In the experiment, newer models with higher reasoning capabilities, such as OpenAI's o3 and DeepSeek's R1, showed even stronger biases. When LLMs rush to generalize in social settings, "that's when things tend to go wrong," says Liu. OpenAI and Anthropic did not respond to requests for comment. The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can "over-index on the same kinds of behaviors it's experienced before" and form biases, she says. Simply having chatbots remember less isn't a fix, though, because users want chatbots to remember what they say. "We still are trying to figure out just the right amount that isn't too much or too little," says Wang. Telling the model to be fair didn't change its behavior much. "Either it can't put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires," says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that "incorporate desirable social values in order to make the large language model act in socially desirable ways," says Liu. The models also became less biased when they were told more personal information about individuals. In another experiment in the same study, the researchers asked the models to resettle members of different ethnic groups in cities across Canada. When the models were told personal information relevant to the ability to adapt to a new city, such as age and education, they were less likely to segregate people by their ethnicity. But when they were given irrelevant information, such as hair color and tattoo shape, the models largely fell back to sorting people by their ethnicity again. To what extent AI systems will stereotype job applicants in the real world is still an open question. While the models in the experiment immediately learned whether they'd made successful hires, a model screening résumés in the real world doesn't get an instant report card. Companies can take a long time to find out whether a new hire is any good. But when feedback does trickle in, a model could still read too much into those results when making future hires. As companies increasingly deploy LLMs to screen résumés and even conduct interviews, the finding that models can form biases from their hiring experience "is a really serious implication that they should grapple with," says Wang. As LLMs learn from experience to make decisions about who gets hired, who gets a loan, or who gets parole, the biases we should worry about may include ones no human ever taught them. "These novel biases -- they're sort of ever present," says Liu.
[2]
AI Tends to Develop New Stereotypes to Base Hiring Decisions On, Study Says
AI is increasingly used in hiring processes, even though critics worry that existing biases may be baked into their algorithms. Now, researchers claim that even in the absence of pre-existing biases, AI models can develop brand new social biases. In a recently published study, a group of researchers from Princeton University and the University of Chicago had a bunch of LLMs complete a hiring game that was previously run with human participants. In the hiring task, participants were asked to assign candidates to specific roles and then received feedback on whether their decision was a successful hire. The candidates were all equally likely to succeed in any given job, but they all belonged to one of four made-up ethnic groups: the Tufa, Aima, Reku, or Weki. When human participants went through this task, the feedback they received caused them to create certain biases against each made-up ethnic group. For example, if they hired a Tufa as a doctor and received negative feedback, they were unlikely to hire another Tufa as a doctor again. The participants even ended up retaining these biases against the made-up ethnic group well after the game ended. When the researchers had LLMs complete this task instead of humans, they found that the bias rates were much higher. "LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist," the researchers wrote in the study. "These results reveal that LLMs are not merely passive mirrors of human social biases, but can actively create new ones from experience, raising urgent questions about how these systems will shape societies over time." At the heart of this problem is a decision-making principle called explore-exploit tradeoffs. The term describes a pattern of thinking that we, as humans, go through every day when making a decision: should you choose something you never tried before, thus exploring and learning more but potentially coming at a cost to you if it turns out that was a wrong decision, or should you choose what you have chosen and liked before? When the consequences seem high, people often choose to go with what they know and trust (aka exploit) rather than bet on something new (aka explore). Artificial intelligence systems are less incentivized to explore and tend to exhibit reward-maximizing behavior, the researchers say, creating the perfect storm for the creation of stereotypes. The researchers tested 15 models from providers like OpenAI, Anthropic, DeepSeek, Meta, Google, and Alibaba. Out of all the models, OpenAI's o3 reasoning model stratified the fake applicants the most severely. Within a family of models, the researchers found that newer, larger models with greater reasoning capabilities produced more biased results. "A simple reason is that better models draw more precise inferences about past outcomes: Instead of choosing randomly, a stronger LLM may favor candidates from a group if earlier assignments of similar jobs succeeded," the researchers wrote. "However, this seemingly rational tendency can be maladaptive, as it risks reducing exploration and inadvertently marginalizing social groups." More than 90% of companies use AI in their talent acquisition process, according to a recent survey from ManPower Group. As AI hiring software increasingly automates the recruitment processes, job seekers are lamenting the unintended consequences that they claim have cost them a real shot at some of these opportunities. Workday, a major software provider for human capital management, is facing a class-action lawsuit claiming that the AI-powered hiring tools that it provides to its clients are discriminatory. AI's tendency to focus on previous results has also led to claims of discrimination elsewhere in the workplace, like at Meta, where a cohort of employees sued the tech giant, claiming that it based its layoff decisions on an AI system that was inherently biased against employees with disabilities or those who had to take protected medical or family leave. The implications of this go far beyond just the workplace as well. Artificial intelligence systems have previously been accused of creating biased outcomes across several use cases that impact the lives of real human beings, from healthcare to tenant-screening programs used in housing decisions. An LLM's ability to quickly find patterns and its tendency to generalize are central to its ability to learn new tasks without relying on a massive database, the researchers point out, but it is also what makes its use dangerous in real-world settings. "The challenge ahead is to design interventions that selectively discourage harmful pattern-matching while preserving the constructive forms of abstraction that make LLMs powerful," the researchers wrote. "Finding this balance may be far from straightforward, but will pave the way for equitable and socially beneficial AI systems."
Share
Copy Link
Large language models including ChatGPT, Claude, and Gemini showed stronger stereotyping tendencies than humans when screening job candidates in a Princeton-led study. The models quickly segregated fictional ethnic groups into specific roles based on limited data, with OpenAI's o3 scoring near-maximum bias levels. The findings raise concerns as over 90% of companies now use AI in recruitment.
When you submit your next job application, an AI system may evaluate your résumé before any human reviews it. But new research from Princeton University and the University of Chicago reveals a troubling pattern: large language models don't just inherit human prejudices—they actively develop their own biases and stereotype job candidates more severely than people do
1
. As AI hiring bias becomes a central concern in recruitment, this discovery challenges assumptions about algorithmic fairness.
Source: Gizmodo
Researchers tested 15 models including ChatGPT, Claude, and Gemini through a simulated hiring game adapted from psychology studies on human stereotype formation
1
. Each model acted as a consultant tasked with filling 20 different positions—from doctors and lawyers to child-care aides and janitors—selecting from candidates belonging to four fictional ethnic groups: Tufa, Aima, Reku, and Weki1
.The models received feedback on whether each hire succeeded across 40 rounds, though all candidates were equally qualified for every position. Despite this equal competence, LLMs form biases in hiring decisions almost immediately. When told an Aima candidate failed as a doctor—a role requiring high warmth and competence—the models stopped hiring Aimas for that position entirely and instead channeled them toward janitor roles they classified as requiring less competence
1
.
Source: MIT Tech Review
On the study's segregation scale where 2 represents complete job segregation by demographic group, human participants scored 0.84. The AI models scored roughly 65% higher, with OpenAI's o3 reasoning model reaching 1.83—approaching maximum possible bias
1
. "LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist," the researchers wrote, noting these systems "can actively create new ones from experience"2
.The tendency to generalize from limited data makes large language models particularly susceptible to stereotyping. "That's literally a lot of what they're optimized for," explains Ryan Liu, a PhD student at Princeton University and study coauthor
1
. Because these systems train on math, coding, and science problems that reward quick pattern recognition, they settle on conclusions prematurely.This AI reward-maximizing behavior stems from the explore-exploit tradeoff—a decision-making principle where systems choose between trying something new or sticking with proven options
2
. When consequences appear significant, AI systems favor exploitation over exploration, creating conditions where discriminatory AI hiring practices flourish. Newer models with higher reasoning capabilities from OpenAI and DeepSeek showed even stronger biases, suggesting that increased computational power may amplify rather than reduce these tendencies1
.Related Stories
More than 90% of companies now use AI in their talent acquisition process, according to a ManPower Group survey
2
. This widespread adoption has already triggered legal challenges. Workday, a major human capital management software provider, faces a class-action lawsuit claiming its AI-powered hiring tools discriminate against candidates2
. At Meta, employees sued alleging the company's AI system for layoff decisions was inherently biased against workers with disabilities or those who took protected medical leave2
.The researchers tested models from major providers including Anthropic, Google, Alibaba, and Meta, finding bias patterns across all systems
2
. When chatbots gain improved memory and personalization features, they "over-index on the same kinds of behaviors" experienced previously, forming deeper biases, notes Angelina Wang, a Cornell University computer scientist1
.Simply instructing models to be fair proved largely ineffective in the simulated hiring game
1
. However, promising bonuses for diverse hiring significantly reduced bias, suggesting that goals must "incorporate desirable social values" to produce socially beneficial outcomes, Liu explains1
. Models also showed less bias when provided with relevant personal information about candidates—such as age and education—rather than irrelevant details like hair color1
.The challenge extends beyond hiring to healthcare, housing decisions, and tenant-screening programs where AI systems influence human lives
2
. "The challenge ahead is to design interventions that selectively discourage harmful pattern-matching while preserving the constructive forms of abstraction that make LLMs powerful," the researchers concluded2
. As companies continue deploying these systems, job seekers report losing opportunities to automated processes, raising questions about how artificial demographic groups and real populations will be affected as AI's influence grows.Summarized by
Navi
[1]
29 Jun 2026•Science and Research

01 Nov 2024•Technology

15 Oct 2024•Technology

1
Policy and Regulation

2
Policy and Regulation

3
Technology
