AI tends to develop new stereotypes on which to base hiring decisions, study finds



AI is increasingly being used in hiring processes, although critics worry that existing biases could be built into its algorithms. Now, researchers say that even in the absence of pre-existing biases, AI models can develop entirely new social biases.

In a recently published article studyA group of researchers from Princeton University and the University of Chicago had a group of LLMs complete a recruiting game that was previously run with human participants. In the hiring task, participants were asked to assign candidates to specific roles and then received feedback on whether their decision was a successful hire. All candidates had the same chance of succeeding in any position, but they all belonged to one of four invented ethnic groups: the Tufa, the Aima, the Reku or the Weki. When human participants performed this task, the feedback they received caused them to create certain prejudices against each ethnic group created. For example, if they hired a Tufa as a doctor and received negative feedback, they were unlikely to hire another Tufa as a doctor again. The participants even ended up maintaining these prejudices against the invented ethnic group long after the game ended. When the researchers had LLMs complete this task instead of humans, they found that the rates of bias were much higher.

“LLMs can spontaneously develop new social biases about artificial demographic groups even when no inherent differences exist,” the researchers wrote in the study. “These results reveal that LLMs are not merely passive mirrors of human social prejudices, but can actively create new ones from experience, raising urgent questions about how these systems will shape societies over time.”

At the heart of this problem is a decision-making principle called the exploration-exploitation balance. The term describes a pattern of thinking that we, as humans, go through every day when making a decision: should you choose something you’ve never tried before, exploring and learning more, but potentially at a cost to yourself if it turns out it was the wrong decision, or should you choose what you chose and liked before? When the consequences seem high, people often opt for what they know and trust (aka exploit) rather than go for something new (aka explore). AI systems are less incentivized to explore and tend to exhibit reward-maximizing behavior, researchers say, creating the perfect storm for stereotyping.

The researchers tested 15 models from vendors such as OpenAI, Anthropic, DeepSeek, Meta, Google and Alibaba. Of all models, O3 by OpenAI The reasoning model stratified false applicants more severely. Within a family of models, the researchers found that newer, larger models with greater reasoning capabilities produced more biased results.

“A simple reason is that better models draw more accurate inferences about past results: instead of
“Rather than choosing at random, a stronger LLM may favor candidates from a pool if previous similar job assignments were successful,” the researchers wrote. “However, this seemingly rational tendency can be maladaptive, as it risks reducing exploration and inadvertently marginalizing social groups.”

More than 90% of companies use AI in their talent acquisition process, according to a recent survey of the ManPower Group. As AI recruiting software increasingly automates hiring processes, job seekers lament the unintended consequences that they say have cost them a real shot at some of these opportunities. Workday, a major provider of human capital management software, faces a class action lawsuit alleging that the AI-based recruiting tools it provides to its clients are discriminatory. AI’s tendency to focus on past results has also led to complaints of discrimination elsewhere in the workplace, such as in Goalwhere a group of employees sued the tech giant, alleging that it based its firing decisions on an artificial intelligence system that was inherently biased against employees with disabilities or those who had to take protected medical or family leave.

The implications of this go far beyond the workplace. AI systems have previously been accused of creating biased results in several use cases that impact the lives of real human beings, from health care to tenant selection programs used in housing decisions.

An LLM’s ability to find patterns quickly and its tendency to generalize are critical to its ability to learn new tasks without relying on a massive database, the researchers note, but it’s also what makes its use dangerous in real-world settings.

“The challenge ahead is to design interventions that selectively discourage harmful pattern matching while preserving the constructive forms of abstraction that make LLMs powerful,” the researchers wrote. “Finding this balance may not be easy, but it will pave the way for equitable and socially beneficial AI systems.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *