Skip to main content

Search...

AI Bias: People Adopt Biases from Algorithms

AI bias spreads because people adopt it from algorithms even faster than from humans and rarely notice. Comparative data keeps AI decisions in check.

• • Updated: • 9 min read
Cover of the expert talk on 'AI Bias: People Adopt Biases from Algorithms' with Sam Goetjes and Richard Seidl.

AI bias is the systematic distortion in decisions made by artificial intelligence when the training data reflects social inequalities. People rarely spot it, because they trust AI recommendations much as they trust human judgment. Worse still, when a biased recommendation comes from an algorithm, people adopt the bias faster and more strongly than when it comes from another person.

Key Takeaways

  • Biased recommendations from an algorithm rub off more strongly than biased recommendations from a person: study participants adopted a gender bias faster and more consistently when it came from an automated system.
  • Trust in algorithms follows the same psychological patterns as trust in people, so a flawed AI judgment gets accepted just like the judgment of an experienced colleague.
  • Most participants did not notice a gender bias in the recommendations across 16 consecutive hiring decisions, even though names perceived as female were consistently rated lower at the same level of competence.
  • Without ongoing monitoring against comparison data, a vicious circle forms: people confirm the biased AI judgments, the AI is trained further on those judgments and the bias grows.

AI Recommendations Are Not Neutral Just Because a Machine Made Them

People tend to trust an algorithm’s recommendation more than a person’s, even when both make the same mistake. This tendency to over-rely on automated suggestions is known as automation bias, and an online study by Sam Goetjes shows it clearly. More than 330 working adults took part and were asked to assess job applications for a hiring decision.

The setup was simple. For each application, participants gave a recommendation from 0 to 100 percent. They then saw a supposed expert recommendation, either from an HR employee or from an automated decision system, and were allowed to adjust their own rating. They did this 16 times in a row.

That is where the key finding appeared. When the biased recommendation came from an algorithm, participants took on the bias faster and more strongly than when a person had made it. The common assumption that people start out skeptical of machines did not hold up.

Why Automation Bias Makes People Treat an Algorithm Like a Colleague

People build trust in an algorithm on the same factors as trust in a person. That is the first solid result of the study, and everything else builds on it.

Sam Goetjes took a trust model originally developed for organizations and relationships between people and tested it against interaction with an algorithm. The model held. Not just barely significant, but about as strong as it is for people.

This matches the computer-as-social-actor hypothesis: people treat an algorithm or an AI socially, like someone they are dealing with, and build trust accordingly. In practice, the way you come to trust a new colleague also works on a decision system. Weaknesses included.

Why People Rarely Notice Bias in AI Decisions

Most people do not notice when a system judges in a systematically biased way. In the study, names perceived as female were consistently rated lower than names perceived as male, at the same level of competence. The vast majority of participants did not pick up on the pattern across all 16 applications.

The comparison was carefully built. Eight names perceived as female and eight perceived as male, paired at identical competence levels, validated in an earlier evaluation study. Annika had exactly the same skills as Thomas. Only the gender of the name was different.

There was also a time effect. From the third or fourth application on, participants moved closer and closer to the recommendation they were shown. Once you know a task and slip into an automated routine, you check less and start thinking: maybe the system is right. The more routine the task, the easier it is to sway you.

The Vicious Circle of Human and Machine Bias

The real risk appears when human bias and machine bias confirm each other. You bring your own prejudices with you. If an automated system happens to rate a person the same way, you feel confirmed, even though the lower rating may have nothing to do with competence.

Sam Goetjes is open about her own experience: she noticed a bias in herself just from the photo on an application, even though she had studied the subject. Knowing about bias does not automatically protect you from it.

That self-confirmation closes the loop. The biased rating flows back into the system, no corrected data comes in, and the AI never improves because nobody sees the error. And gender bias does not affect a small minority. It affects half the population.

“It starts with one group of people, maybe with one characteristic. But when that grows, it’s not a minority anymore. It’s half the population that is being disadvantaged.”

(Sam Goetjes)

How Bias Arises in an AI Recruiting Tool

An AI looks for patterns, not for fairness. When the first screening of applications is handed to a system, it looks at what separates successful employees from less successful ones and derives predictions from that.

Some of these patterns seem reasonable, such as experience in the field as a sign of later performance. Others are pure correlation with no real connection. The example from the study: if 75 percent of a successful group happen to enjoy playing soccer, the system may link soccer to job performance.

Distortions like this depend heavily on the training data. The widespread belief that an AI is more objective than a person by default does not hold. It is only as good as the data it was trained on.

Piloting and Monitoring Beat Blind Trust

Anyone introducing an AI decision system needs comparison data, not the hope that errors will get noticed. As the study shows, they do not. The most effective countermeasure is a setup that keeps the system testable at all times.

A practical approach for test managers:

  • Run both in parallel during the pilot: Keep the old process, manual or otherwise, running next to the new AI system for a while and compare the results over time.
  • Evaluate with more than one source: Do not just point to the training data. Cross-check with other tools so you have evidence of quality.
  • Monitor over time: Especially when a system keeps evolving and turns into a black box, check regularly whether it is drifting in directions nobody wants.

Comparison data works better than people doing the comparing, because people can be swayed themselves. A vendor cannot trace in detail what happens inside the box anyway. But the vendor can trace which data was used for training and measure over the long run where the system is heading.

Awareness Is the Lever Quality Assurance Needs

The first step against biased AI decisions is to stop treating them as neutral. Software quickly comes across as an objective, factual tool. People sit back and expect it to be fine. That complacency is the problem.

For testing and quality management, the focus shifts. Training data quality and whether the AI works at all still matter, but those questions are familiar. A second level comes on top: how do people react to the recommendations, and do they adopt biases without noticing?

Behind the business interest that comes first for most companies are real people who are already being disadvantaged by these decisions today. If you start early, invest the effort once and build the system to be observable, you keep bias from growing unnoticed. The effort does not have to be large. Awareness is where it starts.

Frequently Asked Questions

Do people trust recommendations from software less than those from a colleague?

No, the opposite was found to be true. In an online study involving over 330 working professionals, respondents were more likely to accept a biased recommendation, and did so more readily, when it came from an automated decision-making system than when it came from an HR employee. The assumption that people are initially skeptical of machines did not hold up to scrutiny.

Do different principles apply to trust in algorithms than to trust between people?

The same factors apply. A trust model originally developed for organizations and interpersonal relationships also worked in the study for interactions with an algorithm, and with comparable strength. This aligns with the “computer-as-social-actor” hypothesis: People treat a system socially as they would a counterpart, with all the weaknesses that entails.

Does AI judge more objectively than a human because it only evaluates data?

No. AI looks for correlations, not fairness, and it is only as good as its training data. Some of the patterns it finds seem plausible, such as field experience as an indicator of performance. Others are mere correlations: If, by chance, 75 percent of a successful group enjoys playing soccer, the system may link soccer to job performance.

Does the likelihood of being influenced increase the more often you use a recommendation system?

Yes. In the study, starting with the third or fourth application, participants increasingly aligned with the system’s recommendation. When you’re familiar with a task and slip into an automated routine, you check things less carefully and are more likely to assume the system is right. Routine thus reduces the depth of your own scrutiny rather than increasing it.

Does knowledge about biases protect you from following a biased system?

No. The author of the study uses her own example to illustrate that she became aware of her own bias just by looking at a photo in an application, even though she had studied the subject. It becomes critical when human and machine biases happen to align: then the evaluation feels validated, without competence playing a role.

How can we determine whether an AI decision-making system makes systematically biased judgments?

By using comparative data rather than hoping that errors will become apparent on their own. Three approaches: run the old process in parallel with the new system for a while during the pilot phase and compare the results; cross-check the quality using additional tools rather than relying solely on the training data; and continuously monitor whether the system is drifting in undesirable directions.

Is it enough to have humans act as a check on AI recommendations?

Comparative data is better suited for this purpose than human comparators, because humans themselves are susceptible to influence. A provider cannot, in any case, track in detail what happens inside a black box. However, they can track what data was used for training and measure, over the long term, where the system is heading. This adds a second layer to testing and quality management: user reactions to the recommendations.

Share this page

Related Posts