When Performance Ratings Become Data: Finding Bias in Performance Appraisals
When Performance Ratings Become Data: Finding Bias in Performance Appraisals
Introduction
Imagine two employees doing almost the same job and getting very similar results. Both complete their work on time, both work with clients, and both meet their targets. But when their performance reviews come in, one gets a very high rating and the other gets an average rating. So, what happened? Did one person actually perform better? Or did the manager see the two employees differently?
Performance appraisals are supposed to help companies understand how employees are performing. They are also used for promotions, salary increases, bonuses, training and career growth. So, a performance rating can affect a lot more than just a number on an HR form. But the problem is that performance ratings are often given by people, and people can have biases. Sometimes the manager may not even know that the bias is affecting the rating.
So, this creates an interesting HR problem. If performance ratings are supposed to be based on employee performance, can HR use data to check whether the ratings are actually fair?
The Problem
The basic problem is simple. Managers have to judge employee performance, but their judgment may not always be completely objective.
For example, imagine a manager describes a male employee as "confident and a strong leader", while a female employee who behaves in the same way is described as "too aggressive". So, the actual behaviour may be similar, but the words used to describe it are different.
Another employee may receive a comment like "very hardworking and helpful", while another employee gets feedback such as "excellent technical skills and strong business impact". Both comments sound positive, but they are not really talking about the same things.
This becomes a bigger problem when these comments are connected to ratings. So, if one group is regularly praised for leadership and business results while another group is judged more on personality and behaviour, the ratings may not be measuring the same thing.
The problem can become even bigger when performance ratings are used to decide who gets a promotion or a higher bonus. So, a small amount of bias in a performance review can later become a bigger career problem.
Analysis of the Problem
The first thing HR needs to understand is that bias does not always mean that a manager is deliberately trying to treat someone unfairly. A manager may genuinely believe that they are being fair. The problem can come from small judgments that happen without the manager noticing them.
For example, a manager may remember one mistake made by an employee and allow that mistake to affect the entire review. Another manager may give a high rating because one employee recently completed an impressive project. So, two managers can look at similar performance and reach different conclusions.
There can also be a problem with the language used in reviews. If managers write very different types of comments for different groups of employees, HR can actually study those comments. So, performance reviews are not only documents for employees. They can also become HR data.
This is where HR Analytics becomes useful. Instead of HR simply asking, "Are our managers biased?", HR can ask more specific questions.
Are men and women receiving different average ratings? Are different racial or demographic groups receiving different types of feedback? Are some employees receiving more comments about personality while others receive more comments about technical skills? Are some groups being praised for leadership more often? Are these differences connected with promotions or rewards?
So, the performance review itself becomes a source of data.
The Real-Life Case: A Midsized U.S. Law Firm
A very useful real-life case comes from a midsized law firm in the United States. The firm's Diversity and Inclusion director had already looked at a small sample of performance reviews and noticed some possible signs of bias. So, instead of simply assuming that there was a problem, the firm decided to investigate it using data.
The firm worked with the Center for WorkLife Law and conducted an audit of its performance evaluations. At first, most of the reviews looked normal. But when the researchers looked more closely at the data, they found clear differences based on race and gender.
One particularly interesting finding was about leadership. Only 9.5% of people of color received mentions of leadership in their performance evaluations, while the figure was more than 70 percentage points higher for white women. Leadership mentions were also linked with higher competency ratings in the following year. So, the words written in a review could potentially affect how an employee was viewed later.
The researchers also found other patterns. For example, people of color and white women were more likely to have personality-related comments in their reviews. This matters because personality is not always the same thing as job performance. So, an employee could be judged on how they were perceived personally instead of what they actually achieved.
The firm therefore had a very interesting HR problem. The performance review system looked normal on the surface, but the data showed that employees from different groups were not always being evaluated in the same way.
HR Theories, Frameworks, Models and Concepts Reflected in the Problem
One important concept here is the Halo Effect. This happens when one strong quality of an employee influences the manager's entire opinion of that person. So, if an employee is very good at one thing, the manager may assume that they are excellent at everything else too.
The opposite can also happen through the Horns Effect. So, one mistake or negative quality can influence the manager's view of the employee's entire performance.
Another issue is stereotyping. This happens when people make assumptions about someone based on the group they belong to. For example, a manager may unconsciously expect a man to be more assertive or a woman to be more cooperative. So, the same behaviour can sometimes be interpreted differently.
There is also anchoring. This happens when people rely too much on an existing piece of information when making a decision. In performance management, a manager may start with an employee's previous rating and then judge the new performance around that number instead of looking at the current year's evidence carefully. Recent research using data from a multinational company found that managers' ratings could be influenced by employees' previous ratings and by the self-ratings they saw.
Another useful concept is inter-rater reliability. This basically asks whether different managers would give similar ratings when judging the same type of performance. If one manager gives almost everyone a 4 or 5 and another manager gives almost everyone a 2 or 3, the rating system may have a consistency problem.
There is also criterion validity. This asks whether the thing being measured is actually related to the job performance we want to measure. So, if an employee is being rated on "being friendly" when friendliness is not really important for their job, HR should question whether that is a useful performance criterion.
These concepts show that the problem is not simply "managers are biased". The actual performance management system can also make bias easier to happen.
Solution Using HR Theories, Frameworks, Models and Concepts
The first solution is to make performance criteria much clearer. Employees should know exactly what they are being judged on. So, instead of saying "good leadership", the company could define what leadership means for that particular role. It could include things like developing employees, handling difficult situations, meeting team goals and making good decisions.
The second solution is to ask managers to support ratings with evidence. So, instead of writing "excellent employee", the manager should explain what the employee actually did. For example, they could mention a project completed, a customer problem solved, a sales target achieved or a process improved.
This makes the review more fact-based and makes it harder for one general opinion to control the whole rating.
The third solution is to separate performance, potential and personality. So, if a manager thinks an employee needs to improve their communication style, that should not automatically reduce the person's technical performance rating. These things should be looked at separately.
Another solution is to use behaviourally specific feedback. Instead of saying "not a team player", the manager could explain the actual behaviour, such as "did not share project information with the team during the March project". This gives the employee something specific that they can understand and improve.
HR can also use text analytics to study the language used in performance reviews. This is where the topic becomes especially interesting for HR Analytics. Technology can look for patterns in words and feedback. For example, it can check whether women are receiving more personality-based comments while men receive more comments about achievements or technical skills. Deloitte has also described how text analysis and machine learning can be used to identify differences in the language used in feedback.
HR can then compare ratings and comments across different employee groups. So, instead of looking at one employee and saying, "This review seems unfair", HR can look at hundreds of reviews and ask, "Are we seeing the same pattern again and again?"
Finally, HR should measure the results after changing the system. This is important because simply conducting bias training does not prove that bias has been reduced. HR needs to compare the data before and after the intervention.
The Actual Solution Implemented in the Real-Life Case
The law firm did not simply tell managers, "Don't be biased." Instead, it used what the researchers called bias interrupters. The basic idea was to make small changes to the performance management process so that bias had fewer opportunities to affect the final result.
The firm made two important changes for the following year's reviews. First, it redesigned the performance evaluation form. Instead of relying heavily on broad categories, the new form focused more clearly on competencies and required managers to provide evidence to support their ratings. Second, managers received a short workshop that explained common patterns of bias and showed them how to use the new form.
Then, and this is the part I find most interesting, the firm checked the data again.
After one year, the researchers reported a sharp improvement. Women and people of color received more constructive feedback, and differences in the length and complexity of evaluations became much smaller. In the first year, white men's evaluations were longer and more complex. In the following year, the language and word counts were much more similar across groups.
The idea was not to remove performance reviews. Instead, the goal was to make the existing system more objective and evidence-based. This is important because simply removing performance ratings does not automatically remove bias. Bias can still appear in informal feedback, promotion decisions and other parts of HR.
Conclusion
Performance ratings are numbers, but the process behind those numbers is very human.
So, if a company only looks at the final rating, it may miss what happened before that number was created. A rating can be affected by the manager's memory, the words used in the review, the criteria being used, previous ratings and even ideas about how a "good employee" should behave.
This is why I think performance management is becoming an interesting area for HR Analytics. HR does not have to guess whether bias exists. It can actually look for patterns in the data.
So, the question should not only be:
"Who got a high rating?"
It should also be:
"Why did they get that rating?"
And then:
"Are we using the same standards for everyone?"
The real-life law firm case shows that even small changes can make a difference when they are based on actual data. So, maybe the future of fair performance management is not about removing the human side completely. It is about using data to help humans make better and more consistent decisions.
Because at the end of the day, a performance rating should tell us about the work an employee did, and not accidentally tell us more about who the manager thinks the employee is.
```
Comments