Column

When interviewers read the criteria differently, the same answer earns different scores

Even with the same criteria, if people read them differently, a score depends on who gave it. Scores that carry each interviewer’s tendencies cannot simply be pooled and compared, or checked against outcomes after hiring. This piece covers why gaps between human and AI scores help you find where criteria are read differently, and why fixing them matters.

The problem

Criteria that are read differently cause three problems

People can read the same criteria in different ways. Criteria like that cause three problems.

01

Who interviewed decides who passes

The same answer can gain or lose points depending on the interviewer. Near the pass line, that difference changes the result.

02

Pooled scores cannot be compared

When the interviewer or the period changes, so do the interviewer’s tendencies, and the same score means something different. Putting last year’s scores next to this year’s is not a comparison.

03

Outcomes after hiring cannot confirm anything

Even if you match selection scores against reviews after joining, you cannot tell whether the criteria were right or whether interviewers simply read them differently.

How to check against outcomes after hiring is covered in “Whether an interview evaluation was right only shows after the hire”.

Findings

A rating says more about the rater than about the person rated

In 2000, Scullen and colleagues published an analysis of performance ratings after joining. In two data sets of 2,350 and 2,142 managers, each manager was rated by seven people: two bosses, two peers, two subordinates and themselves. Of the variation in ratings, 62% and 53% was explained by each rater’s own tendencies, and only 21% and 25% by how the person rated actually performed. Put another way, more than half the difference in scores came from who did the rating, and only about a quarter from differences in how people actually worked. The scores reflected the raters more than the people being rated.

Interview ratings also differ between interviewers. In a 2013 analysis by Huffcutt and colleagues, the correlation between interviewers’ ratings averaged 0.74 when they sat in the same interview and 0.44 when each interviewed separately. A correlation runs from -1 to 1; the closer to 1, the more the two ratings agree. Put another way, ratings tend to diverge when interviewers see candidates separately and tend to agree when they sit in the same interview. Even so, hearing the same answers side by side does not make two ratings match completely.

The same pattern has been reported in Japanese graduate hiring. In a 2016 study, Suzuki examined ratings from final-round interviews for new graduates at a Japanese software company. Two interviewers, a board director and a division head, each rated candidates from 0 to 8 on five criteria: attitude as a working adult, interpersonal skills, personality fit, sincerity, and potential to deliver results. The rank correlation between the two was highest for personality fit at 0.41 and between 0.23 and 0.25 for attitude, interpersonal skills and sincerity. For potential to deliver results it was 0.04, showing no relationship. In other words, even when rating the same candidates, the two interviewers agreed only weakly on most criteria, and barely at all on the one that looks ahead to performance after joining.

In a 2010 conference paper, Imashiro reasoned that Japanese graduate hiring, which recruits people into the company rather than into a defined role, leaves more room for interviewers to differ than Western hiring for a specific job. Surveying interviewers, she found that those with long sales experience valued some candidate traits differently from those without.

Things the criteria do not mention also find their way into ratings. A 2009 meta-analysis by Barrick and colleagues examined how candidates’ appearance, impression management, and verbal and non-verbal behaviour relate to interviewer ratings. The effect was larger in unstructured interviews, and the authors concluded that what you see in the interview may not be what you get on the job.

Sources: Scullen, S. E., Mount, M. K., & Goff, M. (2000). Understanding the latent structure of job performance ratings. Journal of Applied Psychology, 85(6), 956–970. https://doi.org/10.1037/0021-9010.85.6.956 / Huffcutt, A. I., Culbertson, S. S., & Weyhrauch, W. S. (2013). Employment interview reliability: New meta-analytic estimates by structure and format. International Journal of Selection and Assessment, 21(3), 264–276. https://doi.org/10.1111/ijsa.12036 / Barrick, M. R., Shaffer, J. A., & DeGrassi, S. W. (2009). What you see may not be what you get: Relationships among self-presentation tactics and ratings of interview and job performance. Journal of Applied Psychology, 94(6), 1394–1411. https://doi.org/10.1037/a0016532 / Suzuki, T. (2016). Inter-rater reliability of recruiting interviews: Experimental study on evaluation components [in Japanese]. Japan Journal of Human Resource Management, 17(1), 69–91. https://doi.org/10.24592/jshrm.17.1_69 / Imashiro, S. (2010). Why do interviewers differ in what they look for? Effects of interviewers’ job experience [in Japanese]. Paper presented at the 2010 annual conference of the Japanese Group Dynamics Association. https://www.recruit-ms.co.jp/research/thesis/pdf/2010jgda.pdf

Why it goes unnoticed

Differences in how people read criteria are hard to see

Differences between human ratings get mixed up with each interviewer’s leniency or strictness, so the part that comes from reading the criteria differently is hard to separate out. Few candidates are rated by several interviewers. And one interviewer’s scores alone do not show how that interviewer reads the criteria.

What AI brings

AI reads the criteria as written, which makes it a reference point for the gaps

AI scores using only the wording of the criteria and what was said in the record. It does not use manner or appearance, and it does not bring in standards the criteria do not state. When people and AI score against the same criteria, the AI score, being the criteria applied as written, becomes a reference point for comparing human scores. Generative AI scores are not always exactly the same each time, but clear answers barely move.

How it happens (an illustrative example)

Took initiative (L5: action, who was involved and outcome are all described / L3: describes an action, but not the outcome)

At my last job, when enquiries suddenly rose, I put the handling steps together myself and shared them with the team.

AI score

3

Describes an action but not the outcome, so it fits L3.

Human score

4

No figures were given, but from the way they spoke, they seemed to have got results.

The human reads it one level higher, on a guess the criterion does not mention.

Reading the gaps

Read the gaps by how they show up

When human and AI scores differ, the way the gap shows up tells you what is going on.

01

Humans score higher or lower across the board

If human scores are higher, or lower, on every criterion, the issue is the interviewer’s leniency or strictness, not the criteria.

02

Gaps on particular criteria only

If gaps keep appearing on some criteria only, their descriptions may allow different readings, which makes them candidates for review.

03

Gaps for particular candidates only

If gaps appear for some candidates only, people may be picking up things the record does not capture, such as manner or atmosphere. That is material for checking whether the grounds for the rating are sound.

Why aligning matters

When readings align, scores build up into an asset

Fix the gaps starting from the criteria as written. If a human judgement has a reason behind it, write that standard into the criteria. If it does not, bring the human scoring in line with the criteria. Once readings align, you can do the following.

01

The same yardstick, whoever interviews

The pass line keeps its meaning when interviewers change. New interviewers can score the same way by reading the same criteria.

02

Scores can be compared across periods

The same score carries the same meaning, so the more scores you gather, the more you have to judge from.

03

Outcomes after hiring can test the criteria

The smaller the differences in how interviewers read the criteria, the easier it becomes to check whether the criteria themselves were right.

Rewritten criteria can be tested on past interviews. See “Rewrite your criteria, then test them on past interviews”.

What makes the comparison possible

Because people and AI use the same criteria, the gaps can be compared

Human and AI scores can be compared only when the human scoring sheet and the AI criteria are written in the same words. Where the service fixes the criteria, the AI criteria and your own scoring sheet end up in different words, and cannot be set side by side on one yardstick.

In Proomi you write the criteria in your own words. People can score against the same criteria, so human and AI scores can be viewed side by side, criterion by criterion.

What to compare

Three checks for whether human and AI scores can be compared

When comparing AI interview services, these three checks tell you whether you can bring human and AI scoring into line.

01

Can people score against the same criteria as the AI?

If the criteria differ, the gap between human and AI scores cannot be compared.

02

Is the reasoning behind AI scores kept?

With the reasoning, you can see which reading of which remark caused the gap.

03

Can human and AI scores be viewed side by side per criterion?

A total score alone does not show which criterion the gap is on.

Q&A

Frequently asked questions

Which is right, the human score or the AI score?

Neither is right by default. The AI score is the criteria applied as written. When it differs from the human score, your company decides where the standard should sit, and that decision goes into the wording of the criteria.

Should gaps be eliminated?

There is no need to get them to zero. A gap that appears for particular candidates only may be a sign that people are noticing something the record does not capture, which may be worth a person checking. What you want to align is the reading of the criteria.

How many interviews before a pattern shows?

One gap at a time cannot tell you whether it is chance. Look across a good number of interviews, criterion by criterion, at whether the gaps keep pointing the same way.

Read next

Once you fix where criteria are read differently, test them on past interviews

Once you rewrite the parts that are read differently, you can test them on past interviews. Checking whether scores tell candidates apart at the same time makes it easier to decide which criteria to fix first.

Rescoring

Rewrite your criteria, then test them on past interviews →

Re-score past interviews with rewritten criteria and see how the scores move.

Distribution

More criteria do not mean more evidence →

Why scores bunch up, checked against public civil-service rating data.

Related pages

How it applies, by situation

Differences in reading show up the same way in multi-site hiring and in interviewer training.

Standards

Standardised evaluation →

Set the criteria first and narrow the spread between interviewers.

Sites

Multi-site hiring →

Keep one hiring standard as your sites multiply.

Practice

Interview practice →

Practise interviewing as often as you like, with AI as the candidate.

Show us your evaluation criteria as they are

We will look with you at which of your criteria are most likely to be read differently by different people. Free, about 30 minutes.