Column

More criteria do not mean more evidence

Six criteria do not give you six pieces of evidence. Using public data on national civil servants’ performance ratings, we show why scores bunch up and what to check when you choose an AI interview service.

Public data

In the national civil service’s performance ratings, the lowest grades are barely used

Japan’s Cabinet Bureau of Personnel Affairs surveys how often each rating grade is given in the national civil service. The grades changed from five levels to six in October 2022, so the old scale and the new scale can be compared. On the new six-level scale, rank-and-file staff received “Very good” (優良) 48.0% of the time and “Good” (良好) 33.2% of the time in the competency rating. The two lowest grades, “Somewhat insufficient” and “Insufficient”, come to just 0.5% between them.

National civil servants (rank-and-file staff): share of each rating grade

The grades changed from five levels to six in October 2022. For each type of rating, the old scale and the new scale are shown side by side. Higher grades are on the left, lower on the right.

Competency rating

Old scale · 5 levelsOct 2018 – Sep 2019

Two most-used grades (A, B): 90.4% / two lowest (C, D): 0.4% / effective levels 2.3

New scale · 6 levelsOct 2023 – Sep 2024

Two most-used grades (Very good, Good): 81.2% / two lowest: 0.5% / effective levels 2.7

Performance rating

Old scale · 5 levelsOct 2019 – Mar 2020

Two most-used grades (A, B): 88.4% / two lowest (C, D): 0.5% / effective levels 2.4

New scale · 6 levelsOct 2024 – Mar 2025

Two most-used grades (Very good, Good): 77.8% / two lowest: 0.5% / effective levels 2.8

See the figures
Competency · old scaleS 9.1%A 53.2%B 37.2%C 0.4%D 0.0%
Competency · new scaleOutstanding 1.2%Excellent 17.1%Very good 48.0%Good 33.2%Somewhat insufficient 0.4%Insufficient 0.1%
Performance · old scaleS 11.2%A 52.1%B 36.3%C 0.4%D 0.1%
Performance · new scaleOutstanding 1.8%Excellent 20.0%Very good 46.4%Good 31.4%Somewhat insufficient 0.4%Insufficient 0.1%

Source: Cabinet Bureau of Personnel Affairs, Cabinet Secretariat, “Operation of the revised personnel evaluation system” (March 2026) and “Survey of rating distribution, October 2018 to March 2020” (September 2020), both in Japanese. Document (PDF). Grade names are our translation: Outstanding (卓越して優秀), Excellent (非常に優秀), Very good (優良), Good (良好), Somewhat insufficient (やや不十分), Insufficient (不十分). The old-scale performance rating is shown for the same months (October to March) as the new scale. Effective number of levels = 1 ÷ ( p12 + p22 + … + pk2 ), where pi is the share given grade i and k is the number of grades. It is our own calculation using a standard statistical method. The old and new scales cover different periods and different staff, so the difference is a before-and-after comparison. The survey is a sample drawn by ministry and agency, with shares scaled up to the population. Figures are rounded, so totals may not equal 100%.

This survey is about personnel evaluation, not interviews. Even so, it gives an official, public figure for a familiar pattern: even when many levels are offered, raters actually use only some of them.

Offer six levels, and raters use only some of them.

That said, these figures alone cannot tell us whether the staff really are alike or whether the raters simply rate alike. That is where the problem in the next sections begins.

Background

After the change, the number of levels actually used rose by about 0.4

On the aim of adding a grade, the Cabinet Bureau’s review report (2021) cites raising the ability to tell people apart by increasing the number of categories. The same report also names as an issue that raters do not read the grades in the same way: grade B is meant to mean “normal” but is taken as below normal.

The result shows in the numbers. Counting the levels actually used, allowing for how unevenly the shares fall (the “effective number of levels”), gives 2.3 on the old scale and 2.7 on the new one for the competency rating, and 2.4 and 2.8 for the performance rating. After the change, more levels are in use.

But the rise is about 0.4, short of the one level that was added (1.0). The new scale offers six levels, yet roughly three are effectively used. The share held by the two lowest grades has barely moved either: from 0.4% to 0.5% for the competency rating, and 0.5% before and after for the performance rating.

These figures do not say why the two lowest grades go unused. Raters may find it hard to give a low rating, or there may genuinely be few low performers. Either way, a level that goes unused is not evidence.

The number of levels offered and the number of levels used are different things.

Applied to interview evaluation, it is not enough to ask how many criteria and levels you have set. You need to check how many of them are actually used.

Applied to hiring

Interview scores can bunch up too, for three possible reasons

When interview scores bunch up, the reason falls into one of three. The distribution alone cannot tell you which.

01

Level descriptions are abstract

If a level reads “strong communication skills”, describing an impression rather than behaviour, both AI and interviewers will tend to play safe and pick a middle score.

02

No question tests that criterion

The criterion is set, but no question in the interview tests it. With no evidence to go on, scores bunch up.

03

The applicants are alike

If applicants were already narrowed at the CV stage, small differences in score are the correct result. In that case the criterion is not the problem.

There is also the halo effect: with human interviewers, one impression can colour the score on every criterion. If the effect is strong, six criteria work in effect as one.

How to check

Whether a criterion works shows in the distribution, not the average

The average score on a criterion can be the same whether or not that criterion tells candidates apart. The two criteria below both average 3.0 across 20 candidates.

Example: scores given to 20 candidates on a 1 to 5 scale (the number above each bar is the number of candidates; the standard deviation is calculated over all 20)

Criterion Ascores spread widely

12345
  • Average score3.0
  • Standard deviationhow spread out the scores are1.14
  • Share of the most common scorehow many candidates got it30%

Differences between candidates show in the scores

Criterion Bscores bunched at 3

12345
  • Average score3.0
  • Standard deviationhow spread out the scores are0.84
  • Share of the most common scorehow many candidates got it60%

Differences between candidates do not show in the scores

What matters is the distribution, not the average. The smaller the standard deviation and the higher the share of the most common score, the less the scores separate candidates.

The same average does not mean the criterion tells candidates apart.

When people score, differences between interviewers get mixed into the scores. Even where the scoring sheet has the same format, the level of detail, the content and how generous or strict each interviewer is will differ. For the same answer, a generous interviewer may give 4 and a strict one 2, so you cannot tell whether the spread in scores reflects the candidates or the interviewers. Putting everyone’s scores on a common scale is work you must do before you can check.

AI interviews make this check easier. The scoring instructions are the same for every candidate, so differences between interviewers do not get into the scores, and every candidate’s score on every criterion is recorded on the same scale. AI has its own tendencies in how it scores, but those tendencies are less likely to change from one candidate to the next. When you choose an AI interview service, also check that you can pull out the scores for every candidate on every criterion as they stand.

Choosing

What sets services apart is whether you can fix a criterion that does not work

When a criterion turns out not to work, what you fix is the description of the criterion or the interview question. Afterwards, re-score the same role and the same applicant pool, and check whether the spread of scores has widened.

Changing the weights will not fix it. Weights move only the total: multiply a criterion on which everyone scores 3 by a weight, and everyone still has the same score.

So when you choose an AI interview service, look at whether you can rewrite the definition of a criterion yourself, not at how many criteria it has. If the service fixes the criteria, you cannot repair one that turns out not to work.

For each service, who decides the criteria is something we checked one by one against public information and set out in our comparison article.

Comparison

AI interview services: nine compared →

Nine major services, sorted into four types by what each one optimises for.

Q&A

Frequently asked questions

How many criteria should I set?

There is no fixed right number. Rather than the number, check whether each criterion you have set actually tells candidates apart.

Does the same thing happen outside interviews?

Yes. The national civil service ratings in this article are one example. You can also check whether performance ratings after hiring bunch up in a few levels.

How many records do I need before I can check?

As a guide, around 100 per role. With a few dozen, one person’s score can shift the share of the most common score by several points, so you cannot tell a real difference from chance.

Read next

Telling candidates apart is not the same as being right

Once you have checked that a criterion tells candidates apart, the next step is to check whether its scores went with how people did after joining. Criteria that did not can be rewritten and tested on past interviews.

Validation

Whether an interview evaluation was right only shows after the hire →

Check hiring evaluations against results after the hire.

Rescoring

Rewrite your criteria, then test them on past interviews →

Re-score past interviews with rewritten criteria and see how the scores move.

Related pages

How it applies, by situation

The number of criteria and how scores spread matter most when setting standards and in graduate hiring.

Standards

Standardised evaluation →

Set the criteria first and narrow the spread between interviewers.

Graduates

Graduate hiring →

See every applicant against one standard, and improve it each year.

Show us your evaluation criteria as they are

We will go through which of your current criteria tell candidates apart, and which do not. Free, about 30 minutes.