
Column
Whether an interview evaluation was right only shows after the hire
Research shows structured interviews predict job performance well on average. An AI interview whose questions and criteria are set in advance counts as a structured interview too. But whether your own evaluation is right is a separate matter. This piece sets out why it is worth checking, and what you need in place to check it.
Terms
What structured interviews and predictive power mean
A structured interview is one where the content of the questions, their order and the criteria for evaluation are fixed in advance and applied to every candidate under the same conditions. Whether the interviewer is a person or an AI, the same questions and the same criteria are used. An AI interview whose questions and criteria are set in advance counts as a structured interview.
Predictive power is how well a score at selection tells you how someone will perform after joining. Research calls it predictive validity and expresses it as the correlation between the selection score and later performance. A correlation runs from −1 to 1. The closer to 1, the more likely it is that people who scored high at selection are rated highly after joining. The closer to 0, the less the selection score has to do with later performance.
Findings
Structured interviews average 0.42 in predictive power, but vary widely by setting
Personnel psychology has long studied the predictive power of each hiring method. In an analysis published in 2022 in the Journal of Applied Psychology, Sackett and colleagues found that structured interviews had the highest predictive power of the methods examined, with an average correlation of 0.42 with later job performance. The research covers structured interviews conducted by people. Whether an AI interview reaches the same level is something you need to check for yourself.
But 0.42 is an average. The same analysis says predictive power differs widely by company and job, even for structured interviews. Used in ten workplaces, nine are expected to reach 0.18 or more, and the remaining one to fall below 0.18.
The difference can be read like this. Rank candidates by selection score, highest first, and split them exactly in half. Call the upper half the “high-scoring group” and the lower half the “low-scoring group”. In a workplace where predictive power is 0.42, about 64% of the high-scoring group go on to perform above average after joining, against about 36% of the low-scoring group. People who scored high clearly tend to perform well afterwards. Where predictive power is 0.18, the figures are about 56% and 44%. The gap is slight, and the score alone is a poor guide to later performance.
Share of people who perform above average after joining, when candidates are split in half by selection score, highest first
Predictive power 0.42average
People who score high clearly tend to perform well after joining
Predictive power 0.18one workplace in ten falls below this
The gap is slight, and the score alone is a poor guide to later performance
The dashed line is 50%. If the score had nothing to do with later performance, both groups would sit at 50%.
Source: Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107, 2040–2068. https://doi.org/10.1037/apl0000994. The shares for the two groups are our own conversion, assuming a symmetric, bell-shaped distribution of scores and performance, and are a guide only.
Why they differ
Even among “structured interviews”, predictive power can change with what is in the interview
The paper points out that structured interviews differ in content. For candidates with no job knowledge or experience, for example, interviews may use situational questions (“what would you do in this situation?”); for candidates who do have it, they may ask about past behaviour (“what did you do?”). The paper says one cannot expect the same level of predictive power from the two, because the content of the interviews is quite different.
The paper also accepts that structured interviews have a high average predictive power, but notes that the spread between settings is large and a low value is a real risk. It says it would be desirable to identify the factors that explain that spread and give more specific estimates by job and by interview design.
A further point: when scores bunch into a narrow range, the link between score and performance looks smaller than it really is.
What follows is that an interview’s predictive power depends on which job it is used for, what it contains and how it is scored. Using a structured interview does not by itself raise predictive power. You need to line up your own selection results against outcomes after joining, make them visible, and check whether the interview really does predict.
Note: this section is based on the discussion in Sackett, Zhang, Berry, & Lievens (2022).
Making use of it
With an AI interview, matching records to outcomes after joining can raise the accuracy of evaluation
An AI interview records, right after the interview, a score and its reasoning on each criterion for every candidate, in the same format. That record can be used as it stands to match against outcomes after joining.
Matching shows which criteria were tied to success after joining. Criteria that were tied to it give you grounds to weigh them more heavily in the next selection. Criteria that were not are a prompt to revisit their description or the question. Getting a score and being right are different things, so adding this check lets you raise the accuracy of evaluation step by step.
The outline of the check is simple. Split candidates on each criterion into a high-scoring group and a low-scoring group, and see whether their outcomes after joining differ. If they do, that criterion was right.
Who owns the criteria
To check, the selection criteria must connect to the criteria used after joining
Suppose selection looks at “communication”, while the post-hire performance review has only “proposal skills” and “negotiation skills”. Line the two up and you cannot say which criterion was right. Matching works when the selection criteria connect to the criteria used after joining: defined in the same words, or mapped to one another.
If you define the criteria yourself, the same definitions can be used on current employees. If you use a framework designed on the service side, you have to ask again for a version designed for employees, and end up with two separate yardsticks, one for hiring and one for evaluation. The moment there are two yardsticks, the premise for matching is gone.
If hiring criteria and performance criteria are written in the same words, the two become one yardstick.
Limits of the check
No firm conclusion in the first year
There are three easy ways to misread the result.
01
You cannot see what happened to those you rejected
Data exists only for those who joined. Someone you rejected on a low score may well have performed. The range you can check is narrow from the start, so differences come out smaller.
02
Small numbers let chance in
If only a handful join each year, one resignation can change the conclusion. A single year tells you little; it only becomes readable once you have kept records for three years.
03
Placement and managers get mixed in
Performance after joining is not decided by the person alone. With a different placement or direct manager, the same person is rated differently.
All three shrink if you gather more data or compare like with like. The other way round, when you start checking decides when you can judge.
Q&A
Frequently asked questions
Is there any point if only a few people join each year?
A single year gives no conclusion. Even so, keeping each criterion’s score and the outcome after joining is worth doing in itself. Three years of records are enough to judge.
Can I use resignation as the measure?
You can, but it cannot settle things alone. People leave for many reasons besides whether the criteria were right. Some leave for reasons beyond their control, such as family circumstances or illness, so a resignation cannot simply be read as a failure of selection. Whether someone is still there is the quickest measure to get, so it suits a first measure; swap it out when performance reviews are available.
Should the hiring side see performance review results?
It depends on how you run it. They do not need to be passed to recruiters in a form that identifies individuals; totalling by criterion before matching is enough. The aim is to test the criteria, not to re-assess individuals.
Read next
Criteria that did not hold up can be rewritten and tested on past interviews
Criteria whose scores did not go with outcomes after joining can be rewritten, then tested by re-scoring past interviews. Before asking whether an evaluation is right, also check that its scores tell candidates apart.
Rewrite your criteria, then test them on past interviews →
Re-score past interviews with rewritten criteria and see how the scores move.
More criteria do not mean more evidence →
Why scores bunch up, checked against public civil-service rating data.
AI interview services: nine compared →
Nine major services, sorted into four types by what each one optimises for.
Related pages
How it applies, by situation
Checking against results after the hire looks different for mid-career hiring, graduate hiring and goal reviews.
Show us your evaluation criteria as they are
We will check with you whether your selection scores went with how people performed after joining. Free, about 30 minutes.