
Solutions
How to choose an AI interview tool
Decide what you want to optimise before you count features. Five layers and six questions show which type of tool fits your organisation.
Context
How far the AI goes
Candidates open the URL they are sent and take the interview without waiting for an interviewer to be free. The AI asks questions along evaluation criteria defined in advance, follows up on the answers, and records an evaluation with its reasoning. What this page is about is how you decide what to hand it.
A person makes the hiring decision. The AI goes as far as the evaluation and the reasoning behind it. Where that line sits differs between services, and it is worth confirming before you choose.
Five layers
Where AI interview tools actually differ
The top two layers can be chosen on price. The bottom three cannot.
01
Execution
Can it run the interviews? Round-the-clock availability, languages, candidate experience.
02
Evaluation
Can it score? Criteria, scores, reports.
03
Ownership of the criteria
Who decides what the evaluation axes are?
04
Validation
Did the people hired on those criteria actually perform?
05
Continuity
Can the evaluation made at hiring carry through to placement and development?
Vendors compete on the top two layers, and that is where they are hardest to tell apart. If the goal is throughput, the top two are enough. For mid-career and specialist hiring, or hiring that spans several departments and roles, the bottom three are what matter. Ownership of the criteria is starting to become a selection axis; validation and continuity are not yet.
Four types
They split by what they optimise
AI interview services fall into four types, according to what they optimise. They are listed by type, not by service.
A
Shared-framework
Bring high-volume screening onto one common evaluation frame.
B
Low-cost, flat-rate
Let every applicant take an interview without worrying about volume.
C
Sourcing and pool
Supply candidates from a pool.
D
Criteria-ownership
Assess against your own criteria, and turn those criteria into an asset.
Proomi is this type
Which of the nine leading services falls into which type was judged one by one against what each vendor publishes. It is set out in AI interview services: nine compared.
Note that you can choose the criteria-owned type without having the criteria written. We can build them with you, from the definitions your performance reviews already use and from what your interviewers actually look at. What separates the types is not who writes the first draft, but whether the finished criteria are yours and whether you can rewrite them afterwards.
Six questions
What to check before you buy
Customising the evaluation criteria (Q2) is possible with every type. The answers start to diverge at Q3. Six questions whose answers depend on the type.
Q1What does this type optimise?
Whether you want more throughput or better evaluation changes which service fits. Start a feature comparison without settling this and you end up choosing by the length of the comparison table.
Q2Can the evaluation criteria be tailored to us?
“Customisable” can be said of every type. What to check is whether you can define the framework itself, or only adjust the weighting inside the service’s standard one. Ask whether there is a limit on the number or naming of criteria, and whether different roles can have different criteria.
Q3Can current employees be measured on the same scale as candidates?
Most services end at selection and never measure the people already employed. When you can, you can check with real data which criteria your high performers actually score well on.
Q4Can we confirm that people hired on these criteria actually performed?
Do you rely on the service’s statistical validation, does the service hand back the data and leave validation to you, or can you check it against your own selection outcomes and results? This is what the validation layer means in practice.
Q5When AI and human evaluations diverge, can we notice and correct it?
They will diverge. The question is whether there is a way to notice, and whether you can fix it yourself once you have. Waiting for the service to update its model and revising your own criteria definitions are very different degrees of freedom.
How to test rewritten criteria on past interviews is covered in “Rewrite your criteria, then test them on past interviews”.
Q6Can one scale be used across roles and languages?
Can you set criteria per role? Can you run Japanese and English without maintaining separate standards? If the tool only works within a standard framework or templates, it breaks down as roles multiply.
Standardisation
How to keep evaluation aligned
Standardised evaluation means you can confirm the same criteria are applied to everyone, verify that those criteria match your actual hiring decisions, and correct them when they drift. It does not mean lining up the evaluation items. It comes in three stages.
1
Are the same criteria applied to everyone?
If the interview record and the reasoning are kept for every case, this can be checked from the tool’s own data.
2
Do the criteria match our hiring decisions?
Set them against pass-or-fail outcomes and you can see which criteria were actually tied to the decision.
3
Did the criteria hold through to retention?
Set them against post-hire tenure and you can see which criteria predicted people staying.
Stage one shows up in the score distribution per criterion. The smaller the standard deviation, the more the AI is giving the same score. Separate whether the criterion is vaguely worded, whether the questions fail to probe it, or whether the candidates really are alike, and the design can be corrected.
Show the numbers
| Criterion | 1 pt | 2 pt | 3 pt | 4 pt | 5 pt | Std. dev. | Modal share |
|---|---|---|---|---|---|---|---|
| Intent to stay | 12 | 63 | 68 | 30 | 2 | 0.88 | 39% |
| Communication | 4 | 19 | 64 | 64 | 24 | 0.94 | 37% |
| Adaptability | 7 | 36 | 67 | 49 | 16 | 0.99 | 38% |
| Logical thinking | 9 | 33 | 79 | 37 | 17 | 0.99 | 45% |
| Initiative | 13 | 29 | 62 | 48 | 23 | 1.10 | 35% |
| Domain knowledge | 19 | 44 | 54 | 36 | 22 | 1.18 | 31% |
The figure on each bar is the share taken by the most common score. Criteria are sorted by standard deviation, smallest first.
None of this is about making the AI more accurate. It is about checking whether your own hiring criteria are right.
Next to read
Which type each of the nine falls into
The five layers and six questions above, applied to the nine leading services. Each was placed against what its vendor publishes, and they are set side by side on questions, criteria and verification. No ranking.
AI interview services: nine compared →
Nine major services, sorted into four types by what each one optimises for.
Show us your evaluation criteria as they are today
We will walk through the six questions and tell you which type fits. Free, about 30 minutes.