Column

Rewrite your criteria, then test them on past interviews

Evaluation criteria are not set once and left alone. An AI interview keeps the full transcript of every interview, so after you rewrite your criteria you can re-score past interviews against the new ones. This piece covers why AI interviews make this possible when interviews run by people do not, and what it can and cannot do.

With human interviewers

Changing the criteria does not easily change how past interviews are scored

The longer you run a selection process, the more often you want to revisit the criteria: when you start to see what the people doing well after joining have in common, or when you notice that everyone scores about the same on one criterion.

With human interviewers, re-scoring past interviews against new criteria is not simple. Someone has to rewatch the recordings one by one, which takes time. And an interviewer who knows how the hires turned out tends to be swayed by that. So new criteria usually come into use from the next interview, and whether they were good criteria only becomes clear later still.

With AI interviews

The full transcript stays, so past interviews can be re-scored against rewritten criteria

An AI interview keeps the full transcript of every interview. After you rewrite the criteria, that record can be scored again against the new ones. Scoring rests only on what was said in the record, so knowing how a hire turned out does not affect it.

Re-scoring gives you scores for the same interview before and after the rewrite. You can see whose scores changed and how, and whether the gaps between candidates widened or narrowed, without waiting for the next round of hiring.

Our test

Rewriting the criteria and re-scoring widened the gaps between the same four candidates

In October 2026 we set up four candidate personas for a customer success role, each answering with a different level of detail, wrote persona-based interview transcripts for them and tested this. The questions were identical; only how specific the answers were changed. Candidate A answers concretely, with figures and the sequence of events; Candidate D barely answers. Candidates B and C fall in between.

We first scored them against vague criteria such as “Does the candidate communicate well?” and “Does the candidate show initiative?”. We then rewrote the same four areas into criteria that describe what a 5-point answer and a 3-point answer look like, and re-scored the same transcripts. We used Proomi for the scoring and scored each 20 times.

Example of the criteria used in the test (the “initiative” area)

Before

Does the candidate show initiative?

No level descriptions

After

Took initiative

L5Describes an action taken without waiting for instructions in an unexpected situation, who was involved, and the outcome
L4Between L5 and L3
L3Describes an action, but not the outcome
L2Between L3 and L1
L1Says nothing about such an experience (insufficient information)

The other criteria were rewritten in the same way and re-scored.

Average scores when the same four interviews were re-scored against rewritten criteria (out of 5, 20 runs each)

Before: vague criteria

Candidate A
4.62
Candidate B
2.92
Candidate C
2.17
Candidate D
1.45

Gap between A and D: 3.2. D’s highest total matches C’s lowest

After: criteria with level descriptions added

Candidate A
4.80
Candidate B
2.83
Candidate C
1.57
Candidate D
1.01

Gap between A and D: 3.8. The ranking never changed in 20 runs

Each average is the mean of the four criteria scores, averaged again over 20 runs. The transcripts are persona-based, and each one was scored 20 times.

Before the rewrite, Candidate D averaged 1.45. Candidate D only said “I want to learn all sorts of things”, yet scored 2 on the growth criterion in 12 of 20 runs. When the levels of a criterion are not defined, even thin answers tend to pick up points.

After the rewrite, Candidate D’s average fell to 1.01 and the gap to Candidate A widened from 3.2 to 3.8. The gap between Candidates B and C, where a pass line is often drawn, also widened, from 0.75 to 1.26. Before the rewrite, Candidate D’s highest total matched Candidate C’s lowest; after it, the four candidates’ ranking did not change once in 20 runs.

When scores bunch up, adding criteria does not add evidence. That is covered in “More criteria do not mean more evidence”.

How generative AI behaves

Answers that split the score show where to rewrite the criteria

Generative AI writes by choosing each next word based on probabilities. So scoring the same transcript against the same criteria will not always give exactly the same score. Clear answers barely move, though. Scores split when an answer fits the descriptions of two neighbouring levels.

Take the answer “I make a point of getting in touch earlier now”. There is reflection, but it is not specific about what changed or how. It fits the level “reflects, but what changed is vague”, and can also be read as one level lower.

In other words, an answer that splits the score marks a place where the level descriptions do not quite tell answers apart. Human interviewers and AI alike split on answers like this. With AI, the same answer can be re-scored as often as you like, so you can find which answers split on which criteria across every interview. When you rewrite criteria, these are the places to start.

Which level a split answer is treated as is decided by a person, after reading the reasoning and the transcript. The AI’s score is input to that judgement.

What past interviews cannot test

You can change how answers are read, not what was asked

Testing on past interview records comes with two caveats.

01

Only what was asked can be tested

Scoring rests only on what was actually said, so past interviews can only test what was asked in them. Changes to questions take effect from the next interview.

02

Check rewritten criteria on different interviews

Rewriting and checking on the same interviews over and over produces criteria that fit by chance. With enough records, split past interviews in two: rewrite on one half, check on the other. With few, check in the next round of hiring.

Re-scoring does not raise predictive power by itself. What it does is put the material for raising it in your hands without waiting for the next round. The limits of checking against outcomes after hiring are set out in “Whether an interview evaluation was right only shows after the hire”.

What makes this possible

Because you write your own criteria, you can rewrite them and test

Past interviews can be re-scored against new criteria because you write the criteria themselves in your own words. Where the service fixes the criteria, you may be able to change the weights, but not what the criteria say. If you cannot rewrite what they say, the way the same interview is read does not change either.

Writing in your own words means the level descriptions can carry your industry’s language and the behaviour of the people who do well at your company. For a customer success role at a SaaS company, for example, it could look like this:

Example of levels written in your own words (customer success at a SaaS company)

Data-driven customer retention

L5Can describe, with figures, spotting early signs of churn in usage data such as active rates and health scores, acting on them, and turning that into renewals or upsells
L4Has acted on usage data, but cannot speak to results such as renewals
L3Has responded to customer feedback, but without using usage data
L2Has customer-facing experience, but limited to answering enquiries as they come in
L1Says nothing about such an experience (insufficient information)

Your industry’s language and the behaviour of the people who do well at your company can go straight into the level descriptions.

In Proomi you define your own criteria. We carry out the re-scoring of past interviews against rewritten criteria with you, as part of our support.

What to compare

Three checks for whether you can rewrite criteria and test them

When comparing AI interview services, these three checks tell you whether you can keep revising your criteria as you go.

01

Can you write what the criteria say, in your own words?

Changing the weights alone does not change how the same answer is read.

02

Are the full transcript and the reasoning for each score kept?

Without the record, past interviews cannot be re-scored. With the reasoning, you can trace a changed score back to what was said.

03

Can past interviews be re-scored against rewritten criteria?

Check what is covered, including support from the provider.

Q&A

Frequently asked questions

How often should we revisit the criteria?

There is no set frequency. Good times are when you notice a criterion where scores bunch up, when outcomes after hiring have built up, and when the work in a role changes.

Can we revise the questions at the same time?

Changes to questions take effect from the next interview. Anything not asked in past interviews does not become a score when re-scored. Rewritten criteria can be tested on past interviews, but revised questions have to wait for the results of the next interviews.

How many interviews do we need to test this?

As a guide, around 100 per role to see how scores spread. With a few dozen, one person’s score can move the result, and you cannot tell the effect of the rewrite from chance.

Read next

Also check whether scores tell candidates apart, and whether evaluations were right

Before rewriting criteria, checking whether your current scores tell candidates apart, and whether they went with outcomes after joining, makes it easier to decide which criteria to revisit first.

Distribution

More criteria do not mean more evidence →

Why scores bunch up, checked against public civil-service rating data.

Validation

Whether an interview evaluation was right only shows after the hire →

Check hiring evaluations against results after the hire.

Show us your evaluation criteria as they are

We will check with you whether your criteria can be rewritten and tested on past interviews. Free, about 30 minutes.