# The Shortlist Submission Standard: When a List Is Ready

*You will be able to grade any shortlist against fixed criteria, candidate by candidate and as a whole, and decide whether it is ready to present or needs one more pass.*

- Canonical URL: https://www.refolk.ai/guides/shortlist-submission-standard
- Pillar: Recruiting and sourcing
- Format: Standard
- Published: 2026-08-16
- Last reviewed: 2026-08-16
- Reading time: 16 min

You have sourced and screened candidates, and now you have to decide whether the list is ready to send. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring, and it gives you a gradeable definition of done: fixed criteria you can apply candidate by candidate and to the list as a whole, plus a pass/fail checklist two people would score the same way. The goal is to stop shortlists bouncing back from the hiring manager.

Public pages argue endlessly about the right number of candidates. That argument is a distraction. The failure that actually returns a shortlist is misalignment with the brief, and no count fixes that. What follows is the standard I would adopt as team policy.

## Why the "right number" is the wrong question

The count almost never decides whether a shortlist bounces. Misalignment with the brief does. HR.com's Future of Talent Acquisition report found that 34% of organisations cite misalignment with hiring managers as a core talent-acquisition challenge, and that is the mechanism behind a returned list: a manager rejects it because it misses what the role actually requires, not because it had nine names instead of six.

Look at how little the published guidance agrees on size. If the number mattered as much as the internet suggests, the internet would have converged on one.

| Source | Shortlist size | Notes |
|---|---|---|
| toggl / logicmelon | 10-20 | "best practice" |
| cs-recruiters | 5-8 | per position |
| hyring | 4-6 | "sweet spot", >8 = fatigue |
| expectllc | 3-5 (up to 10-12 high-volume) | "golden rule" |

Every one of these is defensible. A useful working rule is to shortlist roughly 10% of applicants adjusted for the pool, so 8 to 12 for 100 applicants, then cut to 4 to 8 for the manager-facing list. Fewer than 3 leaves too few comparison points and risks a weak hire; more than 8 creates interview fatigue and slows the process. But treat that as the last constraint you apply, not the first.

> A returned shortlist almost never means too many names. It means the list missed the brief.

The number-of-candidates question is answerable in one line: size to the brief and the pool, land in 4 to 8 for most roles, and move on. The rest of this document is about the part the count hides.

## What a submission-ready shortlist is

A shortlist is submission-ready when every must-have in the brief is covered by at least one named candidate, every candidate carries the evidence a manager needs to decide without a follow-up, and the whole thing has been scored by a shared instrument and signed by a second reviewer. That is the definition of done. Everything below is how to verify each clause.

There is no single public standard for the exact fields a submission must carry, so this standard constructs one from what applicant-tracking systems and scorecard practice converge on. A candidate record is complete when it holds four things.

| Field | What it holds | What proves it |
|---|---|---|
| Competencies | The 5-7 job-relevant criteria from the brief | Each maps to a must-have or desirable |
| Rating | A score on each competency | Points-based: 1 full, 0.5 partial |
| Evidence note | What the candidate said or did | The raw material behind the score |
| Recommendation | Advance, hold, or reject with reason | The manager can act on it directly |

The evidence note is the load-bearing field and the one most often skipped. A rating without a note is an opinion; a rating with a note is a claim a second reviewer can check. If a competency is scored but you cannot write one sentence of what the candidate actually did to earn that score, the record is not ready.

> **Rule:** The no-follow-up test
>
> A candidate record is submission-ready only when a hiring manager could accept or reject that candidate using the record alone, without asking you a single clarifying question.

## Grade each candidate against the brief

Grade one candidate at a time against the scorecard, not against the other candidates and not against your gut. A candidate passes when every must-have criterion is either met or partially met with an evidence note, and no must-have is unaddressed.

The mechanism is a points-based scoring matrix built directly from the locked brief. Most teams run it this way: if a candidate fully meets a criterion they score one point, if they partially meet it they score half a point, and you add the points to rank. Focus the scorecard on the 5 to 7 areas most critical for the role. More than that and the scan slows without adding signal; fewer and you lose the ability to distinguish candidates.

The reason to insist on the scored instrument is not process hygiene. It is validity. Interviews scored without structure have a predictive validity of just .20, barely better than random selection. Add a structured scorecard with behavioral anchors and that jumps to .51, and the most rigorous methods reach .57.

**.20 to .57 - Predictive validity, unstructured judgment to rigorous structured scoring**

The scorecard is not paperwork; it more than doubles how well the process predicts on-the-job success.

That gap is also why a shared scorecard makes two graders converge. When both reviewers score the same behavioral anchors, private judgment stops driving the outcome. The scorecard replaces "she felt strong" with "she scored 1 on payments experience because she shipped X." That sentence is gradeable. The feeling is not.

### What a criterion looks like when it lies

Every criterion you score can produce a false positive, and the standard is only as good as your ability to spot one.

- A title proves seniority until it does not. A "Senior Engineer" at a company where everyone is senior is not the same signal as the same title at a company that gates the level. Score demonstrated scope, not the word.
- A degree or certification proves a credential, not capability. Research on hidden workers shows a large share of employers lose qualified candidates because filters optimise for credentials over what people can actually do.
- Tenure proves time served, not impact. Read what shipped, not how long they stayed.

When a criterion is credential-shaped, write the evidence note in terms of demonstrated ability. If you cannot, downgrade the score.

## Grade the list as a whole

A list passes composition when every must-have is covered by a named candidate, the slate meets a minimum-of-two threshold for any group you are tracking, and no single current employer dominates. These are three separate checks, and a list can fail composition while every individual candidate passes.

The first check is coverage. Map each must-have to at least one named candidate before you submit. A list of eight that looks full but leaves a must-have with zero qualifying candidates is the classic number-chasing failure: it reads as complete and staffs nobody.

The second check is the slate cliff. This is the composition rule with the hardest evidence behind it, and it is a cliff rather than a slope. A single underrepresented candidate on a shortlist is statistically near-invisible to the decision. Two changes everything: with two women finalists the odds of hiring a woman were 79.14x greater, and with two minority finalists the odds of hiring a minority candidate were 193.72x greater. This is why the standard encodes a minimum of two, not a "consider diversity" line. One is not a slate; it is a token that the status quo defeats.

#### From applicants to a submitted shortlist

| Stage | Figure | Note |
| --- | --- | --- |
| Applicants | 100 | about 14% reach interview per industry survey |
| Longlist | 15-25 | meet essential criteria, scored against the matrix |
| Shortlist | 4-8 | cut to cover every must-have |
| Final round | 2-3 | comparison points for the decision |

*Volume narrows sharply, so the composition checks happen on a small, high-stakes set.*

The third check is employer concentration. There is no documented public control for this, so it is a manual read: if every candidate comes from one current employer or one near-identical background, you have built an echo, not a comparison set. A manager reviewing a slate where everyone worked at the same place gets no real choice. Spread the current employers.

The reason these checks land on you rather than a system is structural. In Refolk's index, dedicated sourcers are a thin layer: about 1.1 sourcers exist per 100 recruiters in the US. In most teams the person composing the shortlist is also the one who screened it, which is exactly the single-reviewer condition the calibration pass exists to break.

| Market | Recruiter / Tech Recruiter / TA profiles | Share of the two-market total |
|---|---|---|
| United States | 112,661 | 93.9% |
| United Kingdom | 7,283 | 6.1% |

When you need to build a slate that already satisfies these composition rules rather than retrofitting it, ask for it directly. [Refolk](/) lets you specify employer diversity, must-have and desirable criteria, and underrepresented-group coverage in the query itself, so the list arrives pre-shaped against the standard instead of being cut down after the fact.

I ran this search: `Senior backend engineers in Berlin with Go and Kubernetes who have shipped payments systems at two different companies` - [see the full result list](https://www.refolk.ai/s/7wpnz62er1).

*Returns candidates already spread across employers, so the slate clears the employer-concentration check before you score it.*

## The procedure

Run these eight steps in order. The first two happen before you touch the pool; the composition and calibration checks are where a list is made or broken.

#### From brief lock to logged submission

1. **Lock the brief with the hiring manager** - Run a 30 to 60 minute intake and agree must-haves versus desirables in writing, along with the business goal behind the hire. Done means a signed criteria list both of you can point to later.
2. **Build the scorecard from the brief** - Convert the locked criteria into 5 to 7 job-relevant items with a rating scale. Done means a scorecard whose fields match the brief one for one, ready for evidence notes.
3. **Screen the pool to a longlist** - Scan applicants against the essential criteria, roughly 7 to 15 seconds per resume and 2 to 3 minutes for a closer read, down to a longlist of 15 to 25 who meet the must-haves. Done means a longlist selected against the matrix, not intuition.
4. **Score and rank against the matrix** - Apply a points-based method: one point for fully meeting a criterion, half a point for partial, then total and rank. Done means a ranked list where every position is backed by an evidence note.
5. **Compose the shortlist against the brief** - Cut to 4 to 8 candidates, confirm every must-have is covered by at least one named person, and check slate composition against the two-in-the-pool threshold and employer concentration. Done means the list covers all must-haves and no single employer dominates.
6. **Run a second-reviewer calibration pass** - Have a peer or panel of at least two people re-score and resolve disagreements in a short calibration meeting. Done means no candidate advanced on a single reviewer's judgment.
7. **Assemble the submission pack** - For each candidate attach competencies, ratings with evidence notes, and a recommendation on next steps. Done means a manager could decide without asking a follow-up question.
8. **Submit and log the rationale** - Send the dated pack and record why each candidate was advanced or held for compliance and audit. Done means the submission is out and the reasoning is written down.

> **Tip:** Front-load the brief lock
>
> The 30 minutes spent signing the criteria list at step one is the cheapest insurance against a step-eight bounce. The 34% misalignment failure is almost always a step-one shortcut showing up at the end.

## How this goes wrong

Most bounced shortlists trace to one of eight failure modes. Each has a false positive - a list that looks ready and is not - and each has a check that catches it. This is the part of the standard that earns its keep.

**Number-chasing.** Hitting a target count while a must-have has zero qualifying candidates. The false positive is a full-looking list of eight that no one can actually do the job. Check: map every must-have to at least one named candidate before submitting.

**Intuition masquerading as scoring.** Speeding through resumes and picking the ones that feel right. That feeling is usually pattern recognition biased toward candidates who resemble people previously hired. The false positive is a ranked list that looks scored but was sorted by gut. Check: every ranking has an evidence note.

**Single-reviewer bias.** One grader's mental scorecard driving the whole list. Single-reviewer decisions carry the highest bias risk, and with roughly 1.1 sourcers per 100 recruiters the screener and composer are usually the same person. Check: two reviewers signed and scores calibrated.

**Credential filter, not capability.** Looks-right-on-paper candidates that bounce at manager review. Rising volume makes this worse: recruiters manage 56% more open requisitions than in 2022, and the initial scan is 7 to 15 seconds, which is exactly how credential-over-capability filtering creeps in. Check: score demonstrated ability, not titles or degrees.

**The "only one" slate.** A lone underrepresented or atypical candidate is statistically near-invisible. The false positive is a slate you call diverse that has exactly one. Check: the two-in-the-pool threshold is met.

**Employer-clone concentration.** Every candidate from one current employer or one identical background. There is no documented public control, so this is a manual read. Check: no single current employer over-represented.

**Missing decision fields.** A pack that forces the manager to ask follow-ups. The false positive is a tidy-looking record missing an evidence note or a recommendation. Check: each candidate has competencies, rating, evidence, and recommendation.

**Stale list.** A slow submission loses candidates: half of surveyed applicants turned down an offer because the process took too long. The false positive is a strong list assembled two weeks ago whose top names are gone. Check: the submission is dated and candidates confirmed still active.

#### What to do with a candidate on the shortlist

Horizontal axis runs from Weak or missing evidence to Strong documented evidence. Vertical axis runs from Must-have not covered to Must-have covered.

| Quadrant | What it means |
| --- | --- |
| Cut, and note the gap | No evidence and no coverage, this candidate is padding the count |
| Hold for verification | Covers a must-have but the score is unbacked, get the evidence before submitting |
| Reject for this role | Well documented but does not meet a must-have, right person wrong brief |
| Submit | Covers a must-have with strong evidence, this is what the shortlist is for |

*Two axes decide the call: does the evidence back the score, and is a must-have covered.*

> **Watch out:** The scorecard does not overrule the brief
>
> A candidate can score highly on your 5-7 criteria and still miss a must-have, because a strong score on other axes cannot backfill a criterion nobody meets. Coverage of every must-have is a gate, not a weighted factor.

## The pass/fail checklist

Run this before the list leaves your hands. It is written so two graders applying it to the same shortlist reach the same verdict. If any item fails, the list is not ready and you make one more pass.

#### Submission-ready shortlist

- [ ] The must-have and desirable criteria are signed off in writing by the hiring manager.
- [ ] The scorecard has 5-7 job-relevant criteria, each mapped to a criterion in the brief.
- [ ] Every must-have criterion is covered by at least one named candidate on the list.
- [ ] Every candidate's ranking is backed by an evidence note, not intuition.
- [ ] Each candidate record carries competencies, ratings, evidence notes, and a recommendation.
- [ ] A manager could decide on each candidate without asking a follow-up question.
- [ ] The slate meets the two-in-the-pool threshold for any group being tracked.
- [ ] No single current employer is over-represented across the candidates.
- [ ] A second reviewer has scored the list and scoring disagreements are resolved.
- [ ] The list is 4-8 candidates for a standard role, sized to the brief and pool.
- [ ] The submission is dated and every candidate is confirmed still active.
- [ ] The rationale for advancing or holding each candidate is logged for audit.

## Adopt it and keep it current

Adopt this as a written policy your team grades against, not a mental habit, because the whole point of a standard is that two people apply it identically. Paste the checklist into your applicant-tracking system as a required gate before the "submit to hiring manager" step, and require the second-reviewer sign-off as a field, not a courtesy.

Two things drift over time and need a re-check cadence. First, the brief: criteria signed for one requisition go stale as the role or market shifts, so re-lock the must-haves at the start of every search rather than reusing last quarter's list. Second, the size band: the 4-8 range is a default, not a rule, and high-volume or executive roles legitimately sit outside it. Where you deviate, write down why in the submission log so the deviation is a decision, not a slip.

The standard's hardest clauses are the composition checks, because they have no system enforcing them and depend on you building a slate that already spans employers and clears the two-in-the-pool threshold. Template a copy-pasteable candidate line and a manager-facing intake so the fields are never improvised.

**Candidate line for the submission pack**

```
Name / current role / current employer:
Must-haves covered: [criterion] - met / partial - [one-sentence evidence note]
Desirables met: [criterion] - [one-sentence evidence note]
Total score (points-based): X of Y
Recommendation: Advance / Hold / Reject - because [reason]
Confirmed active as of: [date]
```

*One block per candidate. Fill every field; an empty field is a failed record.*

**Hiring-manager intake, criteria section**

```
Business goal behind this hire:
Must-haves (a candidate is disqualified without these):
1.
2.
3.
Desirables (nice to have, used to rank):
1.
2.
Slate rules for this role (two-in-the-pool group, employer spread):
Target shortlist size and why:
Signed (manager) / date:
```

*Fill this with the manager and get it signed before sourcing. It becomes the scorecard.*

Keep the checklist open while you compose. The measure of the standard is simple: a list that passes every item should not come back, and if it does, the reason belongs in the checklist as a new item for next time.

## Frequently asked questions

### How many candidates should I shortlist for interview?

There is no single right number, and chasing one is the wrong focus. Published ranges run from 3-5 up to 10-20 per role, with a common sweet spot of 4-6, fewer than 3 giving too few comparison points, and more than 8 causing interview fatigue. Size to the brief and the pool. For 100 applicants, 8-12 is workable. Composition against the must-haves, not the count, is what decides whether the list bounces back.

### What makes a shortlist bounce back from the hiring manager?

The dominant documented reason is misalignment on the brief: 34% of organisations cite hiring-manager misalignment as a core talent-acquisition challenge. Others are credential-over-capability screening, undocumented criteria producing inconsistency between reviewers, and submission packs that force the manager to ask follow-up questions. A list bounces because it misses the agreed must-haves or cannot be decided on without more information, not because it had nine names instead of six.

### What should each candidate record contain before I submit?

Each candidate should carry the job-relevant competencies, a rating on each, an evidence note recording what the candidate actually said or did, and a recommendation on next steps. The test is whether a manager can decide without asking you a single follow-up. If any field is missing or a rating has no evidence behind it, the record is not submission-ready.

### Do I really need a second reviewer if I sourced the list myself?

Yes. Single-reviewer decisions carry the highest bias risk, and in most teams the person composing the shortlist is also the one who screened it, which is precisely that condition. A second grader using the same scorecard is what makes two people grade a case identically. Predictive validity rises from .20 for unstructured judgment to .51 with a structured scorecard, so the calibration pass is not optional overhead.

### What is the two-in-the-pool rule and does it belong in a shortlist standard?

Two-in-the-pool means a shortlist should include at least two candidates from any underrepresented group you are tracking, never exactly one. It belongs in the standard because the effect is a cliff, not a slope: a lone underrepresented finalist is statistically near-invisible, but two finalists raise the odds of hiring a woman by 79.14x and a minority candidate by 193.72x. Encode a minimum of two, not a vague nicety.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/shortlist-submission-standard*
