Refolk
StandardRecruiting and sourcing

The Shortlist Submission Standard: When a List Is Ready

You will be able to grade any shortlist against fixed criteria, candidate by candidate and as a whole, and decide whether it is ready to present or needs one more pass.

16 min readLast reviewed August 16, 2026Read as Markdown

You have sourced and screened candidates, and now you have to decide whether the list is ready to send. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring, and it gives you a gradeable definition of done: fixed criteria you can apply candidate by candidate and to the list as a whole, plus a pass/fail checklist two people would score the same way. The goal is to stop shortlists bouncing back from the hiring manager.

Public pages argue endlessly about the right number of candidates. That argument is a distraction. The failure that actually returns a shortlist is misalignment with the brief, and no count fixes that. What follows is the standard I would adopt as team policy.

Why the "right number" is the wrong question

The count almost never decides whether a shortlist bounces. Misalignment with the brief does. HR.com's Future of Talent Acquisition report found that 34% of organisations cite misalignment with hiring managers as a core talent-acquisition challenge, and that is the mechanism behind a returned list: a manager rejects it because it misses what the role actually requires, not because it had nine names instead of six.

Look at how little the published guidance agrees on size. If the number mattered as much as the internet suggests, the internet would have converged on one.

SourceShortlist sizeNotes
toggl / logicmelon10-20"best practice"
cs-recruiters5-8per position
hyring4-6"sweet spot", >8 = fatigue
expectllc3-5 (up to 10-12 high-volume)"golden rule"

Every one of these is defensible. A useful working rule is to shortlist roughly 10% of applicants adjusted for the pool, so 8 to 12 for 100 applicants, then cut to 4 to 8 for the manager-facing list. Fewer than 3 leaves too few comparison points and risks a weak hire; more than 8 creates interview fatigue and slows the process. But treat that as the last constraint you apply, not the first.

A returned shortlist almost never means too many names. It means the list missed the brief.

The number-of-candidates question is answerable in one line: size to the brief and the pool, land in 4 to 8 for most roles, and move on. The rest of this document is about the part the count hides.

What a submission-ready shortlist is

A shortlist is submission-ready when every must-have in the brief is covered by at least one named candidate, every candidate carries the evidence a manager needs to decide without a follow-up, and the whole thing has been scored by a shared instrument and signed by a second reviewer. That is the definition of done. Everything below is how to verify each clause.

There is no single public standard for the exact fields a submission must carry, so this standard constructs one from what applicant-tracking systems and scorecard practice converge on. A candidate record is complete when it holds four things.

FieldWhat it holdsWhat proves it
CompetenciesThe 5-7 job-relevant criteria from the briefEach maps to a must-have or desirable
RatingA score on each competencyPoints-based: 1 full, 0.5 partial
Evidence noteWhat the candidate said or didThe raw material behind the score
RecommendationAdvance, hold, or reject with reasonThe manager can act on it directly

The evidence note is the load-bearing field and the one most often skipped. A rating without a note is an opinion; a rating with a note is a claim a second reviewer can check. If a competency is scored but you cannot write one sentence of what the candidate actually did to earn that score, the record is not ready.

Grade each candidate against the brief

Grade one candidate at a time against the scorecard, not against the other candidates and not against your gut. A candidate passes when every must-have criterion is either met or partially met with an evidence note, and no must-have is unaddressed.

The mechanism is a points-based scoring matrix built directly from the locked brief. Most teams run it this way: if a candidate fully meets a criterion they score one point, if they partially meet it they score half a point, and you add the points to rank. Focus the scorecard on the 5 to 7 areas most critical for the role. More than that and the scan slows without adding signal; fewer and you lose the ability to distinguish candidates.

The reason to insist on the scored instrument is not process hygiene. It is validity. Interviews scored without structure have a predictive validity of just .20, barely better than random selection. Add a structured scorecard with behavioral anchors and that jumps to .51, and the most rigorous methods reach .57.

.20 to .57
Predictive validity, unstructured judgment to rigorous structured scoring
The scorecard is not paperwork; it more than doubles how well the process predicts on-the-job success.

That gap is also why a shared scorecard makes two graders converge. When both reviewers score the same behavioral anchors, private judgment stops driving the outcome. The scorecard replaces "she felt strong" with "she scored 1 on payments experience because she shipped X." That sentence is gradeable. The feeling is not.

What a criterion looks like when it lies

Every criterion you score can produce a false positive, and the standard is only as good as your ability to spot one.

  • A title proves seniority until it does not. A "Senior Engineer" at a company where everyone is senior is not the same signal as the same title at a company that gates the level. Score demonstrated scope, not the word.
  • A degree or certification proves a credential, not capability. Research on hidden workers shows a large share of employers lose qualified candidates because filters optimise for credentials over what people can actually do.
  • Tenure proves time served, not impact. Read what shipped, not how long they stayed.

When a criterion is credential-shaped, write the evidence note in terms of demonstrated ability. If you cannot, downgrade the score.

Grade the list as a whole

A list passes composition when every must-have is covered by a named candidate, the slate meets a minimum-of-two threshold for any group you are tracking, and no single current employer dominates. These are three separate checks, and a list can fail composition while every individual candidate passes.

The first check is coverage. Map each must-have to at least one named candidate before you submit. A list of eight that looks full but leaves a must-have with zero qualifying candidates is the classic number-chasing failure: it reads as complete and staffs nobody.

The second check is the slate cliff. This is the composition rule with the hardest evidence behind it, and it is a cliff rather than a slope. A single underrepresented candidate on a shortlist is statistically near-invisible to the decision. Two changes everything: with two women finalists the odds of hiring a woman were 79.14x greater, and with two minority finalists the odds of hiring a minority candidate were 193.72x greater. This is why the standard encodes a minimum of two, not a "consider diversity" line. One is not a slate; it is a token that the status quo defeats.

From applicants to a submitted shortlist

  1. Applicants
    100

    about 14% reach interview per industry survey

  2. Longlist
    15-25

    meet essential criteria, scored against the matrix

  3. Shortlist
    4-8

    cut to cover every must-have

  4. Final round
    2-3

    comparison points for the decision

Volume narrows sharply, so the composition checks happen on a small, high-stakes set.

The third check is employer concentration. There is no documented public control for this, so it is a manual read: if every candidate comes from one current employer or one near-identical background, you have built an echo, not a comparison set. A manager reviewing a slate where everyone worked at the same place gets no real choice. Spread the current employers.

The reason these checks land on you rather than a system is structural. In Refolk's index, dedicated sourcers are a thin layer: about 1.1 sourcers exist per 100 recruiters in the US. In most teams the person composing the shortlist is also the one who screened it, which is exactly the single-reviewer condition the calibration pass exists to break.

MarketRecruiter / Tech Recruiter / TA profilesShare of the two-market total
United States112,66193.9%
United Kingdom7,2836.1%

When you need to build a slate that already satisfies these composition rules rather than retrofitting it, ask for it directly. Refolk lets you specify employer diversity, must-have and desirable criteria, and underrepresented-group coverage in the query itself, so the list arrives pre-shaped against the standard instead of being cut down after the fact.

The procedure

Run these eight steps in order. The first two happen before you touch the pool; the composition and calibration checks are where a list is made or broken.

From brief lock to logged submission

  1. Lock the brief with the hiring manager
    Run a 30 to 60 minute intake and agree must-haves versus desirables in writing, along with the business goal behind the hire. Done means a signed criteria list both of you can point to later.
  2. Build the scorecard from the brief
    Convert the locked criteria into 5 to 7 job-relevant items with a rating scale. Done means a scorecard whose fields match the brief one for one, ready for evidence notes.
  3. Screen the pool to a longlist
    Scan applicants against the essential criteria, roughly 7 to 15 seconds per resume and 2 to 3 minutes for a closer read, down to a longlist of 15 to 25 who meet the must-haves. Done means a longlist selected against the matrix, not intuition.
  4. Score and rank against the matrix
    Apply a points-based method: one point for fully meeting a criterion, half a point for partial, then total and rank. Done means a ranked list where every position is backed by an evidence note.
  5. Compose the shortlist against the brief
    Cut to 4 to 8 candidates, confirm every must-have is covered by at least one named person, and check slate composition against the two-in-the-pool threshold and employer concentration. Done means the list covers all must-haves and no single employer dominates.
  6. Run a second-reviewer calibration pass
    Have a peer or panel of at least two people re-score and resolve disagreements in a short calibration meeting. Done means no candidate advanced on a single reviewer's judgment.
  7. Assemble the submission pack
    For each candidate attach competencies, ratings with evidence notes, and a recommendation on next steps. Done means a manager could decide without asking a follow-up question.
  8. Submit and log the rationale
    Send the dated pack and record why each candidate was advanced or held for compliance and audit. Done means the submission is out and the reasoning is written down.

How this goes wrong

Most bounced shortlists trace to one of eight failure modes. Each has a false positive - a list that looks ready and is not - and each has a check that catches it. This is the part of the standard that earns its keep.

Number-chasing. Hitting a target count while a must-have has zero qualifying candidates. The false positive is a full-looking list of eight that no one can actually do the job. Check: map every must-have to at least one named candidate before submitting.

Intuition masquerading as scoring. Speeding through resumes and picking the ones that feel right. That feeling is usually pattern recognition biased toward candidates who resemble people previously hired. The false positive is a ranked list that looks scored but was sorted by gut. Check: every ranking has an evidence note.

Single-reviewer bias. One grader's mental scorecard driving the whole list. Single-reviewer decisions carry the highest bias risk, and with roughly 1.1 sourcers per 100 recruiters the screener and composer are usually the same person. Check: two reviewers signed and scores calibrated.

Credential filter, not capability. Looks-right-on-paper candidates that bounce at manager review. Rising volume makes this worse: recruiters manage 56% more open requisitions than in 2022, and the initial scan is 7 to 15 seconds, which is exactly how credential-over-capability filtering creeps in. Check: score demonstrated ability, not titles or degrees.

The "only one" slate. A lone underrepresented or atypical candidate is statistically near-invisible. The false positive is a slate you call diverse that has exactly one. Check: the two-in-the-pool threshold is met.

Employer-clone concentration. Every candidate from one current employer or one identical background. There is no documented public control, so this is a manual read. Check: no single current employer over-represented.

Missing decision fields. A pack that forces the manager to ask follow-ups. The false positive is a tidy-looking record missing an evidence note or a recommendation. Check: each candidate has competencies, rating, evidence, and recommendation.

Stale list. A slow submission loses candidates: half of surveyed applicants turned down an offer because the process took too long. The false positive is a strong list assembled two weeks ago whose top names are gone. Check: the submission is dated and candidates confirmed still active.

What to do with a candidate on the shortlist

Must-have coveredMust-have not covered
Cut, and note the gap
No evidence and no coverage, this candidate is padding the count
Hold for verification
Covers a must-have but the score is unbacked, get the evidence before submitting
Reject for this role
Well documented but does not meet a must-have, right person wrong brief
Submit
Covers a must-have with strong evidence, this is what the shortlist is for
Weak or missing evidenceStrong documented evidence
Two axes decide the call: does the evidence back the score, and is a must-have covered.

The pass/fail checklist

Run this before the list leaves your hands. It is written so two graders applying it to the same shortlist reach the same verdict. If any item fails, the list is not ready and you make one more pass.

Submission-ready shortlist

  • The must-have and desirable criteria are signed off in writing by the hiring manager.
  • The scorecard has 5-7 job-relevant criteria, each mapped to a criterion in the brief.
  • Every must-have criterion is covered by at least one named candidate on the list.
  • Every candidate's ranking is backed by an evidence note, not intuition.
  • Each candidate record carries competencies, ratings, evidence notes, and a recommendation.
  • A manager could decide on each candidate without asking a follow-up question.
  • The slate meets the two-in-the-pool threshold for any group being tracked.
  • No single current employer is over-represented across the candidates.
  • A second reviewer has scored the list and scoring disagreements are resolved.
  • The list is 4-8 candidates for a standard role, sized to the brief and pool.
  • The submission is dated and every candidate is confirmed still active.
  • The rationale for advancing or holding each candidate is logged for audit.

Adopt it and keep it current

Adopt this as a written policy your team grades against, not a mental habit, because the whole point of a standard is that two people apply it identically. Paste the checklist into your applicant-tracking system as a required gate before the "submit to hiring manager" step, and require the second-reviewer sign-off as a field, not a courtesy.

Two things drift over time and need a re-check cadence. First, the brief: criteria signed for one requisition go stale as the role or market shifts, so re-lock the must-haves at the start of every search rather than reusing last quarter's list. Second, the size band: the 4-8 range is a default, not a rule, and high-volume or executive roles legitimately sit outside it. Where you deviate, write down why in the submission log so the deviation is a decision, not a slip.

The standard's hardest clauses are the composition checks, because they have no system enforcing them and depend on you building a slate that already spans employers and clears the two-in-the-pool threshold. Template a copy-pasteable candidate line and a manager-facing intake so the fields are never improvised.

Candidate line for the submission pack
Name / current role / current employer:
Must-haves covered: [criterion] - met / partial - [one-sentence evidence note]
Desirables met: [criterion] - [one-sentence evidence note]
Total score (points-based): X of Y
Recommendation: Advance / Hold / Reject - because [reason]
Confirmed active as of: [date]

One block per candidate. Fill every field; an empty field is a failed record.

Hiring-manager intake, criteria section
Business goal behind this hire:
Must-haves (a candidate is disqualified without these):
1.
2.
3.
Desirables (nice to have, used to rank):
1.
2.
Slate rules for this role (two-in-the-pool group, employer spread):
Target shortlist size and why:
Signed (manager) / date:

Fill this with the manager and get it signed before sourcing. It becomes the scorecard.

Keep the checklist open while you compose. The measure of the standard is simple: a list that passes every item should not come back, and if it does, the reason belongs in the checklist as a new item for next time.

Questions practitioners ask

How many candidates should I shortlist for interview?

There is no single right number, and chasing one is the wrong focus. Published ranges run from 3-5 up to 10-20 per role, with a common sweet spot of 4-6, fewer than 3 giving too few comparison points, and more than 8 causing interview fatigue. Size to the brief and the pool. For 100 applicants, 8-12 is workable. Composition against the must-haves, not the count, is what decides whether the list bounces back.

What makes a shortlist bounce back from the hiring manager?

The dominant documented reason is misalignment on the brief: 34% of organisations cite hiring-manager misalignment as a core talent-acquisition challenge. Others are credential-over-capability screening, undocumented criteria producing inconsistency between reviewers, and submission packs that force the manager to ask follow-up questions. A list bounces because it misses the agreed must-haves or cannot be decided on without more information, not because it had nine names instead of six.

What should each candidate record contain before I submit?

Each candidate should carry the job-relevant competencies, a rating on each, an evidence note recording what the candidate actually said or did, and a recommendation on next steps. The test is whether a manager can decide without asking you a single follow-up. If any field is missing or a rating has no evidence behind it, the record is not submission-ready.

Do I really need a second reviewer if I sourced the list myself?

Yes. Single-reviewer decisions carry the highest bias risk, and in most teams the person composing the shortlist is also the one who screened it, which is precisely that condition. A second grader using the same scorecard is what makes two people grade a case identically. Predictive validity rises from .20 for unstructured judgment to .51 with a structured scorecard, so the calibration pass is not optional overhead.

What is the two-in-the-pool rule and does it belong in a shortlist standard?

Two-in-the-pool means a shortlist should include at least two candidates from any underrepresented group you are tracking, never exactly one. It belongs in the standard because the effect is a cliff, not a slope: a lone underrepresented finalist is statistically near-invisible, but two finalists raise the odds of hiring a woman by 79.14x and a minority candidate by 193.72x. Encode a minimum of two, not a vague nicety.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next