Refolk
PlaybookRecruiting and sourcing

The Calibration Batch: Lock a Search Before You Source

You will run a calibration batch that converts a hiring manager's reactions into signed-off non-negotiables, tradeable requirements, and target companies before scaled sourcing.

14 min readLast reviewed September 6, 2026Read as Markdown

Before you spend a week sourcing, you show a hiring manager a small batch of sample profiles and turn their reactions into a search-ready brief. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring. It hands you the countable loop: how many profiles to stage, how to spread them, how to capture reactions, and the exact rule that converts each reaction into a revised must-have, a tradeable requirement, or a target company.

Most intake advice tells you to "review profiles together" and stops there. This is the procedure that follows: a batch-and-decision process you can run start to finish and end with a brief that two recruiters could source against.

Why a calibration batch beats an intake questionnaire

A calibration batch works because it trades on recognition instead of recall. Hiring managers who cannot answer "what are you looking for?" can look at a spread of real profiles and describe exactly what they want. That is the whole mechanism.

Intake questionnaires ask the manager to generate a spec from memory. Most managers cannot, so they fall back on words like "senior" and "strategic" that do not narrow anyone. A batch flips the direction: you present concrete people, they react, and their reactions carry more information than any answer to an abstract question. This is why off-target profiles are more useful than flattering ones. A profile the manager rejects tells you where the real bar sits.

Calibration is the single highest-impact move in active sourcing and the most-skipped one. Before sending a single message, you build a shared definition of "great" against actual profiles. Skip it, and when you and the manager are not sharply aligned on "qualified," every profile you surface is a guess.

22% to 13%
Mis-hire rate under better hiring-manager alignment
Bersin by Deloitte data on alignment automation; the counterweight to calibration feeling like overhead.

The cost of skipping is not abstract. If you discover you and the manager are misaligned after you are already bringing people in for interviews, it costs far more than resolving it at the start. A bad hire runs roughly 30% of first-year salary. Calibration is a one-hour insurance premium against that.

What "done" has to produce

A calibrated search produces written agreement on five things, and a profile detailed enough that two independent recruiters would source similar people. Anything less is a moving target, not a brief.

The load-bearing definition is blunt: if you leave intake without agreement on target companies, non-negotiable skills, compensation reality, interview process, and rejection criteria, you do not have a search brief. You have a moving target. The falsifiable test is reproducibility, not thoroughness. If two recruiters read your profile and would go after different people, it fails. If the profile is only "senior," "strategic," or "culture fit," it fails.

Hold the deliverable to this test at sign-off. It is the check most intake templates never impose, and it is the one that stops a week of misaimed sourcing.

Building the batch: how to spread the profiles

Stage real sample profiles across disparate backgrounds, deliberately spread from clearly strong to clearly off-target. They need not be available candidates. This is an exercise in surfacing the bar, not a shortlist.

The practitioner method is to bring profiles covering disparate backgrounds, experiences, and skills, and let the manager pick and choose their favourite parts of each. No single count is standardised. The adjacent practice of interview-panel calibration uses a tight spread, 1 to 2 profiles per reviewer scored against a rubric. For a sourcing calibration, size the batch by coverage rather than volume: enough profiles that each major requirement axis is represented, and at least one profile you fully expect to score low.

The strong/borderline/off-target spread below is a construction, not a published standard, but it maps cleanly onto what each profile is meant to test.

Profile typeWhat it testsWhat a good reaction looks like
Clearly strongWhether your read of "great" matches theirsManager scores it 4 to 5 and names the parts they value
BorderlineWhere the must-have line actually sitsManager hesitates and articulates the missing piece
Off-targetThe real floor and unstated deal-breakersManager scores it low with a written, specific reason

If every staged profile is strong, the manager never reveals the real bar. Unanimous 5s are a failure signal, not a success. Confirm before the session that your batch contains at least one profile designed to fail.

The rubric: 6 to 10 criteria with anchored scores

Score every profile against a one-page rubric of 6 to 10 criteria, split into must-have versus nice-to-have, each with a description of what a 1, 3, and 5 look like. This is the instrument that makes reactions comparable.

Identify the 6 to 10 skills and behaviours that predict success in the role. Separate must-haves from nice-to-haves. For each criterion, write concrete anchors: what a score of 1 looks like, what a 3 looks like, what a 5 looks like. The 1 to 5 scale with anchored descriptors is the default across calibration practice for a reason. Without anchors, a "3" means nothing and scores drift toward optimism.

Calibration rubric row
Criterion: <e.g. Backend systems at scale>
Type: Must-have | Nice-to-have
Score 1 (below bar): <what disqualifying looks like>
Score 3 (meets bar): <what the floor looks like>
Score 5 (exceptional): <what standout looks like>
Manager score: __  Written reason if below 3: <required>

Repeat for each of your 6 to 10 criteria. Mark each as must-have or nice-to-have before the session, not during it.

Keep the rubric to one page. If you cannot fit it on a page, you have too many criteria and the session will run long without adding precision. The upper bound of 10 exists to force prioritisation.

What a calibration brief is built from

  1. Sign-off agreement
    The five items and the feedback SLA, in writing
  2. Rubric
    6 to 10 must-have and nice-to-have criteria with 1/3/5 anchors
  3. Reactions
    Scores and written reasons captured from the batch
  4. Sample profiles
    The staged spread from strong to off-target
Each layer constrains the one below it; skip a layer and the sourcing query has nothing to hold onto.

Run the loop: the eight-step procedure

Run the calibration as a fixed sequence: draft the rubric, assemble the batch, hold the session, capture reactions independently, convert each reaction, run the sign-off test, decide whether to proceed, and lock the SLA. Here is the whole loop.

The calibration batch, start to finish

  1. Draft the rubric
    Turn the job description into 6 to 10 scored criteria split into must-have versus nice-to-have, each with 1/3/5 anchors. Done means a one-page rubric. Budget 30 to 45 minutes.
  2. Assemble the batch
    Pull real sample profiles across disparate backgrounds; they need not be available candidates. Done means a staged spread from strong to deliberately off-target. Budget 1 to 2 hours.
  3. Run the calibration session
    Book 60 to 75 minutes for the first session. Walk the rubric for 10 to 15 minutes, then have the manager score each profile and name favourite parts of each.
  4. Capture reactions independently, then reveal
    Each reviewer scores against the rubric with no discussion, then reveal side by side. Done means every score below 3 has a written reason.
  5. Convert each reaction
    Apply the pre-close and the trade-off question to turn each accept or reject into a revised must-have, a tradeable requirement, or an added or removed target company. Done means an updated rubric plus a target list.
  6. Run the sign-off test
    Confirm agreement on target companies, non-negotiable skills, compensation reality, and rejection criteria. Apply the two-recruiters-would-source-the-same-people test.
  7. Decide: proceed or second batch
    If the manager rejected the whole spread or refused all trade-offs, stage a second batch or run a time-boxed market test with a checkpoint set in advance.
  8. Set the feedback SLA and scale
    Lock a 24 to 48 hour feedback turnaround before scaled sourcing starts. Done means the SLA is written into the brief.

The first session runs 60 to 75 minutes. Recalibration later runs about 30. The independent-scoring step matters more than it looks: if the manager sees your scores first, they anchor to them and you lose the signal. Score separately, then reveal.

Convert each reaction into a criterion or a company

Turn every reaction into a decision using two questions: a pre-close and a trade-off. The pre-close pins down whether a described candidate is a yes. The trade-off exposes which requirements can actually move.

The pre-close: "If I find a candidate from Company X who has Y qualifications and makes $Z a year, would you hire them?" Go back and forth until they agree. Each yes or no hardens a must-have or reveals a nice-to-have masquerading as one.

The trade-off question: "If we find someone exceptional in two of these three areas, which requirement can move?" That single question exposes whether the manager understands the market or is holding out for a unicorn. It is a market-literacy test, not just an alignment step.

The trade-off question is not about compromise. It reveals whether the manager understands supply.

For target companies, do not collect logos. Interrogate the underlying attribute. Ask what makes a company a good hunting ground: customer base, sales motion, technical environment, regulatory exposure, or team size. Recharacterising "hospital-system experience" as "high-volume, compliance-sensitive operations" multiplies the reachable pool and finds adjacent talent instead of recycling the same three obvious competitors.

The supply numbers make the stakes concrete. Loosening a single title requirement changes reachable supply by an order of magnitude.

TitleProfiles (US)Top employers
Sourcer / Talent Sourcer1,223MongoDB, Gartner, Microsoft
Technical Recruiter20,661Experis, K2 Partnering, Snowflake, Google
Derived multiple16.9x-

In Refolk's index of professional profiles, reframing "Sourcer" to the broader "Technical Recruiter" title in the US moves the reachable pool 16.9x, from 1,223 to 20,661. Geography does the same in reverse. The same sourcer title is 11.2x scarcer across the UK border, which is exactly the kind of supply reality the trade-off question is meant to surface before you commit.

TitleCountryProfiles in indexTop employers
Sourcer / Talent SourcerUnited States1,223MongoDB, Gartner, Microsoft
Sourcer / Talent SourcerUnited Kingdom109Deloitte, AMS, Adecco, Bloomberg
Derived ratioUS : UK11.2x-

When the manager reframes a requirement into an attribute, you want to test the new pool immediately, not a week later. Describing the reframed profile in plain language and getting the reachable set back is exactly where a sourcing tool earns its place.

I built Refolk for this moment: you have just recharacterised a requirement by attribute and you want to see the pool it actually reaches, in plain English, without rebuilding a boolean string.

How this goes wrong

Calibration fails in predictable ways, and each has a check that catches it before you scale. The most common failures are a batch that is too flattering, non-answers dressed as criteria, a manager who refuses all trade-offs, and reactions captured verbally instead of in writing.

Here is the failure catalogue and the check for each.

  • Batch too flattering. If every profile is strong, the manager never reveals the real bar. False positive: unanimous 5s. Check: include off-target profiles and confirm at least one scored low with a written reason.
  • "Culture fit" non-answers. A profile described only as "senior," "strategic," or "culture fit" is not usable. Check: apply the two-recruiters-would-source-the-same-people test.
  • Unicorn manager. Refuses every trade-off, agrees all requirements are essential. Check: run "which of the three can move?"; if none, run a time-boxed market test.
  • Budget-versus-target conflict hidden. A manager who wants a $220,000 candidate for a $160,000 package must decide now: increase comp, reduce scope, accept a step-down, or broaden the market. Check: test the comp band against the agreed must-haves before sourcing.
  • Scorecard drift. Scorecards drift toward optimism and the rubric goes stale as the role evolves. Check: recalibrate on a recurring cadence.
  • Reactions captured verbally. Low-score reasons evaporate. Check: require a written reason for any score below 3.
  • Feedback SLA absent. The manager stalls because "nothing feels quite right" and the role becomes "hard to fill." Check: lock a 24 to 48 hour SLA at sign-off.
  • Calibration skipped entirely. Every profile becomes a guess. Check: no scaled outreach until the sign-off test passes.

The budget conflict deserves special attention because it hides so well. A manager can agree to every must-have in good faith while the comp band cannot buy any of them. Test the band against the agreed must-haves before you source, not after the first three candidates decline.

Lock the feedback SLA and keep the brief current

A feedback SLA is part of the calibration deliverable, not a later add-on. Lock a 24 to 48 hour turnaround at sign-off, in writing, before scaled sourcing begins.

Managers stall on ambiguous searches. The manager delays feedback because nothing feels quite right, then the role becomes "hard to fill." Benchmark feedback already sits high, so a rule set at sign-off prevents the spiral. These are the timing anchors worth holding your SLA against.

SignalValueSource
Lever avg feedback time37 hourslever.co
Recommended feedback SLA2 days / 48 hrslever.co; treegarden.io
CV shortlist SLA (specialist)72 hourstreegarden.io
Paxos resume-review SLA24 hoursshrm.org

Where a calibration batch narrows the picture

  1. Staged sample profiles
    many

    strong to off-target

  2. Scored with written reasons
    fewer

    below-3 reasons captured

  3. Converted to criteria and companies
    fewer

    must-haves, trade-offs, targets

  4. Signed-off brief
    one

    passes the two-recruiter test

The batch converts a wide staged spread into a signed-off brief before any outreach happens.

Keep the brief current after sign-off. Scorecards drift toward optimism and the rubric goes stale as the role profile evolves. Recalibration is cheap: about 30 minutes against a fresh handful of profiles. Run it whenever the market surprises you, whenever the manager rejects three sourced candidates in a row, or whenever the role scope shifts.

Before you call the calibration done, run this check.

Calibration sign-off checklist

  • The rubric has 6 to 10 criteria, each marked must-have or nice-to-have, each with 1/3/5 anchors
  • At least one staged profile scored below 3 with a written reason
  • Every score below 3 across the batch has a written reason attached
  • Each must-have survived a pre-close question and is genuinely non-negotiable
  • The trade-off question was asked and at least one requirement was named as movable
  • Target companies are defined by attribute, not only by logo
  • The comp band was tested against the agreed must-haves
  • Two independent recruiters would source similar people from the profile
  • A 24 to 48 hour feedback SLA is written into the brief

If any item fails, you are not cleared to scale. Stage a second batch or run a time-boxed market test with a checkpoint set in advance. Do not spend a month proving what one more hour of calibration would have settled.

What to do next

Run your next real requisition through this loop before you touch a sourcing tool. Draft the rubric from the job description, spend an hour assembling a spread that includes at least one profile designed to fail, and book 75 minutes with the manager. Score independently, convert every reaction with the pre-close and the trade-off question, and hold the result to the two-recruiter test.

Then measure it. Only 20% of organisations measure quality of hire, so even a rough count of how many sourced candidates the manager rejects, and why, will tell you whether the calibration held. If rejections cluster on a criterion you thought was settled, recalibrate against a fresh batch. The loop is cheap to rerun and expensive to skip.

Questions practitioners ask

How many profiles should I put in a calibration batch?

No count is standardised, but the working method is a small deliberate spread that covers disparate backgrounds so the manager can pick favourite parts of each. Interview-panel calibration, an adjacent practice, uses just 1 to 2 profiles per reviewer. Aim for enough range that at least one profile is deliberately off-target and scores low with a written reason. The point is coverage of the requirement space, not volume.

Should calibration happen before or after sourcing starts?

The active-sourcing camp is firm that it happens before any outreach: build a shared definition of great against real profiles before sending a single message. Executive search runs a formal calibration meeting around 7 to 10 days into sourcing, using early market intelligence. For in-house work I side with before: an uncalibrated search means every profile is a guess. Recalibrate later if the market surprises you.

How do I turn a hiring manager's reaction into a search criterion?

Use a pre-close and a trade-off question. The pre-close is 'if I find a candidate from Company X with Y qualifications at $Z, would you hire them?' and you go back and forth until they agree. The trade-off is 'if we find someone exceptional in two of these three, which requirement can move?' Each answer becomes a revised must-have, a tradeable requirement, or an added or removed target company.

What does a finished calibration produce?

Written agreement on target companies, non-negotiable skills, compensation reality, interview process, and rejection criteria. The falsifiable test is that two independent recruiters could read the profile and source similar people. If the profile is only 'senior', 'strategic', or 'culture fit', it is not usable and you do not yet have a brief.

What if the hiring manager rejects every profile in the batch?

Rejecting the whole spread or refusing every trade-off signals mis-aim. Document the agreed profile, timeline, and risks, then run a limited market test with a checkpoint set in advance. Do not spend a month proving what you already know. If they insist every requirement is essential, that is a unicorn manager and the comp band or scope usually has to move.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next