Finding the Stage That Kills an Outbound Sourcing Funnel
You can isolate the one sourcing stage converting furthest below benchmark and attribute it to data, targeting, message, or speed instead of blaming volume.
Key takeaways
- The fixable leak is the stage furthest below benchmark, not the stage with the largest raw drop-off; because roughly 97% of applicants are eliminated before speaking to a human, the biggest absolute drop is structural and not yours to fix.
- A 35% sourced-to-reply rate paired with a 5% reply-to-screen rate is a targeting failure wearing a messaging success, and industry medians cannot tell the two apart.
- Low-rate stages are statistically invisible at normal recruiter volume: a sub-1% offer stage needs hundreds of candidates to read, so most 'offer problems' are just noise from one bad week.
- Bounce sits upstream of everything and compounds non-linearly, with well-maintained lists under 1.5% while the average sender runs 5.1%, so pull hard-bounce rate before you ever rewrite copy.
- Personalized recruiter messages get 18-25% response versus 5-8% for templates, a 3-4x multiple, and InMails under 400 characters earn 22% higher response than average.
- Fixing sourcing beats adding inbound volume because outbound converts at up to Gem's 8x multiplier, so repairing one sourcing stage compounds through the whole funnel.
Your outbound sourcing is producing few hires, and the instinct is to send more. This guide is for in-house recruiters, sourcers, and founders running their own outreach who want to find the single stage where candidates leak and fix that one stage instead of scaling a broken funnel. I carry one worked funnel from sourced to hire, name the constraint stage against benchmark, attribute it to a specific cause, and mark the wrong turns most recruiters take before they get there.
The public answer for recruiters today is a pile of vendor benchmark lists. They tell you the median reply rate. They never tell you which stage of your own funnel is broken. Sales has lever-by-lever pipeline diagnostics; recruiting has a dashboard with two ends and a black box between them. This is the diagnostic.
Why "send more" is the wrong first move
The constraint in a sourcing funnel is almost never the top, but the top is where recruiters look first. The fixable leak is the stage converting furthest below benchmark, and that stage is usually in the middle.
Here is the reasoning. In the broad hiring funnel, roughly 97% of applicants are eliminated before they ever speak with a human. That is the largest absolute drop by far, and it is structural. It is not a defect you introduced and not one you can message your way out of. If you rank your stages by raw drop-off, you will always land on the biggest one, waste your effort there, and leave the real leak untouched.
Adding volume has the same failure signature. More sends lift the raw number of replies, so the dashboard looks better, but the stage that was converting below benchmark still converts below benchmark. You have made the leak bigger, not smaller. The only honest unit of measurement here is the per-stage rate, not the count.
There is also a compounding argument for fixing sourcing rather than pouring in more inbound. Outbound converts far above inbound: Gem reports an 8x outbound-vs-inbound multiplier, SeekOut puts sourced candidates near 5x, and referrals run 4-10x. Repairing one sourcing stage compounds through the whole funnel. Inbound volume dilutes against that 97% screening loss.
The one funnel I will carry through this guide
I will use a single recruiter's outbound funnel over one measurement window and follow it all the way to the fix. The counts below are the worked example this whole guide runs on.
The recruiter sourced 600 candidates for a senior backend role over four weeks, using LinkedIn and email. The stage counts came out like this:
| Stage | Count | Pass-through |
|---|---|---|
| Sourced / contacted | 600 | - |
| Replied | 210 | 35% |
| Screened | 21 | 10% |
| Onsite | 8 | 38% |
| Offer | 5 | 63% |
| Accepted | 4 | 80% |
Read the two ends first, because that is all most dashboards show. Sourced-to-reply is 35%, well above the 18-25% personalized benchmark. Offer-to-accept is 80%, inside the 70-85% band. By the numbers most teams look at, this funnel is healthy. It produced four hires, which feels fine.
It is not fine, and the next sections show why. The stage that is quietly killing this funnel does not appear at either end.
The worked funnel, sourced to accepted
- 600Sourced / contacted
top of funnel
- 210Replied
35%, above the 18-25% benchmark
- 21Screened
10%, far below the 30-50% band
- 8Onsite
38%
- 5Offer
63%
- 4Accepted
80%
Benchmarks by stage, and how much volume each one needs to read
Rank stages by their gap to benchmark, not by raw drop-off, and only trust stages that have enough events behind them to not be noise. A stage rate you cannot read is not a diagnosis; it is a guess.
There is no single source that publishes all five outbound stages cleanly. You assemble them. Contacted-to-reply clusters above the cold-email baseline: templated recruiter outreach earns 5-8%, LinkedIn Recruiter's platform average sits around 13%, and personalized outreach referencing specific candidate experience reaches 18-25%. The mid and late stages come from applicant benchmarks that still bound a sourced funnel.
Where a sourcing funnel leaks, top to bottom
- Data (bounce)A dead address never replies; upstream of everything and compounds non-linearly
- MessageBrevity and personalization drive contacted-to-reply
- TargetingWrong-fit list shows high reply, low reply-to-screen
- SpeedSlow follow-up loses candidates who already replied
- Comp / processLate-stage; shows up as low offer-to-accept
The sample question matters as much as the benchmark. Lower conversion rates produce noisier data, so lower-rate stages need proportionally more volume to read. Borrowing from A/B-test statistics, about 100 conversions per variation is a common minimum and 250-400 is the safer ballpark, over a 3-4 week window rather than under two weeks.
| Stage | Benchmark | Approx events needed to read |
|---|---|---|
| Contacted-to-reply | 13-25% | ~100-400 replies |
| Screen-to-interview | 30-50% | ~100-400 advances |
| Interview-to-offer | 15-25% | ~100-400 offers |
| Offer-to-accept | 70-85% | ~100-400 accepts |
Apply this to the worked funnel. Sourced-to-reply is at 35% on 210 replies: readable and above benchmark. Reply-to-screen is at 10% against a 30-50% band: a gap of 20 to 40 points, the largest in the funnel by a wide margin. The stages below it - onsite, offer, accept - sit on 8, 5, and 4 events. Those numbers are too small to read. A 63% offer rate on five candidates could flip on a single decision. So the constraint stage is not a matter of judgment. It is reply-to-screen, and everything downstream is noise until more volume flows through.
The seven-step diagnostic
The procedure below takes your own sourced-to-hire numbers and isolates the one stage converting furthest below benchmark, then attributes it to a cause before you spend a day fixing the wrong thing.
Diagnose the constraint stage
- Instrument the funnel end-to-endPull counts at every stage - sourced or contacted, replied, screened, onsite, offer, accepted - plus the date window. Done is a single row of six numbers. Most teams leave a black box in the middle.
- Confirm the sample is large enough to readCheck each stage has enough events to not be noise; ~100 conversions is a common minimum, 250-400 is safer. Done is each stage flagged readable or under-powered.
- Convert to per-stage rates and lay benchmarks alongsideCompute each stage's pass-through and put the public benchmark next to it. Done is a two-column table, yours versus benchmark, with the gap in points. Use a vertical benchmark, not the blended number.
- Isolate the single stage furthest below benchmarkRank by gap size, not raw drop-off. Done is one named constraint stage. Ignore the largest absolute drop; it is usually structural.
- Attribute the constraint to data, targeting, message, or speedRun the differential in order: hard-bounce rate, then reply quality, then reply-to-screen ratio, then time-to-first-follow-up. Done is one cause with evidence.
- Apply the one matching fix and hold everything else constantShip only the fix that matches the cause; change nothing else in the window. Done is the fix live and a clean comparison window open.
- Re-measure over a full cycle before judgingWait a full 3-4 week cycle so the new rate has adequate sample, then compare. Done is a new per-stage rate with enough events behind it.
Attributing the constraint: data, targeting, message, or speed
Once you have the constraint stage, run a differential in a fixed order: data first, then message, then targeting, then speed. Each cause has a distinct signature, and each lies in a distinct way, so the order stops you from fixing a symptom.
Data (bounce). A dead address never replies, so a data-quality leak looks exactly like a message problem from a low reply rate. Cold-email bounce averages 5.1% across all senders while well-maintained lists stay under 1.5%, and many programs aim under 2%. Bounce compounds non-linearly: moving from 1.8% to 2.3% can be the difference between the primary inbox and the spam folder. Pull hard-bounce rate before you touch a word of copy. What it proves: if bounce is elevated, low reply is a list problem. How it lies: a clean bounce rate does not rule out a bad list, only a dead one.
Message (contacted-to-reply below benchmark, bounce clean). Personalization and brevity move this stage. Personalized messages referencing specific experience get 18-25% versus 5-8% for templates, a 3-4x multiple. InMails under 400 characters earn 22% higher response than average, while messages over 1,200 characters run 11% below it. What it proves: a low reply rate on a clean list points at copy. How it lies: a high reply rate can still hide a targeting failure downstream.
Targeting (reply high, reply-to-screen low). This is the trap in the worked funnel. A 35% reply rate paired with a 10% reply-to-screen rate is a targeting failure wearing a messaging success. The message is working; the people are wrong. Industry medians cannot see this because they only report one stage at a time. Only the ratio of the two adjacent stages exposes it. What it proves: replies that never survive a screen mean the list, not the message, is off. How it lies: it disguises itself as a win on any dashboard that shows reply rate alone.
Speed (mid-funnel decay). Slow follow-up loses candidates who already replied, a leak invisible to any copy audit. 70.8% of candidate replies arrive within 7 days, candidates reply in a median 3.9 days, and speed-to-contact research found reps roughly 100x more likely to make contact at 5 minutes versus 30. What it proves: if time-to-first-follow-up lags, the leak is process, not copy. How it lies: it produces the same low downstream numbers a message problem would.
For the worked funnel, the differential is quick. Bounce came back clean, under 2%. The reply rate is above benchmark, so the message is not the problem. Time-to-first-follow-up was same-day. That leaves the reply-to-screen ratio, which is exactly where the 20-40 point gap sat. The cause is targeting. The 600 sourced candidates replied because the outreach was good, but too many of them were never a fit for the role, so they fell out at screen. The fix is not a better message and not more volume. It is a re-qualified list.
A high reply rate can be a targeting failure wearing a messaging success, and no median will ever tell you.
Re-qualifying the list is the step where sourcing definition, not sending, does the work. This is where I use Refolk: I describe the exact person I want in plain English and get a list back that matches, so the reply-to-screen ratio recovers because the people arriving at screen actually fit the brief.
How this diagnostic goes wrong
Most misdiagnoses come from reading counts as rates, reading noise as signal, or benchmarking against the wrong population. These are the failure modes, each with the check that catches it.
- Blaming volume when the constraint is a rate. Sending more lifts raw replies while the leaking stage still converts below benchmark. Check: compute per-stage rates, never counts.
- Reading a stage that is pure noise. A 0.5% offer stage on 40 candidates cannot be read; a "collapsed" offer rate is often one bad week. Check: confirm about 100+ events and a full 3-4 week cycle before you believe a rate.
- Bounce hidden inside reply rate. A low reply rate can be a dead list, not weak copy. Check: pull hard-bounce rate first; average is 5.1% and healthy is under 1.5%.
- Two funnels that look identical on a dashboard. One is good at messaging and bad at qualifying; its inverse is the opposite. Medians cannot tell them apart. Check: split sourced-to-reply from reply-to-screen and read the ratio.
- Benchmarking against the wrong population. SaaS and software outreach sits near 4.77%, the lowest vertical, due to inbox saturation. Check: use a vertical and function benchmark, not the blended number.
- Speed masquerading as a message problem. If follow-ups lag, replies decay regardless of copy. Check: measure time-to-first-follow-up before rewriting anything.
- Open rate as a signal. Apple Mail Privacy Protection inflates open tracking. Check: never diagnose off opens; treat them as directional only.
- Fixing offer-accept as "candidate quality." Sustained offer-accept below ~75% usually means comp misalignment or process latency, not bad candidates. Check: audit comp and cycle time before you touch sourcing.
The subtlest of these is the two-funnel case. Consider one team at 35% sourced-to-reply but 5% reply-to-screen, and another at 10% sourced-to-reply but 40% reply-to-screen. Both can produce the same hire count and look identical on a summary dashboard. The first is a targeting problem; the second is a message problem. If you only ever look at one stage, you will prescribe the exact wrong fix to each.
Why comparison, not the median, is the whole method
The reason recruiting has lacked a funnel diagnostic is that published benchmarks are medians, and a median tells you where you sit, not why. The method here is comparison between adjacent stages within your own funnel, benchmarked against the right population.
Two disciplines this method borrows from. From A/B testing, it takes the sample-size discipline: a stage rate is only a diagnosis when enough events sit behind it, which is why low-rate stages near the offer are usually unreadable at recruiter volume. From cold-email operations, it takes the upstream-first ordering: bounce before copy, because a data leak wears the costume of a message leak.
Note the missing layer, plainly. This guide is built on public benchmarks. I would normally anchor the spine to Refolk's own index figures - the exclusive sourcer-population and title-versus-skill comparison counts that show how a re-qualified list changes the reply-to-screen ratio in practice. Those index queries did not return in the session behind this guide, so I have not asserted numbers I cannot stand behind. If you want to pressure-test the targeting attribution on your own funnel, the check is local and simple: split reply and reply-to-screen for two lists, one broad and one tightly re-qualified, and read the ratios side by side.
One more caution on the source-quality multipliers, because they are easy to abuse. Published multipliers use different denominators: SeekOut's ~5x, Navero's 4-10x, Gem's 8x, and HrPanda's 11x for referrals and 32x for internal transfers are not measuring the same thing. Do not average them.
| Source | Multiplier vs inbound |
|---|---|
| Sourced / outbound (SeekOut) | ~5x |
| Direct sourcing / referrals (Navero) | 4-10x |
| Outbound (Gem) | 8x |
| Employee referral (HrPanda) | 11x |
| Internal transfer (HrPanda) | 32x |
Use these to justify fixing sourcing over adding inbound, which they clearly support. Do not use them as a single conversion figure.
Before you call the diagnosis done
Run this list before you ship a fix. Each item is a check the failure modes above will punish you for skipping.
Diagnosis readiness
- I have one row of six stage counts and a date window, not just top and bottom of funnel.
- Every stage I am reading has roughly 100+ events behind it, over a full 3-4 week cycle.
- I ranked stages by gap to benchmark, not by raw drop-off.
- I benchmarked against my vertical and function, not the blended platform average.
- I pulled hard-bounce rate before considering any message rewrite.
- I read reply and reply-to-screen as a ratio, not as two separate wins.
- I checked time-to-first-follow-up before blaming copy.
- I ignored open rates entirely as a diagnostic signal.
- My planned fix changes exactly one lever this window.
What to do next and how to keep it current
Run the diagnostic on your live funnel this week, then re-run it every quarter, because benchmarks and your own list quality both drift. The method is stable; the numbers inside it are not.
Three things to re-check on a cadence rather than treat as fixed. First, the reply benchmark for your vertical, since inbox saturation moves it and SaaS already sits near 4.77%. Second, your bounce rate, because list decay is continuous and a list that was under 1.5% will not stay there without maintenance. Third, offer-to-accept against the 70-85% band, since a sustained slide below ~75% is an early signal of comp or process latency that has nothing to do with sourcing.
The discipline that separates this from a benchmark list is the habit of reading adjacent stages as ratios and fixing exactly one cause per cycle. Do that, and the next time your outbound produces few hires, you will name the stage in an afternoon instead of spending a month sending more into a leak.
Questions practitioners ask
Why isn't my recruiting outreach converting even though I'm sending a lot?
Volume lifts raw reply counts but does nothing for a stage that converts below benchmark. Compute per-stage rates instead of counts, then find the single stage whose gap to benchmark is largest. If replies are healthy but reply-to-screen is weak, the defect is targeting, not send volume. Adding more sends into a broken mid-funnel just enlarges the leak while feeling like progress on the dashboard.
What is a good sourced-to-hire conversion benchmark?
There is no single clean sourced-to-hire number published across all stages; you assemble it from per-stage benchmarks. Useful anchors are contacted-to-reply at 13-25%, screen-to-interview at 30-50%, interview-to-offer at 15-25%, and offer-to-accept at 70-85%. Outbound sourced candidates convert far better than inbound, with published multipliers running from about 5x up to Gem's 8x, so compare your funnel against sourced benchmarks rather than applicant ones.
What reply rate should a recruiter expect?
Templated recruiter outreach gets 5-8%, LinkedIn Recruiter's platform average is around 13%, and personalized messages referencing specific candidate experience reach 18-25%, a 3-4x multiple. LinkedIn also enforces roughly 13% as a floor across 100+ InMails per 14-day period. SaaS-to-SaaS outreach runs lower, near 4.77%, so benchmark against your vertical rather than the blended average.
How many candidates do I need before a stage rate is trustworthy?
Borrowing from A/B-test statistics, about 100 conversions per variation is a common minimum and 250-400 is the safer ballpark, over a 3-4 week window rather than under two weeks. Lower-rate stages produce noisier data and need proportionally more top-of-funnel volume. A sub-1% offer stage needs hundreds of candidates to read, while a 15% reply stage stabilizes quickly.
How do I know if a low reply rate is a data problem or a message problem?
Pull the hard-bounce rate before touching copy. Cold-email bounce averages 5.1% across all senders while well-maintained lists stay under 1.5%, and many programs aim under 2%. If bounces are elevated, you have a data-quality leak dragging every downstream stage, and rewriting your message will not move it. Only after bounce is clean should you audit brevity and personalization.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.