# The Sourcing Funnel Leak Diagnosis Standard

*You can read one search's stage counts, rank each transition against outbound baselines, discard noise and upstream causes, and name the single stage costing you the most qualified candidates.*

- Canonical URL: https://www.refolk.ai/guides/diagnose-sourcing-funnel-leak
- Pillar: Process, data, and compliance
- Format: Framework
- Published: 2026-08-22
- Last reviewed: 2026-08-22
- Reading time: 14 min
- Keywords: where is my sourcing funnel leaking, sourcing funnel conversion benchmarks, sourcing pass-through rate by stage, diagnose recruiting funnel bottleneck, outbound recruiting funnel metrics, response rate vs screen rate

## Key takeaways

- Score outbound funnels against outbound baselines: sourced candidates are roughly eight times more likely to be hired than inbound applicants, so an inbound 3% applicant-to-interview rate flags a healthy sourcing funnel as broken.
- Match reply rate to the recruiting InMail band of 18 to 25%, not the 3.43% cross-industry cold-email floor, or you will misdiagnose a normal stage as a leak.
- A LinkedIn InMail response rate below 13% over 100 or more messages in a 14-day window triggers a send cap, so a response leak mechanically becomes a capacity leak in the next window.
- Set the small-sample gate from the identification pool before outreach: a Rust-in-Germany search starts from 87 people in Refolk's index versus 1,137 for Go-in-US, so the same reply rate is noise in one and signal in the other.
- Rank stages by qualified candidates lost, computed as ratio times volume times downstream cost, not by the lowest percentage alone.
- Screen leaks are the most over-diagnosed stage because latency hides there: with 13 interviews per hire and 44 to 62 day fills, week-three dropouts get coded as screen fails.

You have the stage-by-stage numbers for one search in front of you and you need to decide which single stage to fix first. This guide is for the recruiting and revenue operations people answerable for those numbers, and it gives you a scoring model that reads one funnel against outbound-specific baselines, corrects for small-sample noise and upstream masking, and names the one stage costing you the most qualified candidates.

Most of what ranks for "where is my sourcing funnel leaking" hands you a list of benchmark numbers and stops. Worse, most of those numbers come from inbound application funnels and get applied to outbound sourcing, which is a different machine. This standard is built to do the one thing those pages skip: isolate the true bottleneck in a single funnel and rule out the false ones.

## Why outbound funnels need their own baselines

Score an outbound funnel against inbound benchmarks and you will condemn a healthy funnel. Sourced candidates are nearly eight times more likely to be hired than inbound applicants, so every gate in a sourcing funnel should convert far higher than the applicant-pool rates the listicles quote.

The mechanism is channel yield. Job boards and company marketing generate roughly 90% of all applications but account for only about half of hires. Direct sourcing delivers 11% of hires from just 2.6% of applications, a 4x yield. Your outbound funnel starts from a curated list, not a self-selected applicant flood, so it enters every stage with better-qualified people.

That is why a benchmark like "3% applicant-to-interview" misleads. It is an inbound applicant-pool rate. An outbound funnel that converts at 3% applicant-to-interview is not sick, it is dead. The table below is the reference divergence you score against.

| Stage | Outbound | Inbound | Source |
|---|---|---|---|
| Top-of-funnel passthrough | 67% | 8% | Gem 2022 |
| Ultimately hired | 6% | 1% | Gem 2022 |
| Hire-likelihood multiplier | 8x | 1x | Gem 2026 |

**8x - How much more likely a sourced candidate is to be hired than an inbound applicant**

Gem 2026, up from 5x the prior year, so mis-baselining outbound against inbound gets worse each year.

The gap is widening. The multiplier moved from 5x to roughly 8x in a single year. A team still scoring outbound against inbound norms is drifting further off with every benchmark refresh, so pinning your model to outbound-native numbers is not a one-time fix.

## The five stages and what each transition proves

Define five stages with one unambiguous definition each: identification, outreach-sent, response, screen, and handoff. Each transition between them is a ratio, and each ratio proves something specific about a different part of the machine.

#### The outbound sourcing funnel

1. **Identification** - A named person enters the list from a search
2. **Outreach-sent** - A first message is delivered to that person
3. **Response** - The person replies, positive or negative
4. **Screen** - A recruiter conversation happens
5. **Handoff** - The candidate moves to the hiring team

*Four transitions, each with a distinct root cause when it leaks.*

Here is what each transition tells you, and what it looks like when the number lies to you:

- **Identification to outreach-sent.** Proves your list is actually reachable and worth contacting. If this leaks, contacts are being dropped before send, often to enrichment failures or capacity caps. It lies when a send restriction throttles volume and looks like a list problem.
- **Outreach-sent to response.** Proves the list and the message fit. Score against 18 to 25% recruiting InMail, not the 3.43% cold-email floor. It lies when the list is off-spec, because a bad list produces low response that looks like a copy problem.
- **Response to screen.** Proves responders are the right people and your scheduling holds. It lies when latency between reply and conversation lets warm candidates go cold, which reads as disinterest.
- **Screen to handoff.** Proves your screen is calibrated to the hiring team's bar. It lies more than any other stage, because week-three dropouts caused by process latency get booked here as screen fails.

> **Rule:** Always break out every transition
>
> Two funnels with identical headline reply rates can hide opposite problems: one leaks reply-to-screen, the other sourced-to-reply. Never diagnose from a blended rate.

## Setting the small-sample gate before you trust any rate

A rate computed on too few events is noise, not a leak, and you can set the threshold before outreach even begins from your identification pool size. A 5% response on 40 prospects tells you nothing; the same 5% on 1,100 prospects is a signal worth acting on.

There is no recruiting-specific published minimum, so treat this as a judgement call with statistical grounding rather than a magic number. Conversion is a Bernoulli proportion whose variance is p(1-p), which means lower base rates need larger samples to stabilise. For scale, an A/B test at a 10% minimum detectable effect needs roughly 2,922 conversions and at 5% about 11,141. You will rarely have those volumes in one search, which is exactly why the gate matters: most single-search stage rates are under-powered and should be read as directional, not definitive.

Pool size is knowable up front, and it varies enormously by market and skill. In Refolk's index of professional profiles, a Rust-in-Germany search starts from a very different base than a Go-in-US search.

| Role | Country | Available pool | Ratio vs Germany |
|---|---|---|---|
| Rust Software Engineer | United States | 599 | 6.9x |
| Rust Software Engineer | Germany | 87 | 1.0x |

| Role | Skill | US pool | Ratio vs Rust |
|---|---|---|---|
| Software Engineer (Golang) | Golang | 1,137 | 1.9x |
| Software Engineer (Rust) | Rust | 599 | 1.0x |

Read those together. A thin outreach-to-response rate on Rust in Germany, where the pool is 87, is small-sample noise long before the same rate on Go in the US, where the pool is 1,137, is real. If you know your pool is 87 before you send a single message, you know in advance that any stage rate from that search is under-powered and should not drive a template rewrite.

> **Tip:** Set the gate from the pool, not the outcome
>
> Because pool size is knowable before outreach, decide your minimum trustworthy event count at identification. That stops you from reacting to noise mid-search and rationalising a rewrite after the fact.

Knowing the pool before you commit also changes how you build the list in the first place. [Refolk](/) returns the reachable pool for a plain-English search across the public GitHub graph, public LinkedIn records, and the open web, so you can size a market before you invest outreach in it.

I ran this search: `Senior Rust backend engineers in the US who currently work at infrastructure startups like Oxide.` - [see the full result list](https://www.refolk.ai/s/mdge9wkpnk).

*Returns a sized, reachable pool so you can set your small-sample gate before outreach begins.*

## The scoring procedure

Run these seven steps in order on one search. The first four turn raw counts into scored transitions, and the last three strip out the false leaks so you name the right stage.

#### Diagnosing the worst stage in one funnel

1. **Define stages and entry/exit criteria** - Give identification, outreach-sent, response, screen, and handoff each one unambiguous definition. Companies that skip this produce funnels that cannot be scored accurately.
2. **Pull raw stage counts for one search** - Export absolute counts, not percentages, at every gate for a single search.
3. **Compute per-stage passthrough** - Divide each stage by the one before it so every transition carries a ratio.
4. **Select outbound-specific baselines** - Pair each transition with an outbound benchmark, matching reply rate to 18 to 25% InMail, not the 3.43% cold floor or 3% applicant-to-interview.
5. **Apply the small-sample gate** - Tag any stage with too few events as under-powered, using a threshold set from the identification pool.
6. **Correct for masking and upstream causes** - For each apparent leak, rule latency and intake misalignment in or out before naming it.
7. **Rank stages by qualified candidates lost** - Score each surviving leak by ratio times volume times downstream cost, then name one stage to fix first.

## How to rank when sources disagree

Rank stages by qualified candidates lost, not by the lowest percentage. Sources genuinely disagree here, and the disagreement is the trap most people fall into.

One school says fix the worst ratio first. Another says fix the most expensive stage. Both are half right. A stage with a dreadful ratio but almost no volume costs you fewer real hires than a stage with a merely poor ratio and heavy volume feeding an expensive downstream loop. Reconcile the two views with a single weighting: ratio times volume times downstream cost.

> A channel generating 500 applications with 2% advancement is worse than one generating 100 with 20%.

That line is the whole argument for ranking by loss. Volume is not health. The 500-application channel looks busy on a dashboard and is quietly the worst performer in the building. Rank every channel and every stage by drop-off weighted by what the survivors are worth, never by intake.

The matrix below is the fast version of the ranking call for a single flagged stage.

#### Which flagged stage to fix first

Horizontal axis runs from Few candidates gated here to Many candidates gated here. Vertical axis runs from Cheap downstream to Expensive downstream.

| Quadrant | What it means |
| --- | --- |
| Minor ratio dip, low volume | Log it, do not fix yet |
| Wide loss, cheap roles | Fix second, quick wins |
| Narrow loss, expensive roles | Watch, fix if it worsens |
| Wide loss, expensive roles | Fix first, this is the bottleneck |

*Weight the ratio by how much traffic it gates and how costly the downstream candidates are.*

## Where this diagnosis goes wrong

Most funnel diagnoses fail not because the arithmetic is hard but because the analyst names a false leak. These are the eight ways it happens, each with the check that catches it.

### Comparing outbound to inbound baselines

A 20% reply rate looks weak against an inflated target. Confirm the baseline is outbound and recruiting-specific, 18 to 25% InMail, not the 3.43% cold-email floor or a 3% applicant-to-interview inbound rate. This is the single most common error because the inbound numbers are the ones that rank in search results.

### Small-sample leaks

A 5% response on 40 prospects is noise. Flag any stage below your pre-set event count. A Rust-in-Germany-sized pool of 87 will produce unstable rates that a US pool of 1,137 would not, and you knew that before you sent anything.

### Latency mis-booked as a screen leak

Candidates coded as screen fails often timed out in week three. Cross the stage loss against elapsed days per stage before you name screening. With 13 interviews per hire, up 42% in three years, and time-to-fill running 44 days overall and 62 for engineering, there is a lot of elapsed time for warm candidates to cool. And 57% of candidates have abandoned a process for being too slow, so this is not a rare edge case.

### Intake misalignment read as a messaging leak

Low response gets blamed on copy when the list is off-spec. Audit whether the people who did respond match the brief before you rewrite templates. The 4x yield of direct sourcing shows list quality dominates outcomes, so rewriting copy on a wrong list moves nothing.

### InMail restriction masking response quality

Below 13% response triggers a send cap, which lowers volume, which depresses absolute responses further. Separate response rate from send capacity when you read the numbers, or a response leak will look like a capacity leak and you will fix the wrong one.

### Volume mistaken for health

A busy channel is not a good channel. Rank channels by drop-off, not by how many people they put into the top of the funnel.

### Cross-stage masking

Identical headline reply rates can hide opposite problems. Always break out every transition. A blended rate is where two leaks cancel out on paper and neither gets fixed.

### Fixing lowest percentage instead of largest loss

The stage with the worst ratio is not always the stage costing you the most hires. Rank by candidates lost, weighting ratio times volume times downstream cost.

> **Watch out:** The InMail double penalty
>
> On LinkedIn Recruiter, a response rate below 13% over 100 or more messages in a 14-day window restricts you to one-to-one InMails for the next 14 days. A response leak mechanically becomes a capacity leak, so the loss compounds into the next window unless you fix the message or the list first.

Note what counts as a response under that policy, because it changes your denominator: any reply, an accepted invitation, or a "Not Interested" all count. A funnel where recipients are actively declining can still be clearing the 13% bar while your real qualified-response rate is far lower. Break "Not Interested" out separately when you score outreach-to-response.

## Two upstream causes that eat the top of the funnel

Before you name any top-of-funnel stage as leaking, rule out the two upstream causes that get mis-attributed to it. Both live above the funnel and both look like a stage problem on a dashboard.

The first is intake misalignment. Metaview lists sourcing yield from intake-spec misalignment as the first of its five leak points. If the brief and the list disagree, your responders will be the wrong people even at a healthy response rate, and no template will fix that. The check is to read the profiles of everyone who replied and ask whether they match the brief.

The second is latency, which sits between stages rather than in sourcing itself. Sourcing is rarely where the days go; the gap between a profile arriving and someone deciding on it is the largest single block of dead time in most processes. A candidate who would have passed your onsite but dropped out waiting counts as a screening loss in your dashboard when the root cause was elapsed time.

**Root-cause note per flagged stage**

```
Stage: __________
Measured ratio: ____%   Outbound baseline: ____%
Events at this stage: ____   Under-powered? Y / N
Ratio gap vs baseline: ____ points
Latency check: median days in this stage = ____ (name if > baseline)
Intake check: % of responders matching brief = ____%
Capacity check: any send restriction active? Y / N
Downstream cost weight (low / med / high): ____
Loss score = ratio gap x events x cost weight = ____
Verdict: TRUE LEAK / NOISE / UPSTREAM CAUSE
```

*Fill one out for every stage before you rank. If you cannot rule out the upstream cause, do not name the stage yet.*

## Before you call the diagnosis done

Run this checklist before you commit to a stage. It exists to stop you from acting on a false leak, which is the failure mode that wastes the most effort.

#### Diagnosis readiness

- [ ] Every stage has one unambiguous entry and exit definition
- [ ] Counts are absolute integers for one search, not blended percentages
- [ ] Reply rate is scored against 18 to 25% InMail, not the 3.43% cold floor
- [ ] Every stage is tagged trustworthy or under-powered against a pool-set threshold
- [ ] Any screen loss has been crossed against elapsed days to rule out latency
- [ ] Responders have been checked against the brief to rule out intake misalignment
- [ ] Response rate and send capacity have been separated
- [ ] Every transition is broken out, never diagnosed from a blended rate
- [ ] Stages are ranked by ratio times volume times downstream cost, not lowest percentage alone
- [ ] Exactly one stage is named to fix first

## Keeping the model current

The one number in this model that moves is the outbound-to-inbound multiplier, and it moves in a predictable direction. It went from 5x to roughly 8x in a year, which means your inbound-derived baselines drift further off with each benchmark cycle. Re-anchor your outbound targets whenever a fresh benchmark report lands, and treat any inbound number you inherited from a listicle as suspect by default.

The stable parts are the mechanism, not the values. Latency will always hide in the screen stage, small pools will always produce noisy rates, and volume will always tempt you to call a busy channel healthy. Those do not need re-checking; they need discipline. What does need re-checking is your pool size per search, because that sets your small-sample gate, and Refolk gives you that count before outreach so the gate is set from evidence rather than guesswork. Diagnose one funnel this way, name one stage, fix it, then re-run the whole procedure on the next search rather than assuming the same stage leaks everywhere.

## Frequently asked questions

### What is a good response rate for an outbound sourcing funnel?

Score outbound reply against the recruiting-specific band, not a cold-email figure. Recruiting InMail averages 18 to 25%, the highest response of any industry, so that is your reference range for a sourcing funnel. The 3.43% cross-industry cold-email reply rate is the all-industries floor and will make a healthy recruiting funnel look broken. On LinkedIn specifically, staying at or above 13% over 100 or more messages in 14 days also keeps your send capacity intact.

### How many prospects do I need per stage before a conversion rate is trustworthy?

No recruiting-specific published minimum exists, so treat this as a judgement call rather than a fixed number. Conversion is a Bernoulli proportion whose variance is p(1-p), so lower base rates need larger samples. As a scale check, an A/B test at a 10% minimum detectable effect needs roughly 2,922 conversions and 5% needs about 11,141, which tells you handfuls of prospects produce unstable rates. Set your own threshold from the identification pool and tag anything below it as under-powered.

### Should I fix the stage with the worst percentage or the biggest volume loss?

Rank by qualified candidates lost, not by the lowest ratio alone. Sources disagree: one view says fix the worst ratio first, another says fix the most expensive stage. Reconcile them by weighting ratio times volume times downstream cost. A stage with a mediocre ratio but high volume and expensive downstream candidates can cost more real hires than a stage with a terrible ratio and almost no traffic.

### Why does my screen stage always look like the biggest leak?

Screen is the most over-diagnosed stage because latency hides there. A candidate who would have passed your onsite loop but dropped out in week three counts as a screening-stage loss on your dashboard, when the root cause was elapsed time. With 13 interviews per hire and 44 to 62 day fills, this masking is common. Cross the stage loss against elapsed days per stage before you name screening as the leak.

### Can I use inbound applicant funnel benchmarks for my sourcing funnel?

No. Inbound and outbound funnels behave differently, and the gap is widening. Sourced candidates are roughly eight times more likely to be hired than inbound applicants, up from five times the prior year, and outbound top-of-funnel passthrough runs 67% against 8% inbound. Scoring outbound against inbound baselines like 3% applicant-to-interview flags healthy funnels as broken and gets more wrong every year.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/diagnose-sourcing-funnel-leak*
