Turning a Raw Candidate Count Into a Defensible Pool Estimate
You will convert a raw search count into a corrected pool estimate with an explicit confidence grade and a stated reachable subset, defensible line by line.
You ran a search for a role in a location, got a raw candidate count, and now leadership wants to know how many qualified people really exist and how much to trust the number. This guide is for talent-intelligence analysts, strategy and research teams, and operators sizing a hiring market. It gives you a scoring model that turns any raw count into a corrected pool estimate with an explicit confidence grade and a stated reachable subset, defensible line by line in a review.
Most public guidance either sells a black-box pool number from a vendor calculator or explains generic talent-pool management. Neither shows you how to correct a raw count for the errors that inflate or deflate it. This does. It names each correction factor, its direction, and its effect on your confidence grade.
Why a raw count is never the answer
A raw candidate count is wrong in two directions at once, and the two errors do not cancel. Duplicates and stale profiles deflate the true pool while missing title synonyms and adjacent skills inflate your error, so a single uncorrected number is never defensible on its own.
The two error classes are independent. Duplication is a data-hygiene problem: CRM duplication rates reach up to 20 percent, and 70 percent of organizations struggle with duplicate or inconsistent data. Decay is a freshness problem: B2B contact data decays more than 30 percent per year. Neither of those has anything to do with whether you searched for the right titles. Because the classes are independent, you cannot net them off in your head. You correct each one separately, with a direction and a documented range, and you show the line items.
The suppression side is just as real. Recruiter guidance is to include 4 to 6 title variations in every search, because candidate language is inconsistent and a single title can create a false sense of a small market. Skills-based hiring can expand talent pools by 10x because it uncovers qualified workers employers overlooked. So the same raw count that is inflated by duplicates is also deflated by everyone you never searched for.
The job of this framework is to make every one of those movements explicit, so that a reviewer can challenge any single line without knocking over the whole estimate.
The correction model: factors, directions, and ranges
The model has four named correction factors, each with a fixed direction and a documented range. Two deflate the raw count and two recover from it. Score them in that order so your working number always moves in a traceable path.
Here is the full factor set, drawn straight from the documented ranges:
| Factor | Direction | Documented range | Source basis |
|---|---|---|---|
| Duplicate / orphan profiles | Deflate raw | up to ~20% (CRM); millions on LinkedIn | data-quality research |
| Stale / decayed profiles | Deflate raw | ~30%+ per year of decay | data-quality research |
| Missing title synonyms | Recover (add back) | 4 to 6 variants standard | recruiter guidance |
| Adjacent-skill / skills-based | Recover (add back) | up to 10x ceiling | talent research |
The direction column is the discipline. Duplicates and staleness always deflate, so they come off the top. Synonyms and adjacent skills always recover, so they get added back. A reviewer who sees a deduction in the wrong column knows immediately the model was applied carelessly.
The ranges are ceilings and typical values, not point estimates. Dedup of up to 20 percent means you might strip 5 percent from a clean index and closer to 20 percent from a messy CRM export. The 10x skills figure is an upper bound for what expansion can reveal, not a multiplier you apply by default. Applying 10x as a default is one of the fastest ways to lose a room.
The raw-count correction path
- Raw countOne title, one location, logged verbatim
- DeflateRemove duplicates and stale profiles on a stable key
- RecoverAdd synonym and adjacent-skill matches, bounded by the 10x ceiling
- ReconcileCompare two methods into a stated range
The location boundary is your biggest lever
The location boundary moves the base more than any other single correction, and it must be set by function, not by default. It decides who counts as reachable before any other number is applied, because it redefines the denominator itself.
The swing is dramatic. Employers that offer remote work capture 96 percent of the labor market, while those that do not lose out on 58 percent of candidates. That single choice dwarfs a 20 percent dedup or a 30 percent staleness haircut. And it is function-specific: 89 percent of computer and math jobs allow remote work, versus 61 percent in business and finance and 53 percent in arts and media. You cannot apply one remote multiplier across every role.
Relocation adds a smaller, checkable increment on top of the boundary. An estimated 20 million Americans are considering relocating to a different state for job reasons, but up to 40 percent of relocations fail due to culture shock, logistics, and family separation. So a relocation-willing band belongs in the estimate, but stated explicitly and kept modest.
The defensible procedure is: set a metro base, apply a function-specific remote multiplier, then add a modest relocation-willing band you name out loud. Concentration matters here too. In Refolk's index, Germany's Data Engineer pool is roughly twice as concentrated in its top region as the US pool, which means a metro-radius boundary captures a larger share of the German pool than the US pool at the same radius.
| Title | Country | Raw pool | Top region | Top-region share of sample |
|---|---|---|---|---|
| Data Engineer | United States | 29,634 | New York City metro | ~12% (3/25) |
| Data Engineer | Germany | 3,084 | Berlin | ~24% (6/25) |
The US pool is about 9.6 times the German one, but geographic concentration flips the reachability story. A metro boundary drawn around Berlin captures a meaningfully larger fraction of German data engineers than the same boundary around any single US metro captures of the US pool. Size and reachability are different questions, and the boundary is where you separate them.
Adjacent skills, not headcount, explain "small" pools
When a pool looks tiny, the usual cause is a narrow skill or title search, not a genuinely thin market. Candidates self-describe inconsistently, and a single search term understates supply by construction. This is where recovery corrections earn their keep.
The clearest evidence sits in one market. In Refolk's index, the US Software Engineer pool for Go is 2,779, while the same market for Rust is 603.
| Skill | Raw pool | Top employers in sample |
|---|---|---|
| Go | 2,779 | Meta, Google |
| Rust | 603 | Google, Oxide Computer, Meta |
The Go pool is about 4.6 times the Rust pool in the same country at the same moment. A raw Rust search that ignores adjacent systems languages such as C++ or Go reports 603 and calls the market thin, when the real supply of people who could do systems work is several times larger. That is not a rounding error. It is the reason skills-based expansion is documented at up to 10x: it recovers the qualified workers a rigid title or single-skill filter never returned.
number: 4.6x
label: How much larger the US Go engineer pool is than the Rust pool
note: Same country, same index. A single-skill search reports the smaller number and hides the adjacent supply.
Questions practitioners ask
How many candidates are there for a role and location?
There is no single honest answer until you correct the raw count. Start from your search total, deduct duplicates (up to about 20 percent) and stale profiles (about 30 percent per year of decay), then add back candidates you missed through synonyms and adjacent skills. Reconcile a top-down and a bottom-up figure into a range rather than a point. The defensible answer is a band with a confidence grade, not a single number.
What is a realistically reachable subset of a talent pool?
Anchor it to the active-seeker band, roughly 27 to 37 percent of the corrected pool, then add a small passive-response increment you can defend. Published figures put passive talent at about 70 percent of the workforce, so if you count only active people you are working with roughly a third. Remember that active means seeking any role, not interested in yours, so the active share is a floor rather than a target.
How do I correct a candidate count for duplication and title variants?
Deduplicate on a stable identifier such as a hashed member URL, never on a display name or vanity URL that changes when someone renames their profile. For title variants, re-run the search with 4 to 6 synonyms; if the count jumps, your narrow number was understated. Duplicates deflate the raw count while missing synonyms inflate your error, so treat them as separate line items in opposite directions.
How do I assign a confidence grade to a talent supply estimate?
Score four checks: synonym coverage, dedup applied on a stable key, location boundary declared, and agreement between two estimation methods. High confidence needs all four, medium tolerates one soft check, and low means the boundary or the methods disagree. The confidence signal borrows from market sizing: when top-down and bottom-up estimates converge the grade rises, and when they diverge it flags flawed assumptions.
Does the location boundary really change the number that much?
Yes, more than any other single correction. Offering remote work captures 96 percent of the labor market while not offering it loses 58 percent of candidates, and the remote-capable share is function-specific: 89 percent of computer and math jobs allow remote versus 53 percent in arts and media. Because the boundary redefines the denominator itself, declare it explicitly before you quote any pool number.