Reconciling a Headline Talent-Pool Count to a Defensible Number
You can take one inflated pool count and walk it down to a defensible addressable figure with a deduction trail a stakeholder can audit line by line.
You pulled a talent-pool size from a tool, and now you need the real addressable number before it goes in a brief. This guide is for strategy, research, and talent-intelligence analysts who have one inflated count and a stakeholder waiting. It carries a single worked example all the way down, with the real intermediate figures, the deductions in order, and the two wrong turns most analysts take before the number holds up.
Other guides treat pool-sizing as a forward playbook or completeness as a standard. This one runs backwards: one headline number, carried to a defensible figure, with a trail you can hand across the desk. Follow along on your own case as you read.
Why a headline pool count is almost never the real number
A headline pool count is a starting estimate, not a census, and the gap between it and a defensible figure is usually several multiples. The number inflates for reasons that are documented and predictable, which is exactly why you can deduct them in order.
Three things conspire. First, the count above a threshold is an approximation: in the recruiter surface, exact counts appear only below 1,000 results, and above that you see approximate counts. Second, the query itself is loose: a search for a term scans the whole profile, including schools and recommendations, so it pulls in students, blog writers, and managers. Third, the population underneath is decaying: job titles change 25-35% per year, so some share of the "current" matches are already wrong.
None of this means the tool is broken. It means the headline answers a different question than your brief asks. The brief wants "how many people can I plausibly reach for this role." The tool answers "how many profiles mention these terms somewhere." Reconciliation is the work of turning the second into the first, deduction by deduction, with a source beside each drop.
The worked example: US Rust engineers
I will carry one case the whole way: software engineers in the United States who use Rust. In Refolk's index of professional profiles, there are 595 US software engineers with Rust. That is the defensible base I am working toward. The teardown below starts from a loose, inflated headline and walks down to a number in that neighbourhood, showing what each deduction removes.
Two comparisons anchor how sensitive these numbers are to the question you ask. Same role, different market: the US Rust pool is 6.4x the German one. Same market, different skill: the US Go pool is 2.3x the US Rust pool. Neither multiple is a rounding error. They tell you that "size the Rust pool" and "size the Go pool" are not the same brief, and that a country swap changes the answer by more than any staleness haircut will.
| Market | Rust SWE count | Multiple vs Germany |
|---|---|---|
| United States | 595 | 6.4x |
| Germany | 93 | 1.0x |
Counts are from Refolk's index; the multiple is derived from them.
| Skill (US, Software Engineer) | Count | Multiple vs Rust |
|---|---|---|
| Golang | 1,366 | 2.3x |
| Rust | 595 | 1.0x |
Counts are from Refolk's index; the multiple is derived from them.
Read these two tables together before you touch a deduction. If your headline came from a tool and reads in the tens of thousands for "Rust engineers in the US," the distance between that and 595 is the work. Most of it is title definition, not data quality.
The deduction sequence, from headline to defensible
The frame is standard market sizing, translated into talent terms: the total market is all developers, the serviceable market is those who meet your skills, experience, location, and seniority, and the obtainable market is those also reachable now. Each cut below is a known, documented filter, and the order matters because some cuts remove the same people if you are not careful.
One headline count narrowing to reachable-now
- widestRaw headline (loose keyword)
estimate above 1,000
- narrowerTitle-scoped match
the single biggest swing
- narrowerAdjacent populations removed
interns, students, academics
- narrowerStaleness and de-dupe applied
titles decay 25-35%/yr
- narrowestActive / reachable share
~30% active, 4.1% looking
Here are the deduction factors you will apply, each with what it does and where it comes from. Treat the percentages as documented rates, not as the exact answer for your pool; your spot-checks calibrate them.
| Deduction step | Documented factor | Source |
|---|---|---|
| Annual data staleness | 22.5%-70.3%/yr (titles 25-35%) | landbase.com; zoominfo.com |
| Single-platform profile gap | ~30% of engineers absent | recruiter.daily.dev |
| Active-only reachable share | ~30% active (4.1% actively looking) | randstad.com; rally |
| Private-work invisibility (GitHub) | 82% contributions private | pin.com |
The title cut usually dwarfs the rest. In a documented keyword case, switching from the exact phrase "computer games" to the synonym set games OR gaming OR game moved results from 150 to 494, roughly 3x, before any staleness or de-dupe. That is the mechanism to respect: whole-profile keyword scans both over-include noise and under-include people who phrase their title differently. Fix the question before you haircut the answer.
Staleness compounds, which trips up analysts who apply a single flat 30%. At 30% annual title decay, data two years old is about 49% valid and three years old about 34%. If the underlying member records are two years stale on average, a one-year haircut understates the problem badly.
Fix the question before you haircut the answer; the title definition swings the count more than the data ever will.
Run it: the step-by-step reconciliation
Work the deductions in this order, logging a count and a source at every step. The goal is not a single perfect number but a trail where every drop has a reason a stakeholder can see.
From inflated headline to defensible figure
- Capture the raw count and its surfaceRecord the exact query, the tool, and which surface produced it so you can reproduce it. Note whether it is exact or estimated, since approximate counts appear above 1,000 and exact counts only below.
- Tighten title and keyword precisionReplace loose keyword matches with field-scoped titles plus synonyms, because keyword fields scan the whole profile including schools and recommendations. Re-run and log the delta as the title-misfire deduction.
- Exclude adjacent populationsRemove interns, students, bootcamp grads, and academics using narrow NOT terms, testing each term with and without it so you do not drop seniors who merely mention it. Log the exclusion string you used.
- Apply a staleness haircutDiscount for profiles whose current title no longer holds, using a dated, sourced rate, and compound it by the age of the underlying data rather than applying a flat single-year cut.
- De-duplicate and strip dormant accountsAccount for duplicate and phantom records; in 2023 one platform removed more than 121M fake accounts, and cross-surface gaps run to tens of millions. Log a de-dupe and phantom deduction.
- Correct for the single-source coverage gapDecide explicitly whether to gross up for who the platform never had (about 30% of engineers absent) or treat the platform count as the ceiling. State the assumption.
- Split active versus reachableIf the brief needs reachable-now, apply the active share: about 30% active and 4.1% actively looking. Report both the addressable and the reachable figure.
- Document the trail and hand it overBuild one table of step, count, percent removed, and source. Done when a stakeholder can audit every drop without asking you a question.
A note on step two, the title cut: this is where most of your inflation disappears, and it is also where a tool that only matches keywords will fight you. If you can ask for the role in plain language and get title-scoped, de-noised results back, you skip the Boolean tuning entirely.
That is the hard part done in one query. When the role is specified in plain English, Refolk applies the title scoping and the intern, student, and bootcamp exclusions for you, so the number you carry into the staleness and active-share cuts is already the clean base rather than the raw headline.
Where this goes wrong: the failure modes
Most reconciliations fail at one of seven predictable points. Each has a false positive it produces and a cheap check that catches it. This is the section to keep open while you work, because a wrong deduction is harder to spot than a wrong headline.
- Trusting the headline as exact. Above 1,000 the count is an estimate, not a census. The false positive is quoting a rounded approximate number as precise. Check: confirm the count is under the exact-count threshold, or label it estimated.
- Over-excluding with NOT. Stacking NOT intern, junior, student also drops seniors who once held or merely mentioned those terms. Check: compare counts with and without each NOT term, one at a time.
- Loose keyword as title. Whole-profile scans count people who only mention a term, the "data scientist returns students and blog writers" problem. Check: scope to title fields and spot-check 20 profiles by eye.
- Double-counting the coverage gap. Grossing up for the 30% off-platform and then adding a separate developer-account total counts the same engineers twice. Check: define one canonical universe before adding anything back.
- Applying staleness and active-share to the same drop. Stale-title removal and active-versus-total are different cuts; conflating them under-counts. Check: label each deduction by exactly what population it removes.
- Assuming one surface is truth. The same query differs across a product's own surfaces by tens of millions. Check: name the surface beside every number and never mix surfaces in one trail.
- Treating passive as unreachable. The 30% active figure means 30% are actively seeking, not 30% interested in your role. Check: state whether the final number is addressable, reachable, or interested.
Two of these are the wrong turns I see most. The first is the title-as-keyword mistake at the top of the funnel: analysts haircut aggressively for staleness and coverage while leaving a bloated keyword match in place, so they subtract 50% from a number that was 3x too big to begin with. The second is double-counting the coverage gap: grossing up the platform figure by 30% for off-platform engineers, then adding a developer-account total on top, which quietly inflates the market by the exact population you just tried to add back.
When to add back versus when to deduct
Addressable versus reachable: two numbers, not one
A brief usually needs two figures, not one: the addressable pool, meaning everyone who fits, and the reachable-now pool, meaning those you can realistically engage today. Reporting only the larger one oversells the opportunity; reporting only the smaller one hides the market.
The reachable cut is where numbers collapse, because the filters are independent and multiplicative. Start with a platform that misses about 30% of engineers. Of those present, about 30% are active and only 4.1% are actively looking. Stack those and the reachable-now figure can fall under a tenth of the headline. That is not a flaw in the method; it is the honest shape of a talent market.
What each layer of the number means
- Addressableeveryone matching role, skills, seniority, location
- On-platformthe share any single source actually holds
- Current and realafter staleness and phantom-account deductions
- Reachable-nowthe active and open subset you can engage today
Be precise about what the active share means. The roughly 30% active figure is misleading at face value: it means 30% are actively seeking a move, not 30% interested in your particular role. So the reachable-now number is an upper bound on engagement, not a demand forecast. Say so in the brief, in one sentence, next to the figure.
For the Rust example, the defensible addressable base is 595 US software engineers in Refolk's index. If the stakeholder asks "how many can we approach this quarter," you apply the active share to that base and report a far smaller reachable figure, clearly labelled, with the deduction shown. The two numbers belong side by side, never collapsed into one.
Verify before you ship: the reconciliation checklist
Run this before the number leaves your desk. Every item is a thing a stakeholder could challenge, so clearing it means the question is already answered when asked.
Before the number goes in the brief
- The exact query, tool, and surface are recorded and reproducible.
- The count is labelled exact or estimated based on the threshold.
- The title cut is field-scoped and spot-checked against 20 real profiles.
- Each NOT exclusion was tested with and without it, and none drops valid seniors.
- The staleness haircut is compounded by the age of the underlying data, not a flat year.
- De-dupe and phantom-account deductions are logged with a source.
- The coverage gap is added back from exactly one canonical source, or not at all, and the choice is stated.
- Addressable and reachable-now are reported as two separate labelled numbers.
- The final figure is tagged addressable, reachable, or interested, never all three.
- The deduction table shows step, count, percent removed, and source on one page.
The deliverable is that table. One row per deduction, four columns, every drop sourced. A stakeholder should be able to read down it and reconstruct your logic without a meeting.
| Step | Count after | % removed | Source / note | |---|---|---|---| | Raw headline (surface named) | 0000 | - | tool + surface; estimated above 1,000 | | Title-scoped match | 0000 | 00% | title fields + synonyms; spot-checked 20 | | Adjacent populations removed | 0000 | 00% | NOT intern/student/bootcamp; tested each | | Staleness haircut | 0000 | 00% | titles 25-35%/yr, compounded by data age | | De-dupe / phantom accounts | 0000 | 00% | duplicate + dormant removal | | Coverage gross-up (optional) | 0000 | +00% | one source only; universe defined | | Reachable-now (active share) | 0000 | 00% | ~30% active; label as reachable |
Fill one row per step; keep the source column populated for every drop.
Keeping the number current
A reconciled pool count is a snapshot, not a constant, so date it and set a re-check trigger rather than treating it as settled. The underlying population decays on a schedule you can plan around, and the surfaces you pulled from can shift between runs.
Re-run the reconciliation when any of three things changes: the brief's skill or geography (because the base moves before any deduction, as the 2.3x and 6.4x multiples show), the staleness clock (titles change 25-35% per year, so a figure more than a few quarters old deserves a fresh pull), or the surface (if the tool updates, re-capture rather than trust the old headline). Refresh cadence for pool reports is not published as a fixed interval, so set your own and write it on the brief.
When you re-run, keep the old deduction table beside the new one. If the final number moved, you want to see which row moved it: a bigger base, a steeper staleness cut, or a different active share. That side-by-side is what turns a one-off estimate into a figure you can defend quarter after quarter, and it is the difference between a number someone trusts and a number someone has to redo.
Questions practitioners ask
Why is my talent pool number higher in one tool than another for the same query?
The surfaces of a single product can disagree by tens of millions for the same query. One documented analysis found 91M, 140M, and 148M for India across a tool's insights surface, its recruiter surface, and investor reporting, a 57M gap. Same-query runs also differed between a personal account and the recruiter product, reported as 150 versus 43. Always name the surface beside every number and treat none as a census.
How much does a talent pool size get overcounted by loose titles?
No source quantifies an exact inflation percentage, so treat any fixed figure as not established. The directional magnitude is large: a documented keyword case moved from 150 to 494 results, roughly 3x, just by switching from an exact phrase to synonyms. Whole-profile keyword scans count students, blog writers, and managers who merely mention a term, so scope to title fields and spot-check 20 profiles before trusting the count.
Should I add back the people missing from the platform to get a true market size?
Decide explicitly and only once. About 30% of software engineers have no profile on a given platform, so you can gross up for them. But if you also add a separate developer-account total, you will count the same engineers twice. Define one canonical universe first, then add back from exactly one source. State the choice in the deduction trail.
What is the difference between addressable and reachable in a talent pool?
Addressable is everyone who fits the role; reachable-now is the subset you can actually engage today. About 30% of the workforce is active and only 4.1% are actively looking, so the reachable figure can be under a tenth of the headline. The active share is not interest in your specific role either, so label the final number as addressable, reachable, or interested, never all three at once.
How do I handle engineers who look inactive but are strong?
Flag them as missing evidence, not removed population. With 82% of GitHub contributions in private repositories, a capable engineer can appear dormant while shipping constantly. Over-deducting these people is as wrong as over-counting noise. In the trail, separate a line for people the platform never surfaced from a line for people you have deliberately excluded.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.