Choosing a Second Engineering Hub by Countable Talent Depth
You will carry two or three cities through a documented count of reachable engineers for your stack and defend a single-city pick with numbers an exec accepts.
Key takeaways
- In Refolk's index, US Rust engineers (601) and Mexico Python engineers (685) sit within 14% of each other, so a rare stack erases the advantage of a larger country.
- Documented boolean discounts alone cut a raw pool by 17% for competitor employers and another 24% for seniority, so a headline count routinely overstates reachable supply by a third before attrition.
- Mexico's published developer count spans 220,000 to 1.9 million depending on the source, a 5x to 9x spread that is a definition artifact, not a real difference in supply.
- CBRE treats a market above 50,000 tech workers as large, which is the first filter that keeps a too-thin city off your shortlist.
- Austin, Denver, and Raleigh-Durham medians sit within about $5,000 of each other, so the second-hub decision turns on depth and competitor density, not salary.
- Engineering attrition held at 12% in 2025, the lowest of any function, so a deep pool with low churn still yields few passive movers to poach.
You are picking a city for a second engineering team, and the question that actually decides it is not cost and not a vendor ranking. It is how many engineers for your exact stack you could realistically reach and hire against local competition. This guide is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It carries one real mandate through a documented count across three candidate cities, with the raw queries, the intermediate headcounts, and the two wrong turns that inflate a pool, so you can run the same count on your own case and defend a single-city pick.
The mandate I will use: a company hiring a backend team that needs Python depth, with a budget for a nearshore option, choosing between a coastal default, a specialized-stack bet, and a nearshore metro cluster. The method is the same whatever your stack.
Why a generic ranking cannot make this call
A published top-40 tech talent ranking is the wrong instrument for a second-hub decision because it is weighted for the average employer, not for your stack and your competitors. The best-known scorecard uses 13 metrics weighted by importance, with tech-talent concentration metrics given the highest weights because they signify clustering. That is useful for seeding a shortlist. It tells you nothing about how many Rust engineers you could reach in Austin without poaching from two named competitors.
The decision turns on a number no ranking publishes: the effective reachable pool for your role cluster in each candidate city. Getting there means starting from a raw count and discounting it, step by step, for the people who are not really in your cluster, the people locked at competitors, and the people who will not move.
From headline pool to the number that decides
- 55,048Raw pool (profile index)
US Python SWE/backend
- ~45,700After stripping over-inclusion
exact-title, dedupe, current-role only
- ~37,900After removing competitor-locked
minus 17% at direct competitors
- ~28,800Effective reachable at target seniority
minus 24% above-band, weighted by attrition
The funnel figures after the first stage are illustrative of the documented discounts applied in order, not published counts. The point is the shape: the headline overstates reachable supply by about a third before you even reach attrition.
Four sources, four different numbers for the same city
Four public sources yield city and stack headcounts, and each counts a different object, so their numbers diverge by design. You cannot mix them without a rule for which one you carry forward.
| Source | What it counts | What it is good for |
|---|---|---|
| BLS OEWS (SOC 15-1252) | Jobs in an employer survey | The floor; a stable baseline to reconcile against |
| Profile index | Current professional profiles | The ceiling; your actual query surface |
| GitHub Innovation Graph | Accounts geolocated by IP | Activity signal, never a population count |
| CBRE Scoring Tech Talent | Ranked metro labor pools | Shortlisting and the large-market line |
The employer survey counted 1,656,880 US software developers in May 2023, with a relative standard error of 0.7%, which is a tight, trustworthy baseline. A developer platform hosts over 180 million developers globally, geolocated by IP, which overcounts badly for any one city. A self-selected developer survey drew over 65,000 respondents in one year, which is a sample, not a population.
The divergence is largest offshore. Mexico's developer count is published as 220,000 to 371,000 by government and survey data, as 563,000 profiles on one professional network, as 974,500 specialists in a regional report, and as more than 1.9 million in one nearshore report. That is a 5x to 9x spread for the same country in the same period.
Step 2 in practice: shortlisting without wasting a count
Shortlist with the large-market line and location quotient before you count anything, so you do not spend two days counting a city too thin to matter. A market above 50,000 tech workers is categorized as large; below that it is small. Concentration is read through location quotient, the ratio of local to national density: San Jose runs 84.6 developers per 1,000 jobs, a location quotient of 7.75, while Washington state sits at 2.34x national.
For the worked mandate I shortlisted three cities with documented reasons:
- A coastal default (San Francisco Bay Area), because it is the deepest Python pool and the comp ceiling, with an average tech wage of $178,000, about $30,000 above Seattle.
- A specialized-stack emerging hub (Austin), to test whether a narrower stack survives a smaller market. Austin, Denver, and Raleigh-Durham medians sit within about $5,000 of each other, so cost is not the differentiator.
- A nearshore cluster (Guadalajara, Mexico City, Monterrey), because a mid-level engineer there runs about $3,156 per month versus roughly $10,154 for a US counterpart, and Latin American wages average about 38% of US levels.
| Metro | Median or avg SWE base | Source type |
|---|---|---|
| Austin | ~$144,034 | Salary aggregator |
| Denver | ~$142,876 | Salary aggregator |
| Raleigh-Durham | ~$139,122 | Salary aggregator |
These figures are different vintages and methods, so read them as directional, not as one methodology. The lesson stands: the emerging US hubs have collapsed the cost arbitrage against each other, which pushes the decision onto depth and competitor density.
The raw pull: three cities, three stacks, real counts
Pull the raw pool from each source against one identical spec, and record every number with its URL. Here is where the worked mandate produced its most important lesson, straight from Refolk's index.
| Segment | Raw count | Share of US Python pool |
|---|---|---|
| US, Python, SWE/backend | 55,048 | 100% |
| US, Rust, SWE/backend | 601 | 1.1% |
| Mexico, Python, SWE/backend | 685 | 1.2% |
The share column is derived by dividing each segment by 55,048. The counts come from Refolk's index on a single spec with the same title set.
Read the second and third rows together. In Refolk's index, US Rust engineers (601) and Mexico Python engineers (685) are within 14% of each other. A bigger country buys you nothing once the stack is rare. Specialization, not geography, is the binding constraint. If the mandate had required Rust rather than Python, the coastal default would have shrunk to nearshore scale, and the whole site-selection logic would flip.
Counting the pool yourself for your exact spec, in one place, is what turns a vendor ranking into a decision you can defend. Refolk lets you run the role cluster and the city in plain English and get the current-role count back, instead of stitching three incompatible exports together.
The two wrong turns that inflate a pool
Two moves inflate every raw count, and both happened in the worked mandate before I caught them. Naming them is the most valuable part of this standard, because they are invisible in the headline number.
Wrong turn one: trusting the whole-country offshore number. The first nearshore pull used a published figure of more than 1.9 million Mexican developers. Carrying that into the TAM made the nearshore option look bottomless. It is a definition artifact. That 1.9 million counts a different object than the 220,000 government figure. The index count of 685 for Python backend engineers is the number that survives contact with the actual spec. The gap between 1.9 million and 685 is not reality; it is four sources counting four things.
Wrong turn two: stacking narrow filters to zero. When I tried to tighten the US pull with Python plus Senior-seniority plus a strict title filter at once, Refolk's index returned 0. There is no world where the real senior Python pool is zero. Stacking three narrow filters nulled a viable pool. The fix is to relax one filter and re-count, then reapply the missing constraint as a proportion rather than a hard gate. If you take a zero at face value, you write off a city that is actually fine.
Stripping over-inclusion down to real engineers
Apply exact-title matching, NOT operators, and identity dedupe before you trust any count, and expect double-digit percentage shrinkage. Over-inclusion is not a rounding error; it is the difference between a deep-looking city and a reachable one.
The dominant inflator is stale profiles. People forget to put an end date on a previous job when they move, and unclosed past positions are responsible for the majority of false positives in people search. A profile reads as a current engineer when the person left the role years ago. Second is synonym bleed: selecting a title value rather than an exact phrase pulls in the synonyms the system guesses, so an SDE, a SAS Programmer, even a recruiting coordinator lands in the count. Third is whole-profile keyword match, where "Python" counts someone who studied it once at school because the search scanned recommendations and old descriptions instead of the current-role skill field.
Is this count safe to carry forward?
The bottom-right quadrant is where a generic ranking sends you wrong. A large raw pool that is 60% locked inside two local competitors is not reachable. Read employer concentration in the pool before celebrating the headline.
The procedure, start to finish
Run these eight steps in order. One caveat on order: concentration-led scorecards start from clustering weights, but boolean practitioners insist the over-inclusion strip must precede any counting, or every later number is wrong. I side with counting clean first.
Count a second-hub decision end to end
- Define the role cluster and stack preciselyFix titles, required skills, and seniority band into one exact-match query spec before counting. A loose spec guarantees synonym inflation downstream.
- Shortlist two or three candidate citiesUse the 50,000-worker large-market line and location quotient to drop markets too thin to matter. Give each shortlisted city a documented reason.
- Pull a raw pool from each source per cityQuery the profile index, the BLS metro table, and GitHub geographic data against the same spec. Record three numbers from three definitions, each with its URL.
- Reconcile divergenceWhere sources disagree 5x to 9x, carry the employer survey as the floor and the profile count as the ceiling, and state which you take forward.
- Strip over-inclusionApply exact-title matching, NOT operators, and identity dedupe. Expect double-digit shrinkage and a count of current-role, in-cluster people only.
- Discount for reachabilitySubtract the competitor-locked fraction and weight by target-seniority share and annual attrition to get one effective-reachable number per city.
- Score hiring difficultyOverlay active postings and offer-acceptance benchmarks to produce a supply-to-demand ratio per city.
- Recommend one city with the tablePresent raw pool, effective reachable, cost, and difficulty side by side so an exec sees the full chain from headline to decision.
Discounting for reachability and scoring difficulty
The reachable count is the raw pool minus false positives, minus the fraction you cannot poach, adjusted by the slice at your target seniority and annual churn. The discounts are measurable, not vibes. Excluding competitor employers cut one real title search by 17%, and removing seniority higher than the target cut another 24%. Together that is roughly a third gone before attrition even enters.
Then attrition, which cuts both ways. Engineering attrition held at 12% in 2025, the lowest of any function, and a workforce report put software-engineer attrition at 13.2% against about 9% cross-industry. Low churn is good for retention and bad for hiring: a deep pool with low attrition means fewer passive movers to poach. Supply you cannot move is not supply. Separately, 69% of developers stay under two years at one employer, which tells you the movers exist but concentrate at the junior and mid bands.
Score difficulty from demand pressure. Active core-tech postings totaled nearly 490,000 in one February, with newly posted ads up 15% year over year, and metro-level deltas are published, so you can read which of your shortlisted cities is heating up. On the funnel's close side, technical-role offer acceptance runs 73% against 84% for business roles, so plan for more offers than a business team would to land the same headcount.
A deep pool with low churn is a retention story and a hiring problem at the same time.
| Country | Published dev pool (range) | Monthly mid-level salary |
|---|---|---|
| Mexico | 220K to 1.9M | ~$3,156 |
| Brazil | ~500K to 760K | ~$6,000 |
| US (reference) | 1.66M (employer survey) | ~$10,154 |
The pool ranges diverge widely by definition, so treat the wide end as a ceiling only. The salary column is the real reason a nearshore hub stays on the list, but it only matters if the reachable count in the first table can staff a team.
Failure modes to check before you present
Run every count against this list before it reaches an exec. Each failure mode has a false positive it produces and a cheap check that catches it.
- Stale-profile inflation. A "current" title that ended years ago. Check: sample 20 profiles for end dates; expect most of the noise here.
- Synonym bleed. A non-engineer in the count. Check: rerun with exact-phrase quotes and compare the delta.
- Whole-profile keyword match. Someone who studied the skill once. Check: constrain to current-role skill fields, not free text.
- Cross-platform double count. The same engineer on two platforms. Check: dedupe by identity before summing; never add two platform totals.
- Category creep. Counting developers plus QA plus testers (1.9M) when you want developers only (1.66M). Check: lock the occupation code or exact title set.
- Over-filtering to zero. Stacked filters null a real pool. Check: relax one filter and re-count.
- Competitor blindness. A large pool that is mostly locked at local rivals. Check: read employer concentration before celebrating the headline.
Before you call the count defensible
- Each city has three raw numbers from three sources, each with its URL
- You stated which source is the floor and which is the ceiling for every city
- You sampled 20 profiles per city and confirmed current end dates
- Exact-title and NOT operators applied, and identity dedupe run across sources
- Competitor-locked fraction subtracted and target-seniority share applied
- Attrition weighting documented, and no stacked query left sitting at zero
- Final table shows raw pool, effective reachable, cost, and difficulty side by side
Making the recommendation and keeping it current
Recommend one city with the full table visible, so the exec sees the chain from headline to reachable number rather than a single asserted figure. The deliverable is one row per city with raw pool, effective reachable, median cost, and supply-to-demand ratio. The city that wins is rarely the deepest raw pool; it is the best reachable count at an acceptable cost and difficulty.
| City | Raw pool (source) | Effective reachable | Median base | Active postings | Supply:demand | Note | |------|-------------------|---------------------|-------------|-----------------|---------------|------| | City A | ___ (employer survey floor) | ___ | $___ | ___ | ___ | ___ | | City B | ___ (profile ceiling) | ___ | $___ | ___ | ___ | ___ | | City C | ___ (reconciled) | ___ | $___ | ___ | ___ | ___ | Recommendation: ___ because effective reachable ÷ annual openings beats the alternatives at acceptable cost.
Fill one row per shortlisted city; keep every source URL in a footnote so the number is auditable.
Keep the count current by re-running two things. First, the raw pull, because profiles accumulate and the stale-profile share drifts; re-count quarterly during an active build. Second, the difficulty overlay, because postings and the hiring cycle swing hard. One software postings index ran a peak of 225 in early 2022 and sat near 70 in early 2024 and early 2026, against a 2020 baseline of 100. A ratio that looked comfortable at the peak can invert in a cooler market, so re-check the mechanism rather than trusting a number you captured months ago. Document the source and date on every figure, and the recommendation stays defensible long after the first pull.
Questions practitioners ask
Can I just use a vendor top-40 tech talent ranking instead of building my own count?
No, not for a defensible decision. A published ranking weights metrics for the average employer, so it tells you where generic talent clusters, not where your specific stack is deep and reachable. Use the ranking to seed a shortlist and to apply the 50,000-worker large-market line, then build your own count for your exact role cluster. The gap between a generic ranking and your reachable number is exactly where the decision lives.
Why do my raw talent-pool numbers for the same city disagree so much?
Because each source counts a different object. An employer survey counts jobs, a profile index counts accounts, and a developer platform counts signups geolocated by IP. For Mexico, published counts span 220,000 to 1.9 million, a 5x to 9x spread, purely from definition. Carry the employer survey as your floor and the profile count as your ceiling, and never add two platform totals together because you will double-count the same people.
How much should I discount a raw pool to get a reachable number?
Start from documented boolean discounts: excluding competitor employers cut one real search by 17%, and removing seniority above your target cut another 24%. That is roughly a third gone before attrition. Then weight by the share at your target seniority and by annual attrition, which held at about 12% for engineering. There is no single published cutoff, so document each discount so an exec can audit the chain.
Is a nearshore hub actually cheaper enough to matter?
Often yes on salary and no on reachable depth. A mid-level engineer in Mexico runs about $3,156 per month against roughly $10,154 for a US counterpart, and Latin American wages average about 38% of US levels. But if your stack is rare, the reachable pool can be tiny: Mexico Python engineers number 685 in Refolk's index, close to US Rust engineers at 601. Cost arbitrage means nothing against a pool you cannot staff a team from.
What is the single most common mistake that inflates a city's talent pool?
Stale profiles. People forget to add an end date when they change jobs, so an unclosed past position reads as a current role and gets counted. This is responsible for the majority of false positives in people search. Sample 20 profiles for end dates before you trust any raw count, and constrain your query to current-role skill fields rather than scanning the whole profile.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.