Refolk
TeardownMarket and talent intelligence

Choosing a Second Engineering Hub by Countable Talent Depth

You will carry two or three cities through a documented count of reachable engineers for your stack and defend a single-city pick with numbers an exec accepts.

15 min readLast reviewed October 8, 2026Read as Markdown

Key takeaways

  • In Refolk's index, US Rust engineers (601) and Mexico Python engineers (685) sit within 14% of each other, so a rare stack erases the advantage of a larger country.
  • Documented boolean discounts alone cut a raw pool by 17% for competitor employers and another 24% for seniority, so a headline count routinely overstates reachable supply by a third before attrition.
  • Mexico's published developer count spans 220,000 to 1.9 million depending on the source, a 5x to 9x spread that is a definition artifact, not a real difference in supply.
  • CBRE treats a market above 50,000 tech workers as large, which is the first filter that keeps a too-thin city off your shortlist.
  • Austin, Denver, and Raleigh-Durham medians sit within about $5,000 of each other, so the second-hub decision turns on depth and competitor density, not salary.
  • Engineering attrition held at 12% in 2025, the lowest of any function, so a deep pool with low churn still yields few passive movers to poach.

You are picking a city for a second engineering team, and the question that actually decides it is not cost and not a vendor ranking. It is how many engineers for your exact stack you could realistically reach and hire against local competition. This guide is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It carries one real mandate through a documented count across three candidate cities, with the raw queries, the intermediate headcounts, and the two wrong turns that inflate a pool, so you can run the same count on your own case and defend a single-city pick.

The mandate I will use: a company hiring a backend team that needs Python depth, with a budget for a nearshore option, choosing between a coastal default, a specialized-stack bet, and a nearshore metro cluster. The method is the same whatever your stack.

Why a generic ranking cannot make this call

A published top-40 tech talent ranking is the wrong instrument for a second-hub decision because it is weighted for the average employer, not for your stack and your competitors. The best-known scorecard uses 13 metrics weighted by importance, with tech-talent concentration metrics given the highest weights because they signify clustering. That is useful for seeding a shortlist. It tells you nothing about how many Rust engineers you could reach in Austin without poaching from two named competitors.

The decision turns on a number no ranking publishes: the effective reachable pool for your role cluster in each candidate city. Getting there means starting from a raw count and discounting it, step by step, for the people who are not really in your cluster, the people locked at competitors, and the people who will not move.

From headline pool to the number that decides

  1. Raw pool (profile index)
    55,048

    US Python SWE/backend

  2. After stripping over-inclusion
    ~45,700

    exact-title, dedupe, current-role only

  3. After removing competitor-locked
    ~37,900

    minus 17% at direct competitors

  4. Effective reachable at target seniority
    ~28,800

    minus 24% above-band, weighted by attrition

Each stage removes people a raw count wrongly treats as hireable, and the documented discounts alone take out roughly a third.

The funnel figures after the first stage are illustrative of the documented discounts applied in order, not published counts. The point is the shape: the headline overstates reachable supply by about a third before you even reach attrition.

Four sources, four different numbers for the same city

Four public sources yield city and stack headcounts, and each counts a different object, so their numbers diverge by design. You cannot mix them without a rule for which one you carry forward.

SourceWhat it countsWhat it is good for
BLS OEWS (SOC 15-1252)Jobs in an employer surveyThe floor; a stable baseline to reconcile against
Profile indexCurrent professional profilesThe ceiling; your actual query surface
GitHub Innovation GraphAccounts geolocated by IPActivity signal, never a population count
CBRE Scoring Tech TalentRanked metro labor poolsShortlisting and the large-market line

The employer survey counted 1,656,880 US software developers in May 2023, with a relative standard error of 0.7%, which is a tight, trustworthy baseline. A developer platform hosts over 180 million developers globally, geolocated by IP, which overcounts badly for any one city. A self-selected developer survey drew over 65,000 respondents in one year, which is a sample, not a population.

The divergence is largest offshore. Mexico's developer count is published as 220,000 to 371,000 by government and survey data, as 563,000 profiles on one professional network, as 974,500 specialists in a regional report, and as more than 1.9 million in one nearshore report. That is a 5x to 9x spread for the same country in the same period.

Step 2 in practice: shortlisting without wasting a count

Shortlist with the large-market line and location quotient before you count anything, so you do not spend two days counting a city too thin to matter. A market above 50,000 tech workers is categorized as large; below that it is small. Concentration is read through location quotient, the ratio of local to national density: San Jose runs 84.6 developers per 1,000 jobs, a location quotient of 7.75, while Washington state sits at 2.34x national.

For the worked mandate I shortlisted three cities with documented reasons:

  • A coastal default (San Francisco Bay Area), because it is the deepest Python pool and the comp ceiling, with an average tech wage of $178,000, about $30,000 above Seattle.
  • A specialized-stack emerging hub (Austin), to test whether a narrower stack survives a smaller market. Austin, Denver, and Raleigh-Durham medians sit within about $5,000 of each other, so cost is not the differentiator.
  • A nearshore cluster (Guadalajara, Mexico City, Monterrey), because a mid-level engineer there runs about $3,156 per month versus roughly $10,154 for a US counterpart, and Latin American wages average about 38% of US levels.
MetroMedian or avg SWE baseSource type
Austin~$144,034Salary aggregator
Denver~$142,876Salary aggregator
Raleigh-Durham~$139,122Salary aggregator

These figures are different vintages and methods, so read them as directional, not as one methodology. The lesson stands: the emerging US hubs have collapsed the cost arbitrage against each other, which pushes the decision onto depth and competitor density.

The raw pull: three cities, three stacks, real counts

Pull the raw pool from each source against one identical spec, and record every number with its URL. Here is where the worked mandate produced its most important lesson, straight from Refolk's index.

SegmentRaw countShare of US Python pool
US, Python, SWE/backend55,048100%
US, Rust, SWE/backend6011.1%
Mexico, Python, SWE/backend6851.2%

The share column is derived by dividing each segment by 55,048. The counts come from Refolk's index on a single spec with the same title set.

Read the second and third rows together. In Refolk's index, US Rust engineers (601) and Mexico Python engineers (685) are within 14% of each other. A bigger country buys you nothing once the stack is rare. Specialization, not geography, is the binding constraint. If the mandate had required Rust rather than Python, the coastal default would have shrunk to nearshore scale, and the whole site-selection logic would flip.

55,048
US software and backend engineers with Python in Refolk's index
Top metros San Francisco and New York; this is the ceiling the coastal default starts from before any discount.
685
Mexico software and backend engineers with Python in Refolk's index
Top metros Guadalajara, Mexico City, Monterrey; close enough to the 601 US Rust pool to prove the stack matters more than the country.

Counting the pool yourself for your exact spec, in one place, is what turns a vendor ranking into a decision you can defend. Refolk lets you run the role cluster and the city in plain English and get the current-role count back, instead of stitching three incompatible exports together.

The two wrong turns that inflate a pool

Two moves inflate every raw count, and both happened in the worked mandate before I caught them. Naming them is the most valuable part of this standard, because they are invisible in the headline number.

Wrong turn one: trusting the whole-country offshore number. The first nearshore pull used a published figure of more than 1.9 million Mexican developers. Carrying that into the TAM made the nearshore option look bottomless. It is a definition artifact. That 1.9 million counts a different object than the 220,000 government figure. The index count of 685 for Python backend engineers is the number that survives contact with the actual spec. The gap between 1.9 million and 685 is not reality; it is four sources counting four things.

Wrong turn two: stacking narrow filters to zero. When I tried to tighten the US pull with Python plus Senior-seniority plus a strict title filter at once, Refolk's index returned 0. There is no world where the real senior Python pool is zero. Stacking three narrow filters nulled a viable pool. The fix is to relax one filter and re-count, then reapply the missing constraint as a proportion rather than a hard gate. If you take a zero at face value, you write off a city that is actually fine.

Stripping over-inclusion down to real engineers

Apply exact-title matching, NOT operators, and identity dedupe before you trust any count, and expect double-digit percentage shrinkage. Over-inclusion is not a rounding error; it is the difference between a deep-looking city and a reachable one.

The dominant inflator is stale profiles. People forget to put an end date on a previous job when they move, and unclosed past positions are responsible for the majority of false positives in people search. A profile reads as a current engineer when the person left the role years ago. Second is synonym bleed: selecting a title value rather than an exact phrase pulls in the synonyms the system guesses, so an SDE, a SAS Programmer, even a recruiting coordinator lands in the count. Third is whole-profile keyword match, where "Python" counts someone who studied it once at school because the search scanned recommendations and old descriptions instead of the current-role skill field.

Is this count safe to carry forward?

High competitor lockLow competitor lock
Thin and reachable
Viable only for a small team; confirm seniority share before committing
Deep and reachable
The target profile; count reachable, then verify difficulty
Thin and locked
Walk; neither depth nor mobility supports a hub
Deep and locked
The trap; headline looks great, reachable pool is small
Thin raw poolDeep raw pool
Depth means nothing until you know how much of the pool you can actually move.

The bottom-right quadrant is where a generic ranking sends you wrong. A large raw pool that is 60% locked inside two local competitors is not reachable. Read employer concentration in the pool before celebrating the headline.

The procedure, start to finish

Run these eight steps in order. One caveat on order: concentration-led scorecards start from clustering weights, but boolean practitioners insist the over-inclusion strip must precede any counting, or every later number is wrong. I side with counting clean first.

Count a second-hub decision end to end

  1. Define the role cluster and stack precisely
    Fix titles, required skills, and seniority band into one exact-match query spec before counting. A loose spec guarantees synonym inflation downstream.
  2. Shortlist two or three candidate cities
    Use the 50,000-worker large-market line and location quotient to drop markets too thin to matter. Give each shortlisted city a documented reason.
  3. Pull a raw pool from each source per city
    Query the profile index, the BLS metro table, and GitHub geographic data against the same spec. Record three numbers from three definitions, each with its URL.
  4. Reconcile divergence
    Where sources disagree 5x to 9x, carry the employer survey as the floor and the profile count as the ceiling, and state which you take forward.
  5. Strip over-inclusion
    Apply exact-title matching, NOT operators, and identity dedupe. Expect double-digit shrinkage and a count of current-role, in-cluster people only.
  6. Discount for reachability
    Subtract the competitor-locked fraction and weight by target-seniority share and annual attrition to get one effective-reachable number per city.
  7. Score hiring difficulty
    Overlay active postings and offer-acceptance benchmarks to produce a supply-to-demand ratio per city.
  8. Recommend one city with the table
    Present raw pool, effective reachable, cost, and difficulty side by side so an exec sees the full chain from headline to decision.

Discounting for reachability and scoring difficulty

The reachable count is the raw pool minus false positives, minus the fraction you cannot poach, adjusted by the slice at your target seniority and annual churn. The discounts are measurable, not vibes. Excluding competitor employers cut one real title search by 17%, and removing seniority higher than the target cut another 24%. Together that is roughly a third gone before attrition even enters.

Then attrition, which cuts both ways. Engineering attrition held at 12% in 2025, the lowest of any function, and a workforce report put software-engineer attrition at 13.2% against about 9% cross-industry. Low churn is good for retention and bad for hiring: a deep pool with low attrition means fewer passive movers to poach. Supply you cannot move is not supply. Separately, 69% of developers stay under two years at one employer, which tells you the movers exist but concentrate at the junior and mid bands.

Score difficulty from demand pressure. Active core-tech postings totaled nearly 490,000 in one February, with newly posted ads up 15% year over year, and metro-level deltas are published, so you can read which of your shortlisted cities is heating up. On the funnel's close side, technical-role offer acceptance runs 73% against 84% for business roles, so plan for more offers than a business team would to land the same headcount.

A deep pool with low churn is a retention story and a hiring problem at the same time.
CountryPublished dev pool (range)Monthly mid-level salary
Mexico220K to 1.9M~$3,156
Brazil~500K to 760K~$6,000
US (reference)1.66M (employer survey)~$10,154

The pool ranges diverge widely by definition, so treat the wide end as a ceiling only. The salary column is the real reason a nearshore hub stays on the list, but it only matters if the reachable count in the first table can staff a team.

Failure modes to check before you present

Run every count against this list before it reaches an exec. Each failure mode has a false positive it produces and a cheap check that catches it.

  • Stale-profile inflation. A "current" title that ended years ago. Check: sample 20 profiles for end dates; expect most of the noise here.
  • Synonym bleed. A non-engineer in the count. Check: rerun with exact-phrase quotes and compare the delta.
  • Whole-profile keyword match. Someone who studied the skill once. Check: constrain to current-role skill fields, not free text.
  • Cross-platform double count. The same engineer on two platforms. Check: dedupe by identity before summing; never add two platform totals.
  • Category creep. Counting developers plus QA plus testers (1.9M) when you want developers only (1.66M). Check: lock the occupation code or exact title set.
  • Over-filtering to zero. Stacked filters null a real pool. Check: relax one filter and re-count.
  • Competitor blindness. A large pool that is mostly locked at local rivals. Check: read employer concentration before celebrating the headline.

Before you call the count defensible

  • Each city has three raw numbers from three sources, each with its URL
  • You stated which source is the floor and which is the ceiling for every city
  • You sampled 20 profiles per city and confirmed current end dates
  • Exact-title and NOT operators applied, and identity dedupe run across sources
  • Competitor-locked fraction subtracted and target-seniority share applied
  • Attrition weighting documented, and no stacked query left sitting at zero
  • Final table shows raw pool, effective reachable, cost, and difficulty side by side

Making the recommendation and keeping it current

Recommend one city with the full table visible, so the exec sees the chain from headline to reachable number rather than a single asserted figure. The deliverable is one row per city with raw pool, effective reachable, median cost, and supply-to-demand ratio. The city that wins is rarely the deepest raw pool; it is the best reachable count at an acceptable cost and difficulty.

Second-hub decision table skeleton
| City | Raw pool (source) | Effective reachable | Median base | Active postings | Supply:demand | Note |
|------|-------------------|---------------------|-------------|-----------------|---------------|------|
| City A | ___ (employer survey floor) | ___ | $___ | ___ | ___ | ___ |
| City B | ___ (profile ceiling) | ___ | $___ | ___ | ___ | ___ |
| City C | ___ (reconciled) | ___ | $___ | ___ | ___ | ___ |
Recommendation: ___ because effective reachable ÷ annual openings beats the alternatives at acceptable cost.

Fill one row per shortlisted city; keep every source URL in a footnote so the number is auditable.

Keep the count current by re-running two things. First, the raw pull, because profiles accumulate and the stale-profile share drifts; re-count quarterly during an active build. Second, the difficulty overlay, because postings and the hiring cycle swing hard. One software postings index ran a peak of 225 in early 2022 and sat near 70 in early 2024 and early 2026, against a 2020 baseline of 100. A ratio that looked comfortable at the peak can invert in a cooler market, so re-check the mechanism rather than trusting a number you captured months ago. Document the source and date on every figure, and the recommendation stays defensible long after the first pull.

Questions practitioners ask

Can I just use a vendor top-40 tech talent ranking instead of building my own count?

No, not for a defensible decision. A published ranking weights metrics for the average employer, so it tells you where generic talent clusters, not where your specific stack is deep and reachable. Use the ranking to seed a shortlist and to apply the 50,000-worker large-market line, then build your own count for your exact role cluster. The gap between a generic ranking and your reachable number is exactly where the decision lives.

Why do my raw talent-pool numbers for the same city disagree so much?

Because each source counts a different object. An employer survey counts jobs, a profile index counts accounts, and a developer platform counts signups geolocated by IP. For Mexico, published counts span 220,000 to 1.9 million, a 5x to 9x spread, purely from definition. Carry the employer survey as your floor and the profile count as your ceiling, and never add two platform totals together because you will double-count the same people.

How much should I discount a raw pool to get a reachable number?

Start from documented boolean discounts: excluding competitor employers cut one real search by 17%, and removing seniority above your target cut another 24%. That is roughly a third gone before attrition. Then weight by the share at your target seniority and by annual attrition, which held at about 12% for engineering. There is no single published cutoff, so document each discount so an exec can audit the chain.

Is a nearshore hub actually cheaper enough to matter?

Often yes on salary and no on reachable depth. A mid-level engineer in Mexico runs about $3,156 per month against roughly $10,154 for a US counterpart, and Latin American wages average about 38% of US levels. But if your stack is rare, the reachable pool can be tiny: Mexico Python engineers number 685 in Refolk's index, close to US Rust engineers at 601. Cost arbitrage means nothing against a pool you cannot staff a team from.

What is the single most common mistake that inflates a city's talent pool?

Stale profiles. People forget to add an end date when they change jobs, so an unclosed past position reads as a current role and gets counted. This is responsible for the majority of false positives in people search. Sample 20 profiles for end dates before you trust any raw count, and constrain your query to current-role skill fields rather than scanning the whole profile.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next