Mapping a Competitor's Real Location Footprint, Metro by Metro
You can produce a metro-by-function map of a rival's workforce, corrected for remote-to-HQ inflation and vague regions, with a confidence flag on every cell.
You want to know where a rival's people actually sit: which metros hold how much of which function, so you can pick a site, plan a poaching push, or set a pay band. This guide is the outside-in procedure for building that map from public employee signals when you have no access to the target's HRIS. It is written for strategy, research, and talent-intelligence teams, and it corrects for the two failures that quietly break every location count - remote staff defaulting to the headquarters metro, and profiles that only say "Greater X Area" - then puts a confidence flag on every cell so you know which numbers to act on.
The vendor pages that rank for this job pitch location-intelligence tools that assume you already have clean payroll data by city. You do not. What you have is a population of public profiles, each with a self-reported location string of uneven honesty, plus a handful of government datasets that bound a metro but never resolve a single firm. The map is buildable from those, but only if you handle the noise deliberately.
What "location footprint" means and why the naive count is wrong
A location footprint is a grid: functions down one axis, metros across the other, with a headcount and a confidence flag in each cell. The naive way to build it - count profiles, group by the location field - produces a map that is confidently wrong in two specific places.
The first error is remote inflation. Remote workers sit a median 185 miles from their employer's headquarters, versus 10 to 15 miles for fully in-person staff, yet many keep the HQ label on their profile. So the HQ metro cell absorbs people who have never worked there. The useful thing about this error is that it concentrates: it lands almost entirely in one cell, so a single correction fixes most of the map.
The second error is vagueness. LinkedIn derives a location from a postal code and then offers the member a choice - "Denver, Colorado" or "the Greater Denver Area" - and users pick the broader label to appear in more searches. Country-only strings like "United States" are the extreme case. If you collapse "Greater X Area" to the downtown core you over-credit downtown; if you force country-only rows to HQ you compound the first error.
The people you most need to place are exactly the ones whose profiles refuse to say where they are.
These two errors interact, and the interaction is the whole reason this guide exists. In Refolk's index, profiles that surface on a "remote" keyword are three times more likely to give a country-only location than general profiles - 36% versus 12% in matched 25-profile samples. Vagueness and remoteness co-occur. That means the unknown bucket is disproportionately remote, so you cannot distribute it pro-rata across metros. Do that and you re-inject the remote population you just removed.
The three families of location signal and what each proves
There are three families of location evidence, and each proves something different. Mixing them without labelling which is which is how footprints get built on sand.
| Signal family | What it proves | How it misleads |
|---|---|---|
| Self-reported profile location | Where a person says they are | Remote staff default to HQ; "Greater X Area" hides the real city |
| Job-posting office cities | Where a company is hiring | Not where current staff sit |
| Government establishment data (QCEW, LODES) | Jobs physically reported at an address | Cannot isolate one company or one function |
Self-reported profile location is your primary spine because it is the only source that ties a person to a company and a function at the same time. It is also the noisiest, which is why the correction steps below exist. Job-posting cities tell you intent, not reality; keep them out of the count and use them only as a directional cross-check on where a metro is growing. Government data is the opposite: it is real physical employment, covering more than 95 percent of U.S. jobs, but it cannot see a single firm. It bounds your estimate. It never confirms it.
The key discipline is to infer remote status structurally rather than by keyword. Only 215 of 325,147 US Software Engineer profiles in Refolk's index surface on the word "remote." If you search for the word, you miss everyone. Instead, infer it from role type, distance between the stated home and the employer's known offices, and the 36% figure from a 2024 ADP Research Institute analysis showing employees are more likely to report to a manager who lives elsewhere.
Profile supply differs by country, so set your priors first
Before you count anything, calibrate what "big" looks like in each country you are mapping. Profile supply is not uniform, and a rival's cells in a low-supply country will look thin in absolute terms even when the company dominates locally.
| Title | Country | Total profiles | Derived US:country ratio |
|---|---|---|---|
| Software Engineer | United States | 325,147 | 1.00 |
| Software Engineer | United Kingdom | 40,707 | 7.99x |
Totals are from Refolk's index; the ratio is derived (325,147 / 40,707 = 7.99).
The US has roughly eight times the UK Software Engineer supply. If your target runs a London engineering hub and a San Francisco one, the raw profile counts will make London look like a satellite even if it is the larger real team, purely because the underlying population you are sampling from is smaller. Normalise within-country before comparing across borders. The cleanest way is to express each metro cell as a share of that country's mapped total for the function, then compare shares, not raw counts, across the border.
The vague-location problem, measured
The share of profiles that give only a vague region is the single input that determines how much of your map is trustworthy, and it varies sharply by segment. No public source cleanly quantifies the metro-versus-vague split across all profiles, so treat any single published number as not established. Refolk's index gives concrete directional readings from matched samples.
| Segment | Country-only in 25-profile sample | Share | Derived multiple |
|---|---|---|---|
| US Software Engineer (general) | 3 | 12% | 1.0x |
| US Software Engineer (remote keyword) | 9 | 36% | 3.0x |
| UK Software Engineer (general) | 4 | 16% | 1.3x |
Raw counts are from Refolk's index topRegions samples; shares and multiples are derived. These are 25-row samples, so read them as directional, not precise.
Two things follow. First, a general engineering population loses roughly one in eight rows to country-only vagueness, which is manageable if you keep those rows explicit. Second, the moment a population skews remote, that loss triples. The vague bucket is not random missing data - it is structurally weighted toward the people you most need to place. That is why step three below routes country-only and "Remote" to an explicit unknown bucket, and why step four never redistributes that bucket blindly across the metros.
Where profiles fall out of a clean metro count
- 100Raw profiles pulled
Everyone tied to the employer
- 90Deduplicated roster
Duplicates and stale records removed
- 79Resolves to a specific metro
About 12% land in the country-only bucket
- 70Survives remote correction
Modelled remote share moved out of HQ into unknown
Run the map: an eight-step procedure
This is the full procedure, start to finish. It runs about four to five analyst-days for a single mid-size target. Each step names who does it, how long it takes, and what "done" looks like.
Metro-by-function footprint, end to end
- Scope the target and functionsDefine the company entity, the functions to map such as engineering and sales, and the candidate metro list. Done when you have a fixed function taxonomy and a candidate metro set. About half a day.
- Pull the profile populationCollect all public profiles tied to the employer with current title and raw location string. Done when you have a deduplicated roster whose headcount roughly reconciles to a known total. Half to one day.
- Normalise locations to metrosMap each raw string to a metro via postal-code and geo lookup, route "Greater X Area" to its core metro, and route country-only or "Remote" to an explicit unknown bucket rather than silently to HQ. Done when every row carries a metro or an unknown flag. About one day.
- Apply the remote and HQ correctionEstimate how many HQ-metro rows are actually remote using role type and the 185-mile and 36% manager-elsewhere anchors, then move a modelled share out of HQ into unknown. Done when the HQ cell has a stated correction and a residual confidence flag. About half a day.
- Scale the sample to full countsInflate located profiles to the estimated true function headcount per metro while tracking the coverage ratio. Done when you have a metro-by-function grid with counts. About half a day.
- Corroborate against government dataBound each metro's function count with QCEW establishment and employment at county or MSA level, and use LODES block-level workplace counts where needed. Done when each cell has an external upper and lower bound and impossible cells are re-examined. Half to one day.
- Attach confidence flagsGrade every cell by sample coverage, vague-share, and remote exposure. Done when each cell carries a green, amber, or red flag. About half a day.
- Set the re-map cadenceDefine the headcount-delta and staleness thresholds that trigger a refresh. Done when you have a documented monitoring rule. About a quarter day.
On step two, the population pull
Reconcile your roster to a known total early. A published headcount, even a rough one, tells you your coverage ratio - located profiles divided by true headcount - which every later step depends on. If one account is doing deep extraction, budget for it: a single LinkedIn account safely deep-extracts about 50 profiles a day, so a 2,000-person target is weeks of manual collection. This is where structured sourcing earns its place. Refolk lets you ask for the population in plain English and pull the roster with location and title attached, which is the friction step three cannot start without.
On steps four and five, correction order
Sources disagree on whether to correct for remote before or after scaling. Correcting first is cleaner because you scale a corrected base, but it assumes your sample is representative of the true remote share. If your coverage ratio is uneven across metros, scale first per metro, then apply the correction to the HQ cell only. Whichever order you pick, document it, because the two produce different HQ numbers.
Corroborate with government data without over-trusting it
Government establishment data bounds a metro cell; it never resolves one firm. Use it to catch impossible cells - an HQ engineering count that exceeds plausible office employment for the whole county is a clear sign of remote inflation - not to confirm a number.
| Source | Finest geography | Time lag | Coverage |
|---|---|---|---|
| QCEW | County / MSA | ~6 months | >95% of U.S. jobs |
| LODES / OnTheMap | Census block | Annual, through 2023 | UI-covered jobs |
| QWI | County | Quarterly | UI-covered jobs |
Sources: bls.gov/cew; lehd.ces.census.gov; nj.gov LED documentation.
QCEW is your default bound. It publishes establishment, employment, and wage data down to 6-digit NAICS at county level where disclosure rules permit, released roughly six months after each quarter, the fastest source at that detail. When you need sub-county resolution - to see whether jobs cluster in a specific business district - LODES workplace-area characteristics reach census-block detail for most states through 2023. QWI adds detailed industry and person characteristics but offers no geography below county, so it is a supporting read, not a mapping source.
The trap is suppression. Much QCEW county data is withheld to protect the anonymity of individual firms, and a suppressed cell can look like a zero. Enable "show suppressed" and treat suppression as unknown, never as absence. A "zero" that is really a withheld cell will make you conclude a metro has no employment when it may have plenty.
What each layer of evidence can and cannot claim
- Government boundQCEW and LODES cap a metro at plausible physical employment
- Scaled profile countInflated located profiles fill the metro-by-function grid
- Located profile baseRows resolved to a specific metro after normalisation
- Remote correctionHQ cell adjusted down using distance and manager-elsewhere anchors
- Unknown bucketCountry-only and Remote rows held explicitly, never redistributed
How this goes wrong: the failure modes that break the count
Most footprint maps fail in one of seven ways. Each has a false positive it produces and a specific check that catches it. Run these before you flag any cell green.
- Remote defaulting to HQ. Produces an inflated HQ function cell. Check: compare the HQ cell to QCEW establishment employment for that county; if the profile count exceeds plausible office capacity, remote inflation is present.
- "Greater X Area" collapsed to the core city. Produces suburb-heavy metros over-credited to downtown. Check: sample the raw strings and confirm postal-code resolution rather than label matching.
- Country-only or "Remote" silently dropped or forced to HQ. This is the exact population you must not misassign, since remote-keyword profiles are three times more likely to be country-only. Check: keep an explicit unknown bucket and report its size.
- Scaling from a non-representative sample. Produces a metro that over-indexes on public profiles looking larger than it is. Check: report the coverage ratio per metro and flag cells below a coverage floor red.
- Stale locations read as current. A mapped mover still shows the old metro because profiles go materially stale within a year. Check: weight by profile freshness and discount cells built on old snapshots.
- Corroborating against suppressed QCEW cells. Produces a "zero" that is actually withheld. Check: enable "show suppressed" and treat suppression as unknown, not absence.
- Counting the keyword instead of the attribute. Keyword-based remote detection misses almost everyone because only 0.066% of US Software Engineer profiles surface on "remote." Check: infer remote from distance-to-office and role, not the word.
The two that matter most are the first and the seventh, because they compound. If you both under-detect remote workers (by searching the keyword) and default the ones you miss to HQ, your HQ engineering cell can be off by a wide margin while looking perfectly clean. The structural inference in step four is the only defence.
Grade every cell: the confidence flag
Every cell gets a green, amber, or red flag driven by three inputs: sample coverage, vague-share, and remote exposure. A number without a flag is not a finding; it is a guess wearing a finding's clothes. Use this rubric and adjust the thresholds to your own risk tolerance, documenting whatever you choose.
GREEN - Coverage ratio >= 0.6, vague-share < 15%, not the HQ metro OR remote correction applied and residual < 10%. AMBER - Coverage ratio 0.3 to 0.6, OR vague-share 15% to 30%, OR HQ metro with correction applied but residual 10% to 25%. RED - Coverage ratio < 0.3, OR vague-share > 30%, OR HQ metro with no remote correction, OR bounded only by a suppressed QCEW cell.
Apply per cell. A cell drops to the lowest flag any single input triggers.
A green cell is one you can put in front of a site-selection committee or use to set a pay band. An amber cell is directional - fine for prioritising a poaching push, not for a capital decision. A red cell is a placeholder that says "we do not yet know," which is a legitimate and useful thing for a map to say. Because 55% of organisations set geographic pay by city or metro area, a mislabelled cell here has a direct cost, so err toward the lower flag when an input is borderline.
Before you call the footprint done
- Every row resolves to a metro or sits in an explicit unknown bucket - none silently forced to HQ.
- The HQ cell has a stated remote correction and a residual confidence flag.
- Each metro cell reports its coverage ratio, and cells below the floor are flagged red.
- Every cell has a QCEW or LODES bound, and suppressed government cells are treated as unknown, not zero.
- Cross-border cells are expressed as within-country function shares before comparison.
- The unknown bucket's size is reported and has not been redistributed pro-rata.
- Profile freshness is recorded and stale-heavy cells are discounted.
- Every cell carries a green, amber, or red flag with the rule that set it.
Keep the map current
A footprint map is a snapshot with a short half-life, so decide its refresh rule the day you ship it. Public profiles go materially stale within a year if never re-collected, because a meaningful share of professionals change jobs each year, while LinkedIn revises headcount only about twice yearly, typically applied retroactively. Guidance suggests members refresh their profiles at least quarterly, which implies location fields drift on a similar cadence.
No published percentage-over-window rule exists to trigger a re-map, so set your own threshold and document it. A defensible default: a light quarterly refresh that re-pulls the roster and recomputes shares, plus a full re-map triggered by any of three events - a published headcount change beyond a delta you set, a new office announcement in a metro you did not have, or a restructuring that moves a function. Anchor the delta to the coverage you can afford: if a single account extracts about 50 profiles a day, a full re-map of a large target is a multi-week project, so reserve it for genuine change rather than the calendar.
The map earns its keep when it survives contact with a decision. If a site-selection committee, a compensation team, or a recruiting lead can point at a cell, see its flag, and know whether to trust it, you have built the thing the vendor pitches assume you already own. The difference is that you built it from the outside, from public signal, with the two silent failures named and corrected - which is the only version of this map you can actually defend.
Questions practitioners ask
How do I tell if a company's engineers actually work at HQ or are remote?
Do not trust the self-reported location or a keyword search. In Refolk's index only 0.066% of US Software Engineer profiles surface on the word 'remote', so keyword detection misses almost everyone. Infer remote structurally: compare the HQ engineering cell against QCEW establishment employment for that county, and if the profile count exceeds plausible office capacity, remote inflation is present. Remote workers sit a median 185 miles from HQ yet often keep the HQ label.
What does 'Greater Seattle Area' mean and how do I resolve it to a metro?
It is a user choice, not a data gap. LinkedIn derives a location from a postal code and offers the member both the city and the broader 'Greater X Area' label, and many pick the broader one to appear in more searches. Resolve it by postal-code lookup first, then cross-reference the employer's known office cities, and if neither resolves it, route the row to an explicit unknown bucket rather than collapsing it to the downtown core.
Can I use QCEW to confirm how many people a specific company employs in a city?
No. QCEW covers more than 95 percent of U.S. jobs at county and MSA level, but it cannot isolate one company or one function, and much county data is suppressed to protect firm anonymity. Use it to bound a metro estimate, capping a cell at plausible employment, not to confirm one firm's share. Treat suppressed cells as unknown, never as zero.
How often should I re-map a competitor's footprint?
No published percentage-over-window rule exists, so set the threshold yourself. Anchor it to two documented facts: public profiles go materially stale within a year if never re-collected, and LinkedIn revises headcount only about twice yearly, typically retroactively. Guidance suggests location fields drift on a roughly quarterly cadence, so a quarterly light refresh with a full re-map on a material headcount delta is defensible.
Why do my UK metro cells look so thin compared with the US ones?
Because profile supply differs by country. In Refolk's index the US has about eight times the UK Software Engineer supply (325,147 versus 40,707). A rival's UK metro cells will look small in absolute terms even when locally dominant, so always normalise within-country before comparing across borders, or you will misread a strong local hub as a weak one.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.