# Mapping a Competitor's Real Location Footprint, Metro by Metro

*You can produce a metro-by-function map of a rival's workforce, corrected for remote-to-HQ inflation and vague regions, with a confidence flag on every cell.*

- Canonical URL: https://www.refolk.ai/guides/competitor-location-footprint-map
- Pillar: Market and talent intelligence
- Format: Playbook
- Published: 2026-09-28
- Last reviewed: 2026-09-28
- Reading time: 17 min

You want to know where a rival's people actually sit: which metros hold how much of which function, so you can pick a site, plan a poaching push, or set a pay band. This guide is the outside-in procedure for building that map from public employee signals when you have no access to the target's HRIS. It is written for strategy, research, and talent-intelligence teams, and it corrects for the two failures that quietly break every location count - remote staff defaulting to the headquarters metro, and profiles that only say "Greater X Area" - then puts a confidence flag on every cell so you know which numbers to act on.

The vendor pages that rank for this job pitch location-intelligence tools that assume you already have clean payroll data by city. You do not. What you have is a population of public profiles, each with a self-reported location string of uneven honesty, plus a handful of government datasets that bound a metro but never resolve a single firm. The map is buildable from those, but only if you handle the noise deliberately.

## What "location footprint" means and why the naive count is wrong

A location footprint is a grid: functions down one axis, metros across the other, with a headcount and a confidence flag in each cell. The naive way to build it - count profiles, group by the location field - produces a map that is confidently wrong in two specific places.

The first error is remote inflation. Remote workers sit a median 185 miles from their employer's headquarters, versus 10 to 15 miles for fully in-person staff, yet many keep the HQ label on their profile. So the HQ metro cell absorbs people who have never worked there. The useful thing about this error is that it concentrates: it lands almost entirely in one cell, so a single correction fixes most of the map.

The second error is vagueness. LinkedIn derives a location from a postal code and then offers the member a choice - "Denver, Colorado" or "the Greater Denver Area" - and users pick the broader label to appear in more searches. Country-only strings like "United States" are the extreme case. If you collapse "Greater X Area" to the downtown core you over-credit downtown; if you force country-only rows to HQ you compound the first error.

> The people you most need to place are exactly the ones whose profiles refuse to say where they are.

These two errors interact, and the interaction is the whole reason this guide exists. In Refolk's index, profiles that surface on a "remote" keyword are three times more likely to give a country-only location than general profiles - 36% versus 12% in matched 25-profile samples. Vagueness and remoteness co-occur. That means the unknown bucket is disproportionately remote, so you cannot distribute it pro-rata across metros. Do that and you re-inject the remote population you just removed.

**0.066% - Share of US Software Engineer profiles that surface on a "remote" keyword**

In Refolk's index, only 215 of 325,147 US Software Engineer profiles match "remote," so keyword search finds almost no remote workers.

## The three families of location signal and what each proves

There are three families of location evidence, and each proves something different. Mixing them without labelling which is which is how footprints get built on sand.

| Signal family | What it proves | How it misleads |
|---|---|---|
| Self-reported profile location | Where a person says they are | Remote staff default to HQ; "Greater X Area" hides the real city |
| Job-posting office cities | Where a company is hiring | Not where current staff sit |
| Government establishment data (QCEW, LODES) | Jobs physically reported at an address | Cannot isolate one company or one function |

Self-reported profile location is your primary spine because it is the only source that ties a person to a company and a function at the same time. It is also the noisiest, which is why the correction steps below exist. Job-posting cities tell you intent, not reality; keep them out of the count and use them only as a directional cross-check on where a metro is growing. Government data is the opposite: it is real physical employment, covering more than 95 percent of U.S. jobs, but it cannot see a single firm. It bounds your estimate. It never confirms it.

The key discipline is to infer remote status structurally rather than by keyword. Only 215 of 325,147 US Software Engineer profiles in Refolk's index surface on the word "remote." If you search for the word, you miss everyone. Instead, infer it from role type, distance between the stated home and the employer's known offices, and the 36% figure from a 2024 ADP Research Institute analysis showing employees are more likely to report to a manager who lives elsewhere.

## Profile supply differs by country, so set your priors first

Before you count anything, calibrate what "big" looks like in each country you are mapping. Profile supply is not uniform, and a rival's cells in a low-supply country will look thin in absolute terms even when the company dominates locally.

| Title | Country | Total profiles | Derived US:country ratio |
|---|---|---|---|
| Software Engineer | United States | 325,147 | 1.00 |
| Software Engineer | United Kingdom | 40,707 | 7.99x |

Totals are from Refolk's index; the ratio is derived (325,147 / 40,707 = 7.99).

The US has roughly eight times the UK Software Engineer supply. If your target runs a London engineering hub and a San Francisco one, the raw profile counts will make London look like a satellite even if it is the larger real team, purely because the underlying population you are sampling from is smaller. Normalise within-country before comparing across borders. The cleanest way is to express each metro cell as a share of that country's mapped total for the function, then compare shares, not raw counts, across the border.

> **Rule:** Normalise within country before comparing across
>
> Always convert cross-border cells to within-country function shares before comparing them. Raw profile counts embed a country supply bias that will misread a dominant foreign hub as a minor one.

## The vague-location problem, measured

The share of profiles that give only a vague region is the single input that determines how much of your map is trustworthy, and it varies sharply by segment. No public source cleanly quantifies the metro-versus-vague split across all profiles, so treat any single published number as not established. Refolk's index gives concrete directional readings from matched samples.

| Segment | Country-only in 25-profile sample | Share | Derived multiple |
|---|---|---|---|
| US Software Engineer (general) | 3 | 12% | 1.0x |
| US Software Engineer (remote keyword) | 9 | 36% | 3.0x |
| UK Software Engineer (general) | 4 | 16% | 1.3x |

Raw counts are from Refolk's index topRegions samples; shares and multiples are derived. These are 25-row samples, so read them as directional, not precise.

Two things follow. First, a general engineering population loses roughly one in eight rows to country-only vagueness, which is manageable if you keep those rows explicit. Second, the moment a population skews remote, that loss triples. The vague bucket is not random missing data - it is structurally weighted toward the people you most need to place. That is why step three below routes country-only and "Remote" to an explicit unknown bucket, and why step four never redistributes that bucket blindly across the metros.

#### Where profiles fall out of a clean metro count

| Stage | Figure | Note |
| --- | --- | --- |
| Raw profiles pulled | 100 | Everyone tied to the employer |
| Deduplicated roster | 90 | Duplicates and stale records removed |
| Resolves to a specific metro | 79 | About 12% land in the country-only bucket |
| Survives remote correction | 70 | Modelled remote share moved out of HQ into unknown |

*Each stage removes rows that cannot be trusted to a metro, so the located base is smaller than the raw pull.*

## Run the map: an eight-step procedure

This is the full procedure, start to finish. It runs about four to five analyst-days for a single mid-size target. Each step names who does it, how long it takes, and what "done" looks like.

#### Metro-by-function footprint, end to end

1. **Scope the target and functions** - Define the company entity, the functions to map such as engineering and sales, and the candidate metro list. Done when you have a fixed function taxonomy and a candidate metro set. About half a day.
2. **Pull the profile population** - Collect all public profiles tied to the employer with current title and raw location string. Done when you have a deduplicated roster whose headcount roughly reconciles to a known total. Half to one day.
3. **Normalise locations to metros** - Map each raw string to a metro via postal-code and geo lookup, route "Greater X Area" to its core metro, and route country-only or "Remote" to an explicit unknown bucket rather than silently to HQ. Done when every row carries a metro or an unknown flag. About one day.
4. **Apply the remote and HQ correction** - Estimate how many HQ-metro rows are actually remote using role type and the 185-mile and 36% manager-elsewhere anchors, then move a modelled share out of HQ into unknown. Done when the HQ cell has a stated correction and a residual confidence flag. About half a day.
5. **Scale the sample to full counts** - Inflate located profiles to the estimated true function headcount per metro while tracking the coverage ratio. Done when you have a metro-by-function grid with counts. About half a day.
6. **Corroborate against government data** - Bound each metro's function count with QCEW establishment and employment at county or MSA level, and use LODES block-level workplace counts where needed. Done when each cell has an external upper and lower bound and impossible cells are re-examined. Half to one day.
7. **Attach confidence flags** - Grade every cell by sample coverage, vague-share, and remote exposure. Done when each cell carries a green, amber, or red flag. About half a day.
8. **Set the re-map cadence** - Define the headcount-delta and staleness thresholds that trigger a refresh. Done when you have a documented monitoring rule. About a quarter day.

### On step two, the population pull

Reconcile your roster to a known total early. A published headcount, even a rough one, tells you your coverage ratio - located profiles divided by true headcount - which every later step depends on. If one account is doing deep extraction, budget for it: a single LinkedIn account safely deep-extracts about 50 profiles a day, so a 2,000-person target is weeks of manual collection. This is where structured sourcing earns its place. Refolk lets you ask for the population in plain English and pull the roster with location and title attached, which is the friction step three cannot start without.

I ran this search: `People at Stripe whose location says 'Greater Seattle Area' with an engineering or product title.` - [see the full result list](https://www.refolk.ai/s/s8qkwzyk25).

*Returns the exact vague-region rows you must resolve by postal code rather than collapse to a downtown default.*

### On steps four and five, correction order

Sources disagree on whether to correct for remote before or after scaling. Correcting first is cleaner because you scale a corrected base, but it assumes your sample is representative of the true remote share. If your coverage ratio is uneven across metros, scale first per metro, then apply the correction to the HQ cell only. Whichever order you pick, document it, because the two produce different HQ numbers.

> **Watch out:** Never redistribute the unknown bucket pro-rata
>
> The unknown bucket is disproportionately remote (36% country-only among remote-keyword profiles versus 12% general). Spreading it evenly across metros re-injects the remote population you just removed from HQ. Report the bucket's size and leave it as unknown.

## Corroborate with government data without over-trusting it

Government establishment data bounds a metro cell; it never resolves one firm. Use it to catch impossible cells - an HQ engineering count that exceeds plausible office employment for the whole county is a clear sign of remote inflation - not to confirm a number.

| Source | Finest geography | Time lag | Coverage |
|---|---|---|---|
| QCEW | County / MSA | ~6 months | >95% of U.S. jobs |
| LODES / OnTheMap | Census block | Annual, through 2023 | UI-covered jobs |
| QWI | County | Quarterly | UI-covered jobs |

Sources: bls.gov/cew; lehd.ces.census.gov; nj.gov LED documentation.

QCEW is your default bound. It publishes establishment, employment, and wage data down to 6-digit NAICS at county level where disclosure rules permit, released roughly six months after each quarter, the fastest source at that detail. When you need sub-county resolution - to see whether jobs cluster in a specific business district - LODES workplace-area characteristics reach census-block detail for most states through 2023. QWI adds detailed industry and person characteristics but offers no geography below county, so it is a supporting read, not a mapping source.

The trap is suppression. Much QCEW county data is withheld to protect the anonymity of individual firms, and a suppressed cell can look like a zero. Enable "show suppressed" and treat suppression as unknown, never as absence. A "zero" that is really a withheld cell will make you conclude a metro has no employment when it may have plenty.

#### What each layer of evidence can and cannot claim

1. **Government bound** - QCEW and LODES cap a metro at plausible physical employment
2. **Scaled profile count** - Inflated located profiles fill the metro-by-function grid
3. **Located profile base** - Rows resolved to a specific metro after normalisation
4. **Remote correction** - HQ cell adjusted down using distance and manager-elsewhere anchors
5. **Unknown bucket** - Country-only and Remote rows held explicitly, never redistributed

*Read the map from the outside in: government data caps the estimate, profiles populate it, and the correction cleans the HQ cell.*

## How this goes wrong: the failure modes that break the count

Most footprint maps fail in one of seven ways. Each has a false positive it produces and a specific check that catches it. Run these before you flag any cell green.

- **Remote defaulting to HQ.** Produces an inflated HQ function cell. Check: compare the HQ cell to QCEW establishment employment for that county; if the profile count exceeds plausible office capacity, remote inflation is present.
- **"Greater X Area" collapsed to the core city.** Produces suburb-heavy metros over-credited to downtown. Check: sample the raw strings and confirm postal-code resolution rather than label matching.
- **Country-only or "Remote" silently dropped or forced to HQ.** This is the exact population you must not misassign, since remote-keyword profiles are three times more likely to be country-only. Check: keep an explicit unknown bucket and report its size.
- **Scaling from a non-representative sample.** Produces a metro that over-indexes on public profiles looking larger than it is. Check: report the coverage ratio per metro and flag cells below a coverage floor red.
- **Stale locations read as current.** A mapped mover still shows the old metro because profiles go materially stale within a year. Check: weight by profile freshness and discount cells built on old snapshots.
- **Corroborating against suppressed QCEW cells.** Produces a "zero" that is actually withheld. Check: enable "show suppressed" and treat suppression as unknown, not absence.
- **Counting the keyword instead of the attribute.** Keyword-based remote detection misses almost everyone because only 0.066% of US Software Engineer profiles surface on "remote." Check: infer remote from distance-to-office and role, not the word.

The two that matter most are the first and the seventh, because they compound. If you both under-detect remote workers (by searching the keyword) and default the ones you miss to HQ, your HQ engineering cell can be off by a wide margin while looking perfectly clean. The structural inference in step four is the only defence.

## Grade every cell: the confidence flag

Every cell gets a green, amber, or red flag driven by three inputs: sample coverage, vague-share, and remote exposure. A number without a flag is not a finding; it is a guess wearing a finding's clothes. Use this rubric and adjust the thresholds to your own risk tolerance, documenting whatever you choose.

**Cell confidence rubric**

```
GREEN  - Coverage ratio >= 0.6, vague-share < 15%, not the HQ metro OR remote correction applied and residual < 10%.
AMBER  - Coverage ratio 0.3 to 0.6, OR vague-share 15% to 30%, OR HQ metro with correction applied but residual 10% to 25%.
RED    - Coverage ratio < 0.3, OR vague-share > 30%, OR HQ metro with no remote correction, OR bounded only by a suppressed QCEW cell.
```

*Apply per cell. A cell drops to the lowest flag any single input triggers.*

A green cell is one you can put in front of a site-selection committee or use to set a pay band. An amber cell is directional - fine for prioritising a poaching push, not for a capital decision. A red cell is a placeholder that says "we do not yet know," which is a legitimate and useful thing for a map to say. Because 55% of organisations set geographic pay by city or metro area, a mislabelled cell here has a direct cost, so err toward the lower flag when an input is borderline.

#### Before you call the footprint done

- [ ] Every row resolves to a metro or sits in an explicit unknown bucket - none silently forced to HQ.
- [ ] The HQ cell has a stated remote correction and a residual confidence flag.
- [ ] Each metro cell reports its coverage ratio, and cells below the floor are flagged red.
- [ ] Every cell has a QCEW or LODES bound, and suppressed government cells are treated as unknown, not zero.
- [ ] Cross-border cells are expressed as within-country function shares before comparison.
- [ ] The unknown bucket's size is reported and has not been redistributed pro-rata.
- [ ] Profile freshness is recorded and stale-heavy cells are discounted.
- [ ] Every cell carries a green, amber, or red flag with the rule that set it.

## Keep the map current

A footprint map is a snapshot with a short half-life, so decide its refresh rule the day you ship it. Public profiles go materially stale within a year if never re-collected, because a meaningful share of professionals change jobs each year, while LinkedIn revises headcount only about twice yearly, typically applied retroactively. Guidance suggests members refresh their profiles at least quarterly, which implies location fields drift on a similar cadence.

No published percentage-over-window rule exists to trigger a re-map, so set your own threshold and document it. A defensible default: a light quarterly refresh that re-pulls the roster and recomputes shares, plus a full re-map triggered by any of three events - a published headcount change beyond a delta you set, a new office announcement in a metro you did not have, or a restructuring that moves a function. Anchor the delta to the coverage you can afford: if a single account extracts about 50 profiles a day, a full re-map of a large target is a multi-week project, so reserve it for genuine change rather than the calendar.

> **Tip:** Watch the movers, not the whole population
>
> Between full re-maps, monitor only the population most likely to shift the map: senior and function-lead roles and any metro flagged amber. A targeted re-pull of directors and above catches most structural change at a fraction of the collection cost.

The map earns its keep when it survives contact with a decision. If a site-selection committee, a compensation team, or a recruiting lead can point at a cell, see its flag, and know whether to trust it, you have built the thing the vendor pitches assume you already own. The difference is that you built it from the outside, from public signal, with the two silent failures named and corrected - which is the only version of this map you can actually defend.

## Frequently asked questions

### How do I tell if a company's engineers actually work at HQ or are remote?

Do not trust the self-reported location or a keyword search. In Refolk's index only 0.066% of US Software Engineer profiles surface on the word 'remote', so keyword detection misses almost everyone. Infer remote structurally: compare the HQ engineering cell against QCEW establishment employment for that county, and if the profile count exceeds plausible office capacity, remote inflation is present. Remote workers sit a median 185 miles from HQ yet often keep the HQ label.

### What does 'Greater Seattle Area' mean and how do I resolve it to a metro?

It is a user choice, not a data gap. LinkedIn derives a location from a postal code and offers the member both the city and the broader 'Greater X Area' label, and many pick the broader one to appear in more searches. Resolve it by postal-code lookup first, then cross-reference the employer's known office cities, and if neither resolves it, route the row to an explicit unknown bucket rather than collapsing it to the downtown core.

### Can I use QCEW to confirm how many people a specific company employs in a city?

No. QCEW covers more than 95 percent of U.S. jobs at county and MSA level, but it cannot isolate one company or one function, and much county data is suppressed to protect firm anonymity. Use it to bound a metro estimate, capping a cell at plausible employment, not to confirm one firm's share. Treat suppressed cells as unknown, never as zero.

### How often should I re-map a competitor's footprint?

No published percentage-over-window rule exists, so set the threshold yourself. Anchor it to two documented facts: public profiles go materially stale within a year if never re-collected, and LinkedIn revises headcount only about twice yearly, typically retroactively. Guidance suggests location fields drift on a roughly quarterly cadence, so a quarterly light refresh with a full re-map on a material headcount delta is defensible.

### Why do my UK metro cells look so thin compared with the US ones?

Because profile supply differs by country. In Refolk's index the US has about eight times the UK Software Engineer supply (325,147 versus 40,707). A rival's UK metro cells will look small in absolute terms even when locally dominant, so always normalise within-country before comparing across borders, or you will misread a strong local hub as a weak one.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/competitor-location-footprint-map*
