# The Market-Size Count for One Role, From Labor Data to Live Postings

*You can carry your own role and metros through labor data and live posting counts to one defensible number of realistic openings, plus a widen-or-hold call.*

- Canonical URL: https://www.refolk.ai/candidates/guides/market-size-count-one-role
- Pillar: Reading the market
- Format: Teardown
- Published: 2026-09-26
- Last reviewed: 2026-09-26
- Reading time: 17 min

You want one number: how many real, applicable openings exist for your role, and where. Not a national projection you cannot reach, not a job board headline you cannot trust, but a figure you can act on. This guide carries one case, a Data Analyst, all the way through both public labor data and live postings, showing the real queries, the intermediate counts, the wrong turns, and the point where the two numbers disagree. Follow along on your own role.

The short version: there is no single page that answers "how many jobs are there for my role" correctly, because the two authoritative directions measure different things and neither descends cleanly to your metro. You build the answer. What follows is the build.

## Why no ranking page can give you the number

The honest answer is that "openings" means two incompatible things, and job boards report a third thing that is neither. A ranking or dashboard hands you a top-down national projection; a board hands you an inflated advertisement count; nobody reconciles them for a single searcher in a single metro.

Start with the definitions, because the whole exercise depends on them. BLS Employment Projections defines occupational openings as the sum of net occupational employment change and occupational separations. Separations are the projected number of workers permanently leaving the occupation, meaning labor force exits and transfers to other occupations. There is a subtlety that trips everyone: workers who change jobs within the same occupation generate no openings, because there is no net change from that movement. So the BLS number is a modeled 10-year annual flow.

JOLTS is different in kind. It counts positions open on the last business day of the month, where a specific position exists, work is available, the job could start within 30 days, and there is active recruiting from outside the establishment. That is a monthly point-in-time stock, sampled from about 21,000 establishments.

A job-board listing is neither. It is an advertisement with no filled-or-unfilled verification attached to it.

> **Rule:** Three numbers, three meanings
>
> BLS Projections openings is a modeled 10-year annual flow. JOLTS is a monthly point-in-time stock. A board listing is an unverified advertisement. Never treat any two as the same quantity.

This matters before you touch a spreadsheet because the most common mistake is a category error, not an arithmetic one. If you compare a 10-year annual average to a live snapshot and call the result a reconciliation, you have produced a fake number. You can compare them, but only after you state, out loud, the assumption that converts one to the other. No source publishes that crosswalk, so it is your assumption to own.

## The two directions and where each one stops

There are exactly two directions to the answer, and each hits a wall at a different place. Top-down gives you openings but not your metro. Bottom-up gives you your metro but not verified openings. The build is about meeting in the middle honestly.

Here is where each public dataset actually stops.

| Dataset | Measures | Finest geography | Openings? |
|---|---|---|---|
| OEWS | Employment stock and wages | ~530 MSAs | No |
| Projections Central | Avg annual openings | State (some add MSA) | Yes |
| JOLTS | Openings, hires, separations | National plus limited state | Yes (stock) |

Read that table as a set of walls. OEWS is the only source that reliably reaches your metro, and it does not measure openings at all; it measures how many people already hold the job. Projections Central measures openings but is produced by states, so metro detail exists only where a state chooses to model it. Pennsylvania publishes long-term projections for 18 MSAs; Ohio for 8. Most states stop at the state line.

That is the real bottleneck. It is not the occupation, it is the geography. If your state does not model your metro, your metro openings number does not exist to be looked up. You have to build it from live postings. That single fact is why the bottom-up half of this guide is unavoidable rather than optional.

#### What each layer can and cannot tell you

1. **Job boards** - Reach your metro, but every count is an unverified, inflated advertisement stock
2. **OEWS** - Reaches ~530 metros, but measures employed people, not openings
3. **State Projections** - Measures openings, but only some states descend below the state line
4. **BLS national** - Measures openings as a modeled 10-year annual flow, no metro at all

*Openings exist only in the middle two layers, and only one layer reaches your metro.*

## The worked case: one Data Analyst, two metros

I will carry a Data Analyst targeting New York City and Chicago through the whole procedure. The first fork happens immediately, and it is a wrong turn worth showing.

The instinct is to type "Data Analyst" into a projections tool and read off a number. The problem: Data Analyst has no clean, dedicated SOC code. Forcing it into a single code understates openings, and the output looks precise, which is exactly what makes it dangerous. The fix is to bound a range across every SOC where the title plausibly appears rather than committing to one. That is the difference between a defensible figure and a false positive.

The second early temptation is worse. OEWS returns a large, satisfying metro employment number for analyst-type occupations, and it is tempting to read that as demand. It is not. OEWS is a stock of people already employed. Treating it as openings overstates opportunity massively. In this build, OEWS employment is context for sizing the local base and, later, for weighing competition. It is never the openings figure.

> **Watch out:** The two errors that ruin the count on line one
>
> Forcing a no-dedicated-SOC title into one code understates openings behind a precise-looking number. Reading OEWS employment as openings overstates demand by counting people who already have the job.

Now the bottom-up half, which is where the real inflation lives. Suppose the raw headline counts across two or three boards for "Data Analyst" in NYC sum to something in the thousands. That number is nearly meaningless on its own, and here is why in one figure.

#### How a raw board count collapses to reachable roles

| Stage | Figure | Note |
| --- | --- | --- |
| Raw board headline | ~15x | Programmatic syndication duplicates one role across many boards |
| After dedup | down to ~20% | Lightcast-style rule deduplicates up to 80% of collected jobs |
| After ghost strip | minus 18-27% | Remove stale and ghost listings |
| Eligible to you | smaller still | Remove roles failing your authorization, level, and screeners |

*Each stage removes a documented category of noise, and the drop from raw to reachable is more than an order of magnitude.*

The syndication factor is documented: one vendor attributes a 15x duplication increase to programmatic advertising. The dedup is documented too: Lightcast's two-step method removes up to 80 percent of collected jobs, using a 60-day rule where the first appearance is the original, duplicates are dropped for 60 days, and an ad still live after 60 days is counted as a new posting. Academic work separates full duplicates, same title and description, from semantic duplicates, the same position reworded, which is why company-plus-title dedup alone is not enough and you need description-level matching for reposts.

**15x - Duplication attributed to programmatic advertising**

This inflation sits on top of the ghost-job share, so the two compound before you filter for eligibility.

## Stripping ghosts without over-trimming

Ghost and stale postings are the next layer down, and the honest position is that estimates vary widely and none is official. You strip them, but you strip them with a check, because the obvious rule over-trims.

The estimates cluster in a wide band, from three sources using three different methods.

| Source | Estimate | Method |
|---|---|---|
| Greenhouse (2025) | 18-22% | ATS platform data |
| ResumeUp.AI (2025) | 27.4% | Greater-than-30-day listing age on LinkedIn |
| MyPerfectResume (2025) | ~30% | JOLTS openings-minus-hires gap |

The Congressional Research Service is blunt that there are no official statistics on the magnitude of ghost jobs; official opening statistics come from JOLTS. So treat any percentage as a bound, not a measurement. The JOLTS-based framing is instructive on its own: in one June, employers reported 7.4 million openings but made only 5.2 million hires, leaving 2.2 million unfilled. That is roughly a third of the official stock not converting to a hire, before any board noise.

**2.2M - Reported openings that did not become hires in one month**

7.4M openings against 5.2M hires; even the official stock overstates realized hiring by about 30%.

Now the trap. The most convenient ghost filter is a listing-age cut: drop anything older than 30 days. But SHRM reported an average time to fill of 41 days in 2024. A strict 30-day cut therefore kills genuinely slow-to-fill senior roles and biases your fillable count low. Use age as a first pass, then cross-reference the employer's actual hiring velocity and their own career page before you delete a listing. Age alone lies; age plus velocity is defensible.

> Age alone is not evidence a job is dead. A 41-day average fill time means the strict 30-day cut discards real roles.

## Reducing to what you can actually get

This is the load-bearing gap, so I will say it plainly: there is no authoritative, validated public procedure for reducing a deduplicated set to the postings one specific candidate is eligible for. Present this filter as reasoned method, never as a cited standard. Anyone who tells you otherwise is inventing authority.

What you can defend is the mechanism. From your fillable subset, remove roles that fail three tests specific to you:

- **Work authorization.** Roles requiring clearance or a status you do not hold.
- **Level.** Roles gated below or above your band by a hard requirement, not a soft preference.
- **Hard screeners.** Named degrees, licenses, or years-in-tool that function as knockouts.

What survives is your candidate-eligible count. This is the number that actually answers "how many jobs am I actually qualified for," and it is smaller than every earlier count for a reason.

Tailoring and re-scoring every posting to see whether you clear its screeners is where most searchers stall, because doing it by hand for a deduplicated metro list is hours of work. This is the specific friction [Refolk](/candidates) removes: it scores how well your history actually fits each posting, so the eligibility pass becomes a sort rather than a manual read of every listing.

> **Note:** This is your estimate, and that is fine
>
> Because no cited standard exists for eligibility filtering, your candidate-eligible count is a reasoned figure, not a published one. State that when you use it. A number labeled honestly beats a borrowed one dressed as authority.

## The procedure, end to end

Run these eight steps in order for each target metro. Times are rough; the whole pass for two metros is a focused afternoon.

#### From SOC code to one defensible number

1. **Pin the occupation to a SOC code** - Map your title to one 6-digit SOC, since OEWS and Projections both index by SOC. Note adjacent codes that also fit; expect a title like Data Analyst to split across several.
2. **Pull the top-down annual openings figure** - From Projections Central and your state site, record average annual openings for the SOC at the finest geography your state offers. Flag whether it is statewide or metro.
3. **Pull the employment-stock context from OEWS** - Look up metro employment for the SOC to size the local incumbent base. Treat it as context only, never as openings.
4. **Collect the raw live-listing count** - Search each target metro on two or three boards and log the headline count per board. Expect heavy inflation.
5. **Deduplicate the postings** - Collapse by company plus title, then description-level matching for reworded reposts, using a fixed 60-day window. End with one unique-roles number per metro.
6. **Strip stale and ghost listings** - Remove postings past the 30-day age threshold and low-signal employers, but cross-check hiring velocity so slow-to-fill real roles survive.
7. **Apply the eligibility filter** - Remove roles failing your work authorization, level, or hard screeners. Label this as reasoned method, not a cited standard.
8. **Reconcile and decide** - Compare the annual-openings figure to the eligible live snapshot, state your conversion assumption, document disagreement, and land on one figure plus a widen-or-hold decision.

A note on sequence. Some practitioners deduplicate before stripping stale listings; others strip stale first. The order changes your intermediate counts but not the final set, so pick one and keep it consistent across metros so your metro-to-metro comparison stays honest.

## Reconciling top-down and bottom-up

The two numbers will not match, and that is the expected result, not a failure. Your job is to explain the gap, not to erase it.

Here is the reconciliation for the worked case, stated as an assumption rather than a fact. If your state gives you a statewide annual openings figure for the Data Analyst SOC range, and your bottom-up eligible live count for NYC and Chicago is some point-in-time snapshot, you cannot divide one by the other and get truth. What you can say is this: the annual figure is a flow, so on any given day the live stock should be a fraction of it, shaped by how long roles stay open. With a 41-day average time to fill, a role occupies the live stock for roughly six weeks, which is a rough guide to how much of an annual flow is visible at once. Write that assumption down. It is defensible precisely because you named it.

When they disagree sharply, the usual causes are known:

- **The live count is far higher than the flow implies.** Suspect residual syndication or ghost inflation you did not fully strip. Re-check dedup and velocity.
- **The live count is far lower than the flow implies.** Suspect an over-aggressive 30-day cut, or a SOC that split across codes you did not search.
- **Your metro has no modeled openings at all.** Then there is no top-down anchor, and your eligible live snapshot is the whole answer. Say so.

**One-line reconciliation record per metro**

```
Metro: __________
SOC range searched: __________ (+ adjacent: __________)
Top-down annual openings (state/metro): __________  [statewide? y/n]
Raw board count: __________
After dedup: __________   After ghost strip: __________
Candidate-eligible live count: __________
Conversion assumption stated: __________
Gap explanation: __________
Decision: WIDEN / HOLD because __________
```

*Fill one of these per target metro so the final number is auditable months later.*

Once you have two or three metros filled in, you can run the supply search that turns "how many jobs" into "how much competition," which is the number that actually changes your decision.

Ask me this: `Data analysts currently working at companies in the New York City metro area.` - [run the search](https://www.refolk.ai/start?q=Data%20analysts%20currently%20working%20at%20companies%20in%20the%20New%20York%20City%20metro%20area.).

*Returns the live incumbent population for the role in that metro, which is the competition side of your widen-or-hold call.*

## Demand is only half the decision

The number of openings tells you nothing until you set it against supply. A metro with more openings but proportionally more incumbents can be a worse market than a smaller one, which is why the widen-or-hold decision needs both sides on the table.

In Refolk's index, the supply gap between adjacent data roles is large and countable.

| Title | US professionals | Ratio to Data Scientist |
|---|---|---|
| Data Analyst | 62,640 | 2.39x |
| Data Scientist | 26,207 | 1.00x |

**2.39x - More Data Analysts than Data Scientists in the US**

From Refolk's index of professional profiles; a supply figure, meaning competition, not opportunity.

Read that as a warning against counting openings alone. There are 62,640 Data Analyst professionals in the US against 26,207 Data Scientists. If the adjacent role has proportionally fewer incumbents chasing its openings, an eligible-live count that looks smaller can still be the better market. This is exactly the reframe that separates a real market read from a headline: openings on one side, incumbents on the other, and the ratio between them is the decision.

Use the matrix below to place each metro once you have both its eligible live count and its incumbent supply.

#### Widen or hold, per metro

Horizontal axis runs from Few eligible openings to Many eligible openings. Vertical axis runs from Crowded incumbent pool to Thin incumbent pool.

| Quadrant | What it means |
| --- | --- |
| Crowded and thin openings | Widen: add metros or an adjacent SOC, this one is starved |
| Crowded but many openings | Hold cautiously: volume exists but competition is heavy, tailor hard |
| Thin pool, few openings | Watch: low competition but little to apply to, keep as secondary |
| Thin pool, many openings | Hold and go deep: this is your best market, concentrate effort here |

*Plot each metro by eligible openings and by how crowded the incumbent pool is.*

## Where this goes wrong

The failure modes below are the most valuable part of this guide, because each one produces a number that looks right and is not. For each, I give the tell and the check.

- **SOC mismatch.** A title with no dedicated SOC gets forced into one code, understating openings behind a precise-looking figure. Check: list every SOC where the title plausibly appears and bound a range instead of committing to one.
- **Counting OEWS employment as openings.** OEWS is a stock of employed people, not demand; treating it as openings massively overstates opportunity. Check: openings come only from Projections or JOLTS, never OEWS.
- **Trusting headline board counts.** Programmatic syndication inflates by around 15x, so the raw number is nearly meaningless. Check: dedup by company plus title, then description-level, before you believe any board figure.
- **60-day rule artifacts.** A continuously reposted real role can be counted up to six times a year, while a filled role can linger in the stock. Check: verify borderline roles against the employer's own career page.
- **Ghost filter over-trims.** A greater-than-30-day age cut kills legitimately slow senior roles, given the 41-day average fill time. Check: cross-reference hiring velocity, not age alone.
- **Eligibility filter presented as fact.** No cited standard exists for it, so dressing it as authoritative is dishonest. Check: label it plainly as reasoned method.
- **Annual-to-live false equivalence.** Comparing a 10-year annual average to a live snapshot without stating the conversion produces a fake reconciliation. Check: write the conversion assumption down every time.
- **State-metro coverage gap.** Assuming your metro has modeled openings when your state only publishes statewide. Check: confirm on the state projections site before you cite a metro figure.

> **Tip:** The fastest sanity check
>
> If your final eligible-live count is within an order of magnitude of what the annual flow implies at a 41-day fill time, trust it. If it is off by more, one of the eight failure modes is in play; work the list above before you act on the number.

## Keep it current and call the job done

The whole build has a shelf life, because live postings move weekly and projections and ghost-share studies update on their own cycles. Re-run the bottom-up half monthly if you are actively searching; refresh the top-down anchor when your state publishes new long-term projections. The current long-term round covers 2024 to 2034, so note the vintage of whatever you pull.

Before you treat the number as settled, run this check.

#### Before you act on the number

- [ ] One SOC selected, with adjacent codes listed and a bounded range, not a single forced code.
- [ ] Top-down annual openings recorded, flagged as statewide or metro per your state's coverage.
- [ ] OEWS metro employment logged as context only, never entered as an openings figure.
- [ ] Raw board counts deduplicated by company plus title, then by description for reposts.
- [ ] Stale and ghost listings removed with a velocity check, not an age cut alone.
- [ ] Eligibility filter applied and labeled as reasoned method, not a cited standard.
- [ ] Conversion assumption between annual flow and live stock written down explicitly.
- [ ] Incumbent supply pulled for each metro so the decision weighs competition, not just openings.
- [ ] One number per metro and a written widen-or-hold decision with its reason.

The output you keep is not a single national statistic. It is a small table of metros, each with a candidate-eligible live count, an incumbent supply figure, and a one-line decision. That is the thing no ranking page or single board can give you, because it is built for you, at your level, in your metros, and it says out loud where its own numbers are assumptions rather than facts.

## Frequently asked questions

### How many jobs are there for my role, really?

There is no single lookup that answers this honestly. You build two numbers: a top-down annual openings figure from BLS Projections for your SOC code, and a bottom-up count of deduplicated, fillable, eligible live postings in your metros. They measure different things, a modeled annual flow versus a live stock, so you reconcile them by stating the conversion assumption rather than expecting them to match. The defensible answer is the eligible live count, with the annual figure as a sanity bound.

### Why is the number of listings on a job board so much higher than the real openings?

Two inflation sources compound. Programmatic advertising duplicates the same role across boards by a factor cited around 15x, and separately 18 to 27 percent of listings are ghost or stale jobs that will not result in a hire. On top of that, even official openings overstate hiring, with 7.4 million reported openings against 5.2 million hires in one June. A raw headline count can therefore exceed reachable roles by more than tenfold.

### Can I get openings data for my specific city?

Sometimes. OEWS publishes employment for about 530 metros but no openings. Openings come from state projections, and only some states model below the state level, for example Pennsylvania at 18 MSAs and Ohio at 8. If your state only publishes statewide, you cannot look up your metro's openings directly and must build the metro figure from deduplicated live postings instead.

### Is the 30-day age rule a reliable way to remove ghost jobs?

Only partly. A greater-than-30-day age threshold is a common heuristic, but SHRM reported an average time to fill of 41 days in 2024, so a strict age cut discards legitimately open senior roles that are just slow to fill. Use age as a first pass, then cross-reference employer hiring velocity and the company's own career page before dropping a listing.

### How many jobs am I actually qualified for, not just how many exist?

That is the eligibility filter, and no authoritative source publishes a validated procedure for it. Treat it as reasoned method: from your fillable subset, remove roles that fail your work authorization, your level, and any hard screeners. What remains is your candidate-eligible count. Label it honestly as your own estimate rather than a cited standard.

### Should I widen my search or hold to my current role and metros?

Decide with both demand and supply in view. A metro with more openings but proportionally more incumbents is not necessarily better; Refolk's index shows 2.39x as many Data Analysts as Data Scientists in the US, which is competition, not opportunity. If your eligible live count across current metros is thin relative to incumbents, widen by metro or adjacent SOC. If it is healthy, hold and go deeper.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/candidates/guides/market-size-count-one-role*
