Refolk
ReferenceMarket and talent intelligence

The Talent-Market Signal Reference: What Each Data Point Proves

For any public talent-market signal you can name, you will state what it proves, what it does not, and the confound to rule out before you trust it.

16 min readLast reviewed August 3, 2026Read as Markdown

Key takeaways

  • Job postings measure recruitment marketing, not vacancies, and Lightcast removes up to 80% of collected postings as duplicates before any count is defensible.
  • A short median posting duration proves fast removal, not easy hiring; Lightcast bans the metric from time series to avoid a biased reading.
  • In Refolk's index the US ML Engineer pool is roughly 8.5x Germany's (10,530 vs 1,239), so any talent-shortage claim must be geo-scoped or it inverts across borders.
  • In Refolk's index US Go supply is 4.69x US Rust supply (16,014 vs 3,416), so a hard-to-hire read tracks the skill token as much as the underlying market.
  • A postings read carries a 36-hour lag, JOLTS about five weeks, and tenure up to two years, so placing them side by side implies a coherence that does not exist.
  • A profile update is not a confirmed job change; models predicting career moves from profiles top out near 66% accuracy, so corroborate with an employer and date change.

You have a labor-market number in front of you: a posting spike, a supply-demand ratio, a median tenure that dropped. Someone wants it in a leadership deck by end of day, and the deck will imply a conclusion the number may not support. This reference is for strategy, research, and talent-intelligence analysts who need to state, one signal at a time, what the data point proves, what it does not, and the specific confound that most often breaks the reading.

Existing market-research guides tell you how to build a map or size a pool. None define the raw signals themselves, row by row, with the confound attached. This one does. Jump to the signal you are holding, read what it proves and how it lies, then move on.

Why raw talent signals mislead before you interpret them

Most talent signals fail at the source, not at the analysis. The crawl, the baseline, and the refresh cadence each inject error before you form an opinion, so a clean-looking number can already be wrong when it reaches your slide.

Three structural facts sit underneath almost every misread. First, postings are not vacancies. Job postings measure recruitment marketing by employers, and while they correlate with vacancies, many recruitment practices make the relationship imperfect. Second, the crawl is noisy: Lightcast collects from over 220,000 current and historical sources, and its two-step dedup removes up to 80% of collected jobs as duplicates. Third, a long tail of postings never intended to close pollutes both counts and durations. Analysing US NLx data, researchers found 25% of postings are up more than 90 days and 10% more than 180 days, with some employers running evergreen postings to fill multiple vacancies and 20 postings up for the entire 16 years of data.

80%
Share of collected job postings Lightcast removes as duplicates
Headline crawl counts overstate live demand before any interpretation begins.

The practical consequence: a headline crawl count overstates live demand before you touch it. This is why a CSIRO study found a signal-averaging algorithm predicts changes in vacancies significantly more accurately than raw counts of job postings. The raw number is the input to a correction, not the answer.

The signal reference: what each data point proves and how it lies

Each row below names a signal, states the one thing it proves, and names the confound to rule out first. Read the row for the signal you are holding and stop there.

Posting count (absolute). Proves that employers published a given number of listings across the crawled sources in a window. It does not prove the number of open seats. It lies when a repost wave or a newly indexed job board inflates the count without any change in real demand. Rule out: deduplication handling and source-set changes.

Posting index (relative). Proves the percentage change in postings against a fixed baseline. Indeed's index sets February 1 2020 as the pre-pandemic baseline, so a reading of 101 means postings are 1% higher than that day. It lies when a reader treats "up 10%" as month-over-month. Rule out: what the baseline actually is before you quote the number.

Median posting duration. Proves how many days a posting stayed live and accepting applicants in a region, occupation, or company. Lightcast's own example is that Physician Assistant postings tend to stay open for 36 days, and it frames duration as a difficulty signal. It lies in two directions: long duration can be an evergreen pipeline, not scarcity, and short duration can be a cancellation or repost, not easy hiring. Lightcast withholds the metric from time series to avoid biased results. Rule out: reposting and evergreen behavior for the specific employer.

Supply-demand ratio. Proves relative scarcity for a defined role, geography, and window. It lies when the boundary is unstated, because scarcity inverts across borders. In Refolk's index the US ML Engineer pool is roughly 8.5x Germany's. Rule out: whether the geography and skill token are pinned.

Composite hiring-difficulty index. Proves a vendor's blended difficulty proxy. No single public standard exists; each vendor composes its own from supply, duration, and demand components, and the exact weighting is not established publicly for any major index. It lies when it is read as a measurement rather than a proxy. Rule out: which components went in and over what window.

Median employee tenure. Proves the midpoint of how long workers have stayed. It was 3.9 years in January 2024, down from 4.1 in January 2022. It lies when read as attrition, because a younger age mix or a hiring surge lowers the median without more quitting. Rule out: age composition and recent hiring rate.

Profile update or skill add. Proves a member edited their profile. It does not prove a job change. Models predicting career trajectories from profiles top out near 66.1% accuracy, and skill-add rates are up 140% since 2022, meaning freshness partly reflects behavior, not moves. Rule out: a corroborating employer and date change.

SignalProvesConfound to rule out first
Posting countListings publishedDedup and new sources
Posting indexChange vs baselineWhat the baseline is
Median durationDays a role stayed liveReposts, evergreen tail
Supply-demand ratioRelative scarcity in scopeUnstated geography or skill
Median tenureMidpoint of stayAge mix, hiring surge
Profile updateAn edit was madeCorroborating move evidence
The raw number is the input to a correction, not the answer you put on the slide. </pull>

Supply scarcity is a market fact, not a global one

A talent-shortage claim only means something inside a boundary. Change the country or the skill token and the same signal can flip from scarce to abundant, so scope is not a footnote, it is part of the claim.

Two figures from Refolk's index make this concrete. Across borders, the US ML Engineer pool stands at 10,530 current professionals against Germany's 1,239, roughly an 8.5x gap. Within one country, the skill token alone moves supply several-fold: US Go supply is 4.69x US Rust supply.

MarketMatching professionalsTop employer in sampleShare of US pool
United States10,530Meta100%
Germany1,239BASF Management Consulting11.8%

The Germany figure is 11.8% of the US pool, derived as 1,239 divided by 10,530. Read this as relative supply, not demand: it tells you where candidates live, not how many roles are open for them.

SkillMatching professionalsMultiple of Rust pool
Rust3,4161.0x
Go16,0144.69x

The Go multiple is derived as 16,014 divided by 3,416. A "hard-to-hire" read tracks the skill token as much as the underlying market, so a Rust role looks scarcer than a Go role largely because fewer people list Rust. When you count supply yourself rather than trusting a headline scarcity score, Refolk lets you set the exact geography, skill, and seniority the ratio is computed over, which is the boundary a leadership deck will otherwise leave implicit.

8.5x
US ML Engineer pool relative to Germany's in Refolk's index
10,530 vs 1,239 current professionals - a shortage claim inverts across this border.

Cadence and lag: why side-by-side signals lie about the same month

Signals refresh at wildly different rates, so a chart that lines up postings, openings, and tenure implies a coherence that does not exist. Each describes a different month, or a different year, and reading them as one snapshot is the silent error in most decks.

Postings are fastest. Lightcast US postings run back to January 2010 and are updated daily, with a 36-hour window from posting to publishing, though the vendor revisits sites for up to 14 days to catch late listings. Indeed's index is daily but refreshed weekly, expressed as the percentage change since February 1 2020 using a seven-day trailing average. Government supply and demand data lags far more: JOLTS is released monthly, about five weeks after the reference month ends, resting on a sample of 16,000 business establishments. Tenure is slowest of all, collected biennially each January as a supplement to the Current Population Survey.

Signal classRefreshLag to as-of
Online postings (Lightcast)Daily~36 hours
Postings index (Indeed)Weekly7-day trailing avg
JOLTS openingsMonthly~5 weeks
Employee tenureBiennialUp to ~2 years

Dating a reading to its true as-of

  1. Event happens
    A role opens, closes, or a worker moves
  2. Signal captured
    Crawl, survey, or supplement records it after its own lag
  3. Signal published
    36 hours to 5 weeks to 2 years later, depending on class
  4. You read it
    Date it by the event window, not the pull date
A signal's usefulness depends on the lag between the event and when you can see it.

The rule that follows: never place a 36-hour postings read next to a biennial tenure figure and call it a picture of "now." Date each signal by the event window it covers, and if you must show them together, label each with its own as-of.

The seven-step reading procedure

Run every signal through the same seven steps before it reaches a slide. The point is not speed; it is that a reviewer cannot extend your claim past the evidence once the last step is written.

Read any talent-market signal in seven steps

  1. Classify the signal and its source
    Confirm whether the number is a posting count, a duration, a ratio, a composite index, or survey supply data, and name the publisher. Done when source and definition are recorded.
  2. Establish cadence and vintage
    Note the refresh interval and lag: postings ~36 hours to weekly, JOLTS ~5 weeks, tenure biennial. Done when the reading is dated with its true as-of.
  3. Check the denominator or baseline
    Confirm whether the figure is absolute, indexed (Indeed = 100 at Feb 1 2020), or relative (LinkedIn needs 4+ comparators). Done when you can state what up 10% is measured against.
  4. Rule out the signal's primary confound
    For postings, quantify dedup and evergreen inflation. For duration, rule out reposts and cancellations. For tenure, rule out age-mix shifts. Done when the confound is quantified or bounded.
  5. Cross-validate against a second class
    Compare a demand signal against a supply source such as profiles or JOLTS. Done when two independent classes agree in direction.
  6. Test sample sufficiency
    Confirm the cell exceeds publication and comparator floors before slicing by geo or seniority. Done when no suppressed or thin cell drives the claim.
  7. State proof and non-proof
    Write one line: what the data point proves, what it does not, and the confound ruled out. Done when a reviewer cannot extend the claim beyond the evidence.

On step five, practitioners disagree on ordering. Some cross-check supply first, others demand first. Either works; what matters is that two independent classes agree in direction before you commit. On step six, the floors are real: BLS suppresses cells where the base is under 75,000, marking them with a dash, and LinkedIn's hiring-demand comparison only compares within-set once you have four or more regions or industries. Below that, a "high demand" tile is measured against the entire index, not your comparators.

How each reading goes wrong: the eight failure modes

Every talent signal has a characteristic false positive, and the same eight recur across every deck. Each one is a confident number that is unstable for a reason you can check.

Posting spike read as demand growth. A repost wave or a newly indexed source inflates counts without new seats. Check dedup handling and whether sources were added, since up to 80% of raw postings are duplicates.

Short duration read as easy hiring. Roles cancelled, filled internally, or reposted as fresh listings look fast. Check reposting behavior, and note that Lightcast bans duration from time series precisely to avoid this bias.

Long duration read as scarcity. Evergreen and ghost postings, 25% of them up more than 90 days, never intended to close quickly. Check whether the employer runs continuous pipelines.

Indexed number read as absolute. "Postings up 10%" is versus February 1 2020, not last month. Check the baseline before quoting.

Relative demand read as universal. A LinkedIn "high demand" tile is only relative to the comparators in view and flips below four regions or industries. Check the comparison set.

Tenure drop read as rising attrition. A younger age mix or a hiring surge lowers the median without more quitting. Note the split: public-sector median tenure is 6.2 years, nearly twice the private-sector 3.5 years, so a change in the public-private mix alone moves the headline. Check age composition and hires.

Thin-slice reads. A city-by-seniority cell below the sample floor produces a confident but unstable number. Check cell size against the 75,000 base floor.

Profile update read as confirmed job change. A title edit or skill add is not a move, and models here top out near 66% accuracy. Check for a corroborating employer and date change.

Triaging a signal before it goes in the deck

High decision stakesLow decision stakes
Note in passing
Mention with a one-line caveat and move on
Hold until resolved
Do not present - the confound could reverse the call
Quote freely
Use as supporting color, low risk
Lead with it
Put it on the headline slide with the confound stated
Confound unresolvedConfound ruled out
Where a signal sits decides whether you quote it, caveat it, or hold it.

Cross-validation and sample floors: proving direction, not just level

A single signal proves direction only when a second, independent class agrees. Level claims need enough sample to be stable, and both checks are cheap insurance against a number that reverses under scrutiny.

Cross-validation means pairing a demand signal with a supply signal. If postings for a role rise while the candidate pool in that geography is flat and JOLTS openings for the sector climb, the demand read holds. If the pool is growing as fast as postings, the "shortage" is weaker than the count implies. This is where profile-derived supply proxies earn their place: pool size, mobility (members who changed companies in the past 12 months), skill additions, and tenure. Treat them as behavioral, not static: skills data is actively maintained, with the rate of new skill additions up 140% since 2022, so "freshness" reflects members responding to pressure as much as real movement.

Sample sufficiency is the second check. Before you slice by city and seniority, confirm the cell clears the floor. BLS does not publish values where the base is under 75,000, and JOLTS itself rests on 16,000 establishments, so a narrow slice inherits wide uncertainty. A confident city-level seniority number is often a thin cell in disguise.

The proof-and-non-proof line for a deck footnote
This [signal] from [publisher, as-of date] proves [the narrow claim].
It does not prove [the over-reach a reader will assume].
Primary confound ruled out: [confound] via [check, e.g. dedup applied / cell size 90,000 / two classes agree].
Boundary: [geography, skill token, seniority, window].

Fill in one line per signal you present. If you cannot complete all four clauses, the signal is not ready.

Refolk's index is where the supply half of that cross-check is fastest to run, because you can pin the geography, skill, and mobility filter and get named professionals rather than an aggregate tile. That turns "is the pool growing" from a guess into a count you can defend when the deck is challenged.

Keeping the reference current

This reference stays useful only if you re-check the mechanisms, not the values. The numbers in it will drift; the confounds will not, and the way to re-verify each one is stable.

Re-check three things on a schedule. First, baselines and methods change: Indeed seasonally adjusts using a method the Deutsche Bundesbank developed for daily series, and the FRED text separately cites a January 2021 methodology switch, so confirm which baseline and method your source uses before quoting a change. Second, cadence can shift when a vendor changes its collection window, so re-confirm the lag each time you build a recurring report. Third, sample floors and comparator rules are the guardrails that keep thin slices out of decks, so verify them against the current publication notes rather than memory.

Before this signal goes in the deck

  • The signal is classified and its publisher and definition are written down
  • The reading is dated by its event window, not the pull date
  • The baseline is confirmed - absolute, indexed, or relative with its comparators
  • The primary confound is quantified or bounded, not just named
  • A second, independent signal class agrees in direction
  • Every sliced cell clears the sample floor (BLS suppresses under 75,000)
  • A one-line proof-and-non-proof statement is written and a reviewer cannot extend it

The discipline is simple to state and hard to skip under deadline: name the signal, date it, find its baseline, rule out its confound, corroborate its direction, check the cell size, and write the one line that fences the claim. A number that survives all seven is defensible in front of a leadership team. A number that skips even one is a sales pitch waiting to be exposed.

Questions practitioners ask

What does a spike in job postings mean?

A posting spike means employers increased recruitment marketing, which correlates with demand but is not the same as vacancies. Before you read it as demand growth, check deduplication and whether new job-board sources were added to the crawl, since Lightcast removes up to 80% of collected postings as duplicates. A repost wave or a newly indexed source inflates counts without any real change in open seats.

What does the talent supply-demand ratio actually mean?

It compares available candidates against employer demand for a specific role, geography, and time window, so it is a market fact rather than a global one. In Refolk's index the US ML Engineer pool is about 8.5x Germany's, so the same role can read as abundant or scarce depending on where you draw the boundary. Always state the geography, skill token, and seniority the ratio was computed over.

How is a hiring-difficulty index calculated?

There is no single public standard; each vendor composes its own. Lightcast exposes unique postings, unique companies, and median posting duration as separate demand and scarcity components, with duration framed as the difficulty signal. LinkedIn computes hiring demand relatively, not absolutely. The exact weighting of supply, duration, demand change, and concentration is not established publicly for any major index, so treat any single difficulty score as a proxy, not a measurement.

What does median posting duration prove?

It shows how many days a posting stayed live and accepting applicants, which proxies fill difficulty. A long duration can mean scarcity, but it can also mean an evergreen posting never meant to close, and 25% of NLx postings stay up more than 90 days. A short duration proves fast removal, not easy hiring, since roles get cancelled, filled internally, or reposted as fresh listings.

Does a drop in median tenure mean attrition is rising?

Not on its own. Median tenure fell to 3.9 years in January 2024 from 4.1 in January 2022, but a younger age mix or a hiring surge lowers the median without anyone quitting more. Before you call it rising attrition, check the age composition and the recent hiring rate, because fresh hires reset tenure downward mechanically.

Why can't I chart posting duration over time?

Lightcast withholds median posting duration from its time series to avoid biased results, because evergreen reposts reset the clock and make a trend line an artifact. If you chart duration over time you are producing exactly the reading the source itself refuses to publish. Use duration as a point-in-time difficulty proxy for a role or company, not as a trend.

Read next