Refolk
FrameworkMarket and talent intelligence

The Hiring-Difficulty Score: Rating a Role Against a Market

You will score any role-and-market pair 1 to 5 for hiring difficulty from public supply and demand data, then turn that score into one sourcing decision.

17 min readLast reviewed August 25, 2026Read as Markdown

Before you commit headcount to a plan, you need to know how hard each role will actually be to fill in each market you are considering. This guide is for strategy, talent-intelligence, and research teams who have to make that call from public data. It gives you a vendor-neutral rubric that weighs active demand against available supply, produces a 1-to-5 hiring-difficulty score, and maps that score to one sourcing decision: post and wait, source wide, widen the geography, or build internally.

Most existing research either sizes a talent pool without demand, or reads a competitor's roadmap without supply. Neither tells you whether a specific role is fillable. Scoring difficulty means putting both sides of the ledger on the same scale and computing a ratio you can defend, with thresholds you can point to.

What the hiring-difficulty score measures

The hiring-difficulty score is a 1-to-5 rating of how hard one role is to fill in one market, computed as active demand weighed against available supply and confirmed by compensation and time-to-fill. It is a supply-and-demand ratio, not a pool size, and it always resolves to an action.

A pool size answers "how many people exist." The difficulty score answers "how many people exist relative to how many employers want them right now, and how fast the market is moving." Those are different questions with different answers. A market with 5,000 qualified people and 6,000 open roles is harder than a market with 500 people and 200 open roles, even though the first pool is ten times larger.

Two published instruments already blend these inputs, which tells you the shape is sound. Gartner's TalentNeuron scores hiring difficulty 1 to 10 using posting duration, change in demand, competitive concentration measured by HHI, and relative supply. Lever's Recruiter Pressure Index combines job openings, quits, and unemployment into a single 0-to-100 measure. This guide keeps the same inputs but publishes the thresholds so you can compute the score yourself from open job-posting and profile data, without a platform license.

10.5x
How much the same role's supply varies by market
Refolk's index holds 5,348 ML-engineer profiles in the US against 510 in Germany, the same role, two countries.

That 10.5x gap is the whole reason to score per market rather than globally. A role that is a 2 in one country can be a 5 next door.

The four inputs and what each one proves

Four public signals drive the score. Two set the ratio, two confirm it. Each proves something specific, and each lies in a specific way, so learn both.

  • Active demand (unique job postings). Proves how many employers are hiring for the role right now. It lies when reposts inflate the count, so one aggressive employer looks like a hot market.
  • Available supply (profile pool). Proves how many qualified people exist in the market. It lies when the pool is inert, full of people not open to moving, so a big number hides a slow fill.
  • Compensation premium (year over year). Proves whether employers are paying up to win the role, which confirms scarcity the pool count can miss. It lies when a global average is read as a local one.
  • Time-to-fill by seniority. Proves how long the market actually takes to close the role. It lies when a median across all levels hides that senior roles run far longer.

The ratio comes from demand and supply. Compensation and time-to-fill are the corroborators: they either confirm the ratio's reading or push the score one band. If the ratio says balanced but the role takes 90 days and pay is climbing, the ratio is missing something, usually an inert pool.

The four layers of a difficulty score

  1. Decision
    Post, source wide, widen geography, or build
  2. Corroborators
    YoY compensation premium and time-to-fill by seniority
  3. Ratio
    Unique demand postings weighed against available supply
  4. Definition
    Title synonyms, seniority band, and geography, fixed once
The ratio sets the score; compensation and time-to-fill confirm or nudge it.

Building a demand count you can defend

A defensible demand count is unique postings from the last 30 to 60 days, deduplicated on job ID, description hash, and location. Raw feeds are heavily inflated, so the total posting count is close to useless on its own.

The numbers are stark. Raw job-posting feeds often contain 30-45% duplicates, and a fixed deduplication algorithm removes about 80% of all postings collected for redundancy and stable data. The mechanism is the re-post rule: a posting's re-appearances within 60 days of the original are counted as duplicates. So if a job first posts on March 1st, any copy found elsewhere in the next 60 days is a duplicate; if the same ad reposts daily for a year, it counts six times, not 365. Deduplication itself uses a statistical classifier trained on location, job title, company name, and text similarity, with text similarity detected using shingling.

Two operational details follow. First, use a 30-to-60-day window, anchored by the same 60-day rule: a posting re-found after 60 days is legitimately fresh demand, and anything shorter risks missing a slow-filling senior role. Second, benchmark change against fixed lookbacks. Vendors compare demand to the number of similar postings one month, six months, and twelve months ago, and postings data is reclassified and deduplicated every four weeks. A published, role-level decay rate for posting demand is not established publicly, so use the 60-day window plus month-over-month change as your practical proxy rather than inventing a decay curve.

The thresholds: where balanced ends and tight begins

Read the ratio against a published anchor, not against your intuition. The cleanest public threshold is the BLS unemployed-per-opening ratio: values less than one suggest a tight labour market with more openings than unemployed workers, while ratios larger than one indicate slack. Balanced sits near 1.0 to 1.2, not at 1:1.

SignalValueReading
Unemployed per opening, US Jun 20261.0balanced
Unemployed per opening, 2019 average1.2normal to tight
Unemployed per opening, Jul 20096.5very slack
AI demand-to-supply, global3.2:1extremely tight
AI security / ZK specialty8:1acute

Two things to internalise from this table. The 2019 economy, widely called hot, sat at 1.2 unemployed per opening, and June 2026 sat at 1.0. If you score every market at or near 1.0 as scarce, you will flag normal conditions as emergencies. Treat 1.0 to 1.2 as balanced.

The tight end needs a real ceiling, and the AI benchmark supplies one: AI talent demand exceeds supply by a ratio of 3.2 to 1, with over 1.6 million open positions and only 518,000 qualified candidates globally, and in specialties like AI security or ZK cryptography it can exceed 8:1. That is your anchor for a 5.

Here is the five-band scale. Note that no single published five-band mapping exists; the levers below are documented, but the band thresholds are my synthesis, so calibrate them to your own market data before betting a plan on them.

ScoreSupply-to-demand readingWhat it means
1Slack, well above 1.2 candidates per openingMore qualified people than openings
2Comfortable, roughly 1.0 to 1.2Balanced, healthy pool
3Tightening, roughly 0.5 to 1.0Fewer candidates than openings
4Tight, roughly 0.3 to 0.5Approaching the AI-wide 3.2:1 anchor
5Acute, below 0.3At or past 3.2:1, specialty scarcity

Counting supply without manufacturing scarcity

A defensible supply count uses a broad title plus multiple skill synonyms, counted per market from an open profile index. The single most common way this step fails is filtering on one exact skill tag, which measures tagging density, not real talent.

The evidence from Refolk's index is blunt. The US ML-engineer pool holds 5,348 profiles under a broad title match. Narrow that to one exact skill string and it collapses to 6 profiles, a 99.9% drop driven entirely by tagging sparsity. Those people did not vanish; the filter did. Always match on title plus several skill synonyms, and treat any near-zero pool from a single string as a false negative until you widen it.

Geography is the other half. The same role differs by an order of magnitude across borders, so supply is meaningless until you pin it to a market.

MarketML-engineer profilesTop in-pool employer
United States5,348Meta
Germany510Bosch
US-to-Germany multiple10.5x-

The employer column carries a warning too. Top US employers in the pool are Meta, Apple, and Intel; top German employers are Bosch, Meta, and Accenture. If one company dominates both the postings and the profiles, your market is really one buyer and one bench, and both counts are less liquid than they look.

This is the step where a plain-language index earns its place. Rather than exporting boards and stitching deduplicated postings against a scraped profile set, you can ask directly for the pool in the exact market and shape you defined in step one.

With Refolk, the supply count and the demand-adjacent employer read come from the same query, which keeps the numerator and denominator on comparable definitions instead of two different pipelines.

Confirming the score with pay and time-to-fill

Once you have a ratio-based score, two signals confirm it or move it one band: the year-over-year compensation premium and time-to-fill by seniority. Compensation often tightens before the pool visibly shrinks, so a sharp premium jump is an early warning.

The named source for pay is PwC's Global AI Jobs Barometer, measured year over year from job ads comparing AI-skill roles to similar non-AI roles, based on analysis of close to a billion job ads and thousands of company financial reports from six continents. The trajectory across editions is the signal: the average premium hit 56%, up from 25% the year before, then reached 62%. A sustained double-digit year-over-year jump is a tight-market tell. The trap is that these are global averages; a specific metro may not move at all, so confirm year-over-year movement in your target market before you let pay push the score up.

Time-to-fill confirms from the other direction. A long clock on a role you scored as balanced means the pool is inert.

LevelTime-to-fill (days)
Junior15-25
Mid30-45
Senior / Staff60-90+
Tech median, all levels~48

Nearly 40% of senior-level roles take more than 90 days to fill, and the SHRM benchmark averages 44 days across roles with the technology median around 48. So a senior role at 90 days is normal, not alarming, while a junior role stuck at 60 days is a genuine scarcity signal. Read time-to-fill against the seniority band, never against the all-levels median.

Compensation confirms the scarcity a profile count misses, and it moves before the pool visibly shrinks.

The procedure

Run these seven steps in order. Steps two and three can swap; BLS-style analysis leads with demand while pool-sizing guides lead with supply, and either works as long as you compute the ratio the same way for every market.

Score a role-and-market pair

  1. Define the role-and-market pair precisely
    Fix the job-title synonyms, seniority band, and geography. Done means an explicit title-list and location filter you can run identically against postings and profiles.
  2. Pull the demand count
    Count open postings from the last 30 to 60 days, then deduplicate on job ID, description hash, and location. Done means a unique-posting count reported with its total-vs-unique ratio.
  3. Pull the supply count
    Count matching profiles in the same market using broad title plus skill synonyms, never one exact tag. Done means a candidate-pool figure comparable to the demand count.
  4. Compute the ratio and place it on the scale
    Divide supply by demand and read against the BLS tightness rule and the 3.2:1 AI anchor. Done means a provisional 1-to-5 score with numerator and denominator stated once.
  5. Add corroborating signals
    Layer the YoY compensation premium and time-to-fill by seniority onto the ratio. Done means the score confirmed or nudged one band, with the reason recorded.
  6. Map the score to a sourcing decision
    Translate the band into post-and-wait, source wide, widen geography, or build, with the hiring lead present. Done means one named action with an owner.
  7. Set the re-score date
    Schedule quarterly for AI-exposed roles and semi-annual otherwise. Done means a calendar trigger tied to the pair.

From raw postings to a defensible ratio

  1. Raw postings collected
    100

    before any cleaning

  2. After removing 30-45% duplicates
    60

    same-ad reposts inside 60 days

  3. Unique demand count
    20

    vendor pipelines strip up to 80% of raw

Deduplication is the biggest single lever, stripping most of the raw feed before you compute anything.

From score to sourcing decision

Every score must resolve to one named lever with an owner. A 1-to-5 number with no decision attached is a listicle, not a plan. The two documented levers are the build/buy/borrow framework, where build is best when skills are scarce but stable and you can invest through internships, apprenticeships, and internal academies, and geographic widening, where remote talent removes the geographic bottleneck entirely and lets you widen the funnel and speed the hiring loop.

ScoreReadingSourcing decision
1SlackPost and wait; the pool comes to you
2BalancedPost and wait, with light sourcing on close
3TighteningSource wide; run active outreach across the market
4TightWiden geography; source the role where supply exists
5AcuteBuild internally, or widen geography first if the skill lives elsewhere

The order between 4 and 5 matters. Widening geography is faster and cheaper than building, so try it first: if the 10.5x gap between two markets is real, the role that is a 5 at home may be a 3 abroad, and remote hiring closes the gap without a multi-quarter build. Reserve build for when even the widened market stays tight and the skill is stable enough to justify training on it. Note that 77% of employers plan to upskill workers while 41% plan to reduce workforce as AI automates tasks, so building is a live strategy, but it only pays off when the underlying skill will still matter by the time your first cohort is ready.

Choosing between widen and build

Supply elsewhereSupply local
Volatile, local
Source wide and pay the premium; do not build a fast-changing skill
Volatile, elsewhere
Widen geography; borrow rather than build a moving target
Stable, local
Build internally through academies and apprenticeships
Stable, elsewhere
Widen geography first, then build if the widened market stays tight
Skill volatileSkill stable
Cross scarcity with skill stability to pick the top-band lever.

How this scoring goes wrong

The score fails in predictable ways, and each failure has a public tell you can check. This is the part to read twice, because a confidently wrong score sends a whole hiring plan in the wrong direction.

  • Inflated demand from reposts. One aggressive employer reposting daily reads as a hot market. The tell is a high posting count with one dominant company. Check the total-vs-unique ratio and employer concentration; if one firm owns most postings, discount the demand.
  • Stale supply pool. A big pool that is inert scores as easy but fills slowly. The tell is a low ratio paired with a long time-to-fill. Cross-read posting duration and time-to-fill by seniority before trusting a large pool.
  • Skill-string undercounting supply. Filtering by one exact skill tag can collapse the pool to near zero. The tell is a supply count in the single digits for a role you know is common. Use broad title plus multiple synonyms; Refolk's index went from 5,348 to 6 on a single tag.
  • Wrong window. A 12-month posting count captures filled and seasonal roles as live demand. The tell is a demand count that dwarfs your unique count. Restrict to 30 to 60 days and confirm against month-over-month change.
  • Compensation premium misread. A 62% AI premium is a global average that a specific metro may not share. The tell is a tight score justified only by a national number. Confirm year-over-year movement in the target market before scoring tight.
  • Ratio direction flip. Mixing BLS unemployed-per-opening with AI demand-to-supply inverts the score. The tell is one market scoring backwards against your intuition. State the numerator and denominator once and apply it identically everywhere.
  • Score without a decision. A number with no mapped action is not usable. The tell is a scored role that no owner is acting on. Every score resolves to one lever with a name attached.

Keeping the score current

Re-score on skills volatility, not on a fixed calendar. The WEF Future of Jobs Report 2025 finds employers expect 39% of key skills to change by 2030, down from 44% in 2023, which works out to roughly 8% of core skills turning over per year and supports re-scoring general roles semi-annually. AI-exposed roles move faster: the skills sought by employers are changing 66% faster in jobs most exposed to AI, so re-score those quarterly. The exact cadence mapping is a synthesis, not a published rule, so treat it as a starting point and tighten it if your market moves.

Two events should trigger a re-score before the calendar does: a sharp year-over-year jump in the compensation premium, and a large swing in the unique posting count month over month. Both signal the ratio is shifting under you.

Before you call any score final, run this checklist.

Before you publish a hiring-difficulty score

  • The role-and-market pair has an explicit title-synonym list and location filter, applied identically to demand and supply.
  • The demand count is unique postings from a 30-to-60-day window, reported with its total-vs-unique ratio.
  • The supply count uses broad title plus skill synonyms, not a single exact tag, and is scoped to the same market.
  • The ratio's numerator and denominator are stated once and applied the same way to every market compared.
  • The score is confirmed or nudged by YoY compensation and time-to-fill read against the correct seniority band.
  • Employer concentration is checked; a single dominant firm in postings or profiles is flagged and discounted.
  • The score resolves to exactly one lever (post, source wide, widen geography, or build) with a named owner.
  • A re-score date is set: quarterly for AI-exposed roles, semi-annual otherwise.

A score you keep current is worth more than a precise score you compute once. The market moves, the pool ages, and the premium climbs; the discipline is not the first number but the trigger that tells you to compute the next one.

Questions practitioners ask

What is a good supply-to-demand ratio for a role?

Read it against the labour-market baseline, not against 1:1. The US sat at 1.0 unemployed per opening in June 2026 and 1.2 in 2019, so treat roughly 1.0 to 1.2 candidates per opening as balanced. Below that the market is tight; well below one, approaching the 3.2:1 demand-to-supply seen in AI roles, is acute. Always state which way you divide, because BLS ratios and AI-demand ratios run in opposite directions.

How do I count real hiring demand from job postings?

Count postings from the last 30 to 60 days, then deduplicate on job ID, description hash, and location, and report the unique count. Raw feeds run 30-45% duplicate and vendor pipelines strip up to 80% of collected postings, so the un-deduplicated total can overstate demand by roughly 2x. The 60-day rule matters: a re-posted ad only counts again as fresh demand after 60 days.

How often should I re-score a role's hiring difficulty?

Tie the cadence to skills volatility. The WEF projects 39% of core skills changing by 2030, about 8% per year, which supports re-scoring general roles every six months. AI-exposed roles see skills shift 66% faster, so re-score them quarterly. A compensation premium moving sharply year over year is a mid-cycle trigger to re-score early, regardless of the calendar.

Why did my talent pool nearly disappear when I filtered by skill?

You almost certainly filtered on a single exact skill tag, which measures tagging density, not real supply. In Refolk's index the ML-engineer pool dropped from 5,348 to 6 profiles when narrowed to one exact tag, a 99.9% collapse. Use a broad title plus several skill synonyms instead, and treat any near-zero pool from a single string as a false negative until you widen the query.

When should the score tell me to build talent instead of hire?

Build when the ratio is tight, the skills are scarce but stable, and you can invest over months rather than weeks. That is the top of the scale, a 5, where post-and-wait and source-wide both fail because the pool is too small at any price. Widen geography first if the skill exists elsewhere, since remote hiring removes the geographic bottleneck; build internally when even the widened market stays tight.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next