The Hiring-Difficulty Score: Rating a Role Against a Market
You will score any role-and-market pair 1 to 5 for hiring difficulty from public supply and demand data, then turn that score into one sourcing decision.
Before you commit headcount to a plan, you need to know how hard each role will actually be to fill in each market you are considering. This guide is for strategy, talent-intelligence, and research teams who have to make that call from public data. It gives you a vendor-neutral rubric that weighs active demand against available supply, produces a 1-to-5 hiring-difficulty score, and maps that score to one sourcing decision: post and wait, source wide, widen the geography, or build internally.
Most existing research either sizes a talent pool without demand, or reads a competitor's roadmap without supply. Neither tells you whether a specific role is fillable. Scoring difficulty means putting both sides of the ledger on the same scale and computing a ratio you can defend, with thresholds you can point to.
What the hiring-difficulty score measures
The hiring-difficulty score is a 1-to-5 rating of how hard one role is to fill in one market, computed as active demand weighed against available supply and confirmed by compensation and time-to-fill. It is a supply-and-demand ratio, not a pool size, and it always resolves to an action.
A pool size answers "how many people exist." The difficulty score answers "how many people exist relative to how many employers want them right now, and how fast the market is moving." Those are different questions with different answers. A market with 5,000 qualified people and 6,000 open roles is harder than a market with 500 people and 200 open roles, even though the first pool is ten times larger.
Two published instruments already blend these inputs, which tells you the shape is sound. Gartner's TalentNeuron scores hiring difficulty 1 to 10 using posting duration, change in demand, competitive concentration measured by HHI, and relative supply. Lever's Recruiter Pressure Index combines job openings, quits, and unemployment into a single 0-to-100 measure. This guide keeps the same inputs but publishes the thresholds so you can compute the score yourself from open job-posting and profile data, without a platform license.
That 10.5x gap is the whole reason to score per market rather than globally. A role that is a 2 in one country can be a 5 next door.
The four inputs and what each one proves
Four public signals drive the score. Two set the ratio, two confirm it. Each proves something specific, and each lies in a specific way, so learn both.
- Active demand (unique job postings). Proves how many employers are hiring for the role right now. It lies when reposts inflate the count, so one aggressive employer looks like a hot market.
- Available supply (profile pool). Proves how many qualified people exist in the market. It lies when the pool is inert, full of people not open to moving, so a big number hides a slow fill.
- Compensation premium (year over year). Proves whether employers are paying up to win the role, which confirms scarcity the pool count can miss. It lies when a global average is read as a local one.
- Time-to-fill by seniority. Proves how long the market actually takes to close the role. It lies when a median across all levels hides that senior roles run far longer.
The ratio comes from demand and supply. Compensation and time-to-fill are the corroborators: they either confirm the ratio's reading or push the score one band. If the ratio says balanced but the role takes 90 days and pay is climbing, the ratio is missing something, usually an inert pool.
The four layers of a difficulty score
- DecisionPost, source wide, widen geography, or build
- CorroboratorsYoY compensation premium and time-to-fill by seniority
- RatioUnique demand postings weighed against available supply
- DefinitionTitle synonyms, seniority band, and geography, fixed once
Building a demand count you can defend
A defensible demand count is unique postings from the last 30 to 60 days, deduplicated on job ID, description hash, and location. Raw feeds are heavily inflated, so the total posting count is close to useless on its own.
The numbers are stark. Raw job-posting feeds often contain 30-45% duplicates, and a fixed deduplication algorithm removes about 80% of all postings collected for redundancy and stable data. The mechanism is the re-post rule: a posting's re-appearances within 60 days of the original are counted as duplicates. So if a job first posts on March 1st, any copy found elsewhere in the next 60 days is a duplicate; if the same ad reposts daily for a year, it counts six times, not 365. Deduplication itself uses a statistical classifier trained on location, job title, company name, and text similarity, with text similarity detected using shingling.
Two operational details follow. First, use a 30-to-60-day window, anchored by the same 60-day rule: a posting re-found after 60 days is legitimately fresh demand, and anything shorter risks missing a slow-filling senior role. Second, benchmark change against fixed lookbacks. Vendors compare demand to the number of similar postings one month, six months, and twelve months ago, and postings data is reclassified and deduplicated every four weeks. A published, role-level decay rate for posting demand is not established publicly, so use the 60-day window plus month-over-month change as your practical proxy rather than inventing a decay curve.
The thresholds: where balanced ends and tight begins
Read the ratio against a published anchor, not against your intuition. The cleanest public threshold is the BLS unemployed-per-opening ratio: values less than one suggest a tight labour market with more openings than unemployed workers, while ratios larger than one indicate slack. Balanced sits near 1.0 to 1.2, not at 1:1.
| Signal | Value | Reading |
|---|---|---|
| Unemployed per opening, US Jun 2026 | 1.0 | balanced |
| Unemployed per opening, 2019 average | 1.2 | normal to tight |
| Unemployed per opening, Jul 2009 | 6.5 | very slack |
| AI demand-to-supply, global | 3.2:1 | extremely tight |
| AI security / ZK specialty | 8:1 | acute |
Two things to internalise from this table. The 2019 economy, widely called hot, sat at 1.2 unemployed per opening, and June 2026 sat at 1.0. If you score every market at or near 1.0 as scarce, you will flag normal conditions as emergencies. Treat 1.0 to 1.2 as balanced.
The tight end needs a real ceiling, and the AI benchmark supplies one: AI talent demand exceeds supply by a ratio of 3.2 to 1, with over 1.6 million open positions and only 518,000 qualified candidates globally, and in specialties like AI security or ZK cryptography it can exceed 8:1. That is your anchor for a 5.
Here is the five-band scale. Note that no single published five-band mapping exists; the levers below are documented, but the band thresholds are my synthesis, so calibrate them to your own market data before betting a plan on them.
| Score | Supply-to-demand reading | What it means |
|---|---|---|
| 1 | Slack, well above 1.2 candidates per opening | More qualified people than openings |
| 2 | Comfortable, roughly 1.0 to 1.2 | Balanced, healthy pool |
| 3 | Tightening, roughly 0.5 to 1.0 | Fewer candidates than openings |
| 4 | Tight, roughly 0.3 to 0.5 | Approaching the AI-wide 3.2:1 anchor |
| 5 | Acute, below 0.3 | At or past 3.2:1, specialty scarcity |
Counting supply without manufacturing scarcity
A defensible supply count uses a broad title plus multiple skill synonyms, counted per market from an open profile index. The single most common way this step fails is filtering on one exact skill tag, which measures tagging density, not real talent.
The evidence from Refolk's index is blunt. The US ML-engineer pool holds 5,348 profiles under a broad title match. Narrow that to one exact skill string and it collapses to 6 profiles, a 99.9% drop driven entirely by tagging sparsity. Those people did not vanish; the filter did. Always match on title plus several skill synonyms, and treat any near-zero pool from a single string as a false negative until you widen it.
Geography is the other half. The same role differs by an order of magnitude across borders, so supply is meaningless until you pin it to a market.
| Market | ML-engineer profiles | Top in-pool employer |
|---|---|---|
| United States | 5,348 | Meta |
| Germany | 510 | Bosch |
| US-to-Germany multiple | 10.5x | - |
The employer column carries a warning too. Top US employers in the pool are Meta, Apple, and Intel; top German employers are Bosch, Meta, and Accenture. If one company dominates both the postings and the profiles, your market is really one buyer and one bench, and both counts are less liquid than they look.
This is the step where a plain-language index earns its place. Rather than exporting boards and stitching deduplicated postings against a scraped profile set, you can ask directly for the pool in the exact market and shape you defined in step one.
With Refolk, the supply count and the demand-adjacent employer read come from the same query, which keeps the numerator and denominator on comparable definitions instead of two different pipelines.
Confirming the score with pay and time-to-fill
Once you have a ratio-based score, two signals confirm it or move it one band: the year-over-year compensation premium and time-to-fill by seniority. Compensation often tightens before the pool visibly shrinks, so a sharp premium jump is an early warning.
The named source for pay is PwC's Global AI Jobs Barometer, measured year over year from job ads comparing AI-skill roles to similar non-AI roles, based on analysis of close to a billion job ads and thousands of company financial reports from six continents. The trajectory across editions is the signal: the average premium hit 56%, up from 25% the year before, then reached 62%. A sustained double-digit year-over-year jump is a tight-market tell. The trap is that these are global averages; a specific metro may not move at all, so confirm year-over-year movement in your target market before you let pay push the score up.
Time-to-fill confirms from the other direction. A long clock on a role you scored as balanced means the pool is inert.
| Level | Time-to-fill (days) |
|---|---|
| Junior | 15-25 |
| Mid | 30-45 |
| Senior / Staff | 60-90+ |
| Tech median, all levels | ~48 |
Nearly 40% of senior-level roles take more than 90 days to fill, and the SHRM benchmark averages 44 days across roles with the technology median around 48. So a senior role at 90 days is normal, not alarming, while a junior role stuck at 60 days is a genuine scarcity signal. Read time-to-fill against the seniority band, never against the all-levels median.
Compensation confirms the scarcity a profile count misses, and it moves before the pool visibly shrinks.
The procedure
Run these seven steps in order. Steps two and three can swap; BLS-style analysis leads with demand while pool-sizing guides lead with supply, and either works as long as you compute the ratio the same way for every market.
Score a role-and-market pair
- Define the role-and-market pair preciselyFix the job-title synonyms, seniority band, and geography. Done means an explicit title-list and location filter you can run identically against postings and profiles.
- Pull the demand countCount open postings from the last 30 to 60 days, then deduplicate on job ID, description hash, and location. Done means a unique-posting count reported with its total-vs-unique ratio.
- Pull the supply countCount matching profiles in the same market using broad title plus skill synonyms, never one exact tag. Done means a candidate-pool figure comparable to the demand count.
- Compute the ratio and place it on the scaleDivide supply by demand and read against the BLS tightness rule and the 3.2:1 AI anchor. Done means a provisional 1-to-5 score with numerator and denominator stated once.
- Add corroborating signalsLayer the YoY compensation premium and time-to-fill by seniority onto the ratio. Done means the score confirmed or nudged one band, with the reason recorded.
- Map the score to a sourcing decisionTranslate the band into post-and-wait, source wide, widen geography, or build, with the hiring lead present. Done means one named action with an owner.
- Set the re-score dateSchedule quarterly for AI-exposed roles and semi-annual otherwise. Done means a calendar trigger tied to the pair.
From raw postings to a defensible ratio
- 100Raw postings collected
before any cleaning
- 60After removing 30-45% duplicates
same-ad reposts inside 60 days
- 20Unique demand count
vendor pipelines strip up to 80% of raw
From score to sourcing decision
Every score must resolve to one named lever with an owner. A 1-to-5 number with no decision attached is a listicle, not a plan. The two documented levers are the build/buy/borrow framework, where build is best when skills are scarce but stable and you can invest through internships, apprenticeships, and internal academies, and geographic widening, where remote talent removes the geographic bottleneck entirely and lets you widen the funnel and speed the hiring loop.
| Score | Reading | Sourcing decision |
|---|---|---|
| 1 | Slack | Post and wait; the pool comes to you |
| 2 | Balanced | Post and wait, with light sourcing on close |
| 3 | Tightening | Source wide; run active outreach across the market |
| 4 | Tight | Widen geography; source the role where supply exists |
| 5 | Acute | Build internally, or widen geography first if the skill lives elsewhere |
The order between 4 and 5 matters. Widening geography is faster and cheaper than building, so try it first: if the 10.5x gap between two markets is real, the role that is a 5 at home may be a 3 abroad, and remote hiring closes the gap without a multi-quarter build. Reserve build for when even the widened market stays tight and the skill is stable enough to justify training on it. Note that 77% of employers plan to upskill workers while 41% plan to reduce workforce as AI automates tasks, so building is a live strategy, but it only pays off when the underlying skill will still matter by the time your first cohort is ready.
Choosing between widen and build
How this scoring goes wrong
The score fails in predictable ways, and each failure has a public tell you can check. This is the part to read twice, because a confidently wrong score sends a whole hiring plan in the wrong direction.
- Inflated demand from reposts. One aggressive employer reposting daily reads as a hot market. The tell is a high posting count with one dominant company. Check the total-vs-unique ratio and employer concentration; if one firm owns most postings, discount the demand.
- Stale supply pool. A big pool that is inert scores as easy but fills slowly. The tell is a low ratio paired with a long time-to-fill. Cross-read posting duration and time-to-fill by seniority before trusting a large pool.
- Skill-string undercounting supply. Filtering by one exact skill tag can collapse the pool to near zero. The tell is a supply count in the single digits for a role you know is common. Use broad title plus multiple synonyms; Refolk's index went from 5,348 to 6 on a single tag.
- Wrong window. A 12-month posting count captures filled and seasonal roles as live demand. The tell is a demand count that dwarfs your unique count. Restrict to 30 to 60 days and confirm against month-over-month change.
- Compensation premium misread. A 62% AI premium is a global average that a specific metro may not share. The tell is a tight score justified only by a national number. Confirm year-over-year movement in the target market before scoring tight.
- Ratio direction flip. Mixing BLS unemployed-per-opening with AI demand-to-supply inverts the score. The tell is one market scoring backwards against your intuition. State the numerator and denominator once and apply it identically everywhere.
- Score without a decision. A number with no mapped action is not usable. The tell is a scored role that no owner is acting on. Every score resolves to one lever with a name attached.
Keeping the score current
Re-score on skills volatility, not on a fixed calendar. The WEF Future of Jobs Report 2025 finds employers expect 39% of key skills to change by 2030, down from 44% in 2023, which works out to roughly 8% of core skills turning over per year and supports re-scoring general roles semi-annually. AI-exposed roles move faster: the skills sought by employers are changing 66% faster in jobs most exposed to AI, so re-score those quarterly. The exact cadence mapping is a synthesis, not a published rule, so treat it as a starting point and tighten it if your market moves.
Two events should trigger a re-score before the calendar does: a sharp year-over-year jump in the compensation premium, and a large swing in the unique posting count month over month. Both signal the ratio is shifting under you.
Before you call any score final, run this checklist.
Before you publish a hiring-difficulty score
- The role-and-market pair has an explicit title-synonym list and location filter, applied identically to demand and supply.
- The demand count is unique postings from a 30-to-60-day window, reported with its total-vs-unique ratio.
- The supply count uses broad title plus skill synonyms, not a single exact tag, and is scoped to the same market.
- The ratio's numerator and denominator are stated once and applied the same way to every market compared.
- The score is confirmed or nudged by YoY compensation and time-to-fill read against the correct seniority band.
- Employer concentration is checked; a single dominant firm in postings or profiles is flagged and discounted.
- The score resolves to exactly one lever (post, source wide, widen geography, or build) with a named owner.
- A re-score date is set: quarterly for AI-exposed roles, semi-annual otherwise.
A score you keep current is worth more than a precise score you compute once. The market moves, the pool ages, and the premium climbs; the discipline is not the first number but the trigger that tells you to compute the next one.
Questions practitioners ask
What is a good supply-to-demand ratio for a role?
Read it against the labour-market baseline, not against 1:1. The US sat at 1.0 unemployed per opening in June 2026 and 1.2 in 2019, so treat roughly 1.0 to 1.2 candidates per opening as balanced. Below that the market is tight; well below one, approaching the 3.2:1 demand-to-supply seen in AI roles, is acute. Always state which way you divide, because BLS ratios and AI-demand ratios run in opposite directions.
How do I count real hiring demand from job postings?
Count postings from the last 30 to 60 days, then deduplicate on job ID, description hash, and location, and report the unique count. Raw feeds run 30-45% duplicate and vendor pipelines strip up to 80% of collected postings, so the un-deduplicated total can overstate demand by roughly 2x. The 60-day rule matters: a re-posted ad only counts again as fresh demand after 60 days.
How often should I re-score a role's hiring difficulty?
Tie the cadence to skills volatility. The WEF projects 39% of core skills changing by 2030, about 8% per year, which supports re-scoring general roles every six months. AI-exposed roles see skills shift 66% faster, so re-score them quarterly. A compensation premium moving sharply year over year is a mid-cycle trigger to re-score early, regardless of the calendar.
Why did my talent pool nearly disappear when I filtered by skill?
You almost certainly filtered on a single exact skill tag, which measures tagging density, not real supply. In Refolk's index the ML-engineer pool dropped from 5,348 to 6 profiles when narrowed to one exact tag, a 99.9% collapse. Use a broad title plus several skill synonyms instead, and treat any near-zero pool from a single string as a false negative until you widen the query.
When should the score tell me to build talent instead of hire?
Build when the ratio is tight, the skills are scarce but stable, and you can invest over months rather than weeks. That is the top of the scale, a 5, where post-and-wait and source-wide both fail because the pool is too small at any price. Widen geography first if the skill exists elsewhere, since remote hiring removes the geographic bottleneck; build internally when even the widened market stays tight.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.