RefolkCandidates
StandardReading the market

The Market Pay Estimate Standard, and the Numbers That Fail It

After this you can grade a compiled market pay estimate pass or fail on six criteria, so two people grading the same figures reach the same verdict.

15 min readLast reviewed September 13, 2026Read as Markdown

You have pulled pay figures for your role from three or four places, and now they disagree. This guide is for the job seeker who needs to decide whether the picture is solid enough to anchor a target and an ask, or whether it is a pile of numbers dressed up as a conclusion. It gives you a pass/fail grading rubric with six criteria, so you can stop second-guessing, commit to a target, and know exactly which figures to throw out.

Most pages on this topic review each source one by one and end with "use it as one input." That advice never tells you when your own compiled number is done. This standard does. It defines done, and it names the numbers that fail it.

What "trustworthy" means for a compiled pay estimate

A trustworthy market pay estimate is one that passes all six criteria at once: a real sample floor, at least two independent source types, current-or-aged recency, a single-market geography match, a content-based level match, and a clean separation of base from total comp. Miss any one and the estimate fails, because each criterion closes a different way the number lies.

The reason for pass/fail rather than a weighted score is that these failures are not additive. A stale figure is not "mostly right." A total-comp number compared against a base target is not "close." Each defect can move your anchor by more than the negotiation gap you are fighting over. So the grade is binary, and a single fail sinks the estimate until you fix it.

67%
Share of employer-posted ranges that actually contain reported pay
Glassdoor's own analysis; 22% of listings sat above the real number and 11% below, a one-in-three miss rate.

That 67% figure is the whole case for the independence criterion. A number that misses one time in three cannot be adopted from a single source. It has to be corroborated. Everything below turns that instinct into a rule two people can apply the same way.

A stale figure is not mostly right, and a total-comp number is not close. Each defect fails the grade outright.

The six criteria, and what each one proves

The six criteria are sample floor, source independence, recency, geography match, level match, and base-versus-total clarity. Each one exists because a specific, documented distortion lives in the gap it covers. Here is what each proves, and what it looks like when it lies.

  • Sample floor. Proves the median is not an accident of two people. It lies when a confident-looking median actually rests on one or two reports. Glassdoor's floor is three or more submissions before it shows a median at all.
  • Source independence. Proves the number was seen by more than one measurement method. It lies when three "sources" all re-scrape the same crowdsourced pool. Independence requires at least one payroll, survey, or government source distinct from crowdsourced.
  • Recency. Proves the figure reflects the current market. It lies when a five-year-old submission sits next to fresh data with no expiry date. Un-aged staleness can understate pay by 5 to 15 percent.
  • Geography match. Proves the number belongs to your market. It lies when a national figure is applied to a high-cost metro. Differentials run 40 to 60 percent between markets.
  • Level match. Proves the benchmark is your job, not its name. It lies when the title matches but the scope does not. Matching a senior role to a mid-level benchmark is the primary source of benchmarking error.
  • Base-versus-total clarity. Proves you are comparing like with like. It lies when a "high" number is total comp measured against a base-only target. Crowdsourced tools layer bonuses, profit-sharing, and commissions into total pay.

Why one source is never enough

No single source clears the independence criterion, because each source type has a documented direction of bias. Grading is a matter of pulling from types that fail in different directions and reconciling them, not of finding the one honest site.

The three source types bias predictably. Crowdsourced self-report rounds up and never retires stale points. Employer-posted ranges skew low and truncated. Government series lag and report broad categories that understate niche or senior roles. Put a crowdsourced figure next to a government figure and the truth is bracketed between two known errors, which is the entire point of triangulation.

Source typeRecency lagBias direction
Crowdsourced self-reportStale points never retiredRounds up; overstates via total-comp
Employer good-faith rangeCurrent at postingSkews low (~28% spread vs 40-60% norm)
Government (BLS OEWS)12 to 18 monthsBroad category; understates niche and senior

The government option is stronger than most job seekers assume. The BLS Occupational Employment and Wage Statistics program covers over 830 occupations across all US states and more than 500 metro areas, free, drawn from employer payroll records rather than self-report. Its weakness is lag: it updates annually 12 to 18 months behind, and reports broad occupational categories rather than company-specific figures. As of May 2025 the BLS median hourly wage across all US occupations was $24.51, which is a useful anchor for a broad role and useless for a specialized one.

The gold standard is HR-reported survey data from providers like Culpepper, WTW, Radford, or Korn Ferry, which comes directly from employers. Most job seekers cannot buy it, so the realistic independent pair is one crowdsourced tool plus one government series. If you can reach someone who builds these benchmarks, one conversation replaces hours of reconciliation.

How employer-posted ranges narrow toward a usable figure

  1. All posted ranges
    100%

    Starting pool

  2. Ranges containing real pay
    67%

    The rest miss high or low

  3. In-range pay below the median
    over 60%

    Of the 67%, most sits low

Of every posted range, two-thirds contain real pay and most of that real pay sits in the lower half of the band.

The range-accuracy problem, in numbers

Employer-posted ranges are the most tempting source and the most misread. They feel official, but the numbers say a posted range is a weak input that has to be discounted before it counts. Two distortions do the damage: the miss rate and the median trap.

MetricValue
Reported pay within posted range67%
Actual pay below range22%
Actual pay above range11%
In-range pay below range medianover 60%

Read the last row carefully, because it is the one that costs people money. Even when a range is accurate, over 60 percent of real pay falls below the range median. Anchoring your ask at the posted midpoint systematically overshoots what most people in that role actually receive. The mechanism is that employers pad the top of the band, so the midpoint is not the center of a bell curve - it sits above where most pay lands.

There is a quieter distortion too. The recommended high end of a range should sit 40 to 60 percent above the minimum, but an analysis of more than 12 million listings found the actual spread is only about 28 percent. That compression means the advertised top is often below the real ceiling. Meanwhile, under half of US job postings include salary information at all, and only about 10 percent note an exact pay level, so you are often reconciling a handful of truncated bands.

Grade a compiled estimate in seven steps

The procedure below takes a pile of figures to a pass/fail verdict in about half a day of focused work. Run it in order; each step assumes the previous one is done, and the last step is only meaningful once the first six have been executed.

From a pile of figures to a graded verdict

  1. Define the role by content
    Write out responsibilities, scope, decision authority, and reporting level, then confirm an 80% or more content match to your benchmark rather than a title match. Done when you can point to a benchmark job description whose content matches at least 80% of yours.
  2. Fix geography and level
    Pin a single metro or market and one level, and filter every source to both. Done when no figure in your pile comes from a different location or level, because differentials can move the same role 40 to 60 percent between markets.
  3. Pull two independent source types
    Get at least one crowdsourced figure and one government or survey figure, and separate base from total comp. Done when a base number and a total-comp number sit side by side, not merged.
  4. Check the sample floor
    Confirm each load-bearing median rests on enough reports, using Glassdoor's floor of three or more submissions and its confidence badge as the reference. Done when no anchor number rests on one or two reports.
  5. Date-stamp and age every figure
    Record each source's effective date and age forward anything older than nine to twelve months to a single reference date. Done when all figures sit at one date; replace anything over a year old rather than aging it.
  6. Reconcile spread against truncation
    Compare each employer-posted spread to the 40 to 60 percent above minimum norm, discounting bands near 28 percent as skewed low and absurdly wide bands as non-informative. Done when you know which posted ranges to trust and which to drop.
  7. Grade pass or fail
    Apply the six criteria across sample, independence, recency, geography, level, and base-versus-total clarity, failing the estimate on any single miss. Done when two people grading the same evidence would reach the same verdict.

Two of these steps carry more weight than the rest. Defining the role by content is where most estimates are lost before they begin, and dating every figure is where a clean-looking pile silently rots. Give those two the most time.

Building the resume that matches the role you just defined by content, and tailoring it to each posting so your claimed scope lines up with the level you are benchmarking, is exactly the friction Refolk removes - it writes from your own history and scores how well you actually fit each posting, which keeps your target and your applications pointed at the same level.

Content match and geography, done precisely

Match on content, not title, and pin one market before you compare anything. These two criteria - level match and geography match - are where a pile of figures that look consistent quietly turns out to be measuring different jobs in different places.

For level match, aim for a benchmark job description whose content matches 80 percent or more of yours. Not a perfect match, and never a title match. Match by responsibilities, scope, skills, decision-making authority, and reporting relationships. The dossier is blunt that a senior-to-mid mismatch is the primary source of benchmarking error, so if the title says "analyst" but the scope says "leads a team," grade against the scope.

For geography, the number that should scare you is the differential: the same role in San Francisco can carry a 50th-percentile rate 40 to 60 percent above an equivalent secondary Midwest market. A national figure applied to a high-cost metro is not conservative, it is wrong by up to 60 percent. Filter every source to one metro before you reconcile.

Where a benchmark figure sits before you trust it

Right level or contentWrong level or content
Wrong content, wrong market
Discard; it measures a different job elsewhere
Right content, wrong market
Apply the 40-60% differential before using
Wrong content, right market
Re-match on scope, not title, to 80%
Right content, right market
Usable input; corroborate for independence
Wrong geographyRight geography
A figure is only usable when it matches both your job content and your market; the other three quadrants need work first.

How this goes wrong: eight false positives

The most valuable part of a standard is the list of things that pass a lazy look and fail a real one. Every entry below is a false positive: a number that looks trustworthy and is not. Check each one explicitly before you grade an estimate as passing.

  • Median built on a thin sample. A confident-looking median that rests on one or two reports. Check the three-or-more-submission floor and the confidence badge before trusting it.
  • Base-versus-total confusion. A "high" number that is total comp compared against a base-only target. Separate the figures, because tools layer bonuses and commissions into total pay.
  • Range treated as a bell centered on the median. Aiming at the midpoint when over 60 percent of pay sits below the range median. Anchor to reported medians, not posted-range midpoints.
  • Stale point used as current. A five-year-old submission with no expiry sitting next to fresh data. Date-stamp and discard or age it; understatement can reach 5 to 15 percent.
  • Title match instead of content match. Same title, different scope. Enforce the 80 percent content-match rule, since a senior-to-mid mismatch is the primary benchmarking error.
  • Wrong geography. A national number applied to a high-cost metro. Filter to a single market; differentials run 40 to 60 percent.
  • Fake independence. Three "sources" that all re-scrape the same crowdsourced pool. Require at least one payroll, survey, or government source distinct from crowdsourced.
  • Trusting a very wide posted range. Reading a band like $50,000 to $250,000 as informative. Discount any range far outside the 40 to 60 percent norm as non-informative.

Aging, and why the calendar downgrades your data

Aging is the practice of moving a figure forward to a common effective date because surveys are published on a lag. It matters because staleness is a silent downgrade, not a rounding error, and it usually runs in the direction that costs you money.

Data collection closes months before publication, and organizations often use a survey for a full year after it appears, so published data may lag the analysis date by 6 to 18 months. Government series lag 12 to 18 months by design. Left un-aged in a fast-moving market, that gap can understate current compensation by 5 to 15 percent - frequently larger than the negotiation gap a job seeker is arguing over.

The working rules conflict slightly across sources, and the standard resolves the conflict conservatively: age figures between roughly 9 and 18 months forward to your reference date, and replace anything more than a year old rather than aging it, since one source holds that data over a year old should be swapped for a newer source entirely. When in doubt, replace rather than stretch.

Figure log for a single graded estimate
Source name | Source type (crowd/employer/gov/survey) | Base or Total | Sample size / confidence | Effective date | Metro | Level | Aged figure at reference date
Example row:
Glassdoor | crowdsourced | Base | 12 reports, high confidence | 8 months ago | Denver metro | Senior | aged +4% to today
Government series | gov OEWS | Base | full occupation | 14 months ago | Denver metro | (broad) | aged +6% to today

One row per figure. Fill every column before grading; a blank column is a fail on that criterion.

Keep the estimate current, and know where the data is scarce

A graded estimate is a snapshot, not a permanent verdict. Re-grade it when any source crosses its recency threshold, when you change target market or level, or when a new government release lands. The BLS OEWS cycle is the clearest calendar to track: it publishes annually, and knowing your latest release date tells you exactly when your government anchor has gone stale.

Where the standard is thinnest is outside the United States, and it is worth saying plainly. Reliable, methodologically aged, content-matched benchmarks are produced by a scarce profession that is heavily concentrated in one country.

MarketCountRatio vs UK (derived)
United States4,64843.0x
Canada2772.6x
United Kingdom1081.0x

In Refolk's index of professional profiles there are 4,648 US professionals with compensation-analyst or manager titles, against 277 in Canada and 108 in the UK - roughly a 43x US-to-UK gap. The practical consequence: outside the US, fewer practitioners produce aged, content-matched benchmarks, so non-US readers lean harder on lagging government series and should weight the recency criterion even more carefully.

number: 4,648
label: US professionals with compensation-analyst or manager titles in Refolk's index
note: Against 277 in Canada and 108 in the UK - the people who build the aged, content-matched benchmarks you are trying to grade.

Questions job seekers ask

Is Glassdoor salary data accurate enough to anchor my ask?

Not on its own. Glassdoor is self-reported crowdsourced data with no mechanism to correct people who round up and no way to retire a five-year-old submission. It passes only when the median rests on three or more submissions, carries a high confidence badge, and is corroborated by a government or survey source. Use it as one of at least two independent inputs, never as the sole anchor.

How accurate are posted salary ranges in job listings?

By Glassdoor's own analysis, employer-posted ranges contain real reported pay about 67 percent of the time. In 22 percent of listings actual salaries fell below the range and in 11 percent they were above it. Even inside accurate ranges, over 60 percent of real pay sits below the range median, so a posting is a corroborating input, not a verdict.

How old is too old for salary data?

Replace anything more than about a year old rather than aging it, and age figures between roughly nine and eighteen months forward to a common reference date. Published survey data often lags six to eighteen months, and government series lag twelve to eighteen. Left uncorrected in a fast market, that staleness can understate current pay by 5 to 15 percent.

How do I cross-check salary data sources properly?

Require at least two genuinely independent source types, meaning one crowdsourced and one government or survey figure, not three sites that all re-scrape the same crowdsourced pool. Match on job content at 80 percent or more, filter to a single metro and level, and separate base from total comp before you compare anything.

What is the single most common benchmarking mistake?

Matching on title instead of content. A senior role matched to a mid-level benchmark is the primary source of benchmarking error. Match on responsibilities, scope, decision authority, skills, and reporting relationships, and aim for 80 percent or more description content match, not a perfect title match.

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Read next