Refolk
StandardMarket and talent intelligence

The Pay Benchmark Standard: When a Salary Range Is Ready to Post

You can grade any draft pay range post-ready or not-yet against six fixed criteria and produce a one-page evidence trail that survives challenge.

18 min readLast reviewed August 8, 2026Read as Markdown

This is the standard for deciding one thing: whether a salary range you built from public and market data is defensible enough to post on a requisition or hand to Finance. It is written for compensation analysts, comp leads, and the talent-intelligence teams who feed them, and it turns a fuzzy judgment call into a gradeable checklist. Read it and you can grade any draft range post-ready or not-yet against six fixed criteria, and produce the one-page evidence trail that survives a legal or Finance challenge.

Most compensation guides tell you how to run a benchmarking process. Almost none tell you when to stop and whether the number you produced can be trusted. That gap is the actual moment of risk. Under pay-transparency law the range goes on a public posting and, in the EU, the burden of proof can shift to the employer. The point of a standard is that two analysts grade the same draft the same way, instead of ending at "buy our data."

What "post-ready" means for a salary range

A range is post-ready when it clears six criteria at once: a tight job match, a sample floor, a real source blend, uniform aging, a documented percentile, and a written evidence trail. Miss any one and the number stops meaning anything, even if the arithmetic is clean.

The reason to fix all six in advance is that the failures do not announce themselves. A range can look precise because titles align, clear the five-company floor while hiding a bimodal distribution, and carry a confident aged figure built on a four-year-old edition. Each of those passes a casual read. The standard exists so a second grader catches them before the number is public.

Here are the six criteria, each stated so two people would grade the same case the same way.

CriterionPost-ready looks likeNot-yet looks like
Job match70-80% content and scope overlap, documentedTitle match only, no scope note
Sample floor5+ companies per cut, none over 25%Suppressed cut used silently, or n unknown
Source blend2-3 sources, method families differOne pool resold twice, treated as two
AgingEvery source aged to one effective dateMixed dates, or aging on a stale edition
PercentileTarget set and documented before pullPercentile backed into after the offer
DocumentationOne-page trail with rationale and sign-offNumber with no written reasoning

Each criterion has a signal that tells you it is lying, covered in the failure-modes section. Grade against the definition, not against how confident the spreadsheet feels.

4,756
US compensation practitioners in Refolk's index
The pool of analysts, managers, and partners this standard is written for, and the people who peer-review a draft range.

Criterion one: how tight the job match must be

Match on content, not title, and require a 70 to 80% overlap in core responsibilities before you pull any data. A proper survey match compares job duties, responsibilities, scope, and required skills, and the good rule of thumb is a 70 to 80% match on core responsibilities.

The framing that keeps analysts honest is this: benchmarking answers "what does a data engineer with these skills, at our company size, in our city, paid at our chosen position against the market, make." Drop any one of those qualifiers and the number stops meaning anything. Scope is the qualifier people skip most, and it is decisive. A Director of Operations at a 200-person startup has a fundamentally different scope than one at a 10,000-person enterprise, and no survey will warn you that you matched across that gap.

The tell that a job match has failed is a title that maps cleanly but a scope that does not. A Vice President at a small startup might be equivalent to a Director or even a Senior Manager at a large corporation. If the responsibilities, decision-making authority, and headcount owned do not overlap by 70 to 80%, the match is cosmetic. Re-grade on scope, then on skills, then only last on title.

Inconsistent job documentation is the upstream cause of most unreliable benchmarks. Each role should map to a standardized profile capturing scope, responsibilities, required qualifications, and decision-making authority. If two analysts cannot look at the profile and agree it is the same job, the comparator set is already broken.

Criterion two: the sample floor and what breaks it

Every reported statistic must draw on at least five companies, none of them exceeding 25% of the weight, on data at least 90 days old. This is the antitrust safe-harbor floor, adopted across the survey industry, and it is the hardest line in the standard.

Stated in full, the safe harbor requires that data be at least 90 days old, that each statistic reported have at least five companies reporting data, that data be aggregated so no single company can be identified, and that no single company represent more than 25% of any statistic. ERI applies a tighter global variant: no data are reported for any job at any level where fewer than five companies match, or three outside the US. When a cut falls below the floor, reputable providers suppress it rather than publish it. Your job is to notice the suppression and respond by broadening the peer group, blending an additional source, or dropping to a higher job-family aggregate.

The subtler trap is skew. When salaries are concentrated toward one end of the scale or clumped at multiple points, a larger sample is needed than a normal distribution would require. A five-company cut can clear the floor and still be unreliable if it is bimodal. Check the shape of the distribution, not just the count.

The floor also travels badly across markets, and this is where thin comparator supply catches teams by surprise.

Sample floor by market thickness and match tightness

Tight match (scope and skill)Loose match (family level)
Thin market, loose match
Often clears five companies; watch for skew before trusting it
Deep market, loose match
Clears easily; tighten the match to earn precision
Thin market, tight match
Highest risk of a suppressed cut; broaden peers or blend a source
Deep market, tight match
The target state; tight match with the floor cleared
Thin comparator marketDeep comparator market
The same methodology can pass in a deep market and collapse below the floor in a thin one at the same seniority and city.

In Refolk's index, US Software Engineer supply runs at 16.1x Germany's (348,402 profiles versus 21,693). A US cut that easily clears five employers can drop below the floor when transplanted to a smaller market at the same seniority and city filter. The same methodology then yields a defensible range in one geography and an un-postable one in another.

Role / marketProfiles in Refolk's indexRatio
Software Engineer, United States348,40216.1x Germany
Software Engineer, Germany21,693baseline
Compensation practitioners, US4,756reviewer pool

Before you post a scarce-skill cut, check whether the comparator population actually exists at the seniority and location you filtered to. Sizing that supply pool directly is faster than discovering the floor breach after the survey suppresses the cut. Refolk answers that question in plain English, so you can test market thickness before you commit to a source.

Criterion three: the source blend and how to reconcile it

Blend at least two sources across different method families, then reconcile them to one effective date. Cross-referencing at least two sources, ideally one with a broad baseline like the BLS OEWS and one with more current real-time data, gives you a more complete and reliable picture.

There are three method families, each with a known weakness. Published survey data is rigorous but lagging. Real-time platforms are fresh but sometimes limited in coverage. Custom peer-group analysis is precise but small-sample. Blending two or three sources with appropriate weighting produces more reliable benchmarks, and the reason it works is that the families fail in different directions.

Method familyStrengthWeakness
Published surveyRigorous, validatedLags 6 to 18 months
Real-time platformCurrentCoverage can be thin
Custom peer setPrecise to your peersSmall sample

The false pass here is a single-source blend wearing a multi-source costume. Two products that both resell the same self-reported pool are one source, not two. Confirm the underlying methods differ before you claim you blended. And treat free self-reported websites with suspicion: they are often unreliable, based on self-reported data that lacks validation. They can serve as a sanity check, never as one of your two load-bearing sources.

Named survey providers you will encounter include Radford (Aon), Mercer, Willis Towers Watson, and Salary.com. Record which methodology each uses and its effective date, because you cannot age or weight what you have not dated.

Criterion four: aging every source to one date

Age every source to a single common effective date using a 2 to 3% annual factor, prorated by month. If you use data from multiple surveys but do not age each to a common date, you are comparing apples to oranges.

The mechanics are simple. Take the forecasted annual market movement, say 3.0%, divide by 12 to get one month's movement (0.0025), count the months from the source's effective date to your target date, and multiply. Apply that uplift to the source figure. Do it for every source so they all land on the same date before you blend or weight them.

Aging a source to a common effective date

  1. Read effective date
    Find when the source's data collection actually closed
  2. Set annual factor
    Choose a 2 to 3% market-movement rate
  3. Prorate by month
    Divide the annual factor by 12 and count months to target
  4. Apply uplift
    Multiply the source figure and record the aged value
Every source lands on one date before any blending, or the comparison is meaningless.

The critical limit: aging is a decay function, not a fix. Surveys are published on a lag, data collection closes months before publication, and organizations often use an edition for a full year after that. By the time a team is using it, the data may reflect market conditions from 12 to 24 months ago. Applying an aggressive factor to a four-year-old edition produces a confident wrong number; even aggressive aging factors cannot fully account for the market changes. Check the original effective date, not the aged one, and retire editions that are simply too old to rescue.

Aging cannot save a stale source; it only tells you how far the number has already drifted.

This bites hardest in fast-moving roles. With movement at 2 to 3% a year and lags of 12 to 24 months, a stale source understates competitive pay precisely for the roles most likely to be posted under scrutiny.

Criterion five: choosing the percentile with a written rationale

Anchor the midpoint to a chosen percentile and write the rationale before you pull data, not after the offer. P50 is the most commonly used reference point and the default market anchor for most organizations, and a 2023 SHRM survey found 87.6% of HR professionals use percentile data for compensation decisions.

The standard band mapping puts the minimum near P25, the midpoint at P50, and the maximum at P75 to P90.

Band pointTypical percentileWho it fits
MinimumP25Qualified but less experienced hire
MidpointP50Fully competent, experienced employee
MaximumP75-P90Exceptional or tenured

The percentile is a decision, so it needs a reason. A target above P50 should be justified by talent scarcity or an explicit pay philosophy, not chosen to make a preferred number look market-aligned. The failure is backing into a percentile after you already know the offer you want to make. That reverses the logic and destroys defensibility. Set the target, document why, then pull the data.

Under EU rules the percentile alone will not carry the decision. Market rates can inform a decision through a talent-scarcity factor, but should be one input among several, not the sole rationale for a pay difference. The rationale that survives challenge names skills, effort, responsibility, and working conditions, with market data as one supporting input.

Criterion six: the documentation that defends the range

The range is only as defensible as its written trail. Clearly outline the steps taken, including data sources used, criteria for selecting comparable positions, and any adjustments made, so the documentation can explain the decision to any stakeholder who challenges it.

This is where the real legal risk lives, and it is the reason a standard beats a process. Under EU Directive 2023/970 the burden of proof can shift: if a worker presents facts suggesting discrimination, the employer has to demonstrate no breach took place. Absent evidence, the default assumption goes against you. An undocumented but correct range is therefore more dangerous than a documented approximate one. Member States have until 6 July 2026 to implement the Directive, and a pay gap above 5% is defensible if it is objectively justified and documented.

The Directive raises the documentation bar from a spreadsheet to a mapped philosophy: a written pay philosophy mapping every criterion to the four legal standards of skills, effort, responsibility, and working conditions, with compensation bands (min, midpoint, max) per family and level, documented regional adjustments, and sign-off from legal counsel.

US law adds a separate constraint through the good-faith standard. Ranges must span from the lowest to the highest compensation the employer actually believes it will offer, and cannot include open-ended phrases like "$30,000 and up." Employer-size thresholds decide whether a posting must carry a range at all.

StateEmployees to trigger range-in-posting
Colorado / DC1 in-state worker
New York4
New Jersey10
California / Illinois / Washington15
Hawaii50

Penalties are real: Colorado runs $500 to $10,000 per posting, and EU corporate penalties are capped at €14,000. As of 2026, 14 states plus DC have pay-transparency laws, and eleven require a salary range in the posting itself.

One-page pay-range evidence trail
ROLE / PROFILE ID: ___  (scope + level, 70-80% content match confirmed: Y/N)
COMPARATOR CUT: location ___  company size ___  industry ___
SOURCES (2-3):
  1. ___ | method family ___ | effective date ___
  2. ___ | method family ___ | effective date ___
SAMPLE FLOOR: each cut >= 5 companies (Y/N)  no employer > 25% (Y/N)  skew checked (Y/N)
AGING: annual factor ___%  common effective date ___
PERCENTILE TARGET: P___  rationale (scarcity / philosophy): ___  set before pull (Y/N)
BAND: min ___  midpoint ___  max ___
LEGAL / EU MAPPING: skills / effort / responsibility / conditions documented (Y/N)
SECOND-GRADER VERDICT: post-ready / not-yet
SIGN-OFF: legal ___  finance ___  date ___

Fill every field. If a field is blank, the range is not post-ready.

The procedure: from draft range to post-ready

Run these eight steps in order. Steps one through six build the range, step seven is the independent re-grade, and step eight produces the evidence trail. The whole cycle for one role runs a few hours of analyst time plus a short review.

Grading a draft range to post-ready

  1. Fix the job match
    Map the role to a standardized profile capturing scope, level, responsibilities, and decision-making authority. Done means a documented 70-80% content match, not a title match.
  2. Set the comparator cut
    Lock location, company size, and industry before any data pull. Done means every qualifier is written down.
  3. Select and blend sources
    Choose at least two source types across the three method families. Done means 2-3 named sources, each with its methodology and effective date recorded.
  4. Check the sample floor
    Confirm each cut clears the five-company floor and no employer exceeds 25% weight. Done means suppressed cuts are flagged, not silently used.
  5. Age all data to one date
    Apply a 2-3% prorated factor to a common effective date. Done means every source carries the same effective date.
  6. Choose the percentile and build the band
    Anchor the midpoint to a chosen percentile with a written rationale, then set min and max. Done means the target is justified, not backed into.
  7. Grade against the checklist
    A second analyst independently re-grades all six criteria. Done means two graders reach the same post-ready or not-yet verdict.
  8. Produce the evidence trail
    Put sources, dates, factors, and rationale on one page and record legal or Finance sign-off. Done means the trail can survive a challenge.

One ordering note where the sources disagree: some methodologies age the data after matching, others treat aging as inseparable from source selection. Either works, provided every source ends up on the same effective date before you blend. Do not let the ordering debate become an excuse to skip the common-date discipline.

How this goes wrong: failure modes and false positives

Most bad ranges pass a casual read. These are the specific ways a range looks post-ready and is not, and the check that catches each one.

Failure modeWhat it looks likeThe check
Title match as content matchRange looks precise because titles alignRe-grade on scope; a VP at a startup may equal a Senior Manager
Skew under the floorCut clears 5 companies but is bimodalInspect distribution shape, not just n
Aging masks stalenessConfident figure from a 4-year-old editionCheck the original effective date, not the aged one
Fake blendTwo products, one self-reported poolConfirm the method families differ
Range gamed wideLegally compliant but useless bandCheck the width against the real target
Market rate as sole rationaleEU range justified only by market dataRationale must name skills, effort, responsibility, conditions
Percentile chosen after the offerTarget reverse-engineered to fitConfirm the target was set before the data pull
Unvalidated free dataSelf-reported website used as a sourceTreat as sanity check only, never load-bearing

The wide-band failure deserves a second look because it is the most tempting. Posting $40,000 to $400,000 for a role with a real target of $90,000 to $110,000 may satisfy the letter of the law but can trigger regulator attention. There is an upside to this constraint: because open-ended and artificially wide bands are prohibited, the posting law indirectly forces the sample discipline the rest of this checklist demands. You cannot hide a weak benchmark behind a wide band, so a tight, honest band is both the legal and the analytical answer.

Keeping the standard current

A range that was post-ready last quarter can drift out of readiness without anyone touching it, because its sources age and the law moves under it. Re-grade on a schedule, not only when someone challenges a number.

Two things decay on their own. First, source freshness: with market movement at 2 to 3% a year and survey lags of 12 to 24 months, a range built on last year's edition understates competitive pay in scarce roles first. Re-pull and re-age before you repost a role in a fast-moving market. Second, the legal surface: Member States implement the EU Directive by 6 July 2026, and US coverage stands at 14 states plus DC with eleven requiring a range in the posting. Re-check the employer-size thresholds and posting requirements for every state you hire in before you rely on a stored answer.

Post-ready sign-off

  • The job match is a documented 70-80% content and scope overlap, not a title match
  • Every cut clears five companies with no employer over 25%, and skew was checked
  • Two or three sources are named, across different method families, each dated
  • All sources are aged to one common effective date; no edition is beyond rescue
  • The target percentile was set and documented before the data pull
  • The rationale names skills, effort, responsibility, and working conditions
  • The band width matches the real offer intent, with no open-ended top
  • A second analyst reached the same post-ready verdict independently
  • The one-page evidence trail is complete with legal or Finance sign-off

When all nine boxes are checked and two graders agree, the range is post-ready. Anything short of that is not-yet, and not-yet means broaden the peer group, blend another source, or wait for a fresher edition before the number goes public.

Questions practitioners ask

How many pay data sources should I blend for a defensible range?

At least two, and ideally across different method families. The accepted minimum is cross-referencing two sources, one with a broad public baseline like the BLS OEWS and one with more current real-time data. Blending two or three sources with appropriate weighting produces a more reliable benchmark. Two products that both resell the same self-reported pool count as one source, not two, so confirm the underlying methods actually differ.

What is the minimum sample size for a salary benchmark?

The widely adopted antitrust safe-harbor floor requires at least five companies reporting each statistic, data at least 90 days old, aggregation so no single company is identifiable, and no single company representing more than 25% of any statistic. ERI applies at least five matching companies, or three outside the US. Below the floor, suppress the cut and broaden the peer group rather than publish it.

How do I age salary survey data to a current date?

Apply an annual market-movement factor, typically 2 to 3%, prorated by month to a common effective date. Divide the forecasted annual movement by 12, count the months to age, and multiply. Age every source to the same date before blending. Aging is a decay function, not a fix: it cannot rescue a four-year-old edition, and data may already reflect market conditions from 12 to 24 months ago.

Which percentile should I use for a salary range?

P50 is the most commonly used reference point and the default market anchor for most organizations. A common mapping sets the band minimum at P25, the midpoint at P50, and the maximum at P75 to P90. Whatever you choose, set and document the target percentile before the data pull, justified by talent scarcity or pay philosophy. A percentile backed into after the offer defeats defensibility.

Is market rate enough to justify a pay range under EU law?

No. Under EU Directive 2023/970, market rates can inform a decision through a talent-scarcity factor, but must be one input among several, not the sole rationale for a pay difference. The burden of proof shifts to the employer, so you need gender-neutral criteria mapped to skills, effort, responsibility, and working conditions, plus salary bands, audit trails, and a clear rationale for every outcome.

How wide can a posted salary range be?

US good-faith standards require the range to span from the lowest to the highest compensation the employer actually believes it will offer, with no open-ended phrases like $30,000 and up. Posting $40,000 to $400,000 for a role targeted at $90,000 to $110,000 may satisfy the letter of the law but can trigger regulator attention. The rule effectively forces the sample discipline this checklist demands.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next