# The Pay Benchmark Standard: When a Salary Range Is Ready to Post

*You can grade any draft pay range post-ready or not-yet against six fixed criteria and produce a one-page evidence trail that survives challenge.*

- Canonical URL: https://www.refolk.ai/guides/pay-benchmark-post-ready-standard
- Pillar: Market and talent intelligence
- Format: Standard
- Published: 2026-08-08
- Last reviewed: 2026-08-08
- Reading time: 18 min

This is the standard for deciding one thing: whether a salary range you built from public and market data is defensible enough to post on a requisition or hand to Finance. It is written for compensation analysts, comp leads, and the talent-intelligence teams who feed them, and it turns a fuzzy judgment call into a gradeable checklist. Read it and you can grade any draft range post-ready or not-yet against six fixed criteria, and produce the one-page evidence trail that survives a legal or Finance challenge.

Most compensation guides tell you how to run a benchmarking process. Almost none tell you when to stop and whether the number you produced can be trusted. That gap is the actual moment of risk. Under pay-transparency law the range goes on a public posting and, in the EU, the burden of proof can shift to the employer. The point of a standard is that two analysts grade the same draft the same way, instead of ending at "buy our data."

## What "post-ready" means for a salary range

A range is post-ready when it clears six criteria at once: a tight job match, a sample floor, a real source blend, uniform aging, a documented percentile, and a written evidence trail. Miss any one and the number stops meaning anything, even if the arithmetic is clean.

The reason to fix all six in advance is that the failures do not announce themselves. A range can look precise because titles align, clear the five-company floor while hiding a bimodal distribution, and carry a confident aged figure built on a four-year-old edition. Each of those passes a casual read. The standard exists so a second grader catches them before the number is public.

Here are the six criteria, each stated so two people would grade the same case the same way.

| Criterion | Post-ready looks like | Not-yet looks like |
|---|---|---|
| Job match | 70-80% content and scope overlap, documented | Title match only, no scope note |
| Sample floor | 5+ companies per cut, none over 25% | Suppressed cut used silently, or n unknown |
| Source blend | 2-3 sources, method families differ | One pool resold twice, treated as two |
| Aging | Every source aged to one effective date | Mixed dates, or aging on a stale edition |
| Percentile | Target set and documented before pull | Percentile backed into after the offer |
| Documentation | One-page trail with rationale and sign-off | Number with no written reasoning |

Each criterion has a signal that tells you it is lying, covered in the failure-modes section. Grade against the definition, not against how confident the spreadsheet feels.

**4,756 - US compensation practitioners in Refolk's index**

The pool of analysts, managers, and partners this standard is written for, and the people who peer-review a draft range.

## Criterion one: how tight the job match must be

Match on content, not title, and require a 70 to 80% overlap in core responsibilities before you pull any data. A proper survey match compares job duties, responsibilities, scope, and required skills, and the good rule of thumb is a 70 to 80% match on core responsibilities.

The framing that keeps analysts honest is this: benchmarking answers "what does a data engineer with these skills, at our company size, in our city, paid at our chosen position against the market, make." Drop any one of those qualifiers and the number stops meaning anything. Scope is the qualifier people skip most, and it is decisive. A Director of Operations at a 200-person startup has a fundamentally different scope than one at a 10,000-person enterprise, and no survey will warn you that you matched across that gap.

The tell that a job match has failed is a title that maps cleanly but a scope that does not. A Vice President at a small startup might be equivalent to a Director or even a Senior Manager at a large corporation. If the responsibilities, decision-making authority, and headcount owned do not overlap by 70 to 80%, the match is cosmetic. Re-grade on scope, then on skills, then only last on title.

> **Rule:** Content match before data pull
>
> Fix scope, level, company size, location, and industry in writing before you pull a single number. A match documented after the fact is not a match, it is a rationalization.

Inconsistent job documentation is the upstream cause of most unreliable benchmarks. Each role should map to a standardized profile capturing scope, responsibilities, required qualifications, and decision-making authority. If two analysts cannot look at the profile and agree it is the same job, the comparator set is already broken.

## Criterion two: the sample floor and what breaks it

Every reported statistic must draw on at least five companies, none of them exceeding 25% of the weight, on data at least 90 days old. This is the antitrust safe-harbor floor, adopted across the survey industry, and it is the hardest line in the standard.

Stated in full, the safe harbor requires that data be at least 90 days old, that each statistic reported have at least five companies reporting data, that data be aggregated so no single company can be identified, and that no single company represent more than 25% of any statistic. ERI applies a tighter global variant: no data are reported for any job at any level where fewer than five companies match, or three outside the US. When a cut falls below the floor, reputable providers suppress it rather than publish it. Your job is to notice the suppression and respond by broadening the peer group, blending an additional source, or dropping to a higher job-family aggregate.

The subtler trap is skew. When salaries are concentrated toward one end of the scale or clumped at multiple points, a larger sample is needed than a normal distribution would require. A five-company cut can clear the floor and still be unreliable if it is bimodal. Check the shape of the distribution, not just the count.

The floor also travels badly across markets, and this is where thin comparator supply catches teams by surprise.

#### Sample floor by market thickness and match tightness

Horizontal axis runs from Thin comparator market to Deep comparator market. Vertical axis runs from Loose match (family level) to Tight match (scope and skill).

| Quadrant | What it means |
| --- | --- |
| Thin market, loose match | Often clears five companies; watch for skew before trusting it |
| Deep market, loose match | Clears easily; tighten the match to earn precision |
| Thin market, tight match | Highest risk of a suppressed cut; broaden peers or blend a source |
| Deep market, tight match | The target state; tight match with the floor cleared |

*The same methodology can pass in a deep market and collapse below the floor in a thin one at the same seniority and city.*

In Refolk's index, US Software Engineer supply runs at 16.1x Germany's (348,402 profiles versus 21,693). A US cut that easily clears five employers can drop below the floor when transplanted to a smaller market at the same seniority and city filter. The same methodology then yields a defensible range in one geography and an un-postable one in another.

| Role / market | Profiles in Refolk's index | Ratio |
|---|---|---|
| Software Engineer, United States | 348,402 | 16.1x Germany |
| Software Engineer, Germany | 21,693 | baseline |
| Compensation practitioners, US | 4,756 | reviewer pool |

Before you post a scarce-skill cut, check whether the comparator population actually exists at the seniority and location you filtered to. Sizing that supply pool directly is faster than discovering the floor breach after the survey suppresses the cut. [Refolk](/) answers that question in plain English, so you can test market thickness before you commit to a source.

I ran this search: `Staff-level backend engineers in the Netherlands who list Kubernetes` - [see the full result list](https://www.refolk.ai/s/2a5xtzd74h).

*Returns the comparator population for a scarce-skill cut so you can see whether it clears the five-employer floor before you post.*

## Criterion three: the source blend and how to reconcile it

Blend at least two sources across different method families, then reconcile them to one effective date. Cross-referencing at least two sources, ideally one with a broad baseline like the BLS OEWS and one with more current real-time data, gives you a more complete and reliable picture.

There are three method families, each with a known weakness. Published survey data is rigorous but lagging. Real-time platforms are fresh but sometimes limited in coverage. Custom peer-group analysis is precise but small-sample. Blending two or three sources with appropriate weighting produces more reliable benchmarks, and the reason it works is that the families fail in different directions.

| Method family | Strength | Weakness |
|---|---|---|
| Published survey | Rigorous, validated | Lags 6 to 18 months |
| Real-time platform | Current | Coverage can be thin |
| Custom peer set | Precise to your peers | Small sample |

The false pass here is a single-source blend wearing a multi-source costume. Two products that both resell the same self-reported pool are one source, not two. Confirm the underlying methods differ before you claim you blended. And treat free self-reported websites with suspicion: they are often unreliable, based on self-reported data that lacks validation. They can serve as a sanity check, never as one of your two load-bearing sources.

Named survey providers you will encounter include Radford (Aon), Mercer, Willis Towers Watson, and Salary.com. Record which methodology each uses and its effective date, because you cannot age or weight what you have not dated.

> **Watch out:** One pool resold twice is still one source
>
> If both of your sources trace back to the same self-reported data, you have single-source risk dressed as a blend. Check the method family, not the vendor name.

## Criterion four: aging every source to one date

Age every source to a single common effective date using a 2 to 3% annual factor, prorated by month. If you use data from multiple surveys but do not age each to a common date, you are comparing apples to oranges.

The mechanics are simple. Take the forecasted annual market movement, say 3.0%, divide by 12 to get one month's movement (0.0025), count the months from the source's effective date to your target date, and multiply. Apply that uplift to the source figure. Do it for every source so they all land on the same date before you blend or weight them.

#### Aging a source to a common effective date

1. **Read effective date** - Find when the source's data collection actually closed
2. **Set annual factor** - Choose a 2 to 3% market-movement rate
3. **Prorate by month** - Divide the annual factor by 12 and count months to target
4. **Apply uplift** - Multiply the source figure and record the aged value

*Every source lands on one date before any blending, or the comparison is meaningless.*

The critical limit: aging is a decay function, not a fix. Surveys are published on a lag, data collection closes months before publication, and organizations often use an edition for a full year after that. By the time a team is using it, the data may reflect market conditions from 12 to 24 months ago. Applying an aggressive factor to a four-year-old edition produces a confident wrong number; even aggressive aging factors cannot fully account for the market changes. Check the original effective date, not the aged one, and retire editions that are simply too old to rescue.

> Aging cannot save a stale source; it only tells you how far the number has already drifted.

This bites hardest in fast-moving roles. With movement at 2 to 3% a year and lags of 12 to 24 months, a stale source understates competitive pay precisely for the roles most likely to be posted under scrutiny.

## Criterion five: choosing the percentile with a written rationale

Anchor the midpoint to a chosen percentile and write the rationale before you pull data, not after the offer. P50 is the most commonly used reference point and the default market anchor for most organizations, and a 2023 SHRM survey found 87.6% of HR professionals use percentile data for compensation decisions.

The standard band mapping puts the minimum near P25, the midpoint at P50, and the maximum at P75 to P90.

| Band point | Typical percentile | Who it fits |
|---|---|---|
| Minimum | P25 | Qualified but less experienced hire |
| Midpoint | P50 | Fully competent, experienced employee |
| Maximum | P75-P90 | Exceptional or tenured |

The percentile is a decision, so it needs a reason. A target above P50 should be justified by talent scarcity or an explicit pay philosophy, not chosen to make a preferred number look market-aligned. The failure is backing into a percentile after you already know the offer you want to make. That reverses the logic and destroys defensibility. Set the target, document why, then pull the data.

Under EU rules the percentile alone will not carry the decision. Market rates can inform a decision through a talent-scarcity factor, but should be one input among several, not the sole rationale for a pay difference. The rationale that survives challenge names skills, effort, responsibility, and working conditions, with market data as one supporting input.

## Criterion six: the documentation that defends the range

The range is only as defensible as its written trail. Clearly outline the steps taken, including data sources used, criteria for selecting comparable positions, and any adjustments made, so the documentation can explain the decision to any stakeholder who challenges it.

This is where the real legal risk lives, and it is the reason a standard beats a process. Under EU Directive 2023/970 the burden of proof can shift: if a worker presents facts suggesting discrimination, the employer has to demonstrate no breach took place. Absent evidence, the default assumption goes against you. An undocumented but correct range is therefore more dangerous than a documented approximate one. Member States have until 6 July 2026 to implement the Directive, and a pay gap above 5% is defensible if it is objectively justified and documented.

The Directive raises the documentation bar from a spreadsheet to a mapped philosophy: a written pay philosophy mapping every criterion to the four legal standards of skills, effort, responsibility, and working conditions, with compensation bands (min, midpoint, max) per family and level, documented regional adjustments, and sign-off from legal counsel.

US law adds a separate constraint through the good-faith standard. Ranges must span from the lowest to the highest compensation the employer actually believes it will offer, and cannot include open-ended phrases like "$30,000 and up." Employer-size thresholds decide whether a posting must carry a range at all.

| State | Employees to trigger range-in-posting |
|---|---|
| Colorado / DC | 1 in-state worker |
| New York | 4 |
| New Jersey | 10 |
| California / Illinois / Washington | 15 |
| Hawaii | 50 |

Penalties are real: Colorado runs $500 to $10,000 per posting, and EU corporate penalties are capped at €14,000. As of 2026, 14 states plus DC have pay-transparency laws, and eleven require a salary range in the posting itself.

**One-page pay-range evidence trail**

```
ROLE / PROFILE ID: ___  (scope + level, 70-80% content match confirmed: Y/N)
COMPARATOR CUT: location ___  company size ___  industry ___
SOURCES (2-3):
  1. ___ | method family ___ | effective date ___
  2. ___ | method family ___ | effective date ___
SAMPLE FLOOR: each cut >= 5 companies (Y/N)  no employer > 25% (Y/N)  skew checked (Y/N)
AGING: annual factor ___%  common effective date ___
PERCENTILE TARGET: P___  rationale (scarcity / philosophy): ___  set before pull (Y/N)
BAND: min ___  midpoint ___  max ___
LEGAL / EU MAPPING: skills / effort / responsibility / conditions documented (Y/N)
SECOND-GRADER VERDICT: post-ready / not-yet
SIGN-OFF: legal ___  finance ___  date ___
```

*Fill every field. If a field is blank, the range is not post-ready.*

## The procedure: from draft range to post-ready

Run these eight steps in order. Steps one through six build the range, step seven is the independent re-grade, and step eight produces the evidence trail. The whole cycle for one role runs a few hours of analyst time plus a short review.

#### Grading a draft range to post-ready

1. **Fix the job match** - Map the role to a standardized profile capturing scope, level, responsibilities, and decision-making authority. Done means a documented 70-80% content match, not a title match.
2. **Set the comparator cut** - Lock location, company size, and industry before any data pull. Done means every qualifier is written down.
3. **Select and blend sources** - Choose at least two source types across the three method families. Done means 2-3 named sources, each with its methodology and effective date recorded.
4. **Check the sample floor** - Confirm each cut clears the five-company floor and no employer exceeds 25% weight. Done means suppressed cuts are flagged, not silently used.
5. **Age all data to one date** - Apply a 2-3% prorated factor to a common effective date. Done means every source carries the same effective date.
6. **Choose the percentile and build the band** - Anchor the midpoint to a chosen percentile with a written rationale, then set min and max. Done means the target is justified, not backed into.
7. **Grade against the checklist** - A second analyst independently re-grades all six criteria. Done means two graders reach the same post-ready or not-yet verdict.
8. **Produce the evidence trail** - Put sources, dates, factors, and rationale on one page and record legal or Finance sign-off. Done means the trail can survive a challenge.

One ordering note where the sources disagree: some methodologies age the data after matching, others treat aging as inseparable from source selection. Either works, provided every source ends up on the same effective date before you blend. Do not let the ordering debate become an excuse to skip the common-date discipline.

## How this goes wrong: failure modes and false positives

Most bad ranges pass a casual read. These are the specific ways a range looks post-ready and is not, and the check that catches each one.

| Failure mode | What it looks like | The check |
|---|---|---|
| Title match as content match | Range looks precise because titles align | Re-grade on scope; a VP at a startup may equal a Senior Manager |
| Skew under the floor | Cut clears 5 companies but is bimodal | Inspect distribution shape, not just n |
| Aging masks staleness | Confident figure from a 4-year-old edition | Check the original effective date, not the aged one |
| Fake blend | Two products, one self-reported pool | Confirm the method families differ |
| Range gamed wide | Legally compliant but useless band | Check the width against the real target |
| Market rate as sole rationale | EU range justified only by market data | Rationale must name skills, effort, responsibility, conditions |
| Percentile chosen after the offer | Target reverse-engineered to fit | Confirm the target was set before the data pull |
| Unvalidated free data | Self-reported website used as a source | Treat as sanity check only, never load-bearing |

The wide-band failure deserves a second look because it is the most tempting. Posting $40,000 to $400,000 for a role with a real target of $90,000 to $110,000 may satisfy the letter of the law but can trigger regulator attention. There is an upside to this constraint: because open-ended and artificially wide bands are prohibited, the posting law indirectly forces the sample discipline the rest of this checklist demands. You cannot hide a weak benchmark behind a wide band, so a tight, honest band is both the legal and the analytical answer.

> **Tip:** The second grader is the point
>
> The whole value of a standard is that two people reach the same verdict. If your reviewer cannot re-derive post-ready from the evidence page alone, the trail is incomplete, not the reviewer.

## Keeping the standard current

A range that was post-ready last quarter can drift out of readiness without anyone touching it, because its sources age and the law moves under it. Re-grade on a schedule, not only when someone challenges a number.

Two things decay on their own. First, source freshness: with market movement at 2 to 3% a year and survey lags of 12 to 24 months, a range built on last year's edition understates competitive pay in scarce roles first. Re-pull and re-age before you repost a role in a fast-moving market. Second, the legal surface: Member States implement the EU Directive by 6 July 2026, and US coverage stands at 14 states plus DC with eleven requiring a range in the posting. Re-check the employer-size thresholds and posting requirements for every state you hire in before you rely on a stored answer.

#### Post-ready sign-off

- [ ] The job match is a documented 70-80% content and scope overlap, not a title match
- [ ] Every cut clears five companies with no employer over 25%, and skew was checked
- [ ] Two or three sources are named, across different method families, each dated
- [ ] All sources are aged to one common effective date; no edition is beyond rescue
- [ ] The target percentile was set and documented before the data pull
- [ ] The rationale names skills, effort, responsibility, and working conditions
- [ ] The band width matches the real offer intent, with no open-ended top
- [ ] A second analyst reached the same post-ready verdict independently
- [ ] The one-page evidence trail is complete with legal or Finance sign-off

When all nine boxes are checked and two graders agree, the range is post-ready. Anything short of that is not-yet, and not-yet means broaden the peer group, blend another source, or wait for a fresher edition before the number goes public.

## Frequently asked questions

### How many pay data sources should I blend for a defensible range?

At least two, and ideally across different method families. The accepted minimum is cross-referencing two sources, one with a broad public baseline like the BLS OEWS and one with more current real-time data. Blending two or three sources with appropriate weighting produces a more reliable benchmark. Two products that both resell the same self-reported pool count as one source, not two, so confirm the underlying methods actually differ.

### What is the minimum sample size for a salary benchmark?

The widely adopted antitrust safe-harbor floor requires at least five companies reporting each statistic, data at least 90 days old, aggregation so no single company is identifiable, and no single company representing more than 25% of any statistic. ERI applies at least five matching companies, or three outside the US. Below the floor, suppress the cut and broaden the peer group rather than publish it.

### How do I age salary survey data to a current date?

Apply an annual market-movement factor, typically 2 to 3%, prorated by month to a common effective date. Divide the forecasted annual movement by 12, count the months to age, and multiply. Age every source to the same date before blending. Aging is a decay function, not a fix: it cannot rescue a four-year-old edition, and data may already reflect market conditions from 12 to 24 months ago.

### Which percentile should I use for a salary range?

P50 is the most commonly used reference point and the default market anchor for most organizations. A common mapping sets the band minimum at P25, the midpoint at P50, and the maximum at P75 to P90. Whatever you choose, set and document the target percentile before the data pull, justified by talent scarcity or pay philosophy. A percentile backed into after the offer defeats defensibility.

### Is market rate enough to justify a pay range under EU law?

No. Under EU Directive 2023/970, market rates can inform a decision through a talent-scarcity factor, but must be one input among several, not the sole rationale for a pay difference. The burden of proof shifts to the employer, so you need gender-neutral criteria mapped to skills, effort, responsibility, and working conditions, plus salary bands, audit trails, and a clear rationale for every outcome.

### How wide can a posted salary range be?

US good-faith standards require the range to span from the lowest to the highest compensation the employer actually believes it will offer, with no open-ended phrases like $30,000 and up. Posting $40,000 to $400,000 for a role targeted at $90,000 to $110,000 may satisfy the letter of the law but can trigger regulator attention. The rule effectively forces the sample discipline this checklist demands.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/pay-benchmark-post-ready-standard*
