Refolk
TeardownMarket and talent intelligence

Benchmarking One Role's Pay From Public Visa Filings

You can build a percentile pay band for one role in one metro from public LCA files and defend it against the offered-wage, SOC, level, and multi-site traps.

17 min readLast reviewed August 30, 2026Read as Markdown

You need to know what one role in one city actually pays, and you need to defend the number to a hiring manager or to leadership without hand-waving. This teardown carries a single benchmark all the way through the public DOL labor-condition disclosure files, with the real record counts and the wrong turns, so a talent-intelligence analyst can follow along on their own case. The worked example is Software Developers in the San Francisco metro for FY2025; the method transfers to any SOC and any metro.

Most pages ranking for "h1b salary data by role and city" are either a search box over the same records or a job-seeker negotiation tip sheet. Neither shows you where a naive median off the raw files lies to you. That is the whole job here.

What the LCA file actually is, and what it is not

The Labor Condition Application disclosure file is the employer's attested and offered wage at filing, published quarterly by the Office of Foreign Labor Certification. It is not payroll. It is not proof anyone was hired. Reading it as "median salary" is the first and most expensive mistake.

The file is built from Form ETA-9035/9035E, and each column maps to a form item in the OFLC record layout. The fields you care about for a pay benchmark are a small subset:

FieldWhat it holdsTrap it hides
CASE_STATUSCertified, Certified-Withdrawn, DeniedCertified-only filtering drops valid offers
SOC_CODE / title2018 SOC classificationWrong SOC pulls the wrong labor market
WAGE_RATE_OF_PAY_FROMOffered wage, low endOffered, not paid; range needs resolving
PW_WAGE_LEVELPrevailing wage level I to IVMixing levels widens the range artificially
WORKSITE fieldsCity, county, ZIP of employmentMay be HQ or highest-cost metro, not yours
TOTAL_WORKERSPositions on the filingRow-count medians double-count big filers

The single most important sentence about this file: the wage columns are the employer's attested prevailing wage and the offered wage, not verified payroll. You cannot tell from the LCA whether the worker was actually paid the advertised figure, because that is private payroll data. The LCA shows what the employer promised, not what it delivered. Every number you ship from this file inherits that limit.

Why a naive median off the raw file is wrong

A naive median treats every row as one person and every offered wage as an actual paycheck. Both assumptions are false, and they push the number in predictable directions. Before touching the file, understand the four forces that bend the result.

The wage columns are offered, not paid. The median you compute is the middle of what employers promised, which by law must equal or exceed the higher of the actual or prevailing wage. That makes it a floor, not a market rate.

Rows are not people. In FY2025, 910,968 certified positions came from 561,505 certified applications, about 1.62 positions per LCA. A single LCA can cover many workers, so an unweighted median lets large filers vote once for a filing that represents dozens of hires. Weight by TOTAL_WORKERS or you double-count the big filers.

1.62x
Positions per certified LCA, FY2025
910,968 certified positions came from 561,505 applications, so row-count medians overstate people by roughly 60 percent.

The wage level sets the percentile. Levels I through IV map to roughly the 17th, 34th, 50th, and 67th percentiles of the wage distribution. An SOC dominated by Level I filings produces a median far below true market pay, and you will not know it unless you look at the level mix.

One occupation drags the pool. Software Developers were 286,287 of the 910,968 FY2025 certified positions, or 31.4 percent. Pool all SOC codes into a "tech" median and that one occupation sets the shape of your answer.

How a raw median goes wrong

  1. Raw rows
    Every LCA row counted once, offered wage read as pay
  2. Weight
    Multiply by TOTAL_WORKERS so big filings do not vote once
  3. Segment
    Split by PW_WAGE_LEVEL before taking percentiles
  4. Localize
    Filter WORKSITE to the target metro, not HQ
Four corrections turn a misleading row-count median into a defensible per-metro, per-level offered-wage band.

The eight-step benchmark, worked end to end

Fix the scope, pull the file, filter deliberately, normalize the wage, weight, then compute and cross-check. Each step below has a done condition so you know when to move on. This is the procedure that matches the numbered method in this guide's summary.

Trace one role's pay range through public LCA filings

  1. Scope the benchmark
    Fix one SOC 2018 code, one metro, and one wage year. Done looks like "SOC 15-1252 Software Developers, San Francisco MSA, FY2025." Pick the SOC by duties, not title, and default to the higher-paid occupation when duties combine.
  2. Download the right disclosure file
    Pull the LCA disclosure Excel for the fiscal year from the OFLC Performance Data page and confirm the record-layout version matches. Done looks like a file that opens with the expected columns. Expect hundreds of megabytes.
  3. Filter CASE_STATUS deliberately
    Decide Certified only versus including Certified-Withdrawn, then record row count before and after. Done looks like a documented delta. FY2025 baseline is 561,505 certified of 594,819 processed.
  4. Isolate the SOC and worksite
    Filter to the target SOC and to worksite city, county, or ZIP, not employer HQ. Done looks like a per-metro subset. Place of employment is where work is actually performed, per Fact Sheet 62J.
  5. Choose the wage column and normalize
    Use WAGE_RATE_OF_PAY_FROM, convert hourly to annual at 2,080 hours, and resolve wage ranges to one figure per row. Done looks like one comparable annual number per row.
  6. Weight by positions and de-duplicate
    Multiply by TOTAL_WORKERS when sizing demand and drop duplicate case numbers. Done looks like a weighted distribution. FY2025 ran about 1.62 positions per certified LCA.
  7. Compute percentiles and cross-check FLAG
    Produce P25, P50, and P75 and compare to the four FLAG prevailing-wage levels for the same SOC, area, and year. Done looks like a median inside the I-to-IV band with any gap explained.
  8. Reconcile and flag omissions
    Note that the offered wage excludes bonus and equity and that LCA volume overstates hires. Done looks like a caveat line on total comp and intent versus hire.

Step 1, scope: the fork on SOC choice

The first fork is where people go wrong before they open a file. Some guides pick the SOC by matching the job title. DOL policy picks by duties, and defaults to the higher-paid occupation when duties are combined. For the worked example I set the target in writing: SOC 15-1252 Software Developers, San Francisco MSA, FY2025. Writing it down is not ceremony; it is the artifact you defend later.

Note the wage year. OFLC implemented the 2018 SOC codes once OEWS and O*NET completed the transition from the 2010 codes, so a stale crosswalk will map you to a defunct code. Confirm the SOC exists in the 2018 set.

Step 3, the CASE_STATUS fork and the first wrong turn

Here is the first wrong turn most analysts take. The instinct is to filter CASE_STATUS to "Certified" and move on, because certified feels clean. In FY2025 that means keeping 561,505 rows out of 594,819 processed. It looks conservative.

It is not. Certification is nearly automatic. The FY2025 certification rate was 94.4 percent, denials were 0.66 percent, and withdrawals 4.9 percent. Certifications are mostly pro forma and are only denied for obvious errors and inaccuracies. So filtering to Certified drops only about 5 to 6 percent of rows, but it silently removes Certified-Withdrawn filings that were valid at the moment they were filed. Those are real offers. Dropping them biases the sample toward completed hires and away from what employers were actually offering.

The fix is boring and correct: run the counts both ways, report the delta, and for a pay benchmark keep Certified-Withdrawn in. The certification filter barely changes the median but it changes what the median means.

Fiscal yearCertifiedDeniedWithdrawnCert. rate
FY2020558,6263,98314,72596.7%
FY2022590,3263,09632,66294.3%
FY2025561,5053,92129,39394.4%

The certification rate sits near 95 percent every year, which is exactly why CASE_STATUS is a weak filter and worksite plus SOC plus level are the strong ones.

Step 4, the worksite fork and the multi-location trap

The second wrong turn is trusting the wage on a multi-location filing. If an employee will work in multiple locations, the LCA must identify each worksite and use the prevailing wage for each area, and the highest applicable prevailing wage governs the entire LCA. So a filing that lists San Francisco and Austin can record the San Francisco prevailing wage even for the Austin seat. Filter to employer HQ and you inherit whichever metro the employer happened to list first.

Fact Sheet 62J defines the place of employment as where work is actually performed. Filter the WORKSITE fields to your target city, county, or ZIP, not the employer's headquarters. For the San Francisco benchmark I keep only rows whose worksite resolves to the SF MSA and I flag multi-site rows for review rather than trusting their recorded wage.

Step 5, the wage column and the unit trap

Use WAGE_RATE_OF_PAY_FROM, the low end of the offered range, as the comparable figure. Convert hourly figures to annual at 2,080 hours, which is the full-time convention under 20 CFR 655.731. Where a filing lists a wage range, resolve it to one number per row and note the rule you used so the choice is reproducible. Done means every surviving row carries one comparable annual figure.

Step 6, weighting and the row-count trap

Multiply each row by TOTAL_WORKERS when you are sizing demand or building a distribution meant to represent people rather than filings. Drop duplicate case numbers first. In FY2025 the file averaged 1.62 positions per certified LCA, so the weighting shifts the median toward the wages large filers offer. Recompute both weighted and unweighted and look at the shift; a large gap tells you a few big filers were setting your unweighted median.

Refolk closes a gap this method leaves open: the LCA file tells you what employers offered, but not who is actually in the market at that level. When you need to sanity-check a Level III/IV band against real senior people in the metro, describe the role in plain English and get the list back.

Reading the wage level, and why level-blind medians lie

The wage level is the single field that decides whether your median means anything. Levels I to IV map to fixed points on the wage distribution for the SOC and area, so a median computed across mixed levels is an average of entry and lead pay and describes nobody.

Level II and Level III wages are set midway between Level I and Level IV, at roughly the 34th and 50th percentiles. The full map:

LevelApprox. percentileTypical role signal
I17thentry, minimal experience
II34thqualified, some experience
III50thexperienced/specialized
IV67thfully competent/lead

This is why the median is a floor, not a market rate. Because the offered wage must equal or exceed the higher of the actual or prevailing wage, and because Level I sits near the 17th percentile, an SOC dominated by Level I filings produces a median far below the pay a hiring manager will actually see for a mid-level hire. Segment by PW_WAGE_LEVEL before you take percentiles, and report the level mix alongside the number so the reader can see whether your sample skews junior.

One field can be blank. Wage level is empty when the employer used an independent survey or a collective bargaining agreement rather than the OES level. The fraction of filings with no level is not publicly established, so I check it locally: count blank PW_WAGE_LEVEL rows in your subset and decide whether to drop or bucket them before percentiles.

Then cross-check. Pull the four prevailing-wage levels for your SOC, area, and wage year from the FLAG Wage Search, which replaced the discontinued FLCDataCenter on July 1, 2024. Your computed median should land inside the Level I to Level IV band. If it falls outside, you have a SOC, worksite, or unit-conversion error to find before shipping.

The LCA median is the floor employers promised, not the market rate a hiring manager will pay.

Where this goes wrong: the failure modes to check before shipping

Seven ways this benchmark misleads, each with the check that catches it. Treat this as the pre-ship review; it is the most valuable part of the method because every one of these produces a clean-looking number that is wrong.

  • Median off raw rows. The false positive is a tidy median that over-weights big filers, because each row counts as one person. Check: recompute weighted by TOTAL_WORKERS, about 1.62 positions per LCA in FY2025, and compare the shift.
  • Offered wage read as actual pay. The false positive is a confident "median salary." The columns are attested and offered, not payroll. Check: label the output "offered wage floor" and state that bonus and equity are absent.
  • Wrong SOC code. Picking the wrong SOC is one of the most common reasons LCAs get challenged later, and it also pulls the wrong labor market into your benchmark. Check: match O*NET duties and compare the same offer across candidate SOC codes.
  • Multi-location filing pinned to the wrong metro. The highest-cost worksite can be the recorded prevailing wage for every seat on the filing. Check: filter WORKSITE to the target city or ZIP, not employer HQ, and flag multi-site rows.
  • Certified-only sampling. Dropping Certified-Withdrawn removes valid-at-filing offers and biases toward completed hires. Check: run counts both ways and report the delta.
  • Level-blind medians. Mixing Level I entry filings with Level IV lead filings widens the range artificially. Check: segment by PW_WAGE_LEVEL before taking percentiles.
  • Intent mistaken for headcount. LCA volume overstates hires; the false positive is an inflated market-size number. Check: reconcile against the USCIS Data Hub petition approvals.

That last one deserves its own paragraph because it is the difference between a pay benchmark and a headcount claim. The LCA counts intentions, not hires. It does not tell you whether the worker was hired, selected in the lottery, or granted a visa, and many LCAs never result in an H-1B hire. The scale of the gap is documented: in FY2019 the top 30 H-1B employers received approval for 371,461 positions on LCAs, 38 percent of the 968,538 positions DOL certified, against far fewer actual USCIS petition approvals that year. The USCIS H-1B Employer Data Hub is the reconciling source, but it publishes no wage figures, no job-level detail beyond classification, and no cross-reference key tying petitions to specific LCAs. So you can use it to sanity-check volume, never to fix a wage.

Keeping the benchmark current and reproducible

A pay benchmark from this file has two clocks, and if you ignore either the number stops being reproducible. State both on every deliverable.

The disclosure clock. OFLC publishes cumulative quarterly and fiscal-year releases on the federal fiscal year, October 1 to September 30, plus historical annual reports. The lag is about six weeks: the Q1 FY2026 release, covering data through December 31, 2025, was announced February 13, 2026. So a benchmark cites a fiscal year and a release quarter, not just a year.

The wage-year clock. Prevailing-wage levels update on a July 1 wage year from BLS OEWS. This silently re-levels your benchmark: an offer that read Level III in June can read Level II in July at the exact same dollar figure, because the underlying distribution moved. An undated FLAG screenshot breaks reproducibility. Record the wage year alongside the fiscal year.

Benchmark provenance line for the slide footer
Role: SOC <code> <title>
Metro: <MSA or county/ZIP set>
Source: OFLC LCA disclosure, FY<year>, <release quarter>
Wage year: FLAG OES, effective July 1, <year>
Sample: <n rows kept> after worksite + SOC filter; CASE_STATUS = <Certified | +Withdrawn>
Weighting: by TOTAL_WORKERS (<n weighted positions>)
Output: P25 / P50 / P75 = <values> (offered wage floor; excludes bonus & equity)

Fill each field from your run so anyone can reproduce the number.

One more thing the file cannot see: supply. Benchmarking a US metro from filings alone ignores the offshore pool that shapes real market pay. In Refolk's index there are 412,840 US profiles with Software Engineer or Software Developer titles and 684,937 in India, a 1.66x ratio, while LCA demand concentrates heavily in California at 22.8 percent of FY2025 positions. When you present a number to leadership, the supply picture is the context that makes it a decision rather than a data point.

CountryProfiles (SWE/Dev titles)Index share vs US
United States412,8401.00x
India684,9371.66x

Before you call the benchmark done

Run this checklist against the number before it leaves your hands. Each item is a specific check, not a topic, and every one maps to a failure mode above.

Pre-ship review for one role-and-metro benchmark

  • Target is written as one SOC 2018 code, one metro, one fiscal year, one wage year
  • SOC was chosen by duties against O*NET, not by job title alone
  • WORKSITE is filtered to the target city, county, or ZIP, not employer HQ
  • Multi-site rows are flagged and not trusted at their recorded wage
  • CASE_STATUS counts are recorded both ways and the delta is reported
  • Wage figures use WAGE_RATE_OF_PAY_FROM, hourly converted at 2,080 hours
  • Percentiles are computed within each PW_WAGE_LEVEL, and blank-level rows are handled
  • The distribution is weighted by TOTAL_WORKERS and duplicate case numbers are dropped
  • The computed median lands inside the FLAG Level I to IV band, or the gap is explained
  • Output is labeled "offered wage floor" with a caveat on bonus, equity, and intent-vs-hire

The finished deliverable is not a single number. It is a per-level, weighted, metro-localized offered-wage band with a provenance line, a supply-side caveat, and an explicit statement of what the LCA file cannot tell you. That is the version a hiring manager cannot pick apart, because you have already found every place it could have been wrong.

To keep the work current, re-pull on the next fiscal-year release for a full-year sample, re-check the FLAG levels after each July 1 wage year, and reconcile any headcount claims against the USCIS Data Hub. The method does not change; only the two clocks do.

Questions practitioners ask

Is H1B salary data by role and city the same as what people actually get paid?

No. The LCA discloses the employer's offered and attested prevailing wage, not verified payroll. It is what the employer promised at filing, not what landed in a paycheck, and it omits bonus, equity, and total cash comp. Treat the output as an offered wage floor for one occupation in one metro, and say so on the slide, because the median can sit well below true market pay when Level I filings dominate.

Should I filter LCA disclosure data to Certified only for a salary benchmark?

Run it both ways and report the delta. Filtering to Certified drops only about 5 to 6 percent of rows given the FY2025 94.4 percent certification rate, but it silently removes Certified-Withdrawn offers that were valid at filing and biases the sample toward completed hires. For a pay benchmark, including Certified-Withdrawn usually gives a truer picture of what employers were offering.

How do I read the H1B wage level in the file?

The PW_WAGE_LEVEL field runs I to IV and maps to roughly the 17th, 34th, 50th, and 67th percentiles of the wage distribution for that SOC and area. Level I signals entry, Level IV signals fully competent or lead. Segment your percentiles by level before you compute them, because mixing Level I entry filings with Level IV lead filings widens the range artificially and makes the median meaningless.

How current is the public LCA data and how do I keep a benchmark reproducible?

OFLC publishes cumulative quarterly and fiscal-year releases on the federal fiscal year, Oct 1 to Sept 30, with roughly a six-week lag; the Q1 FY2026 release through Dec 31 was announced Feb 13. Prevailing-wage levels re-set on a July 1 wage year from BLS OEWS, so the same dollar figure can change level across that date. Record the fiscal year, the release quarter, and the wage year on every benchmark.

Where do I look up the prevailing wage levels to cross-check my median?

Use the FLAG Wage Search for OES prevailing wages by SOC, area, and wage year. FLCDataCenter was discontinued on July 1, 2024, so any workflow pointing there is stale. Pull the four level figures for your target SOC and metro and confirm your computed median lands inside the Level I to Level IV band; a median outside that band signals a SOC, worksite, or unit-conversion error.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next