The RLHF Poach: 8,496 US Experts, 4 Bidders, One 9.7x Pay Gap
Mercor, Surge, Scale, and Handshake are bidding against frontier labs for the same PhDs. The sourcing playbook, with numbers from Refolk's index.
The AI training-data market flipped in 18 months. The buyer used to want 10,000 crowd labelers at $15/hr; now the same buyer wants one board-certified radiologist at $300/hr, and four platforms plus every frontier lab are chasing the same person. If you are sourcing for RLHF, RLVR, or agent-environment work in late 2026, the game is no longer volume. It is credential verification speed against a pool small enough to name.
The market flipped from crowd labelers to credentialed experts
AI training data hiring is now a domain-expert market: STEM PhDs, practicing physicians, senior software architects, and licensed attorneys grade model outputs and build graded environments, and they cost 5 to 10 times what a generalist annotator did in 2024. The Information documented the shift in April 2026, and every major platform (Handshake AI, Mercor, Surge, Scale) has repositioned around it.
The mechanism is reinforcement learning from verifiable rewards. RLVR is a training method where the reward signal comes from graded, checkable tasks (a math proof, a differential diagnosis, a passing unit test) rather than a thumbs-up from a crowd worker. Grading those tasks requires the same credentials as producing them. That single change collapsed the labor supply from millions of clickworkers to a headcount you can actually enumerate.
The pay ladder tells the story. National average AI trainer pay sits at $31/hr as of early 2026 (Mercor's own data). Mercor's per-role ceiling for practicing physicians, M&A attorneys, and senior software architects sits at $200+/hr. Handshake AI is listing radiology and ophthalmology gigs at $300/hr. That is a 9.7x premium against the generalist floor, and the top of the ladder is where all four platforms are now fighting.
Mercor, Surge, Scale, and Handshake are now the same business
The four platforms have converged on one product: verified domain experts, priced hourly, delivered fast. What separates them now is credential graph and vetting speed, not category.
- Mercor (founded 2023 by Thiel Fellows Brendan Foody, Adarsh Hiremath, Surya Midha) pulls LinkedIn, GitHub, and Google Scholar signals within minutes of signup and books an AI-led video interview before a competitor gets to the same profile. It acquired Sepal AI in February 2026 (expert-graded benchmarks) and Deeptune in July 2026 (RL environments for agents).
- Surge AI (Edwin Chen, ex-Google) advertises Fields Medal mathematicians and Supreme Court litigators; partners include Meta, OpenAI, Google, Microsoft, and Anthropic. A May 2025 Clarkson Law Firm class action over 1099 misclassification is still live.
- Scale AI / Outlier took a $14.3B check from Meta in June 2025 for a 49% stake. Founder Alexandr Wang moved to Meta's superintelligence lab, which is precisely why the other three platforms accelerated in the second half of 2025.
- Handshake AI turned a 12-year-old campus recruiting network into a specialist supply engine. Demand tripled after the Meta-Scale news; the AI arm passed a $150M run rate by November 2025, exceeding its original campus business.
Here is the comparable snapshot:
| Segment | Figure | Source |
|---|---|---|
| US pros in AI trainer / research engineer / RLHF titles | 8,496 | Refolk index |
| Frontier-lab direct-employ share (DeepMind + Meta + Anthropic + OpenAI) | 36% of sampled pool | Refolk index |
| Generalist AI trainer average pay | $31/hr | mercor.com |
| Mercor domain-expert ceiling | $200+/hr | aitrainingjobsfinder.com |
| Handshake AI specialist ceiling | $300+/hr | remowork.life |
| Premium vs generalist | ~9.7x | Derived |
| Handshake AI vs Mercor average listing pay | $108 vs $91/hr | aigigjobs.com |
| Handshake AI vs Mercor active listings | 95 vs 241 | aigigjobs.com |
Note the two lines that matter most for a sourcer. Handshake pays more per listing but has fewer of them; Mercor runs higher volume at a lower midpoint. And in the row above both, frontier labs already directly employ more than a third of the sampled pool. The vendor layer is being squeezed from both sides.
The real US pool is 8,496 people, and the labs already have a third
Refolk's index counts 8,496 US professionals currently holding titles matching AI Trainer, Research Engineer, or RLHF specialist patterns. The top employers are Google DeepMind, Meta, Sesame, ElevenLabs, Anthropic, and OpenAI, and they account for roughly 36% of a sampled slice. The "vendor market" you think you are competing in is a minority of the real pool.
Two consequences follow.
- If you are a frontier lab, the highest-signal hires are already sitting inside your closest three competitors under a "Research Engineer" title, not on a Mercor waitlist.
- If you are a vendor or a startup routing through vendors, you are fishing downstream of the labs. The people you want most were poached to W2 with equity six months ago.
The public "RLHF specialist" label is even thinner than it looks. Refolk's index shows only 8 clearly tagged current RLHF profiles across the US, UK, and India combined, with employers like Turing, Outlier, Invisible Technologies, ByteDance, DataAnnotation, and Capgemini. Everyone else doing the work has a different title (Research Scientist, Applied Scientist, Evaluations Engineer, or nothing at all). This is exactly the gap Refolk closes for sourcers: describe the work in plain English (grades math proofs for a frontier lab, has RL publications, based in the Bay Area) and get the right people even when their job title doesn't match your search string.
The bottleneck isn't recruiting, it's vetting speed
The scarce resource in RLHF talent sourcing is not candidates. It is the minutes between "this person exists" and "this person is credentialed and booked." Mercor's structural advantage is scraping LinkedIn, GitHub, and Google Scholar within minutes of signup, then running an AI-led video interview before a competitor sends a first message.
That compression matters because domain experts go effectively exclusive once they accept a project. A cardiologist grading 40 hours a month of clinical case data is not also going to onboard at three other platforms next week. Whoever confirms the credential first books the calendar.
Whoever confirms the credential first books the calendar. Everyone else sends a second message to a busy inbox.
The playbook for a founder or in-house recruiter that wants to compete:
- Pre-join the credential graph before you need it. Do not start from a job title. Start from publications, GitHub commits to specific repos (verifiable-rewards, agent-eval, eval harnesses), and license lookups (state bar, ABMS).
- Filter on intersections, not stacks. A JD plus a CS master's should be sourced from the JD side because the pool is smaller. A PhD plus Mandarin fluency should be sourced from the intersection because that is where the top rate lives.
- Reach out with the specific task, not the platform. "Grade 20 differential diagnoses per week at $260/hr, 1099, start Monday" beats "we're hiring domain experts."
- Own the contract structure. All 14 leading platforms use 1099 arrangements with no health insurance, PTO, or employer payroll tax. A W2 offer with equity is a poach weapon.
Environment engineer is the next scarcity, not RLHF rater
The scarce role in 2026 is the person who builds the graded task, not the person who grades outputs. Mercor's July 2026 Deeptune acquisition, Scale's and Surge's own environment products, and the rise of Mechanize, Fleet AI, HUD, and Prime Intellect all point the same direction: RL environments are eating RLHF.
An RL environment engineer is a domain expert who can author an executable, gradeable task (a coding sandbox with hidden test cases, a medical case simulator with a defensible ground-truth diagnosis, a legal-drafting harness with judgeable output). The credential stack is unusual: a working software engineer who is also a domain practitioner, or a domain PhD who ships code.
This population is even smaller than the RLHF rater pool, and it is scattered across titles that don't mention "environment." Candidates carry labels like "Applied Research Engineer," "Simulation Engineer," "Evaluations Lead," or "Staff Engineer, Agents." If you are hiring for this role in Q4 2026, do not search by title. Search by the intersection: contributed to at least one open-source eval harness, holds a domain credential, and lives in a metro your competitor already saturated.
The 1099 model is the poach opportunity
The vendor labor model is structurally weak against a W2 offer, which is the single biggest reason frontier labs are disintermediating Mercor, Surge, Scale, and Handshake right now. No health insurance, no PTO, no employer tax contribution, and spiky project engagements mean the top decile on every platform is one recruiter conversation away from leaving.
If you run engineering hiring at a lab or a well-capitalized startup, the near-term play is mechanical:
- Pull the vendor's public contributor names and profiles (many are semi-public via case studies, LinkedIn work history, and platform leaderboards).
- Rank by credential rarity, not credential count.
- Approach with a W2 offer that includes equity, benefits, and a specific 6 to 12 month scope tied to a training run.
The upside for the candidate is obvious. The downside for the vendor is that its top rater roster hollows out every quarter, which is why the vendor category is now expanding into environments, benchmarks, and tooling (Sepal, Deeptune, HUD-style products) instead of defending the rater layer.
What to do this week
Stop competing on outreach volume and start competing on time-to-verified-credential. Every hour you spend cross-checking a Google Scholar page against a LinkedIn resume against a GitHub profile is an hour Mercor's pipeline finished 20 minutes ago.
- If you are a founder building an RL environments company, name the intersection you actually need (ophthalmology board certification plus Python, for example) and source from the smaller side.
- If you are an in-house recruiter at a lab, treat vendor contributor rosters as a warm inbound list and lead with W2 and equity.
- If you are a technical recruiter at a vendor, defend the top decile with retention bonuses tied to the training run, not the calendar year.
A joined credential graph across GitHub, LinkedIn, and the open web is the leverage point, and Refolk is built for exactly this shape of query: ask in plain English, get the right people, verify in the same session.
FAQ
Who are the four platforms competing for RLHF talent in 2026?
Mercor, Surge AI, Scale AI (via Outlier), and Handshake AI. Mercor and Handshake are the fastest-growing on the specialist side; Surge has the deepest credential roster (Fields Medalists, Supreme Court litigators); Scale is partially inside Meta after the June 2025 $14.3B, 49% deal. All four are being disintermediated to varying degrees by frontier labs hiring credentialed experts directly to W2 roles.
What does a domain expert AI annotator actually earn?
The floor is around $60/hr on Mercor, the mode for accepted contributors is $90 to $120/hr, and the ceiling reaches $200+/hr on Mercor and $300+/hr on Handshake AI for radiologists, ophthalmologists, M&A attorneys, senior software architects, and research-grade math PhDs. The national generalist AI trainer average is $31/hr, so credentialed specialists command roughly a 9.7x premium at the top of the ladder.
Why is "environment engineer" harder to hire than an RLHF rater?
Because the role requires two credentials in one person: shipping-quality engineering skills plus a real domain credential (MD, JD, physics PhD, or equivalent) needed to define a defensible ground truth. The population is scattered across at least a dozen titles that don't include the word "environment," so title-based Boolean searches miss most of them. Source by the intersection of publications, repo contributions, and license.
How should a small team compete against Mercor's vetting speed?
Do not try to match Mercor on volume. Match it on credential-graph readiness before you have a specific role. Keep a live shortlist of the 50 to 200 credentialed experts in your domain, updated with publications and repo activity, so that when a project opens you send a warm, specific message the same day. A plain-English sourcing tool that returns ranked people across GitHub, LinkedIn, and Google Scholar makes this maintainable for a team of one or two.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.