Anthropic's $280/Hour Marlin Bench: Sourcing 1,000 Hidden Engineers
Anthropic pays 1,000 engineers $280 per task via Snorkel to grade Claude Code. Here is how to source the RLHF shadow bench at Snorkel, Scale, and Mercor.
Business Insider just pulled the curtain back on Project Marlin: Anthropic is paying roughly 1,000 external software engineers about $280 for each one-hour task grading Claude Code, routed through Snorkel AI. Top experts clear $3,000 a week. If you source senior engineers, that leak is not a story about AI training economics. It is a map to a pre-vetted, income-hungry bench that does not show up in your LinkedIn boolean.
What Project Marlin is, and why sourcers should care
Project Marlin is Anthropic's internal name at Snorkel AI for a roughly 1,000-person effort where senior software engineers write prompts, review code, and A/B test Claude Code outputs to fine-tune the model to mimic a professional developer. The pay is about $280 per one-hour task, with top experts earning $3,000+ per week, according to Business Insider.
That rate matters because it is a filter. Snorkel's public legal contractor roles pay $10 to $100 per high-quality task. Scale AI and Mercor pay software engineers up to $110 an hour. Marlin pays roughly 2.5x the going rate for the same category of work. Every engineer on Marlin has already survived Snorkel's own quality bar. If you can identify one, someone else has already done the technical vetting for you.
The catch: contractors sign vendor NDAs, they usually already have a day job, and most do not put "AI Trainer" on their profile. The bench is large and deliberately quiet. Contractors A/B test outputs from two models and pick the preferred code, but they do not even know which model versions they are evaluating. Vendor secrecy is why LinkedIn signal is thin.
The RLHF contractor economics in one table
The shape of the market, pulled from Business Insider, Mercor's Experts page, and Refolk's index of professional profiles: the premium at the top and the volume in the middle are what make this a sourcing target rather than a curiosity.
| Figure | Value | Source |
|---|---|---|
| Snorkel "Marlin" engineer rate | $280 / 1-hour task, $3,000+/week top | Business Insider |
| Scale AI / Mercor engineer rate | Up to $110/hour | Business Insider |
| Marlin premium vs peers | ~2.5x | Derived |
| Mercor avg contracted rate (all experts) | $114/hour | Mercor Experts page |
| Mercor roles created | 349.3K | Mercor Experts page |
| Mercor daily contractor payouts | $4M+ | Mercor Experts page |
| U.S. profiles with RLHF/AI-trainer signal | 2,496 | Refolk's index |
| Share of top-25 who self-title "AI Trainer" | ~40% (10/25) | Refolk's index sample |
Two numbers do the heavy lifting. $4M in daily payouts at Mercor implies a working population in the tens of thousands. And the 2,496 U.S. profiles Refolk's index surfaces with RLHF-adjacent language is a floor: most moonlighters actively hide the work.
Individual Mercor postings already break past FAANG hourly equivalents. A Cybersecurity Research Expert role posts at $200 to $250 an hour. The Machine Learning Engineer Talent Network posts $70 to $250. Senior engineers with 5+ years of experience earn at the top of the range.
The vendors are the moat, not the labs
If you want to source this pool, forget Anthropic, OpenAI, and Google. The sourceable names sit at the middleware layer: Snorkel AI, Scale AI (and its contractor brand Outlier), Mercor, Alignerr, Handshake AI, and Invisible Technologies. That is a finite, nameable list of six to ten companies.
The ones that matter most for engineer-grade RLHF:
- Snorkel AI: Series D in May 2025 at a $1.3B valuation. Spun out of Stanford, originally focused on reducing manual data annotation, now running Marlin for Anthropic. Clients include Google, Mistral, Anthropic.
- Scale AI / Outlier: Outlier is the contractor-facing brand and surfaces as a top current employer for RLHF-adjacent profiles in Refolk's index.
- Mercor: $350M raise at a $10B valuation. Founded by Bay Area high school friends Brendan Foody, Adarsh Hiremath, and Surya Midha, who dropped out of college. Clients include six of the Magnificent Seven, OpenAI, and Anthropic.
- Alignerr and Handshake AI: Emerging RLHF employers that consistently appear in Refolk's index for engineers doing this work outside the Big Three.
Building a sourcing search around the labs will return recruiters, PMs, and researchers. Building it around these vendors returns the actual bench.
Title search is the wrong tool
Only about 40% of the top 25 profiles doing RLHF work in Refolk's index publicly title themselves "AI Trainer." The other 60% call themselves "Software Engineer," "Founder," "Sr. Technical Program Manager," or "Data Annotator," or they omit the work entirely. Vendor NDAs and dual-employment fears push contractors to bury the signal.
A Sales Navigator boolean on "AI Trainer" OR "RLHF" misses the majority of the pool. The signal you actually want is a combination of:
- Employer mentions in About sections, past experience, or descriptions: Snorkel, Scale, Outlier, Mercor, Alignerr, Handshake AI, Invisible Technologies.
- Skill stack that matches the paying work: Python, Rust, C++, TypeScript. Mercor lists these as the most-requested languages for engineer contractors.
- Task-type keywords: "code generation evaluation," "gold-standard code solutions," "bug detection," "architecture review." These are the highest-demand categories on Mercor.
- Experience floor: 5+ years, since that is where the top of the rate range concentrates.
- Side-project cues: a personal GitHub with recent activity plus a day-job title at a non-AI company is often the tell.
You will not get all five signals in one field. This is a stitched profile: a LinkedIn About paragraph plus a GitHub commit history plus a personal site. Which is the gap Refolk closes. Describe the person in plain English ("senior backend engineers in the US who mention Outlier, Snorkel, or Mercor anywhere, with recent Rust or TypeScript activity") and Refolk stitches LinkedIn, GitHub, and the open web into a ranked shortlist. Title filters do not need to be the entry point.
The Mercor rehire event is a live liquidity window
Roughly a week after Mercor's $10B raise, Forbes reported the company canned an AI project involving thousands of contractors, then offered to rehire many of them hours later at a lower hourly rate. That is a live catalyst. Right now, in Q4 2025 and heading into early 2026, a large fraction of this bench is actively unhappy and open to conversations.
Treat it like a post-layoff hiring window, with one important twist: these people are not unemployed. They kept their day jobs. What they lost was the second income, the interesting work, or both. The reachout writes itself:
- Acknowledge the vendor by name. Specificity earns replies.
- Do not pitch a full-time role first. Pitch a conversation about interesting work at a lab or infra company.
- If your client is an AI infra company, lean on the fact that the candidate has already been grading frontier models. Most FAANG engineers cannot claim that resume line.
The window closes the moment Mercor stabilizes rates or a competitor absorbs the bench. Weeks, not quarters.
The engineers grading Claude Code are not open to work. They are open to a better second job, which is a very different conversation.
Why the moonlighter beats the "open to work" cohort
The moonlighter inverts the usual passive-candidate calculus. They are currently employed, technically active every week rather than just in interview loops, and cash-motivated in a way that makes response rates jump. Compared to the standard "open to work" pool, which skews toward recent layoffs and career gaps, the RLHF moonlighter looks like this:
- Has a current W-2 job, usually at a mid-tier tech company.
- Ships code on the side for pay, not for portfolio.
- Has been evaluated by Snorkel, Scale, or Mercor's internal quality systems and cleared a bar you did not have to run.
- Understands frontier model behavior at a level most FAANG engineers do not.
- Is comfortable with async, remote, contract-shaped work, which matters for the AI infra companies hiring right now.
There is a narrative hook worth using in outbound too. Boris Cherny, head of Claude Code at Anthropic, said in January he had not handwritten a line of code in over two months, with Claude submitting 22 pull requests in one day. Kate Jensen, Anthropic's head of revenue, has said fully unlocking Claude's potential requires new evaluation methods incorporating domain experts and human feedback. The people doing that evaluation are the exact bench described here. They know they are training the tool that is changing their day job. That awareness is why they respond to specific, informed outreach and ignore generic pitches.
Where these engineers show up as full-time hires
The 2,496 U.S. profiles Refolk's index surfaces with RLHF-adjacent language cluster in a handful of employers today. Outlier, Alignerr, and Handshake AI surface most often as current or recent employers. That does not mean these people work full-time at those vendors. It means the vendor is where the RLHF signal was strong enough to make it into their profile.
Real full-time destinations for this cohort tend to be:
- AI infra companies hiring evaluation engineers, model behavior researchers, and applied-AI engineers.
- Enterprise AI startups that need engineers who understand both product code and model output quality.
- The frontier labs themselves, when they hire "member of technical staff" roles that require model intuition.
If you are recruiting for any of the above, the brief is: senior engineers, current employer is not a frontier lab, but their skill and keyword footprint includes an RLHF vendor and a modern typed-language stack. That query is awkward to write in Boolean and natural to write as a sentence.
The reachout that lands
Name the vendor, not the model. Compare:
- Weak: "I saw you have AI training experience..."
- Strong: "I saw an Outlier mention in your profile and recent Rust commits. A client of mine is staffing model-evaluation engineers and specifically wants people who have graded frontier code outputs. Worth 15 minutes?"
The specificity signals three things at once: you know the vendor, you looked at their GitHub, and you understand the work. Because the candidate is already earning $200 to $280 an hour on the side, a full-time offer needs to clear that math, or a hybrid contract-to-perm structure has to be on the table. Bring comp benchmarks to the first call.
FAQ
How do I know if a candidate is actually on Project Marlin?
You usually cannot confirm Marlin specifically because of Snorkel's NDA. What you can confirm is Snorkel AI as a current or recent employer, plus a software engineering skill stack (Python, Rust, C++, TypeScript), plus the 5+ years of experience that qualifies for the top of the rate card. If all three line up, the probability they are on Marlin or a comparable project is high enough to justify the reachout. Do not name Marlin in the first message; name Snorkel.
Is it ethical to poach engineers off RLHF platforms?
Yes, with one boundary. You are not asking them to breach an NDA or share client information. You are offering them a job. The contractors are 1099s working nights and weekends for cash; they are the definition of a sourceable candidate. What you should not do is ask them to describe specific projects, name lab clients, or share evaluation prompts. Recruit the person, not the IP.
What if my ATS or CRM already has some of these people?
It almost certainly does, and they are mis-tagged. RLHF work rarely gets logged as a discrete skill in an ATS. Run a full-text search across resume and notes fields for Snorkel, Scale, Outlier, Mercor, Alignerr, Handshake AI, and Invisible Technologies. You will find candidates you already sourced two years ago who are now sitting on 18 months of frontier-model evaluation experience. Re-source your own database before you buy more seats.
Will the $280 Marlin rate hold, or is this a spike?
Treat it as a spike that sets a new floor for evaluation-grade engineers, not a permanent $280/hour market. Anthropic is paying a premium because Claude Code is strategic and quality data is the bottleneck. When the next lab launches a coding model, the same rates will get bid up again. What is durable is the pool: tens of thousands of senior engineers who have now proven they can and will do this work. That pool is the sourcing asset, not the current rate card.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.