18 vs 1,938: Sourcing Anthropic's $1B RL Environment Hire
Anthropic may spend $1B+ on RL environments. The best hires are heavy Claude Code users, not ML PhDs. A GitHub-first sourcing playbook.
The September 2026 HN "Who is hiring" thread (posted Sept 1, id 49522897) is stuffed with asks for engineers to design RL benchmarks, red-team model outputs, and write rubrics for agentic coding. TechCrunch reported on Sept 21, 2025 that Anthropic leadership has discussed spending more than $1 billion in a single year on RL environments. If you are still sourcing this hire by pasting "RL Environment Engineer" into LinkedIn, you will run out of candidates by Thursday.
The title search is the trap
The scarcest AI hire of 2026 is the person who converts a messy real-world workflow into a verifiable task a model can be trained against, and the addressable pool is roughly 108x larger when you source by skill and domain instead of by title. In Refolk's index of US professional profiles, only 18 people self-identify with a narrow "RL Engineer" or "Environments Engineer" title. The top employers on that short list are APQX, NVIDIA, Handshake, Nous Research, CoreWeave, and MathWorks. That is a market you exhaust in one afternoon of InMails.
Widen the query to US software and ML engineers who list Reinforcement Learning as a skill and the number jumps to 1,938. Top employers there: Google, Meta, Waymo, Applied Intuition, DoorDash, Red Hat. The mechanism is boring and structural: the job title is less than 18 months old, so a self-identified specialist population cannot exist yet. The talent exists. It is filed under generic SWE and ML titles.
Why Claude Code power users beat ML PhDs
The practitioners building these environments are on record: domain expertise plus expert prompting matters more than ML credentials. Epoch AI's interviews with environment builders capture the shift bluntly. One neolab researcher put it this way: "You don't necessarily need to be an AI researcher, but perhaps a very heavy Claude Code user, a prompt whisperer like Riley Goodside, can be better at figuring out what the frontier is than an AI researcher." An RL environment founder in the same series was tighter: "Domain knowledge and expert level prompting is more important than ML skills for creating tasks."
The economics back it up. Superannotate's writeup of the Epoch survey pegs a Slack clone at $300K-plus per environment. Most of that budget goes into fidelity, not gradient math. Someone who has lived inside Slack, Salesforce, Stripe, or Bloomberg every day for five years knows which edge cases matter. An ML researcher does not.
Anthropic's own Code RL job description reads like a tooling role, not a research role. It asks for people who have "built coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling." And the team openly credits its work: "We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6."
Environments now replicate Slack, Gmail, Salesforce, and Stripe. The best hires are power users of those products who can code.
The numbers behind the RL environments market
Anthropic's rumored $1B/year spend sits on top of a vendor market already doing about $8.5B in annual revenue, and it concentrates fast. The July 2026 tally at aitraining.jobs counts more than 50 companies selling data and RL environments to frontier labs, with a combined valuation around $100B. More than 75 percent of that revenue sits with four players: Scale, Surge, Mercor, and Handshake.
| Segment | US count | Source | Derived |
|---|---|---|---|
| Narrow "RL / Environments Engineer" title | 18 | Refolk index | Baseline |
| SWE / ML with Reinforcement Learning skill | 1,938 | Refolk index | ~108x the narrow pool |
| Vendors selling data + RL envs to labs | 50+ | aitraining.jobs, July 2026 | ~$8.5B revenue, ~$100B valuation |
| Top-4 share of vendor revenue | >75% | aitraining.jobs | Long tail ~25% across 46+ vendors |
| Anthropic reported annual env spend | >$1B | TechCrunch / The Information | ~5% of OpenAI's 2026 R&D compute |
| Cost per high-fidelity env (Slack clone) | $300K+ | Epoch AI via Superannotate | ~3,300 envs/year at $1B ceiling |
Two takeaways for a sourcer. First, the four-vendor concentration means Scale, Surge, Mercor, and Handshake are your densest talent pools for people who have actually shipped an environment. Handshake shows up independently in Refolk's employer list too, which is a cross-source confirmation that these are real, findable populations, not org-chart theater. Second, the interesting arbitrage is the 46-plus smaller vendors and neolabs (Mechanize, Prime Intellect, the ex-Sepal AI folks now inside Mercor after the early-2026 acquisition, ex-Metis engineers now inside DoorDash) where people churn faster and are not yet on every recruiter's list.
A GitHub-first sourcing playbook
Skip LinkedIn's title filter entirely for this role. The high-signal candidates leave a public trail on GitHub because the tooling forces them to. Search these five surfaces in order:
- Prime Intellect's Environments Hub and Bounty Program. The most open on-ramp in the field. Contributors here are literally being paid to build RL environments in public.
- The VoltAgent/awesome-claude-code-subagents repo. More than 100 specialized Claude Code subagents, including a reinforcement-learning-engineer subagent. Every contributor is self-identifying as a heavy Claude Code user applying it to agentic work.
- Repos containing a
.claude/directory, aCLAUDE.mdfile, or subagent configs. GitHub code search surfaces these fast. These are the "prompt whisperers" Epoch AI is describing. - Codeforces and AtCoder top ratings, cross-referenced with GitHub activity. Cognition staffed the Devin team heavily from competitive programmers. It is a repeatable pattern for anyone hiring coding-agent environment builders.
- Eval harness and sandbox repo contributors. Anyone with commits to an eval harness or code-execution sandbox is already doing the work in Anthropic's JD.
This is exactly the friction Refolk is built for: you describe the person in plain English ("US-based engineer who contributes to Prime Intellect Environments Hub, uses Claude Code heavily on GitHub, and has domain depth in fintech or CRM SaaS") and get a ranked shortlist across GitHub, LinkedIn, and the open web. Boolean strings cannot express "Claude Code power user with Salesforce depth." A plain-English query can.
Treat red teamers, eval engineers, and rubric authors as one pool
The September 2026 HN thread bundles three job clusters into one ask, and you should source them as one candidate pool. That thread contains language like: "our current work would especially benefit from ai-forward senior engineering expertise that can design RL benchmarks, red-team model outputs, and create rubrics to evaluate and enhance the agentic coding capabilities of various frontier models."
Read that sentence carefully. It is three job families collapsed into one requisition:
- RL benchmark design. Environment engineering proper.
- Red-team model outputs. AI red team sourcing territory, historically a security-adjacent function.
- Rubric authoring. Evals, traditionally an ML research task.
The same thread includes an ask from a self-funded independent AI lab in Zurich exploring ARC-AGI3, LLM/RL reasoning, and quant, with post-training led by ex-swissAI Apertus engineers. Smaller labs are fishing this exact pond, which means the candidates you surface for one Anthropic req will convert for four other clients.
Searching all three keyword clusters (RL environments, AI red team, eval rubric) compounds your reachable set. If you were only running an "RL environment engineer" search, you missed the red teamers and the rubric authors. They are the same people this year.
Where to actually look, by employer
Concentrate outreach on the employers where environment work is already happening at scale, then work outward. From Refolk's index and the aitraining.jobs vendor tally, the densest pockets of shippable talent sit in three tiers:
- Tier 1, big-4 vendors: Scale, Surge, Mercor (now including ex-Sepal AI), Handshake. High density, high recruiter competition, but the resumes are pre-qualified.
- Tier 2, neolabs and smaller vendors: Mechanize, Prime Intellect, Turing, Micro1, Invisible Technologies, Alignerr, ex-Metis engineers now inside DoorDash. Lower recruiter density, higher churn, best arbitrage.
- Tier 3, adjacent skill pools: Google, Meta, Waymo, Applied Intuition, Red Hat, NVIDIA, CoreWeave, Nous Research. This is where the 1,938-person RL-skill pool lives under generic SWE titles.
If your target is Anthropic Code RL specifically, tier 3 is the sleeper. Someone who spent four years at Applied Intuition writing simulation harnesses for AV stacks has done, in substance, the exact work Anthropic's JD describes. They just do not carry the title.
What to write in the first message
Open with the specific repo or environment you found them through, not a title pitch. The people you want are getting sixty InMails a month that all say "exciting opportunity in AI." The ones from Anthropic and OpenAI already got read. Yours will not, unless you demonstrate you actually looked.
Three concrete opening moves that work:
- Reference a specific PR or bounty they shipped on Prime Intellect's Environments Hub, and ask what they wish the environment could measure that it currently cannot.
- Point to a
CLAUDE.mdfile in their repo and ask what subagent pattern they landed on after iteration. Prompt whisperers love talking about this. - If they contribute to a red team or eval repo, lead with the specific failure mode their commits target. This signals you understand the work is adversarial, not academic.
FAQ
Is an ML PhD a disqualifier for RL environment engineer roles?
No, but it is no longer a prerequisite and, in some cases, it is a weaker signal than shipping product with Claude Code. Epoch AI's interviews with environment builders explicitly name "very heavy Claude Code users" and domain experts as often outperforming AI researchers at figuring out what tasks matter. If your candidate has both a PhD and heavy Claude Code usage, great. If forced to choose, the field is telling you to pick the Claude Code power user with domain depth.
How do I find Claude Code power users on GitHub?
Search for repos containing .claude/ directories, CLAUDE.md files, and Claude subagent configs, then cross-reference contributors against the VoltAgent/awesome-claude-code-subagents repo and Prime Intellect's Environments Hub. These artifacts are only created by people who use Claude Code as a daily driver, which is exactly the signal Epoch AI's interviewees flagged. LinkedIn will not show you any of this.
Should AI red teamers be sourced separately from RL environment engineers?
Treat them as the same candidate pool for 2026 hiring. The September 2026 HN "Who is hiring" thread bundles "design RL benchmarks, red-team model outputs, and create rubrics" into single requisitions, and the underlying skill set (adversarial thinking plus rubric design plus tooling) is one profile with three names. Running your searches across all three keyword clusters compounds your reachable candidates and matches how frontier labs and neolabs are actually writing the job.
Why is the "RL Environment Engineer" title so thin in candidate databases?
The title is under 18 months old, so the self-identified population is structurally tiny. Refolk's index shows only 18 US profiles carrying a narrow environment/RL engineer title versus 1,938 US software and ML engineers who list Reinforcement Learning as a skill under generic titles at Google, Meta, Waymo, Applied Intuition, DoorDash, and Red Hat. Source by skill plus domain, not by title, and your addressable pool grows about 108x overnight.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.