16 US Engineers Hold This Title. Mechanize Pays $400K.
Anthropic will spend $1B on RL environments. Only 16 US engineers hold the exact title. Here is where the other 4,486 convertible candidates sit.
In September 2025, The Information reported Anthropic had discussed spending over $1 billion on reinforcement learning environments in the following year. By early 2026, a directory of 38 vendors had appeared to sell those environments, and Mechanize (founded April 2025) was quoting software engineers up to $500K to build them. The problem for everyone else: the exact job title is roughly a year old, and almost nobody uses it yet.
What an RL environment engineer actually does
An RL environment engineer builds the simulated worlds - a Salesforce clone, a terminal, a "digital office" - that a frontier model trains inside, plus the verifiers that decide whether the model's actions earned a reward. The role sits between simulation engineering, backend, and adversarial security. It is not an ML research role.
Prime Intellect's Will Brown gave the clearest public walkthrough at the AI Engineer conference in December 2025: authoring the environment, wiring the reward signal, and scaling it up to train frontier coding models. If you want to understand what you are hiring for, that talk is the fastest way in.
The mechanical requirements of the job:
- Author high-fidelity simulations of real software (Slack, Salesforce, Excel, an IDE, a terminal).
- Write verifiers and reward functions that survive an adversarial optimizer trying to cheat them.
- Ship agent-eval harnesses (Terminal-Bench, IDE-Bench, EnterpriseBench, VADER) that produce clean training signal.
- Instrument trajectories at scale so RL training runs actually converge.
Notice what is missing: there is no line item for "publish at NeurIPS." This is a systems and security discipline dressed up in ML clothing.
Who is actually buying, and how fast
Roughly three dozen companies now sell RL environments to frontier labs, and the top of the market is consolidating on a monthly cadence. The RL List directory tracks 38 vendors by funding, team, customers, SOC 2 posture, and focus area. An industry tally circulated in July 2026 counted 50-plus vendors selling data and RL environments to labs at roughly $8.5 billion in combined annual revenue.
The names a sourcer needs on a whiteboard today:
| Company | Role in the market | Signal for sourcers |
|---|---|---|
| Mechanize | Pure-play SF shop, ex-Epoch founders | Posted band $300K to $400K, top-of-band $500K |
| Mercor | Consolidator; bought Sepal AI (Feb 2026) and Deeptune (July 2026) | $10B valuation, ~$2B annualized revenue by June 2026 |
| Prime Intellect | Open-source lab, Karpathy and Founders Fund backed | Community entry point via Will Brown's talks |
| Surge AI | Bootstrapped incumbent, ~$1.2B 2024 revenue | Ships EnterpriseBench; frontier models solve only ~30% |
| Fleet AI | Enterprise-software simulations, open-source Harbor | Python SDK plus platform API |
| AfterQuery (YC W25) | Expert human data and computer-use trajectories | Publishes Terminal-Bench, VADER, FinanceQA, IDE-Bench |
| Invisible Technologies | RL gyms and "mirror worlds" | Raised ~$100M in September 2025 |
| Deeptune (now Mercor) | RL "training gyms" for computer-use and code | $43M Series A led by a16z, March 2026 |
The mechanism worth understanding: Scale AI's static-labeling revenue is roughly halving in 2026 while environments are the growth area. Every account executive, program manager, and staff engineer at a static-labeling vendor is a warm candidate for an environments team, whether they know it yet or not.
The $500K number is misleading. Here is the real band.
Mechanize's own careers page currently lists $300K to $400K for the software engineer role building RL environments for frontier AI labs. The $500K figure in TechCrunch reflects a top-of-band or negotiated total-comp package, not the posted range.
This matters for outreach. Recruiters pitching candidates with "market is $500K" will get called out on the first screen. The credible pitch is:
- Posted band: $300K to $400K base at Mechanize.
- Top-of-band, negotiated: reported up to $500K.
- Comparable hourly contractor ceiling at Scale or Surge: well below the salaried total on a full-year basis.
- Equity: Mechanize has raised ~$9.1M at a ~$500M valuation per Tracxn, so early-hire equity is meaningful but not lottery-ticket.
The interesting arbitrage is that environment spend is tiny relative to compute. Epoch AI estimates OpenAI's 2026 R&D compute at roughly $19 billion, so Anthropic's $1B environment budget is a single-digit percentage of what a leading lab spends on GPUs. A comparatively small spend on the humans who build environments decides whether an enormous compute spend produces a working agent or an expensive disappointment. That ratio is the entire argument for why this hire is underpriced.
The bottleneck is not finding capable people. It is that everyone is searching on a title 99.6% of the qualified pool has never used.
The exact-title pool in the US is 16 people
In Refolk's index of professional profiles, only 16 people in the United States currently list the exact title "Reinforcement Learning Engineer" or "RL Engineer." The pool of senior engineers who list "Reinforcement Learning" as a skill is 4,486. That is a ~280x gap between the "already-doing-it" pool and the "convertible in a quarter" pool.
| Segment | US count | Source |
|---|---|---|
| Exact title "Reinforcement Learning Engineer" or "RL Engineer" | 16 | Refolk index, title match |
| "Reinforcement Learning" as a skill, Senior/Manager/Director | 4,486 | Refolk index, skill match |
| Ratio of skilled pool to exact-titled pool | ~280x | Derived |
| Top employers of exact-titled RL engineers | APQX (5), NVIDIA (2), Handshake (2), Tesla, Nous Research, CoreWeave, MathWorks | Refolk index |
| Top employers of senior RL-skilled talent | OpenAI, Mercor, Amazon, Workday, Analog Devices, South Park Commons | Refolk index |
| Geographic concentration (senior RL skill) | SF/Bay Area (8 of top 10 city buckets), Seattle, LA | Refolk index |
The takeaway: if you are running a boolean search on "RL Environment Engineer" in LinkedIn Recruiter, you are competing with every other founder for a couple of dozen people, most of whom already work at NVIDIA, Tesla, or a frontier lab and are not moving for $400K. The real hire is a conversion play.
Sourcing the convertible pool: five signals that beat the title
Stop sourcing on title. Start sourcing on the signals that predict who can ship an RL environment in their first 90 days.
- Adversarial software engineering. CTF players, ex-security engineers, red-teamers. What makes the role hard is that every part has to survive contact with a capable, adversarial optimizer, so the engineer has to think like an attacker about their own reward function.
- High-fidelity simulation work. Anyone who has built a Salesforce or Slack clone for testing, browser-automation frameworks at scale, or terminal sandboxes.
- Eval and benchmark authorship. Contributors to Terminal-Bench, IDE-Bench, VADER, FinanceQA, or EnterpriseBench, or authors of internal agentic evals at any lab.
- RL as a listed skill plus a shipped product. The 4,486-person pool shrinks fast once you filter for shipped RL work versus a coursework mention.
- Verifier and reward-shaping vocabulary. GitHub commits or blog posts mentioning "reward hacking," "verifier," "grader," "trajectory," or "rollout" in the last 18 months.
This combination is the exact gap Refolk closes for RL environment engineer hiring: you describe the person in plain English ("US-based senior engineer, ships RL or agent evals, has security or CTF background, not currently at a frontier lab") and get a ranked shortlist across GitHub, LinkedIn, and the open web instead of a title-only boolean that returns 16 rows.
Where the candidates actually sit in 2026
The senior RL-skilled pool clusters in six places, and each demands a different first message. SF and the Bay Area dominate (8 of the top 10 city buckets in Refolk's index), followed by Seattle and Los Angeles.
- Frontier-lab alumni networks. Not current staff at OpenAI, Anthropic, or the top labs (they will not move for $400K), but leavers who joined seed-stage startups that are now struggling. Watch second-employer tenure hitting the 12-month mark.
- Mercor's contractor bench. Mercor organizes a network of 30,000+ domain experts. A meaningful slice are engineers doing RL-adjacent contract work who would take a full-time offer.
- Scale AI and Surge AI staff engineers. Static labeling is shrinking. Environments are growing. This is the warmest pool in the market.
- Prime Intellect's community. Contributors to open RL environment repos and attendees of Will Brown's AI Engineer 2025 talk.
- CTF competition rosters. DEF CON CTF, PlaidCTF, Google CTF finalists. Cross-reference with anyone who has published on agent evals.
- Robotics and self-driving refugees. The 2024 to 2025 wave of autonomy layoffs produced engineers who have written reward functions for actual robots. That is the closest adjacent skill on the planet.
For the last four buckets, LinkedIn Recruiter is largely useless. The signal lives in GitHub commits, conference speaker lists, and Discord handles. Sourcing AI training engineers well in 2026 means pulling from all three surfaces in a single query, which is what Refolk is built to do.
Front-run the next acquihire
Mercor bought Sepal AI on February 6, 2026 and announced Deeptune on July 9, 2026: two RL-environment acquisitions in five months. At current pace, one of the 38 vendors in the RL List directory gets absorbed every eight to ten weeks.
The Mercor and Sepal AI template is now the standard playbook: a Series A vendor with a strong lab customer gets acquired, and the senior engineers either vest-and-rest or leave inside 90 days. If you are hiring agentic eval engineers, the highest-leverage move is to identify the next three to five acquisition targets and warm up their senior engineers now, pre-deal.
Signals that a vendor is next:
- Series A closed in 2025 or early 2026, no Series B in sight.
- One or two frontier-lab logos as anchor customers (concentration risk).
- Founders who came from a lab rather than an enterprise-software background.
- Recent SOC 2 push, which usually precedes a strategic conversation.
Sepal AI (YC S24, 20k+ domain experts) fit every one of those before Mercor bought it. So does a meaningful share of the directory today.
FAQ
Is "RL environment engineer" a real job title yet, or a marketing label?
It is a real job title at roughly a dozen companies and a marketing label everywhere else. Mechanize, Mercor, Surge, Fleet, AfterQuery, and Prime Intellect all post the role under some variant of the name. But Refolk's index shows only 16 US engineers use the exact title on their profile today, versus 4,486 senior engineers who list reinforcement learning as a skill. Hire on the skill signal, not the title.
What is the real salary band for reinforcement learning environments jobs?
Mechanize's public posting is $300K to $400K for the software engineer role. TechCrunch reported $500K as a top-of-band or negotiated figure, which is credible but not the posted range. Hourly contractor rates at Scale or Surge do not reach these totals on a full-year basis, which is the whole reason these companies are pulling contractors into full-time seats. Equity at an early-stage vendor like Mechanize (~$500M valuation on ~$9.1M raised per Tracxn) is meaningful but not lottery-ticket.
Why is Mercor buying so many RL environment companies?
Mercor reached roughly $2B in annualized gross revenue by June 2026 on the back of expert-network and evaluation work, and RL environments are the fastest-growing line item in frontier-lab budgets. Buying Sepal AI (February 2026) and Deeptune (July 2026) gave Mercor both training-gym infrastructure and a bench of engineers who had already shipped for labs. Treat every Series A vendor in the space as a potential source of senior candidates within 90 days of any deal.
Where do I find agentic eval engineers who are not already at a frontier lab?
The highest-yield pools are: recent leavers from OpenAI, Anthropic, and DeepMind whose second employer has stalled; Scale AI and Surge AI staff engineers hit by the static-labeling contraction; contributors to open benchmarks like Terminal-Bench, VADER, and IDE-Bench; and robotics or self-driving engineers from the 2024 to 2025 autonomy layoffs who have written reward functions in production. LinkedIn alone will miss most of them. Combining GitHub, LinkedIn, and conference-speaker data in one query is the unlock.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.