Bespoke Labs Raised $40M for RL Environments. The Titled Pool Is 48.
Bespoke Labs' $40M Series A lands in a category where fewer than 20 US engineers self-identify as evaluation engineers. Here's where the real pool hides.
On July 6, 2026, Bespoke Labs closed a $40M round (Series A of $31.75M led by Wing VC with Mayfield and The House Fund, plus an $8.25M seed led by 8VC with Jeff Dean) to build environments where AI agents can safely learn, test, and improve before production. Anthropic has reportedly told The Information it plans to spend over $1B on RL environments in the next year. The problem for every founder chasing this thesis: the engineers who actually build eval harnesses at scale do not, for the most part, call themselves eval engineers.
How small is the AI evaluation engineer pool, really?
In Refolk's index of professional profiles, only 48 people worldwide currently carry an AI-specific evaluation-engineer title, and the US subset a recruiter can actually pitch to Bespoke is closer to 15 to 25. That is not a typo. One well-funded startup could hire the entire global titled market and still have desks left over.
Here is the breakdown by exact title string:
| Cohort | Count | Scope |
|---|---|---|
| "AI Evaluation Engineer" | 15 | US only |
| "LLM Evaluation Engineer" | 3 | US only |
| "Agent Evaluation Engineer" | 1 | US only |
| All AI-eval titles combined | 48 | Worldwide |
| RLHF listed as a skill | 61 | US only |
| RLHF pool vs titled pool (US) | ~3.2x | Derived |
| Bespoke headcount vs global titled pool | 40 / 48 ≈ 83% | Derived |
The last row is the punchline. Bespoke is roughly 40 people today. If they doubled, they would exceed the entire planet's supply of humans who describe themselves this way on LinkedIn.
Why does the title barely exist in 2026?
The category is younger than the résumé cycle. "Evals engineer" as a distinct role is under 24 months old, university curricula have not caught up, and comp bands at frontier labs still route this work through "Research Engineer" or "Member of Technical Staff." Nobody voluntarily downgrades their title to a name the market has not priced yet.
Three mechanisms are keeping the label suppressed:
- Comp gravity. Research engineer pays more at Anthropic, OpenAI, and Meta than any variant of "evaluation engineer" does. Rational operators keep the higher-status string on their profile.
- Team topology. At most labs, evals live inside post-training or safety, not in a standalone org, so managers do not push a new title through HR.
- Recruiter search behavior. Because nobody searches for the title, nobody optimizes their profile for it, which keeps it rare, which keeps recruiters from searching for it. A perfect deadlock.
The practical consequence: sourcing this category by LinkedIn title filter is malpractice. You will surface a handful of people at Mistral AI, 1eye, Ardoq, Mindrift, and DataAnnotation, and you will miss the 3x larger cohort doing the actual work under a different label.
Where does the real pool live?
The real pool is the 61 US professionals who list RLHF as a skill, plus the maintainers of a small number of open-source eval and environment repos. That is roughly 3x the named pool, and it does not overlap cleanly with the titled cohort.
Two data points from the Refolk index sharpen this:
- Top current employers of the titled cohort skew toward data-labeling companies (Surge AI, DataAnnotation, Mindrift) more than toward frontier labs. Founders think they are competing with OpenAI for this profile. They are competing with Scale.
- Geographic center of gravity for named AI eval engineers is India (Chennai, Pune, Jaipur, Kolhapur) and scattered EU cities (Berlin, Oslo, Nancy). Only one of the top ten regions is SF Bay Area. Bespoke is in Mountain View.
Sourcing this category by title is malpractice. The people doing the work are labeled Research Engineer or MTS.
The correct sourcing move is to invert the search: start with the artifact (a commit to Terminal-Bench, a PR to lm-eval-harness, an environment published to Prime Intellect's Environments Hub) and then walk backward to the person, regardless of what their title happens to say this quarter. That inversion is the exact gap Refolk closes: you describe the person in plain English ("engineers who have shipped PRs to SWE-bench or Terminal-Bench in the last 18 months, currently at a data-labeling or frontier lab company") and get a ranked shortlist across GitHub, LinkedIn, and the open web, not a title-only slice of LinkedIn.
Which repos actually map the talent?
Eight open-source projects function as a de facto skills registry for RL environments and agent evaluation talent. If someone has meaningful commit history in two or more of these, they are qualified for a Bespoke, Mechanize, Prime Intellect, or HUD role regardless of what their title says.
- Terminal-Bench (Bespoke contributes here)
- SWE-bench
- TAU-bench
- lm-eval-harness (EleutherAI)
- lighteval
- HELM
- GEPA (Bespoke contributes)
- OpenThoughts (Bespoke contributes)
Add to that the Prime Intellect Environments Hub, which already hosts 2,500+ open-source RL environments. Every contributor there is a lead. That is a directly scrapable talent map, and it is bigger than every AI eval engineer title on LinkedIn combined.
Who is Bespoke actually competing with for these hires?
Bespoke is competing with a short list of similarly-staged, similarly-funded environments companies, plus the frontier labs' internal post-training teams. SemiAnalysis notes that nearly every environments company (Habitat, DeepTune, Fleet, Vmax, Turing, Mechanize, Preference Model, Bespoke, Veris.ai) is seed-stage with fewer than 20 employees serving one to three customers.
The named competitors matter because they will bid for the same tiny pool:
- Mechanize. ~$9M angel round in April 2025, publicly works with Anthropic.
- Prime Intellect. ~$21M total from Founders Fund, Menlo, and Andrej Karpathy. Runs the Environments Hub.
- HUD (YC W25). $16M Series A backed by Exceptional Capital. Customers include DoorDash, UiPath, OpenAI, Anthropic. Team is public: Jay Ram (CEO), Lorenss Martinsons (CPO).
If each of these four companies tries to hire five named eval engineers in 2026, the global titled market is drained by Q3. That is not a competitive dynamic. That is a supply cliff, and it is the reason Wing VC and Mayfield can underwrite an "environments" thesis at all: the moat is not IP, it is which team lands the first ten hires.
What does the archetype actually look like?
The archetype is a research engineer with post-training experience, not a self-described eval specialist. Look at Bespoke's own founders: Mahesh Sathiamoorthy is an ex-Google DeepMind staff research engineer, and Alex Dimakis is a UC Berkeley professor. Neither ever carried the title "Evaluation Engineer." Both do the work.
A hireable profile in this category tends to combine:
- Two or more years shipping post-training or RLHF pipelines at a frontier lab, a data-labeling company, or a university lab.
- Public commit history to at least one benchmark or environment repo from the list above.
- A track record of writing eval harnesses that other teams actually use (a much rarer signal than "ran an eval once").
- Comfort with agent-loop debugging: tool use, browser control, terminal environments.
You will not find that intersection with a title filter. You will find it by starting with the artifacts (papers, PRs, environments published) and reverse-mapping to people. When founders describe the shape of the person to Refolk in plain English rather than as a Boolean string, the shortlist that comes back is closer to the archetype than any LinkedIn Recruiter query I have watched go out this year.
What should a recruiter do this quarter?
Stop searching by title. Build a pipeline off open-source contribution graphs, and treat data-labeling companies as the primary feeder, not the frontier labs. Concretely, in priority order:
- Pull the contributor list for Terminal-Bench, SWE-bench, TAU-bench, lm-eval-harness, HELM, and lighteval. Deduplicate. That is your top-of-funnel.
- Cross-reference against Prime Intellect Environments Hub authors. The overlap set is your A-list.
- Add the 61 US professionals who list RLHF as a skill, weighted toward Surge AI, DataAnnotation, Mindrift, plus AWS and TikTok USDS.
- Look at HUD's public team and the maintainers of GEPA and OpenThoughts. These are named, cross-poachable people.
- Set a geographic default of remote-first. The center of gravity is India and EU. A Mountain View-only search will surface the same 15 profiles every other recruiter is already messaging.
- Message on the artifact, not the title. "I saw your PR to Terminal-Bench" opens replies that "AI Evaluation Engineer opportunity" does not.
The founders who staff up first will win the environments category, and the only way to staff up first is to source from the 3x larger adjacent pool that the title filter hides. Refolk was built for this exact problem: plain-English queries that cross GitHub, LinkedIn, and the open web, so the "Research Engineer at Surge AI who has committed to lm-eval-harness twice this quarter" shows up on the same shortlist as the one person in the US who actually wrote "Agent Evaluation Engineer" on their profile.
FAQ
How many AI evaluation engineers exist globally?
In Refolk's index of professional profiles, 48 people worldwide currently carry an AI-specific evaluation-engineer title ("AI Evaluation Engineer," "LLM Evaluation Engineer," "Agent Evaluation Engineer," or "Evals Engineer"). The US subset is roughly 15 to 25 once you strip out aerospace and hardware "Evaluation Engineer" profiles. The adjacent US pool of professionals listing RLHF as a skill is about 61, roughly 3.2x the named pool, and that is the cohort recruiters should actually target.
Why is Bespoke Labs' $40M round hard to deploy on hiring?
Bespoke closed $40M on July 6, 2026 into a category where the global titled labor market is 48 people and heavily concentrated in India and EU, not the Bay Area. Even doubling headcount from 40 to 80 exceeds the entire self-identified global pool. The capital is only deployable if the company sources by skill and open-source contribution rather than by job title, and if it commits to remote hiring or expensive relocations.
What open-source projects should recruiters use to find agent evaluation talent?
The highest-signal repos are Terminal-Bench, SWE-bench, TAU-bench, EleutherAI's lm-eval-harness, lighteval, HELM, GEPA, and OpenThoughts, plus Prime Intellect's Environments Hub (2,500+ open-source RL environments). Contributors to two or more of these are qualified candidates for Bespoke Labs jobs, Mechanize, Prime Intellect, or HUD roles regardless of what their LinkedIn title says. Sourcing off commit history outperforms title search by a wide margin in this category.
Who is Bespoke Labs competing with for eval engineers?
The direct competitors are Mechanize ($9M angel, works with Anthropic), Prime Intellect ($21M from Founders Fund, Menlo, and Karpathy), and HUD (YC W25, $16M Series A, customers include OpenAI and Anthropic). Indirectly, Bespoke competes with the internal post-training and safety teams at Anthropic, OpenAI, and Meta, and with data-labeling companies like Surge AI, DataAnnotation, and Mindrift, which employ a surprising share of the currently-titled eval engineer cohort.