Refolk
September 10, 2026·9 min read

65 Titles, 183 Bounty Hunters: Sourcing Rubric Writers in Sept 2026

The Sept 2026 HN thread wants rubric writers who red-team RL agents. Only 65 US titles exist. The real pool lives on bounty leaderboards.

AI evals engineer sourcingRL rubric designer hiringLLM red team recruiterAI benchmark engineer jobssourcing eval engineers 2026
65 Titles, 183 Bounty Hunters: Sourcing Rubric Writers in Sept 2026

The September 9, 2026 Hacker News "Who is hiring?" thread has a role in it that did not exist a year ago and still does not have a LinkedIn title. Multiple posters are asking for senior engineers who can "design RL benchmarks, red-team model outputs, and create rubrics to evaluate and enhance the agentic coding capabilities of various frontier models." If you source engineers, this is the most under-indexed pool of 2026, and every hour you spend on Boolean strings is wasted.

The role the HN thread just named

The September 2026 HN "Who is hiring?" thread (item 49522897, posted Sept 9) is the first place multiple frontier labs publicly asked for the same hybrid engineer in the same week: part RL benchmark designer, part LLM red-teamer, part rubric author. It is one job description doing the work of three, and no ATS taxonomy catches it.

The verbatim ask from one September poster: "senior engineering expertise that can design RL benchmarks, red-team model outputs, and create rubrics to evaluate and enhance the agentic coding capabilities of various frontier models." Same poster reported 216 applications to their August thread and roughly 10 offers extended, a 4.6% offer rate, or about 21.6 applicants per offer.

That looks like dense supply. It is not. Most of those 216 are generic ML engineers who saw "RL" and applied. The actual pool, the one with hands-on rubric and red-team work, is much thinner and lives somewhere your ATS cannot see.

Why standard title search fails

Standard sourcing fails here because candidates are not yet self-labeling with any title that matches the work. In my index of professional profiles, an exact-title search across "Evaluation Engineer," "Red Team Engineer," "AI Evaluations Engineer," and "Evals Engineer" in the US returns 65 people, and almost none of them work on LLMs.

Here is the breakdown:

  • GE Aerospace: 4
  • Dell: 2
  • Honeywell, US Navy, Kudelski Security: named cyber and defense employers
  • Meta: 1
  • Wayve: 1

Ninety-five percent of that 65 is cyber and aerospace red-team work, not model evaluation. The frontier-lab version of the role has a title footprint you can count on one hand.

Now layer skills on top. Cross-filtering those titles with LLM or RLHF skill tags returns zero profiles. Adding "rubric" or "benchmark" as a keyword to the title-and-skill query also returns zero. That is not a bug in the index. It is vocabulary lag: the work is a year old, resumes refresh on a two-year cycle, and nobody has typed "rubric designer" into their headline yet.

65
US profiles with an evals or red team title
Only 1 sits at Meta and 1 at Wayve. The rest are cyber and aerospace.

If you are still sending "AI evals engineer" InMails, you are fishing in a pond of 65 mostly wrong people. The pool exists. It just has to be reconstructed from where the work actually happens.

The dataset that makes this concrete

The gap between titled supply and actual working population is roughly 8x, and one public bounty program alone has surfaced more qualified humans than the entire titled pool.

SliceUS countSourceNote
Exact titles: Evaluation / Red Team / AI Evaluations Engineer65Refolk indexMostly legacy cyber; 1 at Meta, 1 at Wayve
Same, filtered to LLM / RLHF skills0Refolk indexTitle + skill co-occur nowhere
Same, plus "rubric / benchmark / red team" keyword0Refolk indexBoolean is dead for this role
Anthropic Constitutional Classifiers red-teamers183anthropic.com~2.8x the entire titled pool
Anthropic Feb 2025 red-team cohort339pasqualepillitteri.it~5.2x the titled pool
HN Sept 2026 poster funnel216 apps / ~10 offershnhiring.com4.6% offer rate

The takeaway: public red-team programs have already surfaced roughly 8x more qualified humans than exist under standard job titles. If you keep sourcing by title, you are competing over the wrong 65 people while 500-plus vetted red-teamers sit on public leaderboards.

Where the real pool lives

The real pool lives in four places, none of which are LinkedIn title fields: bounty leaderboards, benchmark paper author lists, frontier-lab research blog bylines, and GitHub PR histories on active eval repos. Work through them in order.

1. Bounty program leaderboards

Bounty programs are functioning as a de facto skills assessment for this labor market. A HackerOne handle with three resolved Anthropic model-safety reports is a stronger signal than any resume claim, and the labs publish these lists.

  • Anthropic's Constitutional Classifiers experiment: 183 active participants spent over 3,000 hours over two months attempting to jailbreak the model.
  • Anthropic's 2025 red-team: 339 security researchers, 3,700 collective hours, more than 300,000 targeted attack interactions, $55,000 distributed to the four teams that cleared at least one level.
  • Anthropic's ongoing HackerOne program: pays up to $35,000 per new universal jailbreak, runs on HackerOne with public leaderboards.
  • Gray Swan Arena: top 40 overall participants get invited to the private paid engagement network. Past arena sponsors include OpenAI, Anthropic, Google DeepMind, UK AISI.
  • Cross-lab bounty payouts in 2026 range from $200 for entry-level findings to $100,000 for exceptional discoveries, across xAI, OpenAI, Anthropic, Google, Microsoft, Mozilla 0din.

Every one of those programs publishes participants. Every published participant is a candidate who has already been assessed on the exact skill combo your JD is asking for.

2. Benchmark paper authors

SWE-bench Pro contains 1,865 problems from 41 actively maintained repos across public, held-out, and commercial splits. The paper (arXiv:2509.16941) has a long author list, and every name is a person who has actually built, curated, and validated agentic coding evals. That is the profile.

Related shops with named benchmark builders in the same cluster: Morph Labs, Proximal Labs (Frontier-SWE), Terminal-Bench (Merrill, Shaw, Carlini et al.), Scale AI's SWE-bench Pro contributors.

3. Frontier-lab research blog bylines

Anthropic's Frontier Red Team publishes 3 to 8 named contributors on posts like "Assessing Claude Mythos Preview's cybersecurity capabilities" (April 2026) and "Measuring LLMs' ability to develop exploits" (May 2026). Ten posts give you 30 to 80 candidates, pre-ranked by lab and already vetted for the exact hybrid.

METR's March 25, 2026 post red-teaming Anthropic's internal agent monitoring was authored by David Rein. Single named author. Canonical example of the profile. If you cannot find David Rein and 20 people like him in an afternoon, your sourcing stack is broken.

The bounty program is the ATS. HackerOne leaderboards have already ranked your candidates for you.

4. GitHub PR histories on eval repos

Fork history and PR authorship on SWE-bench, SWE-bench Pro, Terminal-Bench, and the smaller lab-published eval harnesses tells you who is actively fixing broken test cases. An OpenAI audit reportedly flagged that around 30% of SWE-bench Pro's public split has broken or overly strict tests. Every commit fixing one of those is a work sample.

This is the exact gap I close: you describe the person in plain English ("engineers who authored PRs on SWE-bench Pro or Terminal-Bench in the last 12 months, based in the US, ideally with a HackerOne handle") and get a ranked shortlist across GitHub, LinkedIn, and the open web. No Boolean, no title filter, no waiting for the taxonomy to catch up.

Why benchmarks decaying is a hiring driver

Every retired benchmark is a new hire req, which is why lab budgets for rubric and eval work are climbing while the titled pool stays flat. OpenAI publicly retired SWE-bench Verified as a frontier signal in February 2026 ("Why SWE-bench Verified No Longer Measures Frontier Coding Capabilities"). SWE-bench Pro is already reported at around 30% broken. The half-life of a frontier eval is now under 12 months.

That has two consequences for sourcing:

  1. Frame the role as "benchmark maintenance engineer," not "benchmark author." Labs need people who can fix, iterate, and re-validate at speed, not just publish a paper once.
  2. Snorkel AI is publicly evangelizing rubric design as a discipline: their rubric series argues rubrics should be treated like models, structured, measured, iterated for inter-rater agreement. The named authors of that series are themselves sourcing targets, and the framing gives you a JD that reads like the actual job.
8x
Bounty-vetted candidates vs titled pool
183 Constitutional Classifiers participants plus 339 in the 2025 red-team cohort dwarf the 65 US profiles with a matching title.

The sourcing playbook that actually works in September 2026

Skip title search entirely and build the list from public work in this order: bounty leaderboards, benchmark author lists, lab red-team blog bylines, GitHub PR histories. Then enrich to LinkedIn last, not first.

Concrete sequence:

  1. Pull the Anthropic HackerOne leaderboard and the Gray Swan Arena top 40. That is your first 200 names.
  2. Pull author lists from SWE-bench Pro (arXiv:2509.16941), Terminal-Bench, Frontier-SWE, and the last six months of METR and Anthropic Frontier Red Team posts. Another 100 to 200 names.
  3. Pull GitHub contributors to the top five actively maintained eval repos with 20-plus commits each in 2026.
  4. De-duplicate. Enrich with current employer, location, and reachability. Rank by number of independent sources (a HackerOne handle plus a paper credit plus a GitHub PR is a 3-signal candidate).
  5. Only now write the outreach. Reference the specific finding, paper, or PR. Skip anything that reads like a template.

The 216-to-10 funnel from the August HN poster is the giveaway. That poster is drowning in generic ML applicants and starving for the actual profile. If you can deliver 20 pre-vetted candidates from bounty leaderboards and benchmark author lists, you are not competing with the other 216 resumes. You are the only inbound that matters.

FAQ

What is a rubric writer or evals engineer, in one sentence?

A rubric writer is an engineer who designs the structured criteria, benchmarks, and adversarial tests used to measure and improve frontier model behavior, typically combining RL benchmark design, LLM red-teaming, and rubric authorship in one role. As of September 2026, it has no standard title, no clean skill tag, and no LinkedIn filter that catches it, which is why bounty leaderboards and benchmark paper author lists are better sourcing surfaces than any ATS.

Where do I find these people if title search returns nothing?

Start with public bounty programs (Anthropic HackerOne, Gray Swan Arena top 40, xAI, OpenAI, Google, Mozilla 0din), then move to benchmark paper author lists (SWE-bench Pro, Terminal-Bench, Frontier-SWE), then frontier-lab research blog bylines (Anthropic Frontier Red Team, METR), then GitHub PR histories on active eval repos. A candidate who shows up in two or more of those sources is worth an outreach; one who shows up in three is a priority hire.

Why did one HN poster get 216 applicants but only 10 offers?

Because 216 applicants applied and roughly 200 of them were generic ML engineers who saw "RL" or "LLM" in the JD and hit send. The offer rate of 4.6% reflects vocabulary mismatch rather than actual supply. Real qualified supply, as measured by profiles matching the title plus LLM or RLHF skills, is closer to zero, and the working population lives on bounty leaderboards where you have to source them directly.

Should I still post the job on HN or LinkedIn?

Post it, but do not expect the inbound to deliver the hire. Use the posting to signal the market and to catch the occasional self-aware candidate, and run outbound against bounty leaderboards and benchmark author lists in parallel. Given that public red-team programs have surfaced roughly 8x more qualified people than exist under standard titles, the hire is going to come from outbound in almost every case.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next