HN's September Thread: 216 to 10, and Why LinkedIn Misses 746
The September 2026 HN "Who is hiring?" thread named a role LinkedIn cannot see. Here is how to source RL rubric authors by artifact, not title.
The September 2026 HN "Who is hiring?" thread hit 254 points and 397 comments, and if you read it carefully you can see the outline of a role that did not exist a year ago. The top posting asked for engineers to "design RL benchmarks, red-team model outputs, and create rubrics to evaluate and enhance the agentic coding capabilities of various frontier models." The same posting quietly disclosed the August funnel: 216 applications, roughly 10 offers. If you are still typing "AI Red Team" into LinkedIn, you are searching for 36 people while 746 sit one query away.
What the September 2026 HN thread actually revealed
The thread surfaced a new senior-IC archetype: the RL rubric author, an engineer who designs evaluation benchmarks and adversarial test suites for agentic coding models. The role has no standardized LinkedIn title, no ATS category, and no CyberSeek code, which is why the funnel numbers on that one posting are so lopsided.
Three things about the thread matter for anyone hiring in Q4 2026:
- The archetype dominated. Rubric authors, agentic-coding evaluators, and frontier-model red-teamers were the plurality of senior-IC roles, not a novelty.
- The posting sold the artifact, not the seat. It described what you would build (RL benchmarks, rubrics, red-team runs) instead of asking for a title match.
- The August cohort's math was public: 216 in, ~10 offers, a 4.6% offer rate on a channel with no recruiter and no ATS.
HNHIRING has indexed 60,299 job ads back to January 2018, which gives you the baseline: this archetype was effectively absent from the thread twelve months ago. It is now the top of the page.
The 216 to 10 funnel is a feature, not a bug
A 4.6% offer rate on a raw HN post is not luck; it is the mechanical result of describing a task instead of a title. When you write "design RL benchmarks and author rubrics for agentic coding evals," the only people who apply are the ones who have done it. Keyword-matching disappears because there are no keywords to match.
Compare that to what happens when you post the same role on LinkedIn as "Senior ML Engineer, Evaluation." You get everyone with "ML" and "eval" in a headline, most of whom have never authored a weighted rubric or run an adversarial red-team pass in their lives. The screen becomes the job.
The lesson generalizes. If you are hiring for any role that predates its title (rubric author, agentic eval engineer, AI red-teamer, RL environment designer), stop writing the JD around a headline and start writing it around the artifact you need built by the end of Q1. Applications drop, quality climbs, and your recruiters stop arguing with LinkedIn's keyword parser.
Why title search misses most of the pool
Title search misses this pool because the work was invented inside frontier labs in the last 12 months, and the practitioners still carry legacy headlines like Research Engineer, MLE, or Applied Scientist. Refolk's index makes the gap concrete.
| Segment | Count | Note |
|---|---|---|
| US: RE / MLE / Applied Scientist with RL skill | 746 | Broad pool by artifact and skill |
| US + UK: profiles with explicit "Red Team" AI title | 36 | Narrow pool by title alone |
| Ratio | ~20.7x | RL-skilled pool vs title-searchable pool |
| Sept 2026 HN thread engagement | 254 pts / 397 comments | news.ycombinator.com/item?id=49522897 |
| App to offer, Aug HN posting | 216 → ~10 (4.6%) | hnhiring.com/locations/remote |
| Active AI/ML security postings, Mar 2026 | 2,500+ | SANS 2026 |
| Red team operator postings YoY | +29.2% | CyberSeek via axis-intelligence.com |
In Refolk's US index, 746 people hold a Research Engineer, ML Engineer, or Applied Scientist title and list Reinforcement Learning as a skill. Top employers in that slice are Meta (5), AWS (2), DoorDash (2), Google DeepMind, and Together AI, concentrated in the SF Bay Area (6) and NYC (4). Meanwhile the US+UK pool with an explicit "Red Team"-family title tied to AI or frontier-model work is just 36 profiles, most of them at AIG, Apple, Meta, US Navy, and White Knight Labs. The title "AI Red Team" is barely populated even where the work is being done.
That is a 20.7x gap between what you can find by title and what actually exists. Every sourcer who filters on title:"AI Red Team" is fishing the 36-person pond while the 746-person pond sits one skill-plus-artifact query away. This is the exact gap Refolk closes: you describe the person in plain English ("US-based research engineers with RL skills who have contributed to an agentic coding benchmark") and get a ranked shortlist that ignores headline noise.
The rubric author is not an ML engineer
The scarce skill here is taxonomy design under adversarial conditions, not model training, which is why the best candidates come from senior IC engineers with test-infrastructure or standards-body backgrounds rather than from RLHF teams. Two recent papers make this concrete.
UpBench specifies the job precisely: "Expert annotators construct between 5 and 20 rubric criteria per job, labeled critical, important, optional, or pitfall." LH-Bench goes further, showing that expert-authored rubrics with binary-observable boundaries reduce scoring ambiguity compared to LLM-authored rubrics with generic proficiency anchors. Translation: a good rubric author is someone who has spent years writing test cases that fail cleanly, not someone who has trained a model.
Signals that a candidate can actually do this work:
- Prior authorship on a public benchmark (Senior SWE-Bench, SlopCodeBench, Terminal-Bench, MLE-Bench, LiveCodeBench, SWE-rebench, UpBench, LH-Bench). Every one has named human authors on arXiv.
- GitHub history that includes property-based tests, fuzzing harnesses, or CI infrastructure at scale.
- Standards-body or spec-writing experience (IETF drafts, W3C, OpenAPI, Bazel rules).
- Public writing that argues about edge cases with specificity, not abstractions.
The scarce skill is taxonomy design under adversarial conditions. That is a QA lead with a PhD, not an RLHF engineer.
None of those signals live in a LinkedIn title. They live in commit histories, arXiv author lists, and benchmark leaderboards, which is where you have to source.
The supply gap is real, and everyone is fishing the same pond
The AI-security and evaluation talent shortage is a supply problem, not a demand problem, and the numbers show every eng leader in Q4 2026 is competing for the same few hundred people. WEF's Global Cybersecurity Outlook 2025 found only 14% of organizations have the skilled talent to meet their cybersecurity objectives. ISC2's 2025 read has AI security as the number-one critical-skills priority at 41%, ahead of cloud security at 36%.
SANS's 2026 report counted more than 2,500 active AI/ML security engineer postings on job platforms by March 21, in a category that barely existed a year earlier. CyberSeek data shows red team operator postings grew 29.2% year over year, with AI Red Team Specialist and AI Security Engineer roles commanding $140K to $250K and facing the most severe supply shortages of any cybersecurity subcategory.
Put those numbers together and you get the shape of the problem. The count of open reqs targeting this archetype is in the low thousands, aimed at a pool of a few hundred qualified practitioners in the US and UK. The candidates know this. That is why the compensation band starts where senior IC comp used to top out.
How to source RL rubric authors by artifact, not title
Source this pool by walking backwards from the artifacts they built: benchmark repos, arXiv co-author lists, HN handles, and specific Discord communities, then enrich to a real name and a reachable email. Title search is the wrong first move and it is often the wrong second move too.
A workable playbook:
- Start with the benchmarks. Pull author lists from Senior SWE-Bench (Snorkel AI), SlopCodeBench, Terminal-Bench, MLE-Bench, LiveCodeBench, SWE-rebench, UpBench, and LH-Bench. Each paper has 3 to 15 human authors. That is your seed set.
- Walk the citation graph. Anyone who cites two or more of those papers in their own work is fluent in the vocabulary. Google Scholar and Semantic Scholar make this trivial.
- Cross-reference GitHub. Contributors to the benchmark repos, plus anyone who has opened a substantive issue on one, are pre-qualified.
- Monitor HN monthly. The archetypes that dominate "Who is hiring" today will hit LinkedIn title normalization in 6 to 9 months. Sourcers who read the thread the day it drops beat competitors by two hiring cycles.
- Watch the right rooms. EleutherAI Discord, ML Collective, arXiv-sanity, and the SWE-Bench leaderboard maintainers are where the practitioners argue about rubric design in public.
- Then enrich. Convert the seed set into named individuals with current employers and reachable channels.
The last step is where most sourcing stacks fall over. You end up with a spreadsheet of GitHub handles and no way to tie them to a person you can contact this week. Refolk collapses that into one query: describe the artifact and the adjacent skills in plain English, and get back the named people with current employer, location, and the best channel to reach them across GitHub, LinkedIn, and the open web.
Named companies and people to start from
Start your outreach with the companies that already employ the archetype at density, because their employees have both the skill and the recent context. In Refolk's top-employer breakdown for the RL-skilled research-engineer slice: Meta, AWS, DoorDash, Google DeepMind, and Together AI. Snorkel AI runs the Benchtalks series and shipped Senior SWE-Bench (authors include Henry Ehrenberg and Alex Shaw). Anthropic staffed Claude Opus 5 evals. Each of these teams is a graph you can traverse: hire one, source the next three from their citation list.
The trap is treating those five employers as your only hunting ground. The 746-person pool spreads across dozens of smaller labs, university groups, and IC engineers embedded in non-AI-first companies who did their benchmark work as a side project. Those are the candidates competitors have not called yet, which is exactly why you should.
FAQ
Where do I actually find the September 2026 HN thread?
The thread lives at news.ycombinator.com/item?id=49522897, and hnhiring.com/locations/remote indexes the individual postings with searchable text. Read both. The HN comment thread is where the debate happens (compensation ranges, remote policies, and the occasional counter-recruiting) and hnhiring is where you extract the postings cleanly for downstream sourcing.
Is "AI Red Team" the same as "RL rubric author"?
No, and conflating them is a common sourcing mistake. AI red-teamers probe deployed models for policy and safety failures. RL rubric authors design the evaluation criteria and adversarial test cases that measure agentic coding capability during training. The skills overlap (adversarial thinking, taxonomy design) but the day-to-day work and the target companies differ. The September thread bundles them because a single senior IC often does both.
What compensation should I expect for this role?
CyberSeek data via Axis Intelligence puts AI Red Team Specialist and AI Security Engineer roles at $140K to $250K, and the RL rubric author archetype tracks similarly given the supply shortage. Frontier labs pay well above that band for the top of the 746-person pool. If your budget is capped at generic senior-IC comp, you will lose every competitive process.
How often should I re-sweep HN "Who is hiring"?
Monthly, on the first business day of the month, right after the thread drops. The archetypes that appear in three consecutive threads are the ones worth building a sourcing playbook around. The ones that appear once are noise. This is how you spot the next rubric-author-shaped role 6 to 9 months before LinkedIn's title taxonomy catches up.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.