Refolk
September 15, 2026·10 min read

HN's September 2026 Hiring Thread: 21.6:1 and 202 Agent Builders

The September 2026 HN "Who is Hiring" thread reveals a 21.6:1 funnel and a shift to agent-builder and RL-eval roles against only 202 US candidates.

hacker news who is hiring september 2026hn hiring thread applicant conversionhiring ai agent builder engineerssourcing rl eval engineershn who is hiring trends
HN's September 2026 Hiring Thread: 21.6:1 and 202 Agent Builders

The September 2026 "Ask HN: Who is Hiring?" thread (item 49522897) went live about a day ago, and it is already the cleanest snapshot of what frontier engineering hiring actually looks like right now. One repeat poster disclosed a 21.6:1 applicant-to-offer ratio from August, and the JDs underneath have quietly stopped asking for "senior full-stack." They are asking for people who can write RL benchmarks, red-team model outputs, and author eval rubrics, a job description that barely had a name eighteen months ago.

If you are a recruiter or a founder reading that thread as market data, the interesting number is not 216 applications. It is the 202 US professionals who could plausibly do the work.

What the September 2026 HN thread actually says

The September 2026 thread is a compositional break from a year ago: the top asks are agent builders and RL-eval engineers, not generalist full-stacks, and the disclosed funnel math looks better than it is. HNHIRING now indexes 60,304 job ads back to January 2018, so the shift is measurable against a real baseline rather than vibes.

Two posts anchor the change:

  • Cider Consulting disclosed 216 August applications yielding roughly 10 offers, a 21.6:1 ratio. Their September ask is explicit: senior engineers who can design RL benchmarks, red-team model outputs, and create rubrics to evaluate frontier models' agentic coding capabilities.
  • Catalyst·Wayfare AI is hiring an "Agent Builder" for end-to-end ownership of individual capability agents: prompts, tools, retrieval, deterministic checks, evals, monitoring, and enabling infrastructure.

Neither of these JDs would have parsed cleanly in September 2025. "Agent Builder" was not a job title anyone was hiring for at volume; "eval rubric author" was a research task inside a lab, not a line item on an HN post. Twelve months later, both are the headline asks.

60,304
HN job ads indexed since January 2018
The baseline that makes the September 2026 compositional shift measurable rather than anecdotal.

Why 21.6:1 is a worse funnel than it looks

The 21.6:1 ratio is not a soft market. It is applicant noise: with only 202 real agent-builder professionals in the US, most of the 216 applications came from full-stack engineers pattern-matching on "senior + AI" without the underlying skill. The qualified-applicant ratio Cider is actually working is probably closer to 5:1.

Here is the mechanic. When a JD asks for RL benchmark design and red-teaming, three groups apply:

  1. The tiny cohort of engineers who have actually shipped agent evals in production.
  2. A much larger group of senior full-stacks who have wired up an OpenAI SDK call and think that counts.
  3. New grads and career-switchers hoping the JD is looser than it reads.

Only group one converts. The other two inflate the denominator, make the poster look popular, and burn a full week of screening time. The recruiter takeaway is that HN thread response volumes are not comparable across role categories the way they were in 2022. A 40:1 ratio on a Rust systems role and a 22:1 ratio on an agent-builder role can hide the same three qualified applicants.

The supply side: 202 agent builders, 36 eval engineers

In Refolk's index of professional profiles, only 202 people in the US currently hold titles like "AI Agent Engineer," "Agent Builder," or "Agentic AI Engineer," against 2,510 with "Senior Full Stack Engineer." That is 12.4x more full-stack supply than agent-builder supply, and the gap gets brutal at the eval layer.

SegmentUS supply (Refolk index)Signal
Senior Full Stack Engineers2,510Baseline "generic" HN-thread ask
Agent Builder / AI Agent / Agentic AI Engineer titles20212.4x scarcer than full-stack
Evaluation / LLM Eval / RL Engineer titles3669.7x scarcer than full-stack
Cider Consulting Aug 2026 funnel216 apps → ~10 offers21.6:1 (from HN post)
Agentic-role geographic concentrationSF Bay + Boston = 6 of top 10 regionsHigh cluster, low mobility
Non-AI-native "Evaluation Engineer" employersGE Aerospace, Hussmann, Toyota, DellHardware-reliability, not LLMs

Two consequences fall out of that table.

First, the SF Bay and Boston clusters hold six of the top ten regions for on-title agentic roles. If you are hiring from Austin, Seoul, or London, you are pulling against relocation gravity, not just competing on comp.

Second, the "Evaluation Engineer" title is a Boolean trap. 33 of the 36 US professionals with that exact title work at GE Aerospace, Hussmann, Toyota Boshoku, DNV, Wayve, and Dell. They evaluate turbine blades and seat frames, not LLM outputs. A recruiter who plugs "evaluation engineer" into LinkedIn Recruiter and sends 200 InMails will get near-zero relevant replies and burn their credibility with the pool.

The compensation ceiling nobody on HN is quoting

The top of this market clears at $500,000, and any HN post pricing the same skill at "senior full-stack comp" is invisible to the top decile. Mechanize, founded in April 2025 by ex-Epoch researchers to build RL environments, is working with Anthropic and has offered $500,000 salaries to environment engineers. Slingshot Aerospace's public RL/eval JD posts $150,000 to $250,000 remote US for evaluation frameworks, benchmarks, and simulation-backed validation of multi-step tool-using workflows.

The HN thread does not price roles at $500k. That means the 21.6:1 ratio hides a 0:1 offer-acceptance ratio for the top ten candidates in the country, because those candidates are already inside Mechanize, Anthropic, or a stealth environment shop and would not read the thread with intent.

The 21.6:1 ratio hides a 0:1 offer-acceptance ratio for the top ten candidates in the country.

For a founder using HN as a sourcing channel, the honest read is: HN gets you the middle 60% of the agent-builder pool, and only if your JD is written well enough to filter the full-stacks out on the first pass. The top and bottom deciles are elsewhere.

Where the real agent builders actually live

The best agent builders often do not carry a tidy "AI Agent Engineer" title, do not maintain a polished LinkedIn, and cannot be found by Boolean-matching a job title that barely existed eighteen months ago. Their real work lives in a GitHub repository, a published MCP server, or a weekend hackathon submission.

That is a sourcing problem, not a JD problem. The signals that actually correlate with capability in September 2026:

  • Merged PRs to open-source agent frameworks (LangGraph, CrewAI, OpenHands, Inspect).
  • A published MCP server, especially one with more than a dozen stars and real issue traffic.
  • Blog posts or repos on reward hacking, verifiable rewards, or eval harness design.
  • Heavy daily use of Claude Code or Cursor, visible through the shape of their commits.
  • Prior domain expertise (compilers, security, math, finance) plus a recent pivot into agent tooling.

None of those signals map to a job title. All of them map to artifacts on GitHub and the open web. This is the exact gap Refolk closes: you describe the person in plain English ("US-based engineer who has shipped an MCP server and writes about eval harnesses") and get a ranked shortlist back, without pretending "Evaluation Engineer" is a useful Boolean.

Top companies employing on-title agent builders in the index include OpenHands, Liberate, Observe.AI, Walmart Global Tech, YouTube, J.P. Morgan, and Amazon. That is a poaching list, not a partnership list. It also tells you the buyer profile has broadened past labs: Walmart and J.P. Morgan are hiring the same title Anthropic is.

How to read the thread as a recruiter, not a candidate

Read the September 2026 thread as a compensation and category map, not a to-do list. The three moves that pay off:

  1. Sort JDs by explicit skill, not company name. Posts that list "reward hacking detection," "eval rubric authoring," or "RL benchmark design" are pricing a scarce skill and will lose their hire to Mechanize-style offers. Those companies are your best acquisition targets in six months.
  2. Cross-reference posters against WARN filings and layoff trackers. A company posting three agent-builder roles on HN while it has not backfilled its infra team is signaling a strategy pivot. That is a founder call, not a candidate call.
  3. Extract the JD language, not the JD. The exact phrasing evolving on HN ("verifiable rewards," "deterministic checks," "capability agent") is the new keyword set for GitHub bio and repo README searches. It will not appear on LinkedIn for another year.

Gartner projects that 40% of enterprise applications will ship task-specific AI agents by the end of 2026, up from less than 5% a year earlier. If that lands even directionally, the 202-person US on-title pool needs to grow by an order of magnitude in twelve months, and it will not. The gap gets filled by retitling: senior backend engineers who happen to use Claude Code daily become "agent engineers" on their next offer letter, and the domain-expert plus heavy-Claude-user profile becomes more valuable than a fresh ML PhD for authoring RL tasks.

The vendor question nobody in the thread asks

If you are a founder posting the JD, the question that separates the real hires from the resume-padders is: how do you detect reward hacking? An environment that can be gamed actively teaches your model the wrong behavior, and undoing that is far more expensive than building the environment correctly the first time.

Handshake AI pivoted from a college-jobs platform to AI data and reached about $1 billion in gross annualized AI-training revenue by April 2026, acquiring Cleanlab specifically to move into evaluation and QA. That is the size of the vendor market forming around this question. If your in-house candidate cannot answer it in one interview, treat the vague response as a serious red flag, exactly the same way you would evaluate a vendor pitch.

The paradigm underneath all of this is reinforcement learning from verifiable rewards (RLVR), where the environment itself provides the signal rather than a learned reward model. A coding task passes or fails its unit tests. A math answer matches or does not. Building environments that verify well without leaking exploits is the actual scarce skill, and it is what the top of the September 2026 HN thread is really asking for, under the surface of the JD.

202
US professionals with an agent-builder title (Refolk index)
Against 2,510 senior full-stack engineers, a 12.4x supply gap and the real ceiling on the September thread's hire rate.

For a small team sourcing against these posts, the practical workflow is: read HN weekly for JD language, then run those phrases as plain-English queries against a candidate index to surface the untitled candidates who actually match. Do not filter by job title. Filter by artifact.

FAQ

How reliable is the 21.6:1 ratio as a benchmark for my own funnel?

It is reliable as an upper bound on apparent conversion, and misleading as a floor on qualified conversion. Cider Consulting's 216 August applications yielded roughly 10 offers, but a large share of those applications were from senior full-stacks miscategorizing themselves against an RL-eval JD. If you strip those out, the qualified funnel is probably closer to 5:1, which means your own hiring plan should assume you can screen aggressively without running out of viable candidates.

Is "Evaluation Engineer" a useful title to search for on LinkedIn?

Not in September 2026. Only 36 US professionals hold that exact title in Refolk's index, and 33 of them work at GE Aerospace, Hussmann, Toyota Boshoku, Dell, DNV, or Wayve, evaluating hardware reliability rather than LLM outputs. Search for artifacts instead: MCP servers, contributions to eval harnesses like Inspect, and public writing on reward hacking or verifiable rewards.

What comp band do I need to compete for the top decile?

You need to be at $500,000 total comp to compete for the top RL environment engineers, because that is what Mechanize is paying against Anthropic-funded budgets. If you cannot clear that, target the middle 60% of the pool: senior backend engineers with heavy Claude Code usage and one or two shipped agent projects, priced in the $150k to $250k band that Slingshot Aerospace and comparable posters are already publishing.

Should I bother posting on HN if I cannot pay frontier comp?

Yes, but treat it as a filter, not a pipeline. Post the JD with explicit skill language (reward hacking, RLVR, rubric authoring) to attract the middle cohort and repel the pattern-matchers, then use the applications as a market survey rather than a hire list. Do the real sourcing on GitHub and against a candidate index, and use HN posters themselves as a target list of companies whose next hire will be looking in six months.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next