Refolk
August 25, 2026·8 min read

Meta's Muse Code Just Made 6 Buyers Chase 6 US Engineers

Meta shipped Muse Code Aug 5, 2026. The US pool of engineers who have post-trained a coding agent sits in the low double digits. Source accordingly.

coding agent engineersMeta Muse Code hiringAI coding agent talent poolsourcing post-training engineersSWE-bench recruiting
Meta's Muse Code Just Made 6 Buyers Chase 6 US Engineers

On August 5, 2026, Meta announced Muse Code, its first terminal-based AI coding agent, powered by Muse Spark 1.2 and positioned squarely against Claude Code, Codex, Cursor, Devin, and Google's Antigravity CLI. That makes Meta the sixth serious buyer chasing the same tiny bench of engineers who have actually shipped a coding agent to production, and the pool is smaller than the number of open reqs at any one of these companies.

The actionable US pool is in the low double digits

The engineers who have genuinely post-trained and evaluated a shipping coding agent number in the single to low double digits in the US, not the thousands your ATS keyword search will return. Refolk's index shows 1,375 professionals globally list "coding agent" in their public profiles, but the head of that distribution is founders and CEOs, not operators. Filter to Research Engineer, Member of Technical Staff, or Post-Training Engineer titles inside the US and the number collapses to six people, five of whom carry the "Member of Technical Staff" title.

That is the practical floor for the "has shipped this" cohort. Every additional Muse Code hire above that floor comes from one of six places: Anthropic, OpenAI, Cursor (Anysphere), Cognition, GitHub, or Google. There is no seventh bucket.

6
US engineers with "coding agent" in profile and an MTS/Research Engineer title
Refolk's index of public professional profiles, filtered to operator-grade titles.

Why the founder-heavy head distorts every search

Roughly seven of the top ten titles in the global "coding agent" cohort in Refolk's index are founder or executive titles: Chief Executive Officer, Founder, Co-Founder and CEO. These people are not accepting Meta offers, even nine-figure ones. They are raising seed rounds off the same resume you want to poach. When a naive Boolean search returns 1,375 profiles and your recruiter says "great, we have a funnel," what they actually have is a list of people building competitors to your product.

The realistic hires are the MTS-level tier underneath. Six in the US operator filter, plus a thin adjacent tier that Refolk's index surfaces at Magic, Athena, and Microsoft AI.

What Meta is really buying with Muse Code

Muse Code is a data-collection play disguised as a product launch, which means Meta's eval-ops team will scale faster than its modeling team. Standard pricing is $1.25 per million input tokens and $4.25 per million output tokens, with a contributor tier priced more than 10x cheaper. Meta is also requiring thousands of internal engineers to use Muse Code weekly. 7,000 active internal users have already generated over 800 fixes that improved model performance.

That is not a research org. That is a labeling and eval-triage pipeline. The hire-able profile Meta actually needs looks less like "ML PhD" and more like "infra engineer with taste for evals and harness design." Which is a different search string, sourced from a different pool.

The job title has already changed

"RLHF engineer" is the wrong keyword in 2026. As Pragmatic Engineer noted, an increasing share of engineering work at Anthropic and Cursor is about building environments for agents to execute more efficiently. The real titles and self-descriptions look like:

  • Agent-environment engineer (sandboxes, container orchestration for agent runs)
  • Harness engineer (tool-use scaffolding, MCP integrations, retry logic)
  • Eval engineer (SWE-bench Verified, SWE-bench Pro, Terminal-Bench rigs)
  • Post-training engineer (SFT, DPO, RL loops specifically over code)
  • Member of Technical Staff (the catch-all at Anthropic, OpenAI, and increasingly Meta MSL)

Recruiters searching "RLHF" will miss the actual talent, who list "harness," "sandbox," "SWE-bench," or "Terminal-Bench" instead. This is the exact gap Refolk closes: describe the person in plain English ("engineers who have built agent harnesses or evaluated coding agents against SWE-bench in production"), and get a ranked shortlist that does not depend on which acronym the candidate happened to type into their headline.

Meta already holds 16% of the adjacent pool, and that's a warning

A broader US search for "RLHF evals" in Refolk's index returns 19 people, and Meta already employs 3 of them. That 16% share is not a strength. It is a signal that the marginal outside hire has to come from a direct competitor, and Meta is fighting gravity when it does.

SignalFire's 2025 State of Talent Report found engineers at OpenAI are 8x more likely to leave for Anthropic than the reverse. The DeepMind-to-Anthropic ratio is 11:1. Meta is not the natural destination for this cohort even at nine-figure comp, which is why Alexandr Wang's hiring blitz since joining Meta Superintelligence Labs in June 2025 has needed multi-million-dollar packages just to compete on offer, not on prestige.

The Cherny/Wu round-trip proves the pool is naming-market small

Boris Cherny and Cat Wu, the original Claude Code leads at Anthropic, left for Anysphere (Cursor) in more senior roles, then were hired back to Anthropic. Two named individuals moved the market twice inside a year. That only happens in a pool where the operator count is smaller than a mid-sized company's engineering all-hands. When you are sourcing coding agent engineers, you should be able to name every credible target on one page.

When two people can round-trip between Anthropic and Cursor and it makes headlines, the pool is not a pool. It is a roster.

The numbers side by side

Here is the dataset every sourcing plan for Muse Code, Claude Code, Codex, Devin, Cursor, or Antigravity should start from.

SegmentCountSource
Global profiles mentioning "coding agent"1,375Refolk's index, unfiltered
US: "coding agent" + Research Engineer / MTS / Post-Training titles6Refolk's index, operator filter
US: "RLHF evals" any title19Refolk's index, adjacent skill
Share of RLHF-evals pool already at Meta3 of 19 (16%)Refolk's index
Top-10 global titles that are founder/CEO~7 of 10Refolk's index title distribution
Cursor total headcount, Aug 2025~300Wikipedia

Notice the shape. The unfiltered number is 229x the operator number. Any sourcing pipeline that starts from the 1,375 and does not aggressively filter will spend its cycles messaging people who are either founders of your future competition or juniors who typed a trendy phrase into their About section.

How to source this cohort in Q4 2026

Skip the keyword-Boolean approach and source on behavioral signals: shipped artifacts, benchmark submissions, and named-team lineage. Here is the sequence that works right now.

  1. Start from the leaderboards, not LinkedIn. SWE-bench Verified, SWE-bench Pro, and Terminal-Bench have public submission logs. Claude Opus 4.7 tops SWE-bench Verified at 87.6% and SWE-bench Pro at 64.3%. The engineers on those submissions self-identify as operator-grade.
  2. Map the incumbents by team, not by company. You do not want "someone from Anthropic." You want the Claude Code sub-team, the Codex post-training group, the Cursor agent-harness pod, the Devin RL loop, or the Antigravity CLI eval crew.
  3. Filter for benchmark literacy in the first screen. A candidate who cannot explain the delta between SWE-bench Verified, SWE-bench Pro, and Terminal-Bench has not shipped a coding agent, regardless of what their resume claims. This is the fastest disqualifier in the funnel.
  4. Search on artifacts, not titles. GitHub commits to public agent harnesses, MCP server contributions, eval rig repos, and blog posts about sandbox design are stronger signals than the words on a LinkedIn profile.
  5. Watch Meta's own MSL page for cross-pollination. Meta already holds 16% of the adjacent pool. Some of those people were hired in the last 12 months and are still gettable by a compelling seed-stage pitch.

Where the non-obvious hires are hiding

Refolk's index surfaces Magic, Athena, and Microsoft AI as quiet employers of the tight "coding agent + MTS/Research Engineer" cohort. These are not the names on Meta's target list this week, which is exactly why they are the ones to map first. The same logic applies to research engineers at second-tier labs who have built internal-only coding agents that never got a product name.

If you are running a seed-stage or Series A shop competing against Muse Code hiring, your edge is not comp. It is speed to the second-tier list before Wang's recruiters get there. Refolk is built for exactly this: describe the person in plain English, get the shortlist ranked by artifact strength across GitHub, LinkedIn, and the open web, and start outreach the same day.

The bench beyond the six

Widen the definition and the pool grows, but the value density drops fast. The signals that still separate operator from tourist at that width:

  • Named team lineage: worked directly on Claude Code, Codex, Cursor agent, Devin, or Antigravity CLI.
  • Public benchmark submission: author or co-author on a SWE-bench, Terminal-Bench, or agent-eval paper or leaderboard entry.
  • Harness-shaped open source: contributor to an agent framework, sandbox runner, or MCP tooling project.
  • Post-training specificity for code: talks or writeups on RL over code, verifier design, or reward modeling for tool-use tasks.

Two of these in combination is a strong yes. One alone is a maybe. Zero is a title match, not a candidate.

FAQ

How many engineers have actually shipped a coding agent?

The unfiltered global count of profiles mentioning "coding agent" in Refolk's index is 1,375, but most of those are founders, executives, or people who added the phrase after a trendy launch. The operator-grade cohort with production experience post-training or evaluating a shipping agent is in the single to low double digits in the US when filtered to Research Engineer, MTS, or Post-Training titles. Treat the 1,375 as a top-of-funnel to filter, not a pool to message.

Is "RLHF engineer" still the right search term?

No, and using it will actively hide the best candidates. The work has moved from generic human-preference labeling to agent-environment engineering: sandboxes, tool harnesses, and eval rigs. Search for "harness," "SWE-bench," "Terminal-Bench," "agent evals," or "sandbox" instead, or describe the behavior in plain English and let a semantic sourcing tool find the profiles regardless of which keyword the candidate happened to use.

Can a seed-stage startup realistically compete with Meta for these engineers?

Yes, on speed and mission, not on comp. Meta is fighting gravity in this cohort, with OpenAI engineers 8x more likely to leave for Anthropic than the reverse. Founders who can articulate a specific, unsolved problem in coding agents (verifier design, long-horizon task evals, sandbox security) and move from first contact to offer inside two weeks will beat Meta's slower nine-figure process for the subset of candidates who care more about the work than the package.

What is the single fastest disqualifier in a screen?

Ask the candidate to explain the difference between SWE-bench Verified, SWE-bench Pro, and Terminal-Bench, and where Claude Opus 4.7's 87.6% and 64.3% scores fall on those. Anyone who has actually shipped a coding agent has an opinion about the failure modes of each benchmark. Anyone who has not will pattern-match on the names and give a generic answer. This one question compresses a 45-minute screen into 90 seconds.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next