Refolk
October 6, 2026·9 min read

Saffron Grades AI-Fluent Devs. MCP Authors Are the Public Shortlist.

Saffron, Contrario, and Cursor lock AI-fluency data inside hiring companies. Here is the public GitHub shortlist that beats them from the outside.

sourcing AI-native engineersMCP server authors GitHubhiring Claude Code developersCursor power users recruitingAI-fluent engineer shortlist
Saffron Grades AI-Fluent Devs. MCP Authors Are the Public Shortlist.

Saffron (YC F2026) launched this month promising to score "who's engineering and who's vibecoding" by recording every prompt, diff, and edit a candidate makes in a browser IDE wired to Claude Code. Contrario is doing a version of the same thing at $6M ARR six months in. Cursor's AI Code Tracking API, which would answer the question definitively, is Enterprise-only, still in Alpha, and started metering its deeper Conversation Insights layer in January 2026. If you source from outside the hiring company, none of that data will ever be in your search index. You need a public substitute, and it already exists.

The attribution race is over, and sourcers lost

The signal you actually want - who writes good code with AI agents - now lives behind enterprise contracts, and no amount of LinkedIn tooling will surface it. The sourcer's job is to find the public artifacts that correlate with it instead.

Here is what is locked up:

  • Cursor AI Code Tracking API: Enterprise plan only, Alpha, response shapes "may change," works only at the workspace root.
  • Cursor Blame: AI-aware git blame distinguishing AI from human attribution. Enterprise only.
  • Conversation Insights: previously free in preview, began charging in January 2026 on inference cost plus a Cursor token fee.
  • Saffron session recordings: candidate-side only, by definition. The company that ran the eval owns the data.
  • Contrario rubrics: 200+ customer accounts including Wispr Flow, Listen Labs, and Slash. None of it leaves the ATS.

Even inside the hiring company, Cursor's own docs admit the attribution underreports: copy-paste from chat and offline edits do not register, only accepted suggestions through the built-in flow. So the "ground truth" is leaky even for the one team that owns it. If you are sourcing AI-native engineers from the outside, closed APIs you cannot scrape are a dead end. MIT-licensed MCP repositories you can clone are not.

Why "AI reliance" scoring kills the obvious shortlist

Shortlists built on raw Cursor or Copilot usage will be adversely selected against the moment Saffron-style evals go mainstream, because the thesis is not "most AI usage wins." Saffron's own tagline is "See who's engineering and who's vibecoding," and the product ships an explicit AI Reliance Scoring metric that treats over-dependence as a negative indicator.

The mechanism is simple. If the market starts paying for judgment about when to use an agent, then counting Copilot acceptances is counting the wrong thing. The engineers who clear a Saffron session are the ones who understand the agent well enough to override it, which is a strictly harder population than "people who installed Cursor."

That is why authorship beats usage as a sourcing signal. Writing an MCP server, committing an AGENTS.md to a serious repo, or configuring Claude Code hooks all require that you have internalized how the agent actually works. Installing Cursor requires a download.

Shortlist on authorship of agent tooling, not usage of it. Authorship requires understanding the agent. Usage requires a download.

The public pool is tiny. Here is how tiny.

In Refolk's index of professional profiles, the self-described AI-native engineering cohort is small enough to shortlist by hand, and the headline-search version of it is almost nonexistent.

Signal / segmentCountSource
US engineers in SWE / ML / AI Engineer titles tagging LLM-agent work217Refolk's index
Global engineers putting "Claude Code / Cursor / AI coding agent" in headline9Refolk's index
Headline-search gap (ratio)~24:1Derived from Refolk's index
MCP servers in GitHub's official partner registry at Oct 2025 launch~40InfoQ
Stars on modelcontextprotocol/servers (community catalog)91,000GitHub
Stars on modelcontextprotocol/registry (central service)7,300GitHub
US AI-recruiting funding, 12 months to Jul 2026 (10 companies)$656MSecondTalent
9
Global engineers who self-tag "Claude Code / Cursor / AI coding agent" in their LinkedIn headline
From Refolk's index. Headline-based sourcing is not a strategy for this cohort, it is a rounding error.

Top employers for the 9: Walmart Global Tech, Launch Potato, Avvale. Top employer among the 217: AWS, with 3. The US-based 217 cluster in SF Bay (4), NYC and metro (5), and Santa Clara (2). This is not a market where Boolean on job titles returns a workable shortlist. It is a market where you have to go to GitHub.

MCP authors are the real public list

The Model Context Protocol repositories give you a public, timestamped, name-attached roster of the engineers who have shipped against the agent-tooling frontier. That is the shortlist.

Start with the people who created the thing:

  • David Soria Parra (@dsp) and Justin Spahr-Summers (@jspahrsummers), the two protocol authors.
  • Working-group leads: Rado Dimitrov (Stacklok), Tadas Antanavicius (PulseMCP), Bob Dickinson (TeamSpark), Preeti Dewani (Ravenmail).

Then widen to the author graph around them. The two anchor repos (modelcontextprotocol/servers at 91k stars and modelcontextprotocol/registry at 7.3k) have public PR authors, reviewers, and approvers. Reviewer status on either is a stronger seniority signal than any job title.

Then widen again to GitHub's official MCP Registry, which launched in October 2025 with roughly 40 partner-vetted servers from Microsoft, GitHub, Dynatrace, Terraform, Cloudflare, Docker, Elastic, Grafana Labs, Figma, Atlassian, Hugging Face, AWS Labs, Upstash (Context7), Mendable (Firecrawl), and Exa Labs. Every one of those servers has a CODEOWNERS file or a short list of commit authors. That is your partner-vetted senior roster.

Then widen one more time to the community discovery layers:

  • Smithery (by Henry Mao)
  • PulseMCP
  • mcp-get
  • The "Awesome MCP Servers" lists maintained by punkpeye, wong2, and appcypher

Each of those lists resolves to GitHub handles you can enrich into real people. This is the exact gap Refolk closes: you describe the person in plain English ("engineers who have authored or substantially contributed to an MCP server in the last 12 months, based in North America, currently at a company under 500 people") and get a ranked shortlist back with GitHub, LinkedIn, and email tied together. The handles-to-humans resolution is the hard part, and it is the part a Boolean on /in/ cannot do.

Three GitHub signals that actually predict AI fluency

The three durable, cross-platform artifacts that an AI-native engineer leaves on GitHub are MCP server authorship, AGENTS.md commits, and Claude Code hook configurations. Each one requires more agent literacy than installing an IDE plugin does.

  1. MCP server authorship or maintainer status. Writing a server against the protocol forces you to think in tool-calls, auth scopes, and schema design. Reviewing one forces you to catch other people's mistakes in that frame. Both are senior behaviors.
  2. AGENTS.md commits on real production repos. AGENTS.md is the emerging convention for telling coding agents how to behave inside a given codebase: what to run, what to avoid, where the tests live. Authoring one means you have thought about agent ergonomics for your team, not just for yourself.
  3. Claude Code hook configurations checked into the repo. Hooks are the mechanism Claude Code uses to run scripts at specific lifecycle points. Committing a hook config means you have moved past prompting into agent plumbing.

A fourth, weaker signal: contributions to the Cursor extensions ecosystem, Continue.dev, or Aider. Weaker because those are more consumer-side, but still a sharper filter than "has Cursor in their LinkedIn skills."

What this means for Saffron, Contrario, and the $656M wave

The AI-recruiting category raised roughly $656M across 10 US startups in the 12 months to July 2026, and most of that capital is chasing proprietary attribution. That is a solvable problem for hiring teams, but it is a structural dead end for sourcers.

The split looks like this:

VendorWhat they measureWhere the data livesSourcer access
Saffron (YC F2026)Prompt-level session eval in browser IDEHiring company's accountNone
Contrario ($6M ARR)AI-fluency rubric scoringCustomer ATSNone
Cursor EnterpriseAccepted-suggestion attributionWorkspace ownerNone
Mercor (>$2B ARR)Marketplace-side evalsMercorNone
GitHub MCP RegistryPartner server authorshipPublicFull
modelcontextprotocol/* reposPR authors, reviewers, approversPublicFull

The right read on the Saffron launch is not "another eval company." It is confirmation that the market now agrees agent fluency is the signal, which means the public proxies for it are about to get expensive. The sourcers who build their MCP-author shortlist in Q4 2026 are going to look prescient in Q2 2027.

$656M
US AI-recruiting funding across 10 startups, 12 months to July 2026
Almost all of it is chasing attribution data that never leaves the hiring company. The public shortlist is uncontested.

How to actually build the shortlist this week

Three concrete passes, in order, to go from "I want AI-fluent engineers" to a working list of named humans with contact details.

Pass 1: harvest the authors

Pull commit authors and PR reviewers from the GitHub MCP Registry's roughly 40 partner servers, plus the top 200 community servers in modelcontextprotocol/servers. Deduplicate on GitHub handle. Expect a few thousand handles, heavily skewed to a few hundred serious contributors.

Pass 2: filter for recency and non-trivial contributions

Drop handles whose only commit is a README typo. Keep handles with at least one merged PR of 50+ lines in the last 12 months. This step is where most "GitHub sourcing" projects die; it is also where Refolk earns its keep, because asking "contributors with 3+ substantive PRs on any MCP server since Jan 2026, not at Anthropic or OpenAI" in plain English is faster than scripting it.

Pass 3: resolve to humans

Join GitHub handle to LinkedIn, current employer, and location. This is where the gap between "I have handles" and "I have a sourcing list" closes. The 9-vs-217 ratio in Refolk's index is the reason you cannot start from LinkedIn and work backward; the headline data is not there. You have to start from the artifact and resolve forward.

A good sanity check: your final list should include handles that do not appear in any "AI engineer" LinkedIn search, because those are the people the rest of the market cannot find. If your shortlist and a LinkedIn Recruiter search return the same names, you have built the wrong list.

FAQ

Is MCP server authorship a strong enough signal on its own?

On its own, no. A toy MCP server written over a weekend is weaker than three substantive PRs on modelcontextprotocol/servers. Treat authorship as the entry filter, then rank on PR size, review activity, and whether the server is listed in the GitHub partner registry or has meaningful stars. The partner-vetted roster of ~40 servers is the strongest tier because Microsoft, GitHub, Dynatrace, Cloudflare, and the rest did the vetting for you.

What about engineers who use Claude Code or Cursor heavily but have no public artifacts?

They exist, and Saffron's product is designed precisely to find them inside a hiring funnel. But you cannot source them from the outside, because the signal is private by construction. The public GitHub approach is a precision-over-recall play: you will miss some AI-fluent engineers, and the ones you find will be disproportionately senior and disproportionately interested in tooling, which is usually the correct bias for the roles that justify this kind of sourcing effort.

How do I avoid the "AI reliance" trap Saffron is scoring against?

Weight authorship over usage, and weight review activity over commit count. An engineer who reviews MCP PRs is demonstrating judgment about agent design, which is exactly what AI Reliance Scoring is trying to isolate. An engineer whose only signal is "high Copilot acceptance rate" is the one that score will flag. Build your shortlist on behaviors that correlate with judgment, not with volume.

Why not just wait for Cursor's AI Code Tracking API to open up?

Because it will not. The API is Enterprise-only, Alpha, and the deeper Conversation Insights layer began charging in January 2026 on inference cost plus a Cursor token fee. Even if the pricing changes, the data is scoped to a single workspace root and never leaves the paying customer. Attribution APIs are a product feature for hiring teams, not a sourcing data source. Plan accordingly.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next