Refolk
October 9, 2026·9 min read

Baseten's $13B Round: 76 Inference Engineers Exist. It Needs 200.

Baseten's $1.5B Series F lit an inference hiring war. Here's where the vLLM, TensorRT-LLM, and SGLang engineers actually live in Q4 2026.

GPU inference engineer sourcingvLLM TensorRT-LLM talentBaseten hiringAI inference infrastructure recruitingmodel serving engineers
Baseten's $13B Round: 76 Inference Engineers Exist. It Needs 200.

Baseten closed a $1.5B Series F at a $13B valuation in June 2026, and the hiring plan behind it still defines Q4. The company is running north of $600M in annualized revenue and processing more than a billion inference calls a day, which does not happen without a lot of hands on vLLM, TensorRT-LLM, SGLang, and CUDA kernels. The problem is that the global pool of people who have actually shipped that stack in production is small enough to fit in one Slack channel.

The pool is smaller than Fireworks' headcount

In Refolk's index of professional profiles, exactly 76 people worldwide carry an explicit title of "Inference Engineer," "ML Systems Engineer," "Model Serving Engineer," or "GPU Engineer." Fireworks AI alone has 189 employees. The literal-title market is smaller than one direct competitor.

That gap is the whole story. Baseten's reported need is on the order of 200 inference hires. Set that against 76 titled candidates and you get roughly 2.6 open seats per living person before Fireworks, Together, Modal, Anyscale, NVIDIA, and the labs bid. The top employers of those 76 are not where you would guess from the fundraising headlines: Intel has 3, Apple has 3, Moreh has 2. There is no concentration to raid.

76
People globally with an "Inference Engineer" title
Refolk's index, October 2026. Baseten alone wants roughly 200 hires in this skill band.

A broader keyword pass ("vLLM," "TensorRT," "inference" anywhere in the profile text) returns 68 more profiles, top employers Amazon, Microsoft, Meta. Narrow it to the Venn of "inference," "GPU," "model serving," and CUDA skill and the index returns exactly 1 profile, at Meta. That is not a typo. The intersection of the four most-asked keywords on a Baseten job description, as actually written in profiles, has one person in it.

Why the title market is this thin

Three mechanics compress it:

  • The title "Inference Engineer" barely existed before 2023. The people doing the work show up as "Software Engineer, AI Infra," "ML Systems Engineer," or no inference word at all.
  • Fireworks itself hired most of its early team under "PyTorch internals," because that is what the founding DNA (ex-Meta PyTorch nucleus under Lin Qiao and Dmytro Dzhulgakov) looked like on paper.
  • The skill is new in half the stack. PagedAttention landed in 2023. RadixAttention and SGLang's scheduler work are newer than most promo cycles.

Sourcers who boolean on title will miss roughly 90% of the real market. The job is to find the people before they have the word.

Where the real shortlist actually lives: GitHub, not LinkedIn

The highest-signal shortlist in this market is public and already de-duplicated: it is the overlap of vLLM and SGLang contributors on GitHub. Both projects require the same rare skillset (PagedAttention and RadixAttention internals, CUDA, scheduler work), so the overlap is a self-selecting elite.

Per inclusion-ai.org, 194 developers have committed code to both vLLM and SGLang. That is about 30% of SGLang's total code contributor base, and it is roughly 2.5x the size of the entire explicit-title pool in Refolk's index. More importantly, every one of those 194 names is skill-verified by merged PRs. No ATS can match that signal.

SegmentCountTop employersSource
Explicit titles ("Inference Engineer" et al.)76Intel, Apple, MorehRefolk's index, Oct 2026
Keyword match on vLLM / TensorRT / inference68Amazon, Microsoft, MetaRefolk's index, Oct 2026
Inference + GPU + model serving + CUDA1MetaRefolk's index, Oct 2026
vLLM + SGLang dual contributors194Red Hat, xAI, Oracle, LinkedIn, Skyworkinclusion-ai.org
vLLM contributors in a single release cycle270PyTorch Foundation ecosystemmlai.qa, Sept 2026
Fireworks AI headcount (competitor baseline)~189-jobsbyculture.com, 2026

vLLM is now the most widely deployed open-source LLM inference engine, at roughly 91,000 GitHub stars, with 270 contributors in a single release cycle as of September 2026. 270 per release is the real recruitable number for in-house recruiters at Baseten, Fireworks, or Together, not the 2,000-plus lifetime committers that look impressive in a sourcing deck. People who ship into a release are people who still care.

How to turn the commit graph into a pipeline

  • Pull the author list for the last three vLLM minor releases and the last three SGLang releases. Dedupe by email and GitHub handle.
  • Enrich each handle with its most recent employer. Red Hat, xAI, Oracle, LinkedIn, and Skywork are the SGLang backbone; vLLM leans on Red Hat and the broader PyTorch Foundation ecosystem.
  • Flag anyone whose commits touch csrc/, attention/, or scheduler/. That is where the hard CUDA and scheduling work happens.
  • Cross-reference MLPerf Inference v6.0 submissions. Every MI355X and MI350X submission ran a PyTorch/ROCm stack, several naming vLLM explicitly. The submitter list is a public leaderboard of teams that actually ship.

This is the exact gap Refolk closes for inference sourcing: you describe the person in plain English ("committed to both vLLM and SGLang in 2026, based in the US or sponsorable, not currently at Fireworks or Together"), and Refolk reconciles the GitHub graph with LinkedIn and the open web into one ranked shortlist, instead of making you run four tools and a spreadsheet.

The Berkeley alumni tree is a ~50-person list

Two advisors generated most of this market, and their academic tree is finite enough to source exhaustively. Woosuk Kwon (vLLM creator) and Lianmin Zheng (SGLang creator) both came out of Ion Stoica's Berkeley lab, the same group that produced Spark and Ray. Sky Computing Lab and the former RISELab are the single densest geographic cluster of senior inference talent that exists.

The practical move is to build a named list of every PhD student, postdoc, and visiting researcher who has co-authored with Kwon, Zheng, or Stoica since 2021. It is roughly 50 people. Half are already at Databricks, Anyscale, or one of the labs. The other half are the people Baseten and Fireworks will spend the next year chasing. Source them before they update their LinkedIn headlines.

The literal-title market has 76 people. The commit graph has 194. The Berkeley tree has 50. Sourcers who only read titles are fishing with the wrong net.

The competitor map: who else is bidding for the same 194 people

Baseten is not the biggest wallet in the room. Fireworks AI closed a $1.505B Series D in July 2026 at a $17.5B post-money valuation, past $1B in annualized revenue, serving more than 40 trillion tokens per day, with over 95% of that volume on customer-specialized models. Together AI announced $800M in Series C funding in July 2026 for $1.3B total raised. Modal and Anyscale are smaller but hire out of the same Berkeley tree.

Here is how to think about the competitive pitch per candidate archetype:

  • Kernel engineers (CUDA, FireAttention-style custom kernels): Fireworks is the natural home and has the biggest cash advantage. Baseten's counter is multi-cloud breadth, since the platform spans multiple cloud providers rather than a single accelerator vendor.
  • Scheduler and serving engineers (PagedAttention, RadixAttention, continuous batching): The vLLM and SGLang dual-contributor pool. The pitch is production scale, which Baseten's 1B+ daily inference calls supports directly.
  • PyTorch internals: Fireworks' founding DNA (Lin Qiao, Dmytro Dzhulgakov). Hard to poach without a hook. The counter-move is Meta and NVIDIA alumni who did not join Fireworks in 2022.
  • ROCm / AMD-side inference: New surface as of MLPerf v6.0. AMD's inference org, OEM teams, and anyone who has shipped a ROCm vLLM path. Underfished.

The round's split-price structure (some investors in at $11B, others at the $13B headline) tells you capital is already pricing in execution risk, meaning talent risk. The firms that staff fastest get the next markup. The firms that staff slowest get written down.

2.6
Open inference seats per living titled candidate
200 reported Baseten hires against 76 titled profiles globally, before Fireworks, Together, Modal, and the labs bid.

Visa sponsorship is the hard constraint nobody writes down

A disproportionate share of the SGLang contributor base sits in the Chinese ML community, with first-class DeepSeek and Qwen support reflected in the project's roadmap and Skywork named in its contribution backbone. For US-based recruiters, this means visa capacity is the binding constraint on roughly half the usable pool.

Three practical implications:

  • Any US employer without an active H-1B transfer and O-1 pipeline is competing for a 90-person market, not a 194-person one.
  • Canadian and UK entities (Toronto, London, Dublin) are a real arbitrage for Baseten-adjacent companies with multi-country footprints. Processing is faster and the candidate pool is less bid-against.
  • Remote-from-origin arrangements with a US payroll are the fastest unlock. Several SGLang contributors already work this way for Oracle and LinkedIn.

Baseten's non-cash pitch has to carry the round

Baseten cannot outbid Fireworks on cash per head. Fireworks is at a higher valuation with a bigger token volume to point at. Baseten's edge, if the round is going to pay back, has to be the architecture pitch: multi-model, multi-cloud, and the customer logos that already run on it.

Three concrete hooks the Baseten recruiting team should be using in cold outreach this quarter:

  1. Multi-cloud breadth. Baseten runs across cloud providers, which matters to any engineer tired of being pinned to one accelerator roadmap. Fireworks' kernel work is NVIDIA-shaped. Baseten's is not.
  2. Scale as a selling point. 1B+ inference calls a day, roughly 1,900% YoY revenue growth to around $600M ARR by March 2026. The hard-problem density per engineer is high.
  3. Customer-engineer warm intros. Baseten's logo wall (Cursor, Mercor, OpenEvidence) is a sourcing surface. Engineers at those companies have already integrated against the stack and understand it well enough to interview without ramp. Pipe them.

For in-house recruiters at Baseten and its competitors, the sourcing math is the same regardless of logo: the pool is 194 people on a good day, visa capacity cuts that in half, and the firm that reconciles GitHub, LinkedIn, and the open web fastest gets the shortlist. Boolean on title returns 76 people and most of them are not available.

FAQ

How many "inference engineers" actually exist?

In Refolk's index of professional profiles, 76 people worldwide carry an explicit title of "Inference Engineer," "ML Systems Engineer," "Model Serving Engineer," or "GPU Engineer" as of October 2026. The top employers are Intel (3), Apple (3), and Moreh (2). The real recruitable pool is larger (the vLLM + SGLang dual-contributor overlap is 194), but anyone sourcing purely on job title is working with a sub-100-person market.

Why is the GitHub commit graph better than LinkedIn for this role?

Because the skills (PagedAttention internals, RadixAttention, CUDA kernel work, scheduler engineering) are too new to show up reliably in titles or self-reported skills. Merged PRs to vLLM and SGLang are a verified, dated, de-duplicated signal that someone has actually done the work. 194 developers have contributed to both projects, and 270 contributed to a single recent vLLM release cycle. Those lists are public and higher-signal than any ATS field.

Who are Baseten's biggest hiring competitors right now?

Fireworks AI ($1.505B Series D at $17.5B, July 2026), Together AI ($800M Series C, July 2026), Modal, Anyscale, NVIDIA's internal inference teams, and the labs (OpenAI, Anthropic, xAI). Fireworks is the most direct cash threat per head, with an ex-Meta PyTorch founding team under Lin Qiao and Dmytro Dzhulgakov. Together is the second-largest check in the space. Modal and Anyscale pull from the same Berkeley tree.

What's the fastest way to build a Baseten-style pipeline in Q4?

Start with the vLLM + SGLang dual-contributor list, cross-reference MLPerf Inference v6.0 submitters, and layer on the Ion Stoica academic tree out of Berkeley Sky Computing Lab. Enrich with current employer and visa status, then rank by whether their recent commits touch attention, kernels, or scheduling code. Tools that reconcile GitHub, LinkedIn, and the open web in one query collapse what used to be a four-tool workflow into a single plain-English prompt.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next