Refolk
August 26, 2026·9 min read

Fireworks' 400 Hires: The Ex-PyTorch Inference Pool Is Tiny

Fireworks, Together, and Baseten raised $3.8B in four weeks. The ex-PyTorch inference pool they are all chasing is measured in dozens.

ex-PyTorch engineers hiringAI inference engineer sourcingFireworks AI hiringvLLM contributor recruitingspecialized inference engineers
Fireworks' 400 Hires: The Ex-PyTorch Inference Pool Is Tiny

On July 16, 2026, Fireworks AI closed a $1.505B Series D at a $17.5B valuation, and CEO Lin Qiao told CNBC she plans to grow the company from ~200 to 600 employees by year-end. Together AI ($800M at $8.3B) and Baseten ($1.5B at a $13B headline) closed in the same four-week window. Four buyers, roughly $3.8B in fresh capital, and a pool of engineers who can actually build production disaggregated inference that is small enough to fit on a single page.

The math problem hiding inside a 400-person hiring plan

Fireworks needs to hire 400 people in about five months, but the intersection of "PyTorch-native" and "production inference serving" that its technical moat depends on is a pool of dozens, not thousands. The headline number is a GTM plan; the constraint is a bench that fits in one room.

Qiao ran the PyTorch team at Meta from July 2015 to September 2022, managing 300+ engineers. That team is the upstream supply for every serious inference startup in the market right now. Fireworks' technical claims - custom CUDA kernels, model sharding, semantic caching, 12x faster than vLLM - are not generic ML work. They require compiler, kernel, and framework internals experience that most "ML engineers" have never touched.

Here is the comparable data, pulled from Refolk's index of professional profiles plus public disclosures:

SegmentCountSource
US engineers, PyTorch skill + Senior/Director/VP seniority, tight filter~1 in highest-signal cutRefolk index
US engineers, PyTorch + title contains "Inference"/"ML Systems"/"Model Serving"/"AI Infrastructure"2 (top employers: Modal, SpaceXAI)Refolk index
Fireworks planned hires by Dec 2026400 (200 to 600)CNBC
Combined raise: Fireworks + Together + Baseten in 4 weeks$3.8BForbes
PyTorch team size Qiao led at Meta pre-2022300+TechFundingNews
Publicly named ex-PyTorch co-founders/leads across the 4 startups<20Founder pages

The Refolk numbers look absurdly low because the strict intersection really is that narrow. Broader adjacent pools (any "ML engineer" who has touched PyTorch) run to the tens of thousands. Those people cannot ship a custom attention kernel on a deadline. That is why the true elite pool is directionally in the dozens, and why the four buyers are structurally cornered.

$190M+
Dollars raised per publicly named ex-PyTorch engineer
Derived from $3.8B across Fireworks, Together, and Baseten divided by fewer than 20 publicly named co-founders and technical leads.

Why "four buyers, one pool" is only 30% true

The four inference startups compete for the same engineers on roughly the top third of roles, and then their needs diverge sharply. Treating them as identical buyers is what leads sourcers to burn outreach on the wrong people.

The differentiation matters because it changes who you should be poaching from where:

  • Fireworks: PyTorch-native optimization, custom CUDA kernels, semantic caching. Needs compiler and kernel engineers.
  • Together AI: Model agnosticism, neocloud economics. Needs distributed systems and scheduler people.
  • Baseten: Truss deployment, VPC, multi-tenant. Needs Kubernetes, compliance, and platform engineers.
  • Groq: LPU hardware. Needs silicon-adjacent engineers, a mostly separate pool.

The narrow overlap - the top ~30% of infra roles at Fireworks, Together, and Baseten - is where the ex-PyTorch inference pool actually gets fought over. The other 70% is differentiated demand, which is good news if you are sourcing for Baseten and bad news if you assumed the inference market was one homogeneous pond.

The 400 reqs are not 400 inference engineers

Qiao told CNBC Fireworks will build "a formidable sales team after years of having customers sign themselves up." Fireworks also hired former Salesforce exec George Hu as President in April 2026. That is a heavy GTM signal. A realistic split of the 400 reqs looks like:

  • 60 to 120 deep-infra roles (the ones that require the true ex-PyTorch/kernel pool)
  • 150+ GTM, forward-deployed, and sales engineering roles
  • The remainder in SRE, security, product, and G&A

The competition for the actual ex-PyTorch tier is smaller in absolute reqs than the "400 new hires" headline implies. It is also 100% concentrated across four buyers, which is a worse dynamic than if it were spread across forty.

NVIDIA is cross-recruiting itself

NVIDIA sits on the cap table of Fireworks, Together, and Baseten - it contributed $150M of Baseten's January 2026 round led by IVP and CapitalG. Which means NVIDIA's own DevRel, TensorRT-LLM, and Inference Microservices teams are the shared feeder pool being drained by NVIDIA's own portfolio companies.

If you are sourcing for any of the four, watch these leading indicators:

  1. TensorRT-LLM team departures on LinkedIn and GitHub commit history.
  2. NVIDIA Inference Microservices (NIM) solutions architects going quiet on public channels.
  3. CUDA MODE Discord contributors who suddenly stop posting internal NVIDIA context.
  4. PyTorch Conference speaker lists from 2024 to 2026, especially anyone with a Meta or NVIDIA affiliation.

NVIDIA's TensorRT-LLM team is one of the largest concentrations of production inference expertise outside of Meta's former PyTorch org. Every departure is a signal that one of the four buyers just closed a role you were competing on.

The Refolk numbers look absurdly low because the strict intersection of PyTorch-native and production inference really is that narrow.

The shadow pool: vLLM, SGLang, and TensorRT-LLM contributors

The most under-sourced group in this market is the top GitHub contributors to open-source inference runtimes, not ex-Meta employees. Fireworks explicitly claims 12x faster than vLLM, which means the moat is beating vLLM engineers, not exclusively hiring ex-PyTorch alumni.

The three communities worth mapping today:

  • vLLM (UC Berkeley Sky Lab): top 50 GitHub contributors by commits to core kernel and scheduling paths.
  • SGLang (LMSYS): contributors who have shipped structured generation and radix attention work.
  • TensorRT-LLM (NVIDIA): external contributors and issue authors with credible production context.

Mapping these three lists by hand takes days. Filtering by "based in US, currently at a company that is not one of the four buyers, seniority >= Senior, has shipped a merged PR in the last 12 months" takes longer. This is the exact gap Refolk closes: describe the person in plain English and get a ranked shortlist across GitHub, LinkedIn, and the open web without stitching filters together in three tools.

The already-placed tier you cannot poach

A meaningful fraction of the ex-PyTorch universe is already off the market at destinations that will not be outbid by a Series D. Sourcing against them is wasted outreach.

Concrete examples of the already-placed tier:

  • Soumith Chintala (PyTorch co-founder, now Thinking Machines + NYU). Not moving.
  • Dwarak Rajagopal (VP AI Engineering, Snowflake; ex-Meta PyTorch Core Frameworks lead, ex-Google AI Frameworks). Hyperscaler-adjacent, comp locked.
  • Dmytro Dzhulgakov (Fireworks CTO, PyTorch core maintainer, co-created ONNX). Founder equity.
  • Dmytro Ivchenko, James Reed, Pawel Garbacki, Chenyu Zhao, Benny Chen (Fireworks early team). All at Fireworks.

Once you strip out the already-placed tier and the current employees of the four buyers, the true addressable ex-PyTorch inference pool is the residual after you subtract every name that will never take your call. That residual is what makes the "dozens, not thousands" framing directionally correct.

Where the 400 reqs will realistically land

The 400 Fireworks reqs will not land where the press release implies. They will split into four rough buckets, each with a distinct sourcing playbook.

Realistic distribution and where the people come from:

BucketRough shareWhere they come from
Deep infra (kernels, compilers, PyTorch internals)15 to 30%Meta PyTorch alumni, NVIDIA TensorRT-LLM, vLLM/SGLang contributors
GTM and sales engineering30 to 40%Salesforce, Databricks, Snowflake, Confluent
Forward-deployed and solutions15 to 20%Palantir, Databricks Mosaic, Anyscale
SRE, platform, security, G&A20 to 30%Broad market

The deep-infra bucket is where the four-buyer squeeze actually bites. Everything else is a hard but normal recruiting problem.

Why EU startups lose this round

Two of Fireworks' named co-founders, Dzhulgakov and Ivchenko, are Ukrainian engineers who relocated to the US. A meaningful fraction of the global PyTorch-core alumni pool sits on non-immigrant US visas. That structurally advantages US-based buyers with mature immigration operations. EU inference startups chasing specialized inference engineers from this pool are competing with one hand tied behind their back.

The sourcing playbook for the next 90 days

If you are hiring against Fireworks, Together, or Baseten right now, work the shadow pool first and the ex-PyTorch tier second. The shadow pool is bigger, less contested, and often more current on the actual production techniques that matter.

A tactical 90-day plan:

  1. Map the top 50 contributors to vLLM, SGLang, and TensorRT-LLM by last-12-months activity. Rank by seniority signal.
  2. Watch NVIDIA departures weekly. Anyone leaving TensorRT-LLM, NIM, or DevRel is a 30-day hot lead.
  3. Skip the already-placed tier. Do not waste cycles on named PyTorch co-founders and staff at Thinking Machines, Snowflake AI, or Databricks Mosaic.
  4. Segment by buyer need. Baseten reqs pull from Kubernetes and multi-tenant infra pools; Fireworks reqs pull from kernel and compiler pools. They are not interchangeable.
  5. Use plain-English search for intersections that structured filters cannot express, like "shipped a merged PR to a top-3 inference runtime in the last year and is currently at a non-inference company." Refolk was built for exactly this shape of query and returns ranked shortlists across GitHub, LinkedIn, and the open web in one pass.

The Fireworks headline reads like a hiring war. The actual constraint is that the pool of ex-PyTorch engineers hiring managers actually want is small enough to name in a spreadsheet. If you can name them before your competitor does, you win the round. If you cannot, no amount of budget fixes the math.

FAQ

How many ex-PyTorch engineers are actually available to hire?

The publicly named tier of ex-PyTorch technical leads across Fireworks, Together, Baseten, and adjacent startups is fewer than 20 people, and most are founders or already placed. The broader addressable pool of engineers who genuinely shipped PyTorch internals or production inference at Meta scale is in the dozens, not thousands. Refolk's tightest filter on PyTorch skill combined with inference-serving titles surfaces only a handful of profiles, which confirms the pool is measured in dozens.

Should sourcers focus on ex-Meta PyTorch alumni or open-source contributors?

Both, but open-source contributors to vLLM, SGLang, and TensorRT-LLM are the more under-sourced group in 2026. Fireworks explicitly benchmarks 12x faster than vLLM, which means the technical moat depends on beating those engineers, and the top GitHub contributors are often not on any recruiter's LinkedIn list. Map the top 50 contributors to each runtime and rank by seniority before touching ex-Meta lists that every competitor is already working.

Why does NVIDIA matter to sourcing against Fireworks?

NVIDIA sits on the cap table of Fireworks, Together, and Baseten, which means its own TensorRT-LLM, NIM, and DevRel teams are the shared feeder pool being drained by its portfolio. Departures from those specific NVIDIA teams are the strongest leading indicator that a competitor closed a role in the ex-PyTorch tier. Track them weekly.

Are all 400 Fireworks reqs going to inference engineers?

No. Fireworks hired ex-Salesforce exec George Hu as President in April 2026 and Qiao explicitly told CNBC the company is building a "formidable sales team." A realistic estimate is that 60 to 120 of the 400 reqs are deep-infra roles, and the rest are GTM, forward-deployed, SRE, and G&A. The four-buyer squeeze on the true ex-PyTorch pool is real, but it is smaller in absolute reqs than the headline number implies.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next