Fireworks Just Raised $1.5B. The Titled US Inference Pool Is 5.
Fireworks AI plans to triple to 600 engineers. The US pool titled "Inference Engineer" is 5. Here is where to actually source them.
Fireworks AI just closed a $1.5B Series D at a $17.5B valuation and told the market the money is going into GPUs and engineers. The founding template (Lin Qiao pulling her old PyTorch lieutenants out of Meta) worked once. It cannot work again, because the pool it drew from is roughly the size of one team offsite.
How small is the actual ex-PyTorch inference pool?
Small enough to fit on a single Zoom grid. In Refolk's index, only 5 US professionals currently carry the literal title "Inference Engineer," "LLM Inference Engineer," or "Model Serving Engineer." Broadening to any US profile that lists "PyTorch inference" in headline or skills returns 34.
That is the entire titled surface area for a company that plans to net-hire around 400 engineers by year end. It is also the pool Together AI, Baseten, NVIDIA's TensorRT-LLM org, and Cursor's Composer team are all raiding at the same time.
The mechanism is worth stating plainly. PyTorch was a framework team, not an inference-serving team. Lin Qiao managed more than 300 engineers at Meta between July 2015 and September 2022, but only a 7-person nucleus (Dmytro Dzhulgakov, Dmytro Ivchenko, James Reed, Benny Chen, Chenyu Zhao, Pawel Garbacki, and Qiao herself) co-founded Fireworks. The Venn overlap of "shipped PyTorch internals" and "ran production model serving at trillion-token scale" is the real bottleneck, and it was already close to exhausted the day Fireworks incorporated.
Why did Fireworks' round make the shortage worse overnight?
Because the round funded three things at once (headcount, GPUs, and enterprise deployment) and every dollar chases the same 34-person keyword pool. Fireworks surpassed $1B in annualized revenue, up 5x year-over-year, and nearly tripled daily token volume from 15 trillion to more than 40 trillion. The company has roughly 200 employees today and plans to triple headcount by year end.
Here is what the recruiter-facing math looks like:
| Slice | Count | Source |
|---|---|---|
| US, exact title "Inference Engineer" / "LLM Inference Engineer" / "Model Serving Engineer" | 5 | Refolk index |
| US, "PyTorch inference" keyword + PyTorch skill | 34 | Refolk index |
| US, "vLLM / TensorRT / inference serving" keyword | 4 | Refolk index |
| Fireworks founding inference core (ex-PyTorch) | 7 | Index Ventures |
| Fireworks current → year-end target headcount | ~200 → ~600 | Company plan |
| Planned net new hires vs. titled US pool | ~400 / 5 | Derived |
The Series C was $250M at a $4B valuation in October 2025. Nine months later the valuation is $17.5B, a roughly 4.4x jump. Compensation bands move with valuation. Anyone on that 34-person list who was slow-playing a conversation in Q1 is now a counteroffer risk for whoever hires them next.
Who is Fireworks actually competing with for these engineers?
Three groups, and the third is the one most recruiters miss. Together AI and Baseten are the obvious direct competitors in the AI inference cloud market. NVIDIA is inside the Series D cap table and expanding its own inference org (Triton, TensorRT-LLM) through acquisitions, so expect counteroffers from the same investor that just wrote the check.
The non-obvious competitor is Cursor. As of last year, roughly half of Fireworks' revenue came from Cursor. Cursor is now scaling its Composer model on external compute and priced Composer 2 at $0.50 per million input tokens and $2.50 per million output tokens, competing directly with Fireworks-hosted models on unit economics. Cursor still accesses Kimi K2.5 through Fireworks' hosted RL and inference platform, so the relationship is not clean, but the hiring signal is: Cursor's ML systems team is targeting the same 34-person pool, and they are doing it from a customer's balance sheet.
The competitive set looks like this:
- Together AI, Baseten: direct inference-cloud competitors, mutual raid targets.
- NVIDIA (Triton, TensorRT-LLM): Series D participant and hiring rival, with the deepest counteroffer capacity.
- Cursor Composer team: former customer building in-house model serving.
- Hyperscaler ML systems (AWS, Microsoft, AMD): where most of the 34-profile pool actually sits today.
Fireworks' biggest customer is also its biggest hiring competitor, and it is paying with Fireworks' own revenue.
Where do you actually find model serving engineers?
Stop searching titles. Search shipped artifacts. The 34-person "PyTorch inference" keyword pool is real, but it undercounts because the skill does not map to any standard LinkedIn title. The people you want are named in commit histories, not job titles.
Five sourcing surfaces that actually convert:
- vLLM contributors: the open-source LLM inference engine. Commit graphs surface engineers currently at AWS, AMD, Samsung Semiconductor, and Scale AI. Refolk's index shows Meta contributes only 2 of the 34 US "PyTorch inference" profiles; the rest are scattered across exactly these companies.
- SGLang contributors: newer, smaller contributor set, and less picked over than vLLM. Fan Zhao's team at Intel (Senior Director, AI Frameworks) publicly optimizes both vLLM and SGLang. That entire team is a raid target that will not appear on any "ex-Meta" filter.
- TensorRT-LLM and Triton contributors: NVIDIA-adjacent, and the engineers who leave NVIDIA for a startup are usually the ones already frustrated with internal roadmap velocity.
- PyTorch Governing Board and core contributor rosters: public list of the operators behind the framework, most of whom are not at Meta.
- The Fireworks / Together / Baseten diaspora itself: three-year-old startups now have their first attrition wave. This is where the next pool comes from.
Describing this cohort in a keyword search is where most sourcing stacks break. This is the exact gap Refolk closes: you describe the person in plain English ("US engineer who has shipped vLLM PRs and worked on batched inference at a hyperscaler") and get a ranked shortlist across GitHub, LinkedIn, and the open web without stitching together three tools.
What signals separate a real model serving engineer from an ML generalist?
Ship history in one of four sub-disciplines, and production scale in at least one of them. "Model serving engineer" is a narrower skill than "ML engineer," which is why title-based sourcing returns 5 people and keyword sourcing returns 34. The people who can actually do the job have one or more of the following on their resume:
- Quantization: INT8, FP8, AWQ, GPTQ. Look for CUDA kernel work or contributions to bitsandbytes, AutoGPTQ, or llama.cpp.
- Batching and scheduling: continuous batching, PagedAttention, prefix caching. vLLM and SGLang PRs are the cleanest signal.
- Speculative decoding and multi-token prediction: Medusa, EAGLE, lookahead decoding contributions.
- Multi-model routing: engineers who have run a fleet of fine-tuned model variants behind one endpoint. More than 95% of tokens Fireworks serves come from models specialized on customers' proprietary data, which means the routing layer is now more valuable than any single model.
That last point reframes the entire hiring plan. When 95% of served tokens hit custom models, you do not need pretraining talent. You need people who can quantize, batch, speculatively decode, and route. Most job descriptions still ask for the wrong thing, which is another reason the titled pool looks like 5.
What does the second-order diaspora actually look like?
It is the highest-signal, lowest-competition source right now, and almost nobody is hunting it systematically. Fireworks, Together AI, and Baseten are all past their first three years. Early engineers vest, some leave, and the ones who leave carry exactly the production model-serving experience the rest of the market is starving for.
The second-order sourcing plan looks like this:
- Fireworks alumni: even at ~200 employees, a vested cohort exists from the pre-Series C era.
- Together AI alumni: earlier stage, longer tenure, and less picked-over.
- Baseten alumni: the most enterprise-focused of the three, and their engineers tend to have real deployment scars.
- Ex-PyTorch engineers not at Fireworks: Soumith Chintala, PyTorch co-founder, is now at Thinking Machines and NYU. The rest of that graph is small and largely mapped, but worth walking.
- Cursor Composer team alumni: not a live pool yet, but the moment priorities shift after Composer's next release, it becomes one.
Walking that graph by hand takes weeks. Describing it in a single query and letting the tool traverse GitHub, LinkedIn, and the open web is where Refolk earns its keep, especially when the recruiter running the search has never shipped a CUDA kernel and cannot judge a commit history on sight.
How should a founder or head of talent actually run this search?
Run three parallel sourcing lanes, not one. The mistake is treating "LLM inference engineers" as a single pool. It is three pools with different half-lives and different close rates.
- Titled pool (5 people, US): fully mapped in an afternoon. Assume every one of them is in conversation with Fireworks, Together, Baseten, and NVIDIA already. Do not build a hiring plan around them.
- Keyword pool (34 people, US): workable in a quarter. Prioritize by commit recency and current employer stability. AWS and Microsoft profiles are more portable than AMD or Samsung Semiconductor profiles.
- Artifact pool (hundreds, global): vLLM, SGLang, TensorRT-LLM, and Triton contributor graphs. This is where the next hire actually comes from, and it is invisible to any title-based search.
For Fireworks AI hiring specifically, the math forces lane three. You cannot triple to 600 engineers by fishing a 5-person pond. For everyone else chasing ex-Meta AI infrastructure talent, the takeaway is quieter and more useful: the "ex-Meta PyTorch alumni" filter that seemed like a cheat code two years ago is now the most crowded query in the market, and the actual talent has already routed around it.
FAQ
How many ex-PyTorch engineers are actually available to hire?
The cohort large enough to have shipped inference at scale is the 7-person founding nucleus at Fireworks plus a long tail of engineers now distributed across OpenAI, Google, and Meta itself. Even generously counting every ex-PyTorch engineer at an inference startup, the pool is well under 100. The framework team was large (300+ under Lin Qiao at Meta) but the subset that shipped production model serving at trillion-token scale is the actual bottleneck, and most of them are already placed.
Is title-based sourcing enough for LLM inference roles?
No. Refolk's index shows only 5 US professionals with the literal title "Inference Engineer," "LLM Inference Engineer," or "Model Serving Engineer." The skill does not map cleanly to any standard LinkedIn title, so title-based sourcing systematically undercounts. Keyword searches on "PyTorch inference," "vLLM," "TensorRT-LLM," and "SGLang" plus GitHub contributor graphs are the only way to build a realistic pipeline.
Why is Cursor a hiring rival if it is also a customer?
Because Cursor is building Composer, an in-house model priced at $0.50/$2.50 per million input/output tokens, and it needs the same model-serving engineers Fireworks needs. Roughly half of Fireworks' revenue historically came from Cursor, and Cursor is now scaling Composer on external compute while continuing to use Fireworks for RL training on Kimi K2.5. That makes Cursor's ML systems team a direct competitor for the 34-person "PyTorch inference" keyword pool, funded partly by Fireworks' own revenue.
Where should recruiters look besides Meta alumni?
The Intel AI Frameworks team under Fan Zhao (publicly optimizing vLLM and SGLang), the vLLM and SGLang contributor graphs on GitHub, NVIDIA's Triton and TensorRT-LLM orgs, and the second-order diaspora of Fireworks, Together AI, and Baseten themselves. Refolk's index shows Meta contributes only 2 of 34 US "PyTorch inference" profiles; the other 32 sit at AWS, Microsoft, AMD, Samsung Semiconductor, and Scale AI, none of which appear in an "ex-Meta" filter.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.