Together AI's $800M Series C: The Real Inference Pool Is 129
Together AI raised $800M on July 1, 2026. The vLLM plus Rust plus CUDA pool is 129 globally, and the contestable slice is under 20.
On July 1, 2026, Together AI announced an $800M Series C at an $8.3B post-money valuation and said it is hiring across engineering, research, product, and GTM. The live JDs describe a stack (Rust axum/tokio service layer, CUDA/Triton kernels, vLLM or TensorRT-LLM internals, speculative decoding) that almost nobody has end-to-end. If you are trying to hire against that spec, the pool is smaller and stranger than the headcount ambition implies.
How big is the actual Together AI inference pool?
The realistic full-stack pool is 129 engineers globally, and the contestable US slice is in the single digits. That is not the headline number Together's own recruiting funnel will chase, but it is the number that matters if you are bidding against them.
In Refolk's index of professional profiles, roughly 1,568 people mention vLLM or inference in their headline or experience. That is the outer pool: anyone with any exposure. When you require both Rust and CUDA as listed skills and vLLM-adjacent inference work in the free text, the count collapses to 129 worldwide. Only 3 of those 129 sit in the San Francisco Bay Area today. The rest are spread across Bengaluru, Toronto, San Jose, London, Lisbon, and Buenos Aires.
The 129 number matters because it reframes the fight. Together AI just raised more capital than any open-source inference company in the cycle and is targeting a 50-fold capacity increase over five years, on annual bookings that surpassed $1.15bn last quarter. But it is bidding into a pool the size of a mid-sized engineering org, most of which is already employed at counter-offer-heavy incumbents.
Why the 129 becomes fewer than 20 in practice
The contestable slice of the 129 is roughly 15 to 20 engineers, because the top employers of this cohort counter-offer aggressively and pay in liquid equity. You are not sourcing against a pool of 129; you are sourcing against the fraction of 129 that is willing to move.
Named concentrations inside the 129, by current employer:
- NVIDIA: 3
- Apple: 2
- Microsoft: 2
- Databricks: 1
- Intel: 1
The remaining 120 are diffuse across a long tail of chip startups, frontier labs, and boutique inference shops. That diffusion is good news (no single competitor owns the pool) and bad news (there is no cluster to raid). Together's own JDs are competing for the same people who read NVIDIA's Principal SWE, AI Inference posting, which lists a base range of $272,000 to $431,250 with liquid stock on top.
Two mechanisms push the contestable number below 20:
- Counter-offer economics. NVIDIA, Apple, and Microsoft can refresh RSUs at spot and match base. A pre-IPO Series C offer, even at Together's $8.3B mark (up 2.5x from the $3.3B Series B roughly 16 months earlier), cannot always clear that.
- Geographic friction. If the JD says SF onsite, you lose the Bengaluru, Toronto, London, and Lisbon density in the 129 immediately. That is where the numeric collapse to single digits happens.
What Together AI's JDs actually require
Together's Machine Learning Engineer, Inference role asks for knowledge of TGI, vLLM, TensorRT-LLM, Optimum, speculative decoding, CUDA and Triton programming, with Rust, Cython, and compilers as nice-to-haves. The Core ML (Turbo) role goes further: modifying production inference systems like SGLang or vLLM serving stacks, and shipping speculative decoding systems such as ATLAS.
Translated to sourcing signals, the stack decomposes as follows:
| Layer | JD language | Sourcing signal |
|---|---|---|
| Serving | Rust, axum, tokio | GitHub Rust repos, not LinkedIn "Rust" skill |
| Kernels | CUDA, Triton, CUTLASS | FlashAttention or paged-attention commits |
| Runtime | vLLM, SGLang, TensorRT-LLM | Contributor graphs on vllm-project, sgl-project |
| Algorithms | Speculative decoding, ATLAS, EAGLE, Medusa | Paper authorship, "draft model" in profile |
| Compilers | Cython, TorchInductor | PyTorch, Triton commits |
The trap is treating this as a keyword-AND search on LinkedIn. Nobody in the 129 wrote their profile that way. The Rust identity in this cohort is adjunct, not primary. Refolk's index returns exactly 1 profile globally where "Rust" and "LLM inference GPU" appear as a primary identity. One.
Why Rust is the false signal, not the bottleneck
Rust is a filter that removes the right people. The engineers Together actually wants are C++/CUDA specialists who picked up Rust for the axum/tokio serving layer over a weekend, and they do not list Rust as a top skill on LinkedIn.
If you search "Rust ML engineer" or "Rust inference engineer," you will find backend engineers who wish they worked on models. If you search FlashAttention contributors, paged-attention authors, CUTLASS committers, and TensorRT-LLM maintainers, then screen for whether the person has shipped a Rust service, you hit the actual pool. Together's own research lineage (FlashAttention, Hyena, FlexGen, RedPajama) is a good starting map of where those contributors cluster.
This is the exact gap Refolk closes for inference sourcing: you describe the person in plain English (someone who has shipped speculative decoding in a vLLM-style stack and can hold their own writing an axum service), and Refolk returns a ranked shortlist across GitHub, LinkedIn, and the open web. The tool does not care whether the candidate remembered to list "Rust" as a skill.
The four open-source graphs that beat LinkedIn here
The highest-signal sourcing surface for Together-style inference engineers is the open-source contribution graph, not LinkedIn skills. NVIDIA's own Principal Inference JD explicitly rewards substantial open-source contributions to vLLM, SGLang, PyTorch, Triton, and NCCL. If NVIDIA screens on the commit graph, so should you.
The four repositories that matter most, in order of signal density:
- vllm-project/vllm. The reference open-source serving stack. Maintainers and repeat contributors are the closest thing to a public list of the 129.
- sgl-project/sglang. SGLang has become the discriminator for speculative decoding work. Contributor overlap with vLLM is small, which means SGLang commits identify a distinct sub-pool.
- NVIDIA/TensorRT-LLM. Contributor list skews toward NVIDIA employees, but the external contributors are prime candidates.
- Dao-AILab/flash-attention and NVIDIA/cutlass. The kernel graph. This is where the CUDA-native cohort lives before they surface on serving projects.
LinkedIn will not tell you who merged a paged-attention patch to vLLM in March. GitHub will. Refolk indexes both so a single query returns the union.
Speculative decoding is the real discriminator
Filtering on vLLM returns 1,568 people. Filtering on speculative decoding shippers returns a fraction of that, and that fraction is the pool Together actually pays for. NVIDIA has reported that TensorRT-LLM support for speculative decoding now provides over 3x the speedup in total token throughput, which is exactly why every serious inference company is bidding for the same authors.
The vocabulary to filter on, in profile text and commit history:
- ATLAS (Together's own speculative decoding system, called out in the Core ML JD)
- EAGLE
- Medusa
- "draft model," "target model"
- "self-speculative"
Anyone whose GitHub history includes a PR touching a draft model implementation in vLLM, SGLang, or TensorRT-LLM is inside the pool that Together, NVIDIA, Fireworks, Baseten, Modal, Anyscale, and Perplexity are all bidding for.
The geography arbitrage Together's JD posture leaves open
The 129 pool has meaningful density in Bengaluru, Toronto, London, and Lisbon, and Together's SF-first JD posture leaves those candidates largely unrecruited. A remote-friendly competitor can hire 3 to 5 of them before Together's recruiters route the profiles.
Only 3 of the 129 are in SF Bay. If you insist on onsite, you are competing for those 3 against NVIDIA, Databricks, Fireworks, and every seed-stage inference startup with a Sand Hill lead. The unit economics of that fight are terrible.
Together is bidding into a 20-person contestable market, not 129, and definitely not 1,568.
Who else is bidding for the same 129
Together AI is the loudest bidder this quarter, but the same 129 profiles are being worked by at least eight other buyers who post nearly identical JDs. The bidding war is not Together versus one incumbent; it is Together versus a rotating cast of eight to ten.
Named competitors for the pool, based on JD overlap:
- NVIDIA (Principal SWE, AI Inference, $272K to $431K base)
- Databricks (post-MosaicML inference team)
- Fireworks AI
- Baseten
- Modal
- Anyscale
- Perplexity
- SGLang team at LMSys
Add Together's own named customers (Cognition, Decagon, Eleven Labs, Cursor, Suno), all of whom are also hiring inference engineers to reduce their Together spend, and the effective bidder count inside the 129 pool is closer to a dozen. Aramco Ventures leading Together's round, alongside Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, March Capital, Pegatron, and SentinelOne's S Ventures, signals Gulf sovereign money flowing into US inference infra, which is a new comp-inflation vector worth pricing into your 2026 offers.
The five moves that beat Together's recruiting funnel
If you cannot outspend NVIDIA and you cannot out-brand Together, you win by sourcing off different signals faster. The five moves that actually work against a pool this small:
- Source off commit graphs, not LinkedIn skills. Start with vllm-project/vllm, sgl-project/sglang, and NVIDIA/TensorRT-LLM contributor lists.
- Screen in Rust, do not filter on it. Look for C++/CUDA engineers who have shipped any Rust service, not for Rust-primary engineers.
- Weight speculative decoding evidence heavily. ATLAS, EAGLE, Medusa, draft-model PRs. This filter shrinks 1,568 to the real pool fast.
- Post remote-friendly. The 126 of 129 who do not live in SF are your addressable market.
- Move in seven days, not thirty. If you can get an offer out inside a week of first contact, you win against a slower incumbent even at lower comp.
Sourcing tools that treat this as a plain-English query, Refolk among them, cut the first move from a week of Boolean iteration to a single prompt. That time saving is what turns a 20-person contestable market from impossible to workable.
FAQ
How many engineers globally match Together AI's inference JD?
Roughly 129 profiles in Refolk's index match the full stack (Rust and CUDA skills plus vLLM-adjacent inference work in their text). About 1,568 profiles mention vLLM or inference at any level, but that outer pool includes many candidates without the kernel or serving depth Together's JDs require. The contestable slice, after accounting for counter-offers from NVIDIA, Apple, and Microsoft, is closer to 15 to 20 people.
Is Rust really required for vLLM engineer roles?
Rust is listed as a nice-to-have on Together's Machine Learning Engineer, Inference JD and shows up in axum/tokio service layer work. In practice it is not the bottleneck skill. Refolk's index returns roughly 1 profile globally where "Rust" and "LLM inference GPU" appear as a primary identity. The right pool is C++/CUDA engineers who picked up Rust for serving code, and they will not surface if you filter on Rust as a top skill.
Where should I source inference engineers besides LinkedIn?
Start with the open-source contribution graph: vllm-project/vllm, sgl-project/sglang, NVIDIA/TensorRT-LLM, Dao-AILab/flash-attention, and NVIDIA/cutlass. NVIDIA's own Principal AI Inference JD explicitly rewards contributions to those repos, which means the highest-signal candidates are already public on GitHub. Refolk indexes GitHub, LinkedIn, and the open web together so a single plain-English query returns the union.
What comp should I expect to pay against Together AI?
Use NVIDIA's Principal SWE, AI Inference range as the anchor: $272,000 to $431,250 base, plus liquid RSUs. Together AI's $8.3B post-money valuation means their equity is illiquid but priced aggressively. To win a candidate away from an incumbent with liquid stock, plan on matching base and offering meaningful early-exercise equity, or on winning speed and mission rather than cash.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.