Together AI's $8.3B Round Is Hunting 194 Engineers, Not Thousands
Together AI just raised $800M at $8.3B to scale open-source LLM inference. The pool that can actually ship the kernels is under 200 engineers worldwide.
On July 1, 2026, Together AI closed an $800 million Series C at an $8.3 billion valuation, led by Aramco Ventures with Vista Equity Partners, General Catalyst, NVIDIA, March Capital, and S Ventures. The press release reads like a hiring boom. The underlying labor market reads like a knife fight over roughly 194 people.
If you are recruiting against this raise, or against the neocloud wave more broadly, the interesting number is not the $800M. It is how few engineers on Earth have actually shipped production code to both vLLM and SGLang, the two open-source inference engines that make Together AI's platform faster than the alternative.
What the $800M is actually buying
Together AI's raise is compute plus a very specific kind of engineer, and the second half is the hard part. The company has committed to over 500 megawatts of capacity to fund roughly 50x growth over five years, but capacity is procurable. Kernel-level inference expertise is not.
A few anchors from the round:
- $800M Series C on July 1, 2026, at $8.3B post-money
- Up from a $305M Series B at $3.3B roughly 16 months earlier, a 2.5x valuation step
- Annual bookings crossed $1.15B last quarter
- Customers include Cursor, Cognition, and Decagon
- Together AI claims 31% more transactions per second than the next fastest open-source engine on production coding-agent workloads
That 31% TPS edge is the tell. It is not a marketing number, it is an artifact of paged-attention tuning, RadixAttention scheduling, custom CUDA kernels, and speculative decoding work. The people who can produce that edge are the same people whose GitHub handles show up in the vLLM and SGLang commit graphs. Everyone else in the neocloud bracket, Baseten ($1.5B at $13B in June 2026), Fireworks, Anyscale, is bidding for the same names.
The vLLM and SGLang pools are smaller than the headlines suggest
The honest count of engineers who can ship the kind of inference code Together AI is paying for is under 200 worldwide, not the tens of thousands implied by GitHub stars. Here is the ladder, from vanity metric to real signal.
vLLM is the Berkeley-originated open-source LLM inference engine, seeded by Woosuk Kwon in Ion Stoica's lab. SGLang is its structured-generation cousin, seeded by Lianmin Zheng out of the same lineage and LMSYS.org. Both projects hit tens of thousands of stars fast. Neither has a large committer base.
| Segment | Count | Source |
|---|---|---|
| vLLM PR-submitters, cumulative | ~2,000 | vLLM project reporting |
| Implied total SGLang code contributors | ~647 | Derived from inclusion-ai.org |
| Developers who shipped code to BOTH vLLM and SGLang | 194 | inclusion-ai.org |
| US profiles matching "vLLM inference engineer" | 334 | Refolk index |
| Global profiles matching "vLLM contributor inference" | 27 | Refolk index |
Two ratios matter more than the raw counts:
- Only about 1 in 10 vLLM PR-submitters (194 of ~2,000) has also shipped SGLang code. That dual-project cohort is the honest proxy for "understands paged attention and RadixAttention scheduling."
- The broad US "vLLM inference engineer" pool of 334 is only 1.72x the elite 194 dual-committer set. There is barely any depth behind the front line.
This is the exact gap Refolk closes for infra recruiters: instead of scraping GitHub for star-counters, you describe the engineer in plain English ("has merged PRs to vLLM in the last 12 months and shipped SGLang code") and get a ranked shortlist tied back to LinkedIn and the open web.
Where the 194 actually live
The elite inference cohort is concentrated in three metros, and two of them are not in the United States. In Refolk's index of professional profiles, the tightest "vLLM contributor inference" query returns 27 people globally, and the geographic split is not what US-centric recruiters expect.
Top regions in the 27-profile cohort:
- San Francisco Bay Area: 3
- Bengaluru: 3
- Hyderabad: 3
The Bay Area does not lead. It ties. Bengaluru and Hyderabad together outweigh it. That inverts the standard sourcing assumption that inference kernel work is a US-only market and it directly reprices remote and offshore packages. A Together AI recruiter posting an on-site SF requisition for this role is competing with three local candidates, not thirty.
The broader US pool of 334 does concentrate more predictably:
- SF Bay Area: 5
- San Francisco proper: 3
- Kirkland and Bellevue corridor: 3
- Top employers: Meta (3), Apple (3), NVIDIA, sgl-project
The Bay Area does not lead the elite inference cohort. It ties Bengaluru. It ties Hyderabad.
The employer graph tells you where to poach
The talent is flowing out of Big Tech into the open-source projects themselves, not the other way around. In the tightest 27-profile cohort, the top current employers are AMD (2), sgl-project, vLLM itself, Google, TikTok, IBM India, and Microsoft AI Innovators Hub. Two of the top four are the OSS projects, meaning they have already poached from the hyperscalers.
That reshapes the competitive set for Together AI:
- vLLM's backbone employer is Red Hat. Any recruiter mapping vLLM contributors should assume Red Hat is a primary current-employer node, then AMD, then the sgl-project and vllm-project GitHub orgs directly.
- SGLang's core contributors sit at xAI, Skywork, Oracle, and LinkedIn. These are not the companies most US recruiters cold-source against for inference work. They should be.
- Named individuals are already public. comaniac (OpenAI) has 77 PRs to vLLM and 17 early PRs to SGLang. CatherineSue (Oracle) went from 4 vLLM bug-fixes to 76 PRs on SGLang as a core contributor. These are the archetypes to build a lookalike search around.
The Ion Stoica lineage is a sourcing shortcut
If you only mine one graph for this hire, mine the UC Berkeley Sky Computing Lab and RISELab alumni network. Woosuk Kwon (vLLM) and Lianmin Zheng (SGLang) both trained under Ion Stoica, the same lab that spawned Spark and Ray. That is not a coincidence, it is a supply chain.
The practical implication for sourcing:
- Sky Computing Lab and RISELab alumni are pre-screened for the exact systems-plus-ML background that produces inference kernel authors.
- LMSYS.org contributors overlap heavily with SGLang. LMArena work is a strong secondary signal.
- Ray Summit vLLM track speakers and vLLM meetup co-hosts (a16z, IBM, Roblox, Cloudflare/BentoML, AWS, NVIDIA per the vllm-project GitHub) are a curated list of people already vetted enough to present publicly.
Recruiters running Boolean strings against LinkedIn will miss most of this because the signals live in GitHub commit graphs, arXiv co-authorship, and Discord roles, not in job titles. That is why Refolk indexes across GitHub, LinkedIn, and the open web at once: you ask for "Stoica lab alumni who have merged inference-related PRs since 2024" and get the answer, rather than three separate exports that never join cleanly.
SGLang is maintainer-constrained, and that is your leverage
Hiring two SGLang core committers meaningfully shifts an open-source project's roadmap, which no amount of Aramco capital buys at vLLM's scale. The tell is response time. Most vLLM issues get responses within 12 hours to 3 days. SGLang typically takes 3 to 5 days. The gap is not laziness, it is a smaller maintainer set relative to inbound load.
For a founder or eng leader, this creates two asymmetric moves:
- If you are Together AI or a direct competitor, hiring SGLang maintainers is worth a premium beyond their market comp because it buys informal roadmap influence on a project your platform depends on.
- If you are a smaller shop, the same math works at your scale. Two hires flip a project's responsiveness enough that your bug reports get triaged first, which compounds into faster production incident recovery.
The catch: SGLang's core contributor set is much smaller than vLLM's. Implied total code contributors are around 647, and the true maintainer subset is dozens, not hundreds. Every one of them is being courted right now.
What this means for the neocloud hiring wave
Every neocloud round in 2026 is chasing the same few hundred people, so speed and specificity beat budget. Together AI's $800M lands on top of Baseten's $1.5B at $13B (June 2026), Fireworks at $17.5B, and a broader wave. Inference is projected to account for roughly two-thirds of all AI compute by the end of 2026, up from about one-third in 2023, so the demand curve steepens from here.
Concrete implications for anyone recruiting against this raise:
- Stop counting stars, start counting merged PRs. GitHub stars are not talent signal. Merged PRs to vLLM or SGLang in the last 12 months is talent signal. Use a threshold (5, 10, 20) and hold it.
- Budget for Bengaluru and Hyderabad packages that beat local Big Tech. The elite pool ties SF at 3 profiles each. If you cannot pay a global comp band, you are recruiting against 3 people in one metro.
- Compete with the projects themselves. sgl-project and vllm-project appear in the top current-employer list. Assume your candidate has a standing offer to work on the OSS full-time.
- Mine the Stoica lineage first. Sky Computing Lab, RISELab, LMSYS.org. These are pre-filtered pools.
- Use plain-English search across sources. Refolk lets you ask for "vLLM committers currently at Red Hat or Oracle, based in India, open to remote US roles" and returns a joined view across GitHub and LinkedIn, which is the query shape this hire actually requires.
The Together AI round is not evidence that this labor market is easy. It is evidence that even $8.3 billion in enterprise value now depends on whether you can convince 20 or 30 specific engineers to take your offer over the one Baseten just made.
FAQ
How many engineers can actually ship production code to vLLM and SGLang?
The honest global count is 194 engineers who have committed code to both projects, per inclusion-ai.org's cross-analysis of the two commit graphs. That is roughly 1 in 10 of the ~2,000 cumulative vLLM PR-submitters and about 30% of SGLang's ~647 implied code contributors. Refolk's tightest inference-contributor query returns 27 profiles globally, which is a stricter definition focused on people currently identifying as inference contributors in their public profiles.
Where is this talent actually located?
In Refolk's 27-profile cohort, the top regions tie at 3 profiles each: San Francisco Bay Area, Bengaluru, and Hyderabad. The broader 334-profile US "vLLM inference engineer" pool concentrates in the SF Bay Area (5), San Francisco proper (3), and the Kirkland/Bellevue corridor (3). Top current US employers are Meta (3), Apple (3), NVIDIA, and sgl-project itself.
Who is Together AI competing against for this hire?
Directly, Baseten (which raised $1.5B at $13B in June 2026), Fireworks ($17.5B), and other neocloud platforms building on open-source inference stacks. Indirectly, the open-source projects themselves, since sgl-project and vllm-project appear as top current employers in the tightest cohort. That means candidates often have a live option to work on the OSS full-time, which most recruiters do not price into their comp bands.
What is the fastest way to source this pool without posting job ads?
Mine three graphs: the vLLM and SGLang GitHub contributor lists filtered by merged PRs in the last 12 months, the UC Berkeley Sky Computing Lab and RISELab alumni network (Ion Stoica's lineage produced both project founders), and LMSYS.org contributors. Ray Summit vLLM track speakers and vLLM meetup co-hosts are already pre-vetted. Refolk collapses these into one plain-English query across GitHub, LinkedIn, and the open web, which is faster than running three separate exports and trying to join them by hand.