Together AI's $800M Buys 500MW. The US CUDA Kernel Pool Is 62.
Together AI's $8.3B Series C funds a 50x infrastructure expansion, but the real bottleneck is a US CUDA kernel pool of 62 engineers.
On July 1, 2026, Together AI closed an $800M Series C at an $8.3B post-money valuation, locked in 500+ MW of compute commitments, and announced the first Y Combinator dedicated GPU cluster. CEO Vipul Ved Prakash said the company will grow its infrastructure roughly 50x over five years and is hiring across engineering and research. The capital is real. The megawatts are real. The engineers who can actually write and tune the kernels that turn those megawatts into revenue are a rounding error.
The real bottleneck isn't $800M or 500 MW, it's 62 people
Together AI's binding constraint is a US pool of roughly 62 CUDA kernel engineers who publicly identify as such, not capital or GPU allocation. In Refolk's index of professional profiles, only 62 US-based people list CUDA as a skill and carry "kernel" in their headline or role. That is the addressable, surfaceable pool for the exact work Together AI just committed 500 MW of silicon to.
A CUDA kernel engineer is a systems programmer who writes and tunes the low-level GPU code (in CUDA C++, PTX, or Triton) that determines whether a matmul runs at 30% or 85% of peak FLOPs. It is the layer between PyTorch and the metal, and it is what separates a neocloud that sells tokens profitably from one that rents Nvidia's margin back to Nvidia.
The working assumption in most neocloud recruiting decks is a "sub-thousand-person" pool. Even that is generous. Scale the 62 public profiles by 10x for engineers who don't advertise the work, and you land under 700 nationwide, competing against Upscale AI ($500M Series A extension at $2B), TensorWave ($350M Series B at $1.55B, 137 total employees), CoreWeave, Lambda, Groq, Nscale, Hydra Host, and every frontier lab.
Why LinkedIn title search returns garbage for this role
LinkedIn title search misses roughly 100% of qualified CUDA kernel engineers because the people who do the work are titled "Member of Technical Staff," "Performance Engineer," or "Principal GPU Software Engineer," not "CUDA Kernel Engineer." When Refolk normalized the 62-person cohort against standard Senior/Manager/Director seniority filters, the count collapsed to zero.
The mechanism is boring and important: title conventions in this niche were set by Nvidia, Apple, and Meta a decade before "kernel engineer" became a recruiter search term. The people doing the work never bothered to rebrand. Consider what a title-first sourcing pass actually catches:
- "CUDA Engineer" at Nvidia: catches applications engineers, not kernel authors.
- "GPU Software Engineer" at Apple: catches Metal driver work, some of it kernel-adjacent.
- "Performance Engineer" at Meta: catches the actual kernel authors, buried in a title shared with SRE and profiling generalists.
- "Member of Technical Staff" at Modular, Anthropic, OpenAI: catches everyone from kernel authors to product engineers with a single string.
This is the exact gap Refolk closes: you describe the person in plain English ("US-based engineers who write CUDA or Triton kernels for LLM inference, currently at a compiler company or neocloud") and get a ranked shortlist that ignores the title field entirely. The tool reads skills, project descriptions, GitHub commits, and paper authorship, which is where kernel work actually leaves fingerprints.
The numbers behind the pool
Here is the full picture from Refolk's US-only index, queried July 2026. The gap between "any CUDA skill" and "actual kernel practitioner" is the story.
| Slice | US profile count | Notes |
|---|---|---|
| Any CUDA skill listed | 16,370 | Broad ceiling; includes CTOs, founders, generalists |
| CUDA + "kernel" in headline/role | 62 | Realistic kernel-writer pool on the open web |
| CUDA + kernel + Senior/Manager/Director filter | 0 | Standard title filters return nothing |
| Triton skill + "gpu kernel" context | 2 | Triton-native cohort is effectively single-digit |
| Ratio of CUDA-any to CUDA-kernel | ~264:1 | Only ~0.38% of "CUDA people" work at the kernel layer |
| Kernel practitioners in SF Bay Area | 4 of 25 sampled (~16%) | Largest cluster; NYC, Boston, Santa Clara, San Jose follow |
The 264:1 ratio is the recruiting math nobody prices in. A Boolean string on LinkedIn for CUDA AND (engineer OR developer) returns 16,370 US results and feels like abundance. The actual pool that can write a fused attention kernel that beats FlashAttention-3 on a specific SKU is 0.38% of that number.
Modular is the chokepoint employer, not Nvidia
The densest non-hyperscaler concentration of CUDA + kernel talent in Refolk's US index sits at Modular, not at any of the neoclouds or frontier labs. Modular ties for the top employer slot in the sample with 3 profiles, ahead of Meta and Apple in the same slice. Microsoft AI, Microsoft, Intel, Vast.ai, General Motors, Roche, and Gulp (YC W25) round out the top-10 employer list.
Why a compiler startup, and not Nvidia? Modular is building Mojo and MAX, a compiler and runtime that competes directly with CUDA's software moat. Every kernel engineer they hire is one who wanted to work on the abstraction layer above CUDA rather than inside it, and they recruited heavily from LLVM and Swift alumni networks. That makes Modular a rare thing in this market: a single company whose entire technical hiring thesis overlaps with what Together AI now needs.
A compiler startup, not a hyperscaler, is now the densest concentration of poachable kernel talent outside Nvidia itself. </pull> If you are running neocloud recruiting for Together AI, Upscale, or TensorWave, Modular is the single highest-yield outbound target in the country. Vast.ai, which shows up in the same cohort, is the second: it is one of the only neocloud-native employers with kernel engineers already on staff, and its comp bands are public enough to know what you need to beat. ## The Nvidia investor-portfolio no-poach ceiling Nvidia's cross-portfolio investing creates an implicit ceiling on kernel-engineer comp escalation between neoclouds, which pushes the real bidding war toward the frontier labs. NVIDIA is a direct investor in Together AI, Upscale AI, TensorWave, Groq, Lambda, Nscale, and Hydra Host. An engineer moving from Together to Upscale isn't really changing employers from Nvidia's point of view.
refolk prompt: US-based engineers who have shipped CUDA or Triton kernels for LLM inference in the last 18 months, currently at Modular, Vast.ai, or an Nvidia-invested neocloud note: You get a ranked shortlist built from GitHub commits, paper authorship, and skill signals, not job titles, with employer tenure and likely-to-move signals attached. slug: jvzkzd9y5x
The mechanism matters because it changes where the real leverage lives. If you are a founder outside the Nvidia portfolio (a frontier lab, a defense-tech shop, an AMD-first cluster like TensorWave in its early days), you have room to price kernel talent honestly. If you are inside it, you are competing with your co-investors for the same 62 people, and the market-clearing price will be set by whoever is willing to break the ceiling. Historically that has been OpenAI, Anthropic, and xAI.
Nvidia's published Lead/Staff CUDA kernel bands ($272K to $431K/yr for 15+ years experience, $152K to $242K for mid/senior) are the floor a neocloud has to beat, not the ceiling. Frontier lab total comp for the same engineers routinely lands 2 to 3x that number once equity is priced.
## Where to actually source the 62 (and the other 600)
The highest-signal sourcing channels for CUDA kernel engineers are open-source repos, academic labs, and alumni networks, not job boards or LinkedIn Recruiter. Six channels that actually work:
1. **vLLM, SGLang, and TensorRT-LLM commit histories.** Refolk shows vLLM contributors already appearing in the Triton-kernel cohort. Every non-trivial PR to a fused kernel is a resume.
2. **FlashAttention and Triton repo issues.** The people filing detailed bug reports against Tri Dao's or OpenAI's code are the pool.
3. **Modular's Mojo Discord and MAX contributor list.** Same three-name cluster the index surfaces, plus the LLVM alumni who followed Chris Lattner over.
4. **Percy Liang's group at Stanford and Ce Zhang's labs at ETH Zürich and University of Chicago.** These are Together AI's co-founder networks; kernel-adjacent systems PhDs come through them directly.
5. **GTC and MLSys paper authorship.** The last two years of MLSys proceedings are a nearly complete roster of kernel-layer researchers who might take an industry role.
6. **Y Combinator W25 batch (specifically Gulp).** Together AI's new YC-dedicated cluster is a warm introduction to the exact founders already writing kernels in-batch.
Refolk indexes all six of those signals as sourcing surfaces, so a plain-English query like "engineers who have committed CUDA kernels to vLLM or SGLang in the last year and are not currently at Nvidia" returns a working list, not a keyword-matched pile. That is what "find anyone, just ask" means in practice for this niche.
## The math on the 50x expansion
500 MW divided by roughly 600 addressable US kernel engineers is about 0.83 MW per engineer, which means Together AI alone needs to hire 1 in every 60 US kernel engineers just to instrument its committed compute. Before Upscale, TensorWave, CoreWeave, or the labs take their share.
```stat
number: 264:1
label: Ratio of CUDA-any to CUDA-kernel practitioners
note: Only ~0.38% of the 16,370 US professionals who list CUDA actually work at the kernel layer.
</stat>
Traditional recruiting math does not close this gap. A funnel that assumes 100 sourced profiles yields 10 conversations yields 1 hire needs 6,200 sourced profiles to close 62 kernel engineers, which is more people than exist in the addressable public pool. The only funnels that work are the ones that start with the actual practitioner list and negotiate from there, which requires knowing exactly who the 62 (or 600) are before you send the first message.
That is the reframe neocloud recruiting needs after Together AI's raise. It is not a capital problem. It is not a GPU allocation problem. It is a naming problem: the 62 people who can do the work do not call themselves what your ATS is searching for, and the recruiters who find them first will decide which neocloud actually ships the 50x.
FAQ
How many CUDA kernel engineers are there in the US?
Refolk's index shows 62 US-based professionals who list CUDA as a skill and carry "kernel" in their headline or role, out of 16,370 people who list any CUDA skill. Scaled generously for engineers who don't advertise the work publicly, the true addressable US pool is under 1,000, which is smaller than a single big-tech ML org.
Why doesn't LinkedIn title search work for CUDA kernel engineers?
Kernel engineers are titled "Member of Technical Staff," "Senior Software Engineer," "Performance Engineer," or "Principal GPU Software Engineer," not "CUDA Kernel Engineer." When Refolk normalized the 62-person cohort against standard Senior/Manager/Director title filters, the count dropped to zero, meaning title-first sourcing misses effectively 100% of qualified candidates.
Which employer has the densest concentration of poachable CUDA kernel talent?
Modular. In Refolk's US index, Modular ties for the top employer slot in the CUDA + kernel intersection with 3 profiles in the sample, ahead of Meta and Apple. A compiler startup, not a hyperscaler, is now the highest-yield outbound target for neoclouds like Together AI, Upscale, and TensorWave.
What sourcing channels actually surface CUDA kernel engineers?
Six channels return real signal: vLLM/SGLang/TensorRT-LLM commit histories, FlashAttention and Triton repo issues, Modular's Mojo contributor list, Percy Liang's and Ce Zhang's academic labs, MLSys and GTC paper authorship, and the Y Combinator W25 batch. All of them require reading project artifacts rather than job titles, which is why plain-English semantic search across GitHub, LinkedIn, and the open web outperforms Boolean strings for this specific role.