Refolk
August 5, 2026·9 min read

Together AI's $800M Series C Wants Kernel Fluency. The Visible Pool Is 6.

Together AI's $800M Series C funds a hiring push for CUDA-fluent inference engineers. The publicly signaled US pool is 6. Here is how to source the rest.

Together AI hiringCUDA inference engineersFlashAttention talent poolsourcing GPU kernel engineersopen source AI researchers
Together AI's $800M Series C Wants Kernel Fluency. The Visible Pool Is 6.

On July 1, 2026, Together AI announced an $800M Series C at an $8.3B valuation, up from $3.3B earlier in 2025, and CEO Vipul Ved Prakash said the plain part out loud: the company is hiring across engineering, research, product, and GTM to grow compute capacity roughly 50x over five years. The chips are the easy problem. The engineers who can write custom CUDA kernels and ship open-model inference at frontier scale are the hard one, and if you are sourcing this role like a normal ML engineer search, you have already lost.

The real hireable pool is measured in dozens, not thousands

The strict intersection of "CUDA-kernel-fluent" and "open-model inference" engineers in the US, self-signaled on public profiles, is 6. That is the floor. The addressable pool, including adjacent contributors and Hot Chips 2025 presenters, is a few hundred. It is not the 20,000 people who list "PyTorch" on a LinkedIn profile.

In Refolk's index of professional profiles, exactly 6 US engineers currently match the strict headline-level intersection of CUDA skill plus inference-kernel keywords. Five of those six sit at NVIDIA, Microsoft, LinkedIn, Linear, or Roche. All are clustered in Santa Clara, the SF Bay, and San Jose. That is not a talent market. That is a dinner party.

6
US engineers publicly signaling CUDA plus inference-kernel work in headline
Refolk's index of professional profiles, strict intersection. The addressable pool with adjacent signals is larger but still small.

Together's own Chief Scientist, Tri Dao, wrote the original FlashAttention paper (NeurIPS 2022) with Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. He is the archetype the company now needs to clone 20 or 30 times over. The archetype barely exists in the wild, and the few who do exist mostly work for Together's customers.

Why the customer roster is the recruiting funnel

Together's named customers are exactly where the kernel-fluent inference engineers already live. That makes the top account list a shadow org chart of the hiring plan, which is a governance vector VCs do not underwrite.

Together has publicly named Cognition, Decagon, Eleven Labs, Cursor, and Suno as customers, and disclosed that Decagon reduced inference costs by approximately 6x after moving to its platform. The mechanism is not magic. It is a small number of engineers who understand how to fuse attention kernels, batch decode, and squeeze the last percentage points out of an H100. That knowledge is portable. The engineer who saved Decagon 6x can save Together's next enterprise customer the same amount, but only one company gets to employ them at a time.

This is the kind of intersection search that generic Boolean strings choke on, and the exact gap Refolk closes: describe the person in plain English ("engineer who has contributed to FlashAttention or SGLang, currently at a Together customer, based in the Bay") and get a ranked shortlist across GitHub, LinkedIn, and the open web in one query.

Poaching from your top accounts to build the product they are paying you for is a churn vector, not a strategy.

The comparable numbers, in one table

Here is the shape of the market in six rows. Each number below comes from either Together's disclosures, the arXiv record for FlashAttention, competing raises reported in 2026, or Refolk's index.

SegmentCountSource
US profiles matching CUDA plus inference-kernel headline signal6Refolk's index (strict intersection)
Share of that pool at NVIDIA, Microsoft, LinkedIn, Linear, or Roche83% (5 of 6)Refolk's index (derived)
FlashAttention library size in lines of code~70,000arXiv 2511.11581
Together AI infra growth multiple over 5 years (stated)~50xTogether / Quartz
Together AI ARR growth from early 2024 to 202638x ($30M to $1.15B)Company disclosure
Competing neocloud raises (Upscale + TensorWave)$850M combinedTechCrunch

Two things jump out. First, 83% of the visible pool works for five specific employers, most of which have larger recruiting budgets than Together. Second, $850M of net-new competing demand (Upscale AI at $500M / $2B, TensorWave at $350M / $1.55B) hit the same 6-to-few-hundred-person market inside roughly a 30-day window. The pool did not grow. The demand curve did. Rate cards for FlashAttention, SGLang, and vLLM contributors have re-priced, and any offer letter drafted off 2025 comp benchmarks will bounce.

The real bottleneck is training-side kernels, not inference

If you only remember one thing: the scarce archetype is the backward-pass kernel engineer, not the inference-serving one. Every job post that says "inference engineer" is fishing in the wrong pond.

FlashAttention-4 was published on March 5, 2026, with preliminary results presented at Hot Chips in August 2025. Per the FA4 release notes, it "currently supports only forward propagation, with backward pass support planned for future releases. This limits its applicability primarily to inference workloads." Read that carefully. The open-source frontier for training-side attention kernels is still open engineering territory. The engineer who can write a numerically stable backward pass for a fused attention kernel on H100 or B200 is a strictly smaller set than "people who can serve LLaMA fast."

Where to actually find them:

  • FlashInfer authors: Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, et al. Their MLSys 2025 paper on customizable attention engines for LLM inference serving is the exact profile.
  • KernelBench authors: Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, Azalia Mirhoseini (arXiv:2502.10517, 2025). Stanford / CRFM is a canonical pipeline.
  • Hot Chips 2025 presenters on attention and inference kernels.
  • SGLang and vLLM top contributors on GitHub, filtered by non-trivial CUDA commit history.
  • CUTLASS and CUDA compiler alumni from NVIDIA who have left in the last 24 months. NVIDIA is now a Together investor, which functionally means alumni are fair game, sitting engineers are not.

Why remote will not save this search

Geographic concentration is a feature of the skill, not a bug of the labor market. This role does not port to Tulsa.

Refolk's top regions for the tight intersection are Santa Clara, SF Bay, and San Jose. Kernel work is built through hardware-adjacent apprenticeships that require physical proximity to test clusters, tape-out engineers, and the specific compiler team who knows why a WMMA instruction stalled on a Tuesday. Remote hires from Warsaw or Waterloo can absolutely do the coding. They cannot, on Slack, absorb the tacit knowledge that lives inside NVIDIA's Santa Clara buildings and the handful of SF-based labs that ship serving stacks at frontier scale.

Practical implication for sourcers: any recruiter promising a "distributed CUDA team" for Together, Upscale, or TensorWave in 2026 is either lying or planning to lose the hires within 12 months. Anchor searches to a 30-mile radius from Palo Alto, then expand grudgingly.

How to actually size the pool for a Together AI hiring push

Break the search into three concentric rings and staff each ring differently. The visible tip is 6 people. The realistic addressable pool is a few hundred. The training pipeline for the next cohort sits at three or four labs.

  1. Ring 1, the visible 6. Direct outreach only, from a founder or the Chief Scientist, not a recruiter. Assume 2 will reply, 1 will take a call, and 0 to 1 will move. Budget accordingly.
  2. Ring 2, adjacent contributors (roughly 200 to 300). FlashAttention, FlashInfer, SGLang, vLLM, and CUTLASS contributors, KernelBench authors, Hot Chips 2025 kernel-track presenters, and ex-NVIDIA CUDA compiler and CUTLASS alumni within 24 months. This is where most of the actual hiring happens. Modal is publicly contributing to FlashAttention-4 and SGLang and hiring off the same contribution graph, so Ring 2 is contested from day one.
  3. Ring 3, the training pipeline. Stanford CRFM and adjacent kernel research groups. New grads and postdocs who have published a kernel paper in the last 18 months. Slower to close, but the only sustainable supply.

The mistake almost every generalist recruiter makes is spending all their time on Ring 2 with LinkedIn Recruiter filters that do not distinguish "CUDA on a resume" from "CUDA in a merged PR to FlashAttention." Refolk was built for exactly this shape of problem: ask for "engineers with merged CUDA PRs to FlashAttention or SGLang in 2025, currently outside NVIDIA, US-based" and get the actual Ring 2 shortlist without hand-scraping GitHub.

What the competing bidders tell you about compensation

Together, Upscale AI, and TensorWave collectively announced roughly $1.65B in fresh capital in a compressed window, all targeting overlapping talent. Any offer benchmarked to pre-2026 comp data is stale.

TensorWave is building on AMD GPU clusters, which sounds like a different pool but actually competes for the same compiler and kernel engineers because ROCm work rewards the same hardware intuition. Upscale AI is chasing the same NVIDIA-alumni cohort. Prakash framed the mission as making intelligence "abundant, not expensive," which is a lovely quote and also a signal that Together plans to compete on scale rather than exclusivity, which puts more pressure on hiring velocity, not less.

Practical comp read: assume the top of Ring 2 is now indexed to frontier-lab research-engineer compensation, not to infra-engineer compensation. The engineers know this. Recruiters catching up to it a quarter late will lose this quarter's best candidates.

FAQ

How many CUDA inference engineers can Together AI realistically hire in the next 12 months?

If Together is disciplined about targeting Ring 2 and treats Ring 1 as a bonus, a plausible plan is 15 to 30 kernel-fluent hires in 12 months, weighted toward ex-NVIDIA CUTLASS and CUDA compiler alumni plus top FlashAttention, SGLang, and vLLM contributors. Anything above 40 requires either poaching directly from the customer roster (Cursor, Cognition, Decagon), which creates real revenue risk, or lowering the bar on what "kernel-fluent" means, which defeats the point of the raise.

Is it actually a conflict for Together to hire from Cursor, Cognition, and Decagon?

Legally, no. Practically, yes. These are named public customers whose inference bills fund Together's ARR growth from $30M to $1.15B. Poaching a lead inference engineer from Decagon, which Together publicly credits with a 6x cost reduction, is the sort of move that gets escalated to a CEO-to-CEO call and can cost a seven-figure ARR contract. Most Together hires from this list will happen quietly, through mutual-friend introductions, not cold outreach.

What is the fastest way to build a defensible Ring 2 list?

Start from the GitHub contributor graphs for FlashAttention, FlashInfer, SGLang, vLLM, and CUTLASS. Filter to merged PRs with non-trivial CUDA changes in the last 18 months. Cross-reference against Hot Chips 2025 and MLSys 2025 author lists. Then enrich with current employer and location. This is doable manually in about a week per organization, or in a single prompt with a tool like Refolk that reads GitHub, LinkedIn, and the open web in one pass.

Does NVIDIA becoming a Together investor change the hiring dynamic?

Yes, in one specific way: it turns sitting NVIDIA engineers into off-limits candidates and NVIDIA alumni into premium ones. Expect a soft détente where Together hires ex-NVIDIA aggressively but does not run outreach into current NVIDIA teams. Recruiters should assume any inbound from a sitting NVIDIA CUDA engineer is either a bluff or a signal that person is already halfway out the door for other reasons.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next

Together AI's $800M Series C Wants Kernel Fluency. The Visible Pool Is 6. · Refolk