Member of Technical Staff, TPU Performance Engineering
Inferact · Singapore, Singapore
- Location
- Singapore, Singapore, Singapore
- Employment
- Full time
- Level
- Staff
- Posted
- 2 months ago
About this role
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.
Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.
Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.
Preferred qualifications:
Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.
Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.
Bonus points if you have:
Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.
Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
Logistics
Location: This role is based in Singapore.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.
As published by Inferact. Applications are handled on their site.
Skills this posting mentions
About Inferact
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
All 19 openings at InferactOne click, then it is written
Apply to Inferact with a resume written for this role.
Queue Member of Technical Staff, TPU Performance Engineering and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Inferact’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Inferact
See all- 5 weeks ago
- 6 weeks ago
- 7 weeks ago
- 2 months ago
- 2 months ago
Member of Technical Staff, CI/CD Infrastructure
San Francisco, California
$200k - $400k/yrStaffEngineering - 2 months ago
Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, California
$200k - $400k/yr2 locationsStaffEngineering
Similar roles elsewhere
See more- 3 days ago
Principal Software Engineer - DevOps / Site Reliability Engineer
Riot GamesSingapore
PrincipalEngineering
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Inferact publishes. Refolk is not the employer and does not handle their hiring. Applications go to Inferact directly.