Member of Technical Staff, CI/CD Infrastructure
Inferact · San Francisco, California
- Location
- San Francisco, California, United States
- Employment
- Full time
- Level
- Staff
- Posted
- 2 months ago
About this role
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference efficient and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware - a position that took years to build.
About the Role
vLLM is growing at a fast pace, and every bit of that growth lands on the CI system. More models, more hardware, more contributors, more ways for things to break. Your job is to advance the CI system so it scales with vLLM’s momentum and unlocks faster development for everyone.
You’ll get to:
Maintain and scale the compute infrastructure that powers CI, release, performance benchmark, accuracy evaluation for vLLM project, across a wide range of models and accelerators including H100/H200, (G)B200/300, AMD MI325/355X, TPU, Intel Gaudi, etc..
Get creative about cutting CI time-to-signal from hours to minutes
Make sure every corner of vLLM code base is well-tested
Keep vLLM releases rock-solid
Build out tooling that helps 3,000+ vLLM contributors move fast
Skills and Qualifications
Minimum qualifications:
Strong experience with Docker, Kubernetes, and containerized build or test environments.
Built CI/CD pipelines from scratch using GitHub Actions, Buildkite, or similar systems.
Familiar with CI design patterns and CI techniques: compute orchestration, handling flaky tests, dependency/environment management, caching, remote execution, test target determination, etc, test coverage, and so on.
Fluent in Python, Bash, Go, or similar for automation and tooling.
Solid fundamentals of Linux, security, networking, storage, package management,.
Bonus points for:
Setting up infrastructure for ML, inference, CUDA, ROCm, or accelerator-heavy workloads.
Running Buildkite at scale, including agents, queues, dynamic pipelines, test sharding, caching, and artifact management.
Operating Kubernetes clusters for CI, batch jobs, test execution, or internal developer infrastructure.
Managing CI/CD in large open-source project
Building dashboards, alerts, runbooks, or tooling for CI observability.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
As published by Inferact. Applications are handled on their site.
Skills this posting mentions
About Inferact
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
All 19 openings at InferactOne click, then it is written
Apply to Inferact with a resume written for this role.
Queue Member of Technical Staff, CI/CD Infrastructure and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Inferact’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Inferact
See all- 5 weeks ago
- 6 weeks ago
- 7 weeks ago
- 2 months ago
- 2 months ago
Member of Technical Staff, TPU Performance Engineering
Singapore, Singapore
S$200k - S$400k/yr2 locationsStaffEngineering - 2 months ago
Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, California
$200k - $400k/yr2 locationsStaffEngineering
Similar roles elsewhere
See morePut this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Inferact publishes. Refolk is not the employer and does not handle their hiring. Applications go to Inferact directly.