Technical Compute Qualification Manager
Together AI · San Francisco
- Location
- San Francisco
- Workplace
- Remote
- Level
- Manager
- Posted
- 4 days ago
About this role
About The Role
Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it carries customer workloads, and this role owns that bar. As Technical Program Manager, Compute Qualification, you will run the process that screens and qualifies prospective compute providers, taking each prospective deployment through a structured evaluation across compute, networking, storage, power, cooling, and operations.
You will coordinate various engineering partners through validation, review provider specifications and test results, and produce clear go/no-go recommendations on whether new capacity meets our standards. It is a high-impact, process-driven role for someone technical enough to know when a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in parallel. "You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads." Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience.
Responsibilities
- Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
- Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
- Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
- Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
- Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , sanity-check performance claims, and surface issues before deeper engineering review.
- Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
- Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows.
Requirements
- 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments.
- Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management.
- Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
- Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
- Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives.
- Willingness to travel to provider and data center sites as needed.
Nice to Have
- Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
- Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
- Experience in AI/HPC cluster design.
- Background working directly with hardware vendors, colocation providers, or cloud capacity providers.
About Together AI
Together AI is an AI-native cloud company building the infrastructure to make AI faster, cheaper, and more accessible. We’re rapidly scaling our GPU footprint: signing our own data center leases, building large-scale clusters, and expanding toward a global owned-infrastructure presence. Our research team has contributed to breakthroughs like FlashAttention, Hyena, and RedPajama, and we co-design across software, hardware, and algorithms to push the frontier of AI efficiency.
Compensation
We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200-250K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our Privacy Policy at https://www.together.ai/privacy
As published by Together AI. Applications are handled on their site.
Skills this posting mentions
About Together AI
Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform - from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.
All 65 openings at Together AIOne click, then it is written
Apply to Together AI with a resume written for this role.
Queue Technical Compute Qualification Manager and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Together AI’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Together AI
See all- Today
- Today
- 3 days ago
- 4 days ago
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Bangalore IndiaRemote
2 locationsStaffEngineering - 4 days ago
- 4 days ago
Similar roles elsewhere
See morePut this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Together AI publishes. Refolk is not the employer and does not handle their hiring. Applications go to Together AI directly.