Staff Product Manager, AI Infrastructure (Fleet Management)
Crusoe · San Francisco, California
- Location
- San Francisco, California, USA
- Employment
- Full time
- Level
- Staff
- Posted
- 2 days ago
About this role
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About the Role:
The Staff Product Manager, Fleet Management owns the product strategy and execution for the systems that keep Crusoe's GPU and CPU fleet running reliably and efficiently at scale. You'll define how Crusoe provisions, monitors, maintains, and optimizes its server fleet supporting AI training and inference workloads, translating operational requirements into scalable product capabilities. This covers the full lifecycle of compute: initial provisioning, firmware and configuration management, burn-in, multi-node testing, health monitoring, maintenance orchestration, repair, and decommissioning.
The role sits at the intersection of infrastructure operations, engineering, and customer experience. Fleet-level systems determine how much of our capacity is sellable, how fast we can bring new clusters online, and how quickly a customer's workload recovers when hardware fails. You'll be defining the problem space as much as solving it, and you should be comfortable operating in ambiguity, with support from senior leadership as the initiative takes shape.
You'll collaborate closely with hardware partners, engineering, infrastructure operations, networking, supply chain, finance, and customer success to ensure cohesive lifecycle management across Crusoe Cloud. As a Staff PM, you'll lead significant cross-functional initiatives, influence technical and product strategy across multiple teams, and be accountable for outcomes that directly impact company goals.
What You'll Be Working On:
Own the vision and product strategy for fleet management, including provisioning, burn-in, lifecycle management, health monitoring, maintenance orchestration, and performance optimization
Drive outcomes across the entire product lifecycle, from discovery through launch and iteration, for fleet management capabilities
Define product direction for systems managing large-scale GPU and CPU infrastructure, ensuring reliability, utilization, and operational efficiency
Lead cross-functional initiatives spanning engineering, infrastructure operations, networking, and customer success to deliver integrated fleet management solutions
Build consensus with tech leads and engineering managers on architecture decisions, tooling investments, and operational process improvements
Identify opportunities to improve fleet utilization, reduce operational overhead, and enhance observability through product innovation
Translate customer-facing reliability and performance impact, gathered through customer success, into product requirements
Mentor other product managers and engineers, sharing expertise in infrastructure operations and complex systems thinking
What You'll Bring to the Team:
6+ years of product management experience delivering infrastructure or platform products, with demonstrated ownership of complex, cross-functional initiatives
Deep expertise in fleet management concepts including provisioning, lifecycle management, monitoring, and operations at scale
Strong understanding of distributed systems, server hardware, and infrastructure automation
Proven ability to solve ambiguous, novel problems requiring research, invention, and creative analysis
Experience driving product vision for complex systems with long-term strategic impact, balancing technical tradeoffs and operational requirements
Track record of building consensus across engineering, operations, and infrastructure teams on technical direction and investment priorities
Strong analytical skills interpreting operational metrics, performance data, and utilization patterns to guide product decisions
Excellent communication skills articulating technical concepts and connecting them to business outcomes
Experience mentoring colleagues in product management or technical roles
Bonus Points
Experience with large-scale GPU infrastructure or AI/ML platform operations
Familiarity with infrastructure automation tools, orchestration systems (Kubernetes, Slurm), or configuration management platforms
Experience with observability platforms, monitoring systems, or incident management workflows
Knowledge of hardware lifecycle management, predictive maintenance, or capacity planning systems
Benefits:
Competitive compensation
Restricted Stock Units
Paid time off & paid holidays
Comprehensive health, dental & vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance, short-term and long-term disability
Professional development & tuition reimbursement
Mental health & wellness support
Commuter benefits (parking & transit)
Cell phone stipend
401(k) Retirement plan with company match up to 4% of salary
Volunteer time off
Compensation Range
Compensation will be paid in the range of up to $215,000-$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
As published by Crusoe. Applications are handled on their site.
Skills this posting mentions
About Crusoe
Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated, purpose-built AI infrastructure solutions trusted by Fortune 500 companies to power their most advanced AI applications. Crusoe is redefining AI cloud infrastructure, with a mission to align the future of computing with the future of the climate. Our AI platform is recognized as the “gold standard” for reliability and performance. Our data centers are optimized for AI workloads and are powered by clean, renewable energy.
All 364 openings at CrusoeOne click, then it is written
Apply to Crusoe with a resume written for this role.
Queue Staff Product Manager, AI Infrastructure (Fleet Management) and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Crusoe’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Crusoe
See all- 2 days ago
Senior Staff Program Lead, Cloud Engineering Operations
San Francisco, California
$230k - $280k/yrStaffEngineering - 3 days ago
- 3 days ago
- 3 days ago
Principal Organizational Development Strategist
Denver, Colorado
$175k - $205k/yr2 locationsPrincipalPeople and HR - 3 days ago
Senior Program Manager, Learning & Development - Tech
San Francisco, California
$160k - $195k/yrSeniorProduct - 4 days ago
Similar roles elsewhere
See more- Yesterday
Software Engineer, Production Inference (Distributed Inference)
Thinking Machines LabSan Francisco, CaliforniaRemote
$350k - $500k/yrMid levelEngineering - Yesterday
- Yesterday
- Yesterday
Senior Software Engineer, Billing Platform
SentrySan Francisco, CaliforniaRemote
$190k - $240k/yrSeniorEngineering - Yesterday
Engineering Manager, Billing Platform
SentrySan Francisco, CaliforniaRemote
$220k - $350k/yrManagerEngineering
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Crusoe publishes. Refolk is not the employer and does not handle their hiring. Applications go to Crusoe directly.