RefolkCandidates
Open nowEngineeringProduct and Design

Staff Product Manager, AI Infrastructure (Fleet Management)

Crusoe · San Francisco, California

Location
San Francisco, California, USA
Employment
Full time
Level
Staff
Posted
2 days ago

About this role

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:

The Staff Product Manager, Fleet Management owns the product strategy and execution for the systems that keep Crusoe's GPU and CPU fleet running reliably and efficiently at scale. You'll define how Crusoe provisions, monitors, maintains, and optimizes its server fleet supporting AI training and inference workloads, translating operational requirements into scalable product capabilities. This covers the full lifecycle of compute: initial provisioning, firmware and configuration management, burn-in, multi-node testing, health monitoring, maintenance orchestration, repair, and decommissioning.

The role sits at the intersection of infrastructure operations, engineering, and customer experience. Fleet-level systems determine how much of our capacity is sellable, how fast we can bring new clusters online, and how quickly a customer's workload recovers when hardware fails. You'll be defining the problem space as much as solving it, and you should be comfortable operating in ambiguity, with support from senior leadership as the initiative takes shape.

You'll collaborate closely with hardware partners, engineering, infrastructure operations, networking, supply chain, finance, and customer success to ensure cohesive lifecycle management across Crusoe Cloud. As a Staff PM, you'll lead significant cross-functional initiatives, influence technical and product strategy across multiple teams, and be accountable for outcomes that directly impact company goals.

What You'll Be Working On:

  • Own the vision and product strategy for fleet management, including provisioning, burn-in, lifecycle management, health monitoring, maintenance orchestration, and performance optimization

  • Drive outcomes across the entire product lifecycle, from discovery through launch and iteration, for fleet management capabilities

  • Define product direction for systems managing large-scale GPU and CPU infrastructure, ensuring reliability, utilization, and operational efficiency

  • Lead cross-functional initiatives spanning engineering, infrastructure operations, networking, and customer success to deliver integrated fleet management solutions

  • Build consensus with tech leads and engineering managers on architecture decisions, tooling investments, and operational process improvements

  • Identify opportunities to improve fleet utilization, reduce operational overhead, and enhance observability through product innovation

  • Translate customer-facing reliability and performance impact, gathered through customer success, into product requirements

  • Mentor other product managers and engineers, sharing expertise in infrastructure operations and complex systems thinking

What You'll Bring to the Team:

  • 6+ years of product management experience delivering infrastructure or platform products, with demonstrated ownership of complex, cross-functional initiatives

  • Deep expertise in fleet management concepts including provisioning, lifecycle management, monitoring, and operations at scale

  • Strong understanding of distributed systems, server hardware, and infrastructure automation

  • Proven ability to solve ambiguous, novel problems requiring research, invention, and creative analysis

  • Experience driving product vision for complex systems with long-term strategic impact, balancing technical tradeoffs and operational requirements

  • Track record of building consensus across engineering, operations, and infrastructure teams on technical direction and investment priorities

  • Strong analytical skills interpreting operational metrics, performance data, and utilization patterns to guide product decisions

  • Excellent communication skills articulating technical concepts and connecting them to business outcomes

  • Experience mentoring colleagues in product management or technical roles

Bonus Points

  • Experience with large-scale GPU infrastructure or AI/ML platform operations

  • Familiarity with infrastructure automation tools, orchestration systems (Kubernetes, Slurm), or configuration management platforms

  • Experience with observability platforms, monitoring systems, or incident management workflows

  • Knowledge of hardware lifecycle management, predictive maintenance, or capacity planning systems

Benefits:

  • Competitive compensation

  • Restricted Stock Units

  • Paid time off & paid holidays

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

Compensation Range

Compensation will be paid in the range of up to $215,000-$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

As published by Crusoe. Applications are handled on their site.

Skills this posting mentions

GPUIncident ManagementProduct Vision

About Crusoe

Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated, purpose-built AI infrastructure solutions trusted by Fortune 500 companies to power their most advanced AI applications. Crusoe is redefining AI cloud infrastructure, with a mission to align the future of computing with the future of the climate. Our AI platform is recognized as the “gold standard” for reliability and performance. Our data centers are optimized for AI workloads and are powered by clean, renewable energy.

All 364 openings at Crusoe

One click, then it is written

Apply to Crusoe with a resume written for this role.

Queue Staff Product Manager, AI Infrastructure (Fleet Management) and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Crusoe’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Crusoe

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Crusoe publishes. Refolk is not the employer and does not handle their hiring. Applications go to Crusoe directly.