RefolkCandidates
Open nowEngineeringCloud Site Operations (DIG)

Staff Data Center Operations Engineer

Crusoe · Denver, Colorado

Location
Denver, Colorado, US
Employment
Full time
Level
Staff
Posted
2 months ago

About this role

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

Crusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org - based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio.

This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as SiteOps' technical presence at headquarters - visible to engineering, procurement, and leadership in a way that a field-based role cannot be.

You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring-ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site.

What You'll Do

Cross-Site Platform Operations & Escalation

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution

  • Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support

  • Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests

  • Root-cause complex hardware issues - PCIe, BMC, thermal, fabric - and produce resolution documentation reusable across the SiteOps org

  • Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages

OEM & ODM Technical Partnership

  • Develop and maintain deep technical relationships with Crusoe's primary hardware partners - currently SuperMicro and HPE, with upcoming ODM’s as growing platforms - at the engineering and field escalation level

  • Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet

  • Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint

  • Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership

Platform Standards & Org Development

  • Own OEM platform technical knowledge at the SiteOps org level - escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites

  • Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms

  • Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures - both for new hire onboarding and ongoing skill development as the fleet and team evolve

  • Contribute to the technician certification program and technical leveling standards across the org

  • Support new site bring-up efforts providing platform readiness and deployment execution expertise

HQ Presence & Cross-Functional Collaboration

  • Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations

  • Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures

  • Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond)

  • Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership

What We're Looking For

Required

  • 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience

  • Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required

  • Deep familiarity with server platform architecture and OEM escalation and RMA processes

  • Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff

  • Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution

  • Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams

  • Strong written communication - comfortable producing escalation documentation, platform runbooks, and leadership reporting

  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)

Preferred

  • Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus

  • Familiarity with ASUS or Quanta server platforms and ODM engagement models

  • Experience with liquid-cooled GPU platforms and CDU integration

  • Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X)

  • Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator

  • Experience contributing to technician certification programs or IC leveling standards within a DC ops organization

Location

This role is based at Crusoe Cloud's headquarters in Denver, CO, with regular travel to our data center sites across the US and internationally. Domestic relocation support is available.

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $150,000 -$170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

As published by Crusoe. Applications are handled on their site.

Skills this posting mentions

Cloud ServicesArtificial IntelligenceODM

About Crusoe

Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated, purpose-built AI infrastructure solutions trusted by Fortune 500 companies to power their most advanced AI applications. Crusoe is redefining AI cloud infrastructure, with a mission to align the future of computing with the future of the climate. Our AI platform is recognized as the “gold standard” for reliability and performance. Our data centers are optimized for AI workloads and are powered by clean, renewable energy.

All 371 openings at Crusoe

One click, then it is written

Apply to Crusoe with a resume written for this role.

Queue Staff Data Center Operations Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Crusoe’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Crusoe

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Crusoe publishes. Refolk is not the employer and does not handle their hiring. Applications go to Crusoe directly.