Senior Staff Program Lead, Cloud Engineering Operations
Crusoe · San Francisco, California
- Location
- San Francisco, California, USA
- Employment
- Full time
- Level
- Staff
- Posted
- 2 days ago
About this role
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About This Role
Crusoe's Cloud Engineering organization is 380 people today and hiring toward 550. When we sell capacity, we make customers two promises: it will be ready on the date we said, and it will work while they use it. At this scale, keeping those promises doesn't happen by accident anymore. This role exists to make sure it happens on purpose.
You'll lead Engineering Operations: the operating system that lets Cloud Engineering leadership agree on outcomes, execute, and keep its promises to customers. You'll own the standing programs that make operations better every week (SLOs, incident follow-up, change safety, capacity delivery tracking, and operational data) and build the systems that feed the weekly executive Engineering Operations review. You don't just report when a metric slips; you find the owner, hold them to the commitment, and make sure a decision happens fast.
You'll start as the first hire in this function, working solo while you prove out the model above. As the mandate holds up, you'll shape how the function grows: what to build, who to bring in, and what to keep lean.
This isn't a role scoped to a single program or a single dashboard. You will be a strategic advisor to leadership on whether teams are actually executing operationally, and the person who makes clear owners, shared metrics, and a steady cadence stick across the org.
The ideal candidate has a technical background (software engineering, SRE, Technical Program Management, or product in an infrastructure context) and is close enough to production systems to hold senior engineering leaders to what they committed.
What You'll Be Working On
Engineering Operations programs
Own the standing Engineering Operations programs: SLO/SLI definition and attainment, incident management and follow-up, change safety (change policy, maintenance readiness, and pre-production checks), capacity delivery tracking, and operational data. Partner with the engineering leaders who own each program and hold the line on outcomes.
Define and track the metrics that show how the org runs, grows, builds, and enables itself: usable capacity, SLO attainment, MTTD and MTTR, customer-found incidents, on-time capacity delivery, milestone slip, change failure rate, follow-up closure rate, and change policy compliance. Data should tell the story before anyone has to ask.
Partner with engineering, SRE, data engineering, and product to keep operational data accurate, with one source of truth for each metric.
Partner with engineering, TPM, product, SRE, data scientists, customer success, and data center operations to ensure operations run smoothly across all functional areas.
Work toward a single pane of glass that makes the operational health of the organization easy to understand at a glance.
Accountability and operating cadence
Own the weekly executive Engineering Operations review: build the systems, data, and pre-reads that feed it, and make sure every risk on the page has an owner, a date, and a next step.
Run the weekly operations planning forum, and partner with Product on a monthly business review that ties operational health to customer outcomes.
Make ownership explicit: clear accountable owners, swimlanes, and articulated outcomes for every program and metric.
Hold teams to what they committed. Drive incident follow-ups, overdue actions, and slipping milestones to closure, and escalate fast when an owner can't fix the problem alone.
Growing the function
Operate as an individual contributor first: prove out the programs, metrics, and cadence before asking for headcount.
Make the case for more capacity as scope outgrows one person, then help hire and onboard the people who join the function.
Own the evolution of the operating system itself, not just the programs inside it: retire process that no longer earns its cost, and keep the overhead on engineering teams as low as possible.
What You'll Bring to the Team
10+ years in software engineering, SRE, technical program management, or a technical product role, close enough to production systems to know what an SLO breach actually means.
A track record of leading org-wide programs at the staff level or above, and of driving accountability across senior engineering leaders without formal authority.
Comfortable starting as a team of one: you don't need a team in place to start driving impact, and you know when the case for headcount is real versus premature.
You've operated in high-growth infrastructure environments where processes are still being built; ambiguity doesn't paralyze you.
You're a natural coordinator who works across teams without formal authority. Engineering, SRE, and product leads trust you because you follow through.
You can turn messy, multi-source data into a clear picture of organizational health, and you know how to pick the few metrics that drive decisions over the many that fill dashboards.
You've designed operating rhythms (weekly reviews, business reviews, incident reviews) that leaders actually use, and you keep them light for the teams that feed them.
Scrappy, low-ego, high-drive. You build the program and tooling yourself when it doesn't exist yet, and you care more about the outcome than the credit.
Bonus Points
Time inside AWS, GCP, Azure, CoreWeave, Lambda Labs, or a similar cloud provider.
Experience running an SLO, incident review, or change management program at a cloud or infrastructure company.
Familiarity with incident.io, Opsgenie, PagerDuty, or similar incident management platforms at scale.
Background in AI/ML infrastructure.
Benefits:
Competitive compensation and equity packages
Restricted Stock Units
Paid time off, paid holidays & leave of absence programs
Comprehensive health, dental & vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance, short-term and long-term disability
Professional development & tuition reimbursement
Mental health & wellness support
Commuter benefits (parking & transit)
Cell phone stipend
401(k) Retirement plan with company match up to 4% of salary
Volunteer time off
Global travel insurance & emergency assistance
Daily meals allowance
Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $230,000 - $280,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
As published by Crusoe. Applications are handled on their site.
Skills this posting mentions
About Crusoe
Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated, purpose-built AI infrastructure solutions trusted by Fortune 500 companies to power their most advanced AI applications. Crusoe is redefining AI cloud infrastructure, with a mission to align the future of computing with the future of the climate. Our AI platform is recognized as the “gold standard” for reliability and performance. Our data centers are optimized for AI workloads and are powered by clean, renewable energy.
All 364 openings at CrusoeOne click, then it is written
Apply to Crusoe with a resume written for this role.
Queue Senior Staff Program Lead, Cloud Engineering Operations and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Crusoe’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Crusoe
See all- 2 days ago
Staff Product Manager, AI Infrastructure (Fleet Management)
San Francisco, California
$215k - $260k/yrStaffEngineering - 3 days ago
- 3 days ago
- 3 days ago
Principal Organizational Development Strategist
Denver, Colorado
$175k - $205k/yr2 locationsPrincipalPeople and HR - 3 days ago
Senior Program Manager, Learning & Development - Tech
San Francisco, California
$160k - $195k/yrSeniorProduct - 4 days ago
Similar roles elsewhere
See more- Yesterday
Software Engineer, Production Inference (Distributed Inference)
Thinking Machines LabSan Francisco, CaliforniaRemote
$350k - $500k/yrMid levelEngineering - Yesterday
- Yesterday
- Yesterday
Senior Software Engineer, Billing Platform
SentrySan Francisco, CaliforniaRemote
$190k - $240k/yrSeniorEngineering - Yesterday
Engineering Manager, Billing Platform
SentrySan Francisco, CaliforniaRemote
$220k - $350k/yrManagerEngineering
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Crusoe publishes. Refolk is not the employer and does not handle their hiring. Applications go to Crusoe directly.