RefolkCandidates
Open nowEngineeringEngineering

Software Engineer, Site Reliability

Hebbia AI · New York, New York

Location
New York, New York, United States
Employment
Full time
Level
Senior
Posted
6 months ago

About this role

About Hebbia

The AI platform for investors and bankers that generates alpha and drives upside.

Founded in 2020 by George Sivulka and backed by Peter Thiel and Andreessen Horowitz, Hebbia powers investment decisions for BlackRock, KKR, Carlyle, Centerview, and 40% of the world’s largest asset managers. Our flagship product, Matrix, delivers industry-leading accuracy, speed, and transparency in AI-driven analysis. It is trusted to help manage over $30 trillion in assets globally.

We deliver the intelligence that gives finance professionals a definitive edge. Our AI uncovers signals no human could see, surfaces hidden opportunities, and accelerates decisions with unmatched speed and conviction. We do not just streamline workflows. We transform how capital is deployed, how risk is managed, and how value is created across markets.

Hebbia is not a tool. Hebbia is the competitive advantage that drives performance, alpha, and market leadership.

The Role

We are looking for a Site Reliability Engineer who thinks like a software engineer first. You will own critical production systems end-to-end, designing, building, and improving them rather than simply operating them. You will write production-quality code that keeps the platform reliable at scale, embed with product
engineering teams to influence architecture from the start, and build the internal tooling that every engineer at Hebbia depends on. This is not a ticket-driven ops role. You will spend most of your time writing code: instrumenting services, eliminating performance bottlenecks, building deployment platforms, and translating incident post-mortems into lasting architectural improvements.

Responsibilities

  • Own critical production services end-to-end, from design and code review through deployment,
    operation, and incident response

  • Profile, benchmark, and rewrite hot paths to eliminate bottlenecks as Hebbia scales

  • Lead incident response and drive post-mortem culture, translating findings into code changes and
    architectural improvements rather than runbooks

  • Design and build observability frameworks from scratch, writing custom instrumentation, alerting
    logic, and debugging tooling that surfaces production issues before customers feel them

  • Define and enforce SLOs across platform services and build the feedback loops that keep
    engineering teams accountable to them

  • Own capacity planning and cost efficiency: model growth, right-size infrastructure, and write
    automation that prevents over-provisioning and resource exhaustion

  • Build robust, well-tested internal platforms and deployment tooling held to the same engineering
    standards as customer-facing code

  • Own and continuously improve CI/CD systems so engineering teams can ship safely and quickly

  • Embed with product engineering teams as a peer software engineer, contributing directly to
    production codebases and co-designing systems for reliability from the start

  • Partner on infrastructure security through threat modeling, hardening, and automated compliance
    tooling

Who You Are

  • 5+ years software development with a track record of writing, shipping, and maintaining production services, not just operating infrastructure

  • Production-grade proficiency in at least one systems or backend language: Go, Python, C++, or Rust

  • Proven experience as a Production Engineer, SRE, or software engineer with a deep infrastructure focus, comfortable owning services end-to-end across the full stack

  • Deep understanding of distributed systems

  • Container orchestration expertise and hands-on experience debugging complex distributed failures in production

  • Working knowledge of OS-level concepts

  • Cloud platform fluency (AWS preferred)

  • Experience in building and maintaining observability stacks

  • Strong CI/CD pipeline expertise and a track record of improving developer velocity without sacrificing safety

  • Background at a company with a Production Engineering or software-focused SRE culture is a strong plus

  • Experience building platforms for AI/ML workloads or high-throughput document processing pipelines is a plus

Compensation

The salary range for this role is $160,000 to $350,000. This range may be inclusive of several career levels at Hebbia and will be narrowed during the interview process based on the candidate’s experience and qualifications. Adjustments outside of this range may be considered for candidates whose qualifications significantly differ from those outlined in the job description.

Life @ Hebbia

PTO: Unlimited

Insurance: Medical + Dental + Vision + 401K

Eats: Catered lunch daily + doordash dinner credit if you ever need to stay late

Parental leave policy: 3 months non-birthing parent, 4 months for birthing parent

Fertility benefits: $15k lifetime benefit

New hire equity grant: competitive equity package with unmatched upside potential

#LI-Onsite

As published by Hebbia AI. Applications are handled on their site.

Skills this posting mentions

PythonGolangDistributed Systems

About Hebbia AI

Hebbia is the leading AI platform for finance. Founded in 2020 by George Sivulka, Hebbia is a generative AI company backed by Andreessen Horowitz, Peter Thiel, and Index Ventures. Investment banks and over 40% of the largest asset managers by AUM use Hebbia’s AI agents to drive investment decisions and automate financial analyst workflows. Users can instantly surface insights over filings, research, and millions of internal documents, enabling citation-backed research, AI-driven document, powerpoint, and spreadsheet generation, and AI driven origination, screening, and diligence. Learn more at hebbia.com.

All 25 openings at Hebbia AI

One click, then it is written

Apply to Hebbia AI with a resume written for this role.

Queue Software Engineer, Site Reliability and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Hebbia AI’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Hebbia AI

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Hebbia AI publishes. Refolk is not the employer and does not handle their hiring. Applications go to Hebbia AI directly.