RefolkCandidates
Open nowEngineering

Senior Software Engineer, Chaos Engineering

Datadog · Paris, France

Location
Paris, France
Level
Senior
Posted
Today

About this role

Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve.

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What You’ll Do:

  • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
  • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
  • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
  • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
  • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
  • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.

Who You Are:

  • You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
  • You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
  • You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
  • You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
  • You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
  • Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.

Datadog values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That’s okay. If you’re passionate about technology and want to grow your experience, we encourage you to apply.

Benefits and Growth:

  • Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
  • Work on reliability systems that operate across Datadog’s production environment and influence how engineering teams design for failure.
  • Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction.
  • Explore practical applications of AI and automation to reliability engineering and operational workflows.
  • Collaborate with engineers across infrastructure, databases, observability, and service teams on complex systems challenges.
  • Mentor other engineers and contribute to technical designs, engineering practices, and platform strategy.
  • Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.

#LI-Hybrid


About Datadog:

Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, Datadog enables businesses to move faster with clarity and confidence. Learn more about #DatadogLife on Instagram, LinkedIn, and Datadog Learning Center.


Equal Opportunity at Datadog:

Datadog is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and other characteristics protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. Here are our Candidate Legal Notices for your reference.

Datadog endeavors to make our Careers Page accessible to all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please complete this form. This form is for accommodation requests only and cannot be used to inquire about the status of applications.

Privacy and AI Guidelines:

Any information you submit to Datadog as part of your application will be processed in accordance with Datadog’s Applicant and Candidate Privacy Notice. For information on our AI policy, please visit Interviewing at Datadog AI Guidelines.

As published by Datadog. Applications are handled on their site.

One click, then it is written

Apply to Datadog with a resume written for this role.

Queue Senior Software Engineer, Chaos Engineering and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Datadog’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Datadog

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Datadog publishes. Refolk is not the employer and does not handle their hiring. Applications go to Datadog directly.