RefolkCandidates
Open nowEngineeringEngineering

Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs · Freiburg (Germany)

Location
Freiburg (Germany)
Level
Staff
Posted
12 months ago

About this role

About Black Forest Labs

We’re the team behind Latent Diffusion, Stable Diffusion, and FLUX - foundational technologies that changed how the world creates images and video. We’re creating the generative models that power how people make images and video - tools used by millions of creators, developers, and businesses worldwide. Our FLUX models are among the most advanced in the world, and we’re just getting started.

Headquartered in Freiburg, Germany with a growing presence in San Francisco, we’re scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.

Why This Role

We're looking for engineers to build and maintain the engine that powers our mission to develop visual intelligence. From maintaining and scaling clusters, to building research platforms to accelerate the rate of innovation, this team operates with large breadth and depth. We build the systems to make multi-week/month long training possible, to orchestrate resources at scale, and at the same time efficiently, enabling the next breakthrough model. If you’re obsessed with distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement, this team would be perfect for you.

What You’ll Work On

  • Maintain research infrastructure, ensuring health, and optimizing components to extract peak performance from the system (both on application, and infrastructure side)
  • Scale infrastructure to meet growing research demands while maintaining reliability and performance
  • Collaborate with research teams to deeply understand their infrastructure needs, and design solutions that balance performance with cost efficiency.
  • Identify and resolve performance bottlenecks and capacity hotspots through deep analysis of distributed systems at scale.
  • Build and evolve telemetry and monitoring systems to provide deep visibility into infrastructure performance, utilization, and costs across our cloud and datacenter fleets.
  • Participate in on-call rotations and incident response to maintain system reliability

Technical Focus

  • Python, Bash, Go
  • Kubernetes
  • Nvidia GPU drivers, and operators
  • OTel, Prometheus

What We’re Looking For

  • Experience building or operating large-scale training platforms
  • Worked with large scale compute clusters (GPUs)
  • Proven ability to debug performance and reliability issues across large distributed fleets
  • Strong problem-solving skills and ability to work independently
  • Strong communication skills and the ability to work effectively with both internal and external partners
  • Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
  • Experience with SLURM
  • Experience building or operating large-scale training platforms

How We Work Together

We’re a distributed team with real offices that people actually use. Depending on your role, you’ll either join us in Freiburg or SF at least 2 days a week (or one full week every other week), or work remotely with a monthly in-person week to stay connected. We’ll cover reasonable travel costs to make this possible. We think in-person time matters, and we’ve structured things to make it accessible to all. We’ll discuss what this will look like for the role during our interview process.

Everything we do is grounded in four values:

  • Obsessed. We are a frontier research lab. The science has to be right, the understanding deep, the product beautiful.
  • Low Ego. The work speaks. The best idea wins, no matter who said it. Credit is shared. Nobody is above any task.
  • Bold. We take the ambitious bet. We ship, we do not wait for conditions to be perfect.
  • Kind. People over politics. We treat each other with genuine warmth. Agency without empathy creates chaos.

If this sounds like work you’d enjoy, we’d love to hear from you.

Base Annual Salary:

EU €100,000 - €230,000 + Equity

US $150,000 - $300,000 + Equity

This role is based in our Freiburg / San Francisco office. We operate a hybrid model and cover reasonable travel costs - relocation is encouraged but not required. We do expect a meaningful in-person presence, and we'll discuss what that looks like for your situation during the process.

As published by Black Forest Labs. Applications are handled on their site.

Skills this posting mentions

Google Cloud PlatformInfrastructureCloud Infrastructure

About Black Forest Labs

We’re the leading frontier AI research lab, continuously building the most advanced technology that shapes the visual understanding of the world. Our team pioneered Stable Diffusion, Stable Video Diffusion, and FLUX.1 and 2 - benchmarks in the evolution of generative AI. Today, these foundations power millions of creations worldwide, from individual artists to enterprise applications. Our most recent models power a wide range of products - turning imagination into reality with precision, speed, and creative control.

All 12 openings at Black Forest Labs

One click, then it is written

Apply to Black Forest Labs with a resume written for this role.

Queue Member of Technical Staff - Research Infrastructure Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Black Forest Labs’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Black Forest Labs

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Black Forest Labs publishes. Refolk is not the employer and does not handle their hiring. Applications go to Black Forest Labs directly.