RefolkCandidates
Open nowData and MLEngineering

Staff Data Engineer

Iterative Health · Cambridge, Massachusetts

Location
Cambridge, Massachusetts, United States
Level
Staff
Posted
4 months ago

About this role

Iterative Health is a healthcare technology and services company powering the acceleration of clinical research to transform patient outcomes.

We built a leading performance-driven network of 100+ sites across the US, Europe, India, and Australia, conducting research directly in the communities where care is delivered across gastrointestinal, hepatology, obesity, and cardiology. By combining deep clinical trial expertise with cutting-edge AI, we connect sponsors' scientific ambitions with high-performing research teams that expedite and expand access to novel therapeutics for patients in need. Today, Iterative Health is headquartered in Cambridge, Massachusetts, and New York City with 250+ employees world-wide.

About the Role

Accelerating clinical research is one of the defining challenges in healthcare. Promising therapies exist that patients can't access because the operational infrastructure to run clinical trials efficiently doesn't exist yet. We're building it. That means designing technology systems that bring order to a fragmented landscape of clinical data sources, automating the operational work that slows trials down, and turning real-world clinical data into a foundation for predictive intelligence.

We're building a uniquely valuable data asset: real-world patient and research data flowing across 80+ trial sites, spanning dozens of EHRs and clinical systems, focused on patient populations that are chronically underserved by existing clinical research infrastructure. Your job is to build the pipelines, data models, and AI infrastructure that make this asset real, from ingestion and normalization through to the systems that power predictions on top of it. You'll own data quality and observability as foundational engineering problems. You'll also have a direct hand in shaping how this data drives our AI strategy, what we model, what we predict, and what becomes possible.

This is an opportunity for someone who wants to be part of a small, fast-moving engineering team at a formative stage. You'll shape what gets built, how decisions get made, and what the team becomes.

Responsibilities

  • Own the data layer and architecture: the models, schemas, and infrastructure decisions that everything downstream depends on
  • Build and operate the pipelines and transformations that move data from ingestion through normalization, enrichment, and into the formats that support analytics, ML training, and production model serving
  • Own data quality and observability: build the systems that make data issues visible and correctable before they compound
  • Partner with ML and engineering teams to identify what's modelable, define training data requirements, and build the data foundations for new predictive capabilities
  • Define how clinical and operational data is governed across the system
  • Evaluate and select the tools and technologies that make up the data stack, with a clear point of view on build vs. buy
  • Help shape the engineering culture of a small, growing team: how technical decisions get made, how problems get debated, what rigor looks like in practice

What We’re Looking For

Required Qualifications

  • 10+ years of experience in data engineering or related roles, with significant time spent building data systems
  • Experience with healthcare data strongly preferred (HL7, FHIR, claims, EHR extracts) or other complex, regulated data domains
  • Deep experience modeling and integrating data from multiple heterogeneous sources with inconsistent schemas and quality
  • Experience applying AI and LLMs to data engineering problems: extraction, normalization, classification, entity resolution
  • Strong understanding of how data infrastructure supports ML workflows from feature engineering to training data pipelines to model serving
  • Fluent in SQL and at least one modern programming language (Python, Java, Scala, Go), with experience across modern data infrastructure - distributed processing, streaming, cloud-native storage, orchestration, and transformation frameworks
  • Have built data systems from early stages, making foundational decisions with incomplete information
  • Naturally raise the quality of the engineering around you through code review, design guidance, and honest technical conversation

Preferred Qualifications

  • Experience building data infrastructure that directly supports ML model training and evaluation
  • Familiarity with clinical trial operations, EDC systems, or life sciences data
  • SOC 2, HIPAA or similar compliance experience baked into engineering practice
  • A track record of building or improving data systems that others had given up on making reliable

New York pay range$200,000-$325,000 USD

At Iterative Health, we’re actively working towards creating an environment that is representative of the diversity of patients our technology serves. We are focused on building an equitable and inclusive culture, and by extension, hiring process. If you require any accommodations to make the application process or interviewing experience more accessible to you, please contact CandidateAccommodations@iterative.health.

As published by Iterative Health. Applications are handled on their site.

Skills this posting mentions

Artificial IntelligenceStorageData Engineering

About Iterative Health

Iterative Health is a healthcare technology and services company powering the acceleration of clinical research to transform patient outcomes. By combining deep expertise in clinical trials with cutting-edge AI, we empower research teams and study sponsors to expand and expedite access to novel therapeutics for patients in need. Today, Iterative Health is based in Cambridge, Massachusetts, and New York City with 250+ employees world-wide.

All 55 openings at Iterative Health

One click, then it is written

Apply to Iterative Health with a resume written for this role.

Queue Staff Data Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Iterative Health’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Iterative Health

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Iterative Health publishes. Refolk is not the employer and does not handle their hiring. Applications go to Iterative Health directly.