RefolkCandidates
Open nowData and MLEngineering & Quality

Principal Data Engineer

Gather AI · Remote (India)

Location
Remote (India)
Workplace
Remote
Level
Principal
Posted
2 months ago

About this role

About Us

Are you ready to build the future of the supply chain? At Gather AI, we're not just creating software; we're pioneering a new era of warehouse intelligence. We've developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining "on-time, in full" delivery.

If you're looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We're leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time.

About the Team

You'll join the Data Platform team at its inception, helping establish the foundation from day one . Today, production and analytical workloads share a single database, and every product team defines its own metrics. This team exists to fix that: designing the warehouse, the transformation layers, and the semantic model that every product and dashboard will build on going forward, in close partnership with product, engineering, and security.

About the Role

Most data platform roles ask you to extend something someone else has already built. This one starts with a blank canvas.

As Principal Data Engineer, you'll architect Gather AI's data foundation from the ground up: separating analytical workloads from live production traffic, building a semantic layer so metrics are defined once and stay consistent everywhere, and linking structured records to real drone imagery and video with full traceability. You'll prove the model end to end on Gather's drone product, then generalize it so every new product extends the foundation instead of rebuilding it, all while working as a Principal-level individual contributor with real influence across engineering, product, and leadership.

What You'll Do

  • Architect a greenfield, multi-layer data warehouse (raw, refined, serving) that separates analytical workloads from production OLTP traffic.
  • Deliver a governed, self-service data-access layer for internal consumers first (Product, CSM, Deployment/Operations, and Leadership) as Phase 1, ahead of customer-facing conversational analytics.
  • Build a semantic and metrics layer so every metric, such as "scan accuracy by site," is defined once in code and stays identical across every dashboard and product, making self-service safe from metric drift.
  • Own the quality bar: 99%+ availability SLA with freshness guarantees, 100% traceability, zero cross-tenant leakage, 99.5%+ pipeline success, and no data loss.
  • Design tenant isolation, per-tenant cost attribution, and schema and row-level RBAC to scale toward hundreds of tenants (300+ target), not today's fleet size.
  • Own data-ingestion correctness at the boundary with the integration/backend team, covering data contracts, schema validation, and pipeline quality, so WMS data lands in the right place, shape, and time across WMS versions.
  • Stand up a data catalog and lineage layer (Purview as the Azure-native fit, DataHub as the open-source alternative) so every consumer can find data, see ownership, and trace lineage when a metric looks wrong.
  • Prove the foundation end to end on Gather's drone product, then generalize it so each new product extends the model instead of rebuilding it
  • Act as the connective tissue between product and ML (3DCC, damage detection). Link structured records to unstructured drone imagery and video with full traceability, and stand up the data-infra readiness for feature stores and annotation pipelines on one trusted foundation.

What You'll Need

  • 10+ years in data engineering, with 3+ years architecting data platforms for data products, analytics, or AI-driven products.
  • Proven experience building a greenfield data warehouse and leading an OLTP to OLAP transition, not just maintaining an existing one.
  • Deep expertise designing multi-layer transformation architectures and reusable frameworks that scale across multiple product areas.
  • Expert SQL and dbt, hands-on ELT and orchestration, and large-scale or streaming data experience.
  • Production experience on a major cloud (Azure preferred, AWS or GCP acceptable), plus infrastructure as code and CI/CD.
  • Track record with data quality, security, governance, and multi-tenancy in production environments.
  • Data transformation and modeling that turns raw multi-source data into refined, serving-ready datasets (raw to refined to serving).
  • Pipeline orchestration and workflow automation for scheduling, dependency management, and reliable execution across data flows.
  • Large-scale and distributed processing of high-volume batch data.
  • Real-time and streaming ingestion that captures and processes event data as it arrives.
  • Semantic and metrics-layer design that defines business metrics once and serves them consistently to every consumer.
  • Serving-layer optimization for fast, low-latency consumption through wide and flattened tables and pre-computed metrics.
  • Cloud data engineering and infrastructure automation that provisions, deploys, and operates the platform reproducibly (cloud-native, infrastructure as code, CI/CD).
  • Data quality, observability, and lineage that ensure trust, freshness, and end-to-end traceability.
  • Security, governance, and multi-tenancy including tenant isolation, access control, and resiliency.
  • Multimodal data integration that links structured records to unstructured image and video (drone captures) with traceability.

Ways of Working

  • Treats data as a product for internal consumers, not just a pipeline feeding dashboards.
  • Comfortable making long-lead architecture calls (platform, isolation model) with incomplete consensus.
  • Strong cross-functional collaborator, works closely with integration/backend, ML, product, customer success teams and internal analytics consumers.

Nice to Have

  • Experience modeling structured data linked to unstructured or blob data such as images, video, or sensor files
  • Experience with feature stores, annotation pipelines, or ML data infrastructure supporting computer vision products.
  • IoT, edge, or device-telemetry background
  • BI or presentation-layer and dashboard design experience
  • Warehousing, logistics, or supply-chain domain knowledge

As published by Gather AI. Applications are handled on their site.

Skills this posting mentions

Data TransformationData EngineeringModeling

About Gather AI

Gather AI is a co-pilot for intralogistics. It combines real-time data from machines and enterprise systems with industry-specific AI to deliver insight when and where it's needed. The result is a platform that helps teams move goods efficiently and prevent disruptions without adding complexity.

All 16 openings at Gather AI

One click, then it is written

Apply to Gather AI with a resume written for this role.

Queue Principal Data Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Gather AI’s form for you.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • 25 sent a week, free
  • No card
  • Nothing sent until you say so

More roles at Gather AI

See all

Similar roles elsewhere

See more

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Listed from the job board Gather AI publishes. Refolk is not the employer and does not handle their hiring. Applications go to Gather AI directly.