Refolk

Top Python Data engineering repositories on GitHub

Pipelines, orchestrators, and ELT/ETL tooling. Filtered to projects whose primary language is Python.

Ranked by stars across 281 Python repositories tagged data-engineering. Refreshed daily.

  1. 1
    apache/superset74,886 · ⑂ 18,375

    Apache Superset is a Data Visualization and Data Exploration Platform

    • superset
    • apache
    • apache-superset
    • data-visualization
    • data-viz
    • analytics
  2. 2
    apache/airflow46,947 · ⑂ 17,898

    Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

    • airflow
    • apache
    • apache-airflow
    • python
    • scheduler
    • workflow
  3. 3
    PrefectHQ/prefect23,899 · ⑂ 2,534

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    • python
    • workflow
    • data-engineering
    • data-science
    • workflow-engine
    • prefect
  4. 4
    airbytehq/airbyte22,117 · ⑂ 5,359

    Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

    • data
    • pipeline
    • data-analysis
    • data-engineering
    • java
    • python
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    Avaiga/taipy19,434 · ⑂ 1,994

    Turns Data and AI algorithms into production-ready web applications in no time.

    • automation
    • data-engineering
    • data-ops
    • data-visualization
    • datascience
    • developer-tools
  7. 6
    dagster-io/dagster16,193 · ⑂ 2,300

    An orchestration platform for the development, production, and observation of data assets.

    • data-pipelines
    • dagster
    • workflow
    • data-science
    • workflow-automation
    • python
  8. 7
    andkret/Cookbook15,434 · ⑂ 2,757

    The Data Engineering Cookbook

    • data-engineer
    • data-engineering
    • big-data
    • best-practices
    • cookbook
  9. 8
    semantica-agi/semantica13,400 · ⑂ 1,514

    Graph-Native Infrastructure for Context and Accountable AI Systems

    • ai
    • ai-governance
    • artificial-intelligence
    • context-engineering
    • context-graphs
    • decision-intelligence
  10. 9
    fivetran/great_expectations11,830 · ⑂ 1,853

    Always know what to expect from your data.

    • pipeline-tests
    • dataquality
    • datacleaning
    • datacleaner
    • data-science
    • data-profiling
  11. 10
    xonsh/xonsh9,651 · ⑂ 743

    🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.

    • xonsh
    • devops
    • iterm2
    • data-engineering
    • security-automation
    • raspberry-pi
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    mage-ai/mage-ai8,826 · ⑂ 990

    🧙 Build, run, and manage data pipelines for integrating and transforming data.

    • machine-learning
    • artificial-intelligence
    • data
    • data-engineering
    • data-science
    • python
  14. 12
    feast-dev/feast7,304 · ⑂ 1,443

    The Open Source Feature Store for AI/ML

    • machine-learning
    • features
    • ml
    • big-data
    • feature-store
    • python
  15. 13
    Zipstack/unstract7,254 · ⑂ 720

    LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

    • ai-agents
    • data-engineering
    • document-ai
    • generative-ai
    • idp
    • json-extraction
  16. 14
    dlt-hub/dlt5,880 · ⑂ 605

    data load tool (dlt) is an open source Python library that makes data loading easy 🛠️

    • data
    • python
    • data-engineering
    • data-lake
    • data-loading
    • data-warehouse
  17. 15
    ruc-datalab/DeepAnalyze4,648 · ⑂ 736

    DeepAnalyze is the first agentic LLM for autonomous data science. 🎈你的AI数据分析师,自动分析大量数据,一键生成专业分析报告!

    • agent
    • agentic
    • agentic-ai
    • chatbot
    • data
    • data-analysis
  18. 16
    aws/aws-sdk-pandas4,119 · ⑂ 750

    pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).

    • python
    • aws
    • pandas
    • apache-arrow
    • apache-parquet
    • data-engineering
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    ploomber/ploomber3,620 · ⑂ 243

    The fastest ⚡️ way to build data pipelines. Develop iteratively, deploy anywhere. ☁️

    • workflow
    • machine-learning
    • data-science
    • data-engineering
    • mlops
    • papermill
  21. 18
    datafold/data-diff2,991 · ⑂ 314

    Compare tables within or across databases

    • database
    • mysql
    • postgresql
    • snowflake
    • rdbms
    • trino
  22. 19
    meltano/meltano2,634 · ⑂ 271

    Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.

    • dataops
    • dataops-platform
    • elt
    • open-source
    • opensource
    • data

Find Python engineers shipping Data engineering

The list above ranks the most-starred public Python repositories tagged with the Data engineering topic, drawn from the public GitHub graph. Across 281 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Data engineering experience.

That overlap is rare. Most Python engineers haven’t shipped Data engineering, and most Data engineering maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.

Refolk turns this list into a search. Ask for Python Data engineering maintainers hiring” or Python engineers shipping Data engineering in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.

How this list is built

Refolk searched GitHub for public Python repositories tagged with the Data engineering topic, ranked them by stargazer count, and kept those with at least 25 stars. The list refreshes once a day.

Last refreshed: Tue, 22 Sep 2026 22:24:04 GMT

Search this list

Need a more specific search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Related lists

See all repository lists.

Or zoom out

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Tue, 22 Sep 2026 22:24:04 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring