Top Python Data engineering repositories on GitHub
Pipelines, orchestrators, and ELT/ETL tooling. Filtered to projects whose primary language is Python.
Ranked by stars across 281 Python repositories tagged data-engineering. Refreshed daily.
- 1apache/superset★ 74,886 · ⑂ 18,375
Apache Superset is a Data Visualization and Data Exploration Platform
- superset
- apache
- apache-superset
- data-visualization
- data-viz
- analytics
- 2apache/airflow★ 46,947 · ⑂ 17,898
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
- airflow
- apache
- apache-airflow
- python
- scheduler
- workflow
- 3PrefectHQ/prefect★ 23,899 · ⑂ 2,534
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
- python
- workflow
- data-engineering
- data-science
- workflow-engine
- prefect
- 4airbytehq/airbyte★ 22,117 · ⑂ 5,359
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
- data
- pipeline
- data-analysis
- data-engineering
- java
- python
- Live search
Find the people behind these repos
Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.
500 free credits on sign-up, no card needed.
- 5Avaiga/taipy★ 19,434 · ⑂ 1,994
Turns Data and AI algorithms into production-ready web applications in no time.
- automation
- data-engineering
- data-ops
- data-visualization
- datascience
- developer-tools
- 6dagster-io/dagster★ 16,193 · ⑂ 2,300
An orchestration platform for the development, production, and observation of data assets.
- data-pipelines
- dagster
- workflow
- data-science
- workflow-automation
- python
- 7andkret/Cookbook★ 15,434 · ⑂ 2,757
The Data Engineering Cookbook
- data-engineer
- data-engineering
- big-data
- best-practices
- cookbook
- 8semantica-agi/semantica★ 13,400 · ⑂ 1,514
Graph-Native Infrastructure for Context and Accountable AI Systems
- ai
- ai-governance
- artificial-intelligence
- context-engineering
- context-graphs
- decision-intelligence
- 9fivetran/great_expectations★ 11,830 · ⑂ 1,853
Always know what to expect from your data.
- pipeline-tests
- dataquality
- datacleaning
- datacleaner
- data-science
- data-profiling
- 10xonsh/xonsh★ 9,651 · ⑂ 743
🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.
- xonsh
- devops
- iterm2
- data-engineering
- security-automation
- raspberry-pi
- Live search
Who is hiring in this space?
I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.
500 free credits on sign-up, no card needed.
- 11mage-ai/mage-ai★ 8,826 · ⑂ 990
🧙 Build, run, and manage data pipelines for integrating and transforming data.
- machine-learning
- artificial-intelligence
- data
- data-engineering
- data-science
- python
- 12feast-dev/feast★ 7,304 · ⑂ 1,443
The Open Source Feature Store for AI/ML
- machine-learning
- features
- ml
- big-data
- feature-store
- python
- 13Zipstack/unstract★ 7,254 · ⑂ 720
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
- ai-agents
- data-engineering
- document-ai
- generative-ai
- idp
- json-extraction
- 14dlt-hub/dlt★ 5,880 · ⑂ 605
data load tool (dlt) is an open source Python library that makes data loading easy 🛠️
- data
- python
- data-engineering
- data-lake
- data-loading
- data-warehouse
- 15ruc-datalab/DeepAnalyze★ 4,648 · ⑂ 736
DeepAnalyze is the first agentic LLM for autonomous data science. 🎈你的AI数据分析师,自动分析大量数据,一键生成专业分析报告!
- agent
- agentic
- agentic-ai
- chatbot
- data
- data-analysis
- 16aws/aws-sdk-pandas★ 4,119 · ⑂ 750
pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).
- python
- aws
- pandas
- apache-arrow
- apache-parquet
- data-engineering
- Live search
Turn any brief into a list like this
I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.
500 free credits on sign-up, no card needed.
- 17ploomber/ploomber★ 3,620 · ⑂ 243
The fastest ⚡️ way to build data pipelines. Develop iteratively, deploy anywhere. ☁️
- workflow
- machine-learning
- data-science
- data-engineering
- mlops
- papermill
- 18datafold/data-diff★ 2,991 · ⑂ 314
Compare tables within or across databases
- database
- mysql
- postgresql
- snowflake
- rdbms
- trino
- 19meltano/meltano★ 2,634 · ⑂ 271
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
- dataops
- dataops-platform
- elt
- open-source
- opensource
- data
Find Python engineers shipping Data engineering
The list above ranks the most-starred public Python repositories tagged with the Data engineering topic, drawn from the public GitHub graph. Across 281 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Data engineering experience.
That overlap is rare. Most Python engineers haven’t shipped Data engineering, and most Data engineering maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.
Refolk turns this list into a search. Ask for “Python Data engineering maintainers hiring” or “Python engineers shipping Data engineering in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.
How this list is built
Last refreshed: Tue, 22 Sep 2026 22:24:04 GMT
Need a more specific search?
Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:
- Python Data engineering maintainers hiringRun
- Python Data engineering contributors in EuropeRun
- Companies shipping Python Data engineeringRun
500 free credits on sign-up, no card needed.
Related lists
- Python · Machine learning
- Python · Deep learning
- Python · Computer vision
- Python · Natural language processing
- Python · LLM
- Python · AI agents
- Python · RAG
- Python · Embeddings
See all repository lists.
Or zoom out
Common questions
How are these repositories ranked?
By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.
How fresh is the data?
The ranking re-renders at least daily. Last refreshed: Tue, 22 Sep 2026 22:24:04 GMT.
Can I find the maintainers and contributors behind these repos?
Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.
Can I use this list for hiring?
That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.