Refolk

Top Python Data science repositories on GitHub

Notebooks, analysis libraries, and data tooling. Filtered to projects whose primary language is Python.

Ranked by stars across 1,211 Python repositories tagged data-science. Refreshed daily.

  1. 1
    apache/superset74,849 · ⑂ 18,356

    Apache Superset is a Data Visualization and Data Exploration Platform

    • superset
    • apache
    • apache-superset
    • data-visualization
    • data-viz
    • analytics
  2. 2
    Asabeneh/30-Days-Of-Python74,191 · ⑂ 13,537

    The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

    • 30-days-of-python
    • python
    • flask
    • github
    • heroku
    • matplotlib
  3. 3
    scikit-learn/scikit-learn67,319 · ⑂ 27,427

    scikit-learn: machine learning in Python

    • machine-learning
    • python
    • statistics
    • data-science
    • data-analysis
  4. 4
    keras-team/keras64,322 · ⑂ 19,801

    Deep Learning for humans

    • deep-learning
    • tensorflow
    • neural-networks
    • machine-learning
    • data-science
    • python
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    pandas-dev/pandas49,761 · ⑂ 20,404

    Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more

    • data-analysis
    • pandas
    • flexible
    • alignment
    • python
    • data-science
  7. 6

    Summer 2027 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.

    • interview-preparation
    • internships
    • jobs
    • university
    • fall-2026
    • github
  8. 7
    apache/airflow46,917 · ⑂ 17,874

    Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

    • airflow
    • apache
    • apache-airflow
    • python
    • scheduler
    • workflow
  9. 8
    streamlit/streamlit45,799 · ⑂ 4,385

    Streamlit — A faster way to build and share data apps.

    • python
    • machine-learning
    • data-science
    • deep-learning
    • data-visualization
    • streamlit
  10. 9
    ray-project/ray43,877 · ⑂ 8,067

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    • ray
    • distributed
    • parallel
    • machine-learning
    • reinforcement-learning
    • deep-learning
  11. 10
    gradio-app/gradio43,586 · ⑂ 3,608

    Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!

    • machine-learning
    • models
    • ui
    • ui-components
    • interface
    • python
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    explosion/spaCy33,910 · ⑂ 4,722

    💫 Industrial-strength Natural Language Processing (NLP) in Python

    • natural-language-processing
    • data-science
    • machine-learning
    • python
    • cython
    • nlp
  14. 12
    eriklindernoren/ML-From-Scratch32,891 · ⑂ 5,489

    Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.

    • machine-learning
    • deep-learning
    • deep-reinforcement-learning
    • machine-learning-from-scratch
    • data-science
    • data-mining
  15. 13
    Lightning-AI/pytorch-lightning31,354 · ⑂ 3,797

    Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.

    • python
    • deep-learning
    • artificial-intelligence
    • ai
    • pytorch
    • data-science
  16. 14
    d2l-ai/d2l-en29,646 · ⑂ 5,130

    Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.

    • deep-learning
    • machine-learning
    • book
    • notebook
    • computer-vision
    • natural-language-processing
  17. 15

    Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

    • python
    • machine-learning
    • deep-learning
    • data-science
    • big-data
    • aws
  18. 16
    reflex-dev/reflex28,894 · ⑂ 1,776

    🕸️ Web apps in pure Python 🐍

    • python
    • framework
    • open-source
    • gui
    • dashboard
    • fullstack
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    plotly/dash24,418 · ⑂ 2,317

    Data Apps & Dashboards for Python. No JavaScript Required.

    • dash
    • plotly
    • data-visualization
    • data-science
    • gui-framework
    • flask
  21. 18
    PrefectHQ/prefect23,884 · ⑂ 2,528

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    • python
    • workflow
    • data-engineering
    • data-science
    • workflow-engine
    • prefect
  22. 19
    sinaptik-ai/pandas-ai23,807 · ⑂ 2,340

    Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.

    • llm
    • pandas
    • ai
    • data-analysis
    • data-science
    • gpt-4

Find Python engineers shipping Data science

The list above ranks the most-starred public Python repositories tagged with the Data science topic, drawn from the public GitHub graph. Across 1,211 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Data science experience.

That overlap is rare. Most Python engineers haven’t shipped Data science, and most Data science maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.

Refolk turns this list into a search. Ask for Python Data science maintainers hiring” or Python engineers shipping Data science in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.

How this list is built

Refolk searched GitHub for public Python repositories tagged with the Data science topic, ranked them by stargazer count, and kept those with at least 25 stars. The list refreshes once a day.

Last refreshed: Sun, 20 Sep 2026 10:50:00 GMT

Search this list

Need a more specific search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Related lists

See all repository lists.

Or zoom out

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Sun, 20 Sep 2026 10:50:00 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring