Refolk

Top Python Data science repositories on GitHub

Notebooks, analysis libraries, and data tooling. Filtered to projects whose primary language is Python.

Ranked by stars across 1,180 Python repositories tagged data-science. Refreshed daily.

  1. 1
    scikit-learn/scikit-learn66,378 · ⑂ 27,087

    scikit-learn: machine learning in Python

    • machine-learning
    • python
    • statistics
    • data-science
    • data-analysis
  2. 2
    Asabeneh/30-Days-Of-Python65,671 · ⑂ 12,291

    The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

    • 30-days-of-python
    • python
    • flask
    • github
    • heroku
    • matplotlib
  3. 3
    keras-team/keras64,094 · ⑂ 19,736

    Deep Learning for humans

    • deep-learning
    • tensorflow
    • neural-networks
    • machine-learning
    • data-science
    • python
  4. 4
    pandas-dev/pandas49,034 · ⑂ 20,019

    Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more

    • data-analysis
    • pandas
    • flexible
    • alignment
    • python
    • data-science
  5. 5
    apache/airflow45,884 · ⑂ 17,264

    Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

    • airflow
    • apache
    • apache-airflow
    • python
    • scheduler
    • workflow
  6. 6
    streamlit/streamlit45,020 · ⑂ 4,291

    Streamlit — A faster way to build and share data apps.

    • python
    • machine-learning
    • data-science
    • deep-learning
    • data-visualization
    • streamlit
  7. 7

    Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.

    • interview-preparation
    • internships
    • jobs
    • university
    • fall-2026
    • github
  8. 8
    gradio-app/gradio42,970 · ⑂ 3,506

    Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!

    • machine-learning
    • models
    • ui
    • ui-components
    • interface
    • python
  9. 9
    ray-project/ray42,953 · ⑂ 7,710

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    • ray
    • distributed
    • parallel
    • machine-learning
    • reinforcement-learning
    • deep-learning
  10. 10
    explosion/spaCy33,675 · ⑂ 4,688

    💫 Industrial-strength Natural Language Processing (NLP) in Python

    • natural-language-processing
    • data-science
    • machine-learning
    • python
    • cython
    • nlp
  11. 11
    eriklindernoren/ML-From-Scratch31,944 · ⑂ 5,347

    Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.

    • machine-learning
    • deep-learning
    • deep-reinforcement-learning
    • machine-learning-from-scratch
    • data-science
    • data-mining
  12. 12
    Lightning-AI/pytorch-lightning31,198 · ⑂ 3,742

    Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.

    • python
    • deep-learning
    • artificial-intelligence
    • ai
    • pytorch
    • data-science
  13. 13

    Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

    • python
    • machine-learning
    • deep-learning
    • data-science
    • big-data
    • aws
  14. 14
    d2l-ai/d2l-en29,023 · ⑂ 5,074

    Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.

    • deep-learning
    • machine-learning
    • book
    • notebook
    • computer-vision
    • natural-language-processing
  15. 15
    reflex-dev/reflex28,579 · ⑂ 1,742

    🕸️ Web apps in pure Python 🐍

    • python
    • framework
    • open-source
    • gui
    • dashboard
    • fullstack
  16. 16
    plotly/dash24,267 · ⑂ 2,299

    Data Apps & Dashboards for Python. No JavaScript Required.

    • dash
    • plotly
    • data-visualization
    • data-science
    • gui-framework
    • flask
  17. 17
    sinaptik-ai/pandas-ai23,591 · ⑂ 2,327

    Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.

    • llm
    • pandas
    • ai
    • data-analysis
    • data-science
    • gpt-4
  18. 18
    matplotlib/matplotlib22,909 · ⑂ 8,359

    matplotlib: plotting with Python

    • matplotlib
    • data-visualization
    • data-science
    • python
    • qt
    • wx
  19. 19
    PrefectHQ/prefect22,655 · ⑂ 2,346

    Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

    • python
    • workflow
    • data-engineering
    • data-science
    • workflow-engine
    • prefect
  20. 20
    recommenders-team/recommenders21,778 · ⑂ 3,326

    Best Practices on Recommendation Systems

    • machine-learning
    • recommender
    • ranking
    • deep-learning
    • python
    • jupyter-notebook
  21. 21
    marimo-team/marimo21,529 · ⑂ 1,144

    A reactive notebook for Python — run reproducible experiments, query with SQL, execute as a script, deploy as an app, and version with git. Stored as pure Python. All in a modern, AI-native editor.

    • notebooks
    • python
    • data-science
    • machine-learning
    • artificial-intelligence
    • data-visualization
  22. 22
    akfamily/akshare20,546 · ⑂ 3,292

    AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库

    • futures
    • financial-data
    • data-science
    • quant
    • fundamental
    • akshare
  23. 23
    ipython/ipython16,724 · ⑂ 4,487

    Official repository for IPython itself. Other repos in the IPython organization contain things like the website, documentation builds, etc.

    • ipython
    • jupyter
    • data-science
    • notebook
    • python
    • repl
  24. 24
    piskvorky/gensim16,443 · ⑂ 4,408

    Topic Modelling for Humans

    • gensim
    • topic-modeling
    • information-retrieval
    • machine-learning
    • natural-language-processing
    • nlp
  25. 25
    dagster-io/dagster15,728 · ⑂ 2,166

    An orchestration platform for the development, production, and observation of data assets.

    • data-pipelines
    • dagster
    • workflow
    • data-science
    • workflow-automation
    • python

Find Python engineers shipping Data science

The list above ranks the most-starred public Python repositories tagged with the Data science topic, drawn from the public GitHub graph. Across 1,180 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Data science experience.

That overlap is rare. Most Python engineers haven’t shipped Data science, and most Data science maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.

Refolk turns this list into a search. Ask for Python Data science maintainers hiring” or Python engineers shipping Data science in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.

How this list is built

Refolk searched GitHub for public Python repositories tagged with the Data science topic, ranked them by stargazer count, and kept those with at least 25 stars. The list refreshes once a day.

Last refreshed: Sun, 21 Jun 2026 11:20:45 GMT

Need a more specific search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

Related lists

See all repository lists.

Or zoom out