Refolk

Top Python Embeddings repositories on GitHub

Models, libraries, and infrastructure for vector representations of text and media. Filtered to projects whose primary language is Python.

Ranked by stars across 325 Python repositories tagged embeddings. Refreshed daily.

  1. 1
    neuml/txtai12,964 · ⑂ 891

    💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows

    • python
    • search
    • nlp
    • semantic-search
    • vector-search
    • txtai
  2. 2
    Embedding/Chinese-Word-Vectors12,230 · ⑂ 2,320

    100+ Chinese Word Vectors 上百种预训练中文词向量

    • chinese
    • chinese-word-segmentation
    • embeddings
    • word-embeddings
    • vectors-trained
    • embedding
  3. 3
    FlagOpen/FlagEmbedding12,177 · ⑂ 916

    Retrieval and Retrieval-augmented LLMs

    • embeddings
    • information-retrieval
    • llm
    • sentence-embeddings
    • text-semantic-similarity
    • retrieval-augmented-generation
  4. 4
    h2oai/h2ogpt11,967 · ⑂ 1,300

    Private chat with local GPT with document, images, video, etc. 100% private, Apache 2.0. Supports oLLaMa, Mixtral, llama.cpp, and more. Demo: https://gpt.h2o.ai/ https://gpt-docs.h2o.ai/

    • chatgpt
    • llm
    • ai
    • embeddings
    • generative
    • gpt
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5

    The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.

    • metric-learning
    • deep-learning
    • computer-vision
    • machine-learning
    • pytorch
    • deep-metric-learning
  7. 6
    MinishLab/semble6,111 · ⑂ 269

    Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read

    • agents
    • code-search
    • embeddings
    • mcp
    • mcp-server
    • model-context-protocol
  8. 7
    shibing624/text2vec4,975 · ⑂ 428

    text2vec, text to vector. 文本向量表征工具,把文本转化为向量矩阵,实现了Word2Vec、RankBM25、Sentence-BERT、CoSENT等文本表征、文本相似度计算模型,开箱即用。

    • similarity
    • nlp
    • text-similarity
    • text2vec
    • word2vec
    • embeddings
  9. 8
    lightly-ai/lightly3,812 · ⑂ 367

    A python library for self-supervised learning on images.

    • deep-learning
    • self-supervised-learning
    • machine-learning
    • computer-vision
    • pytorch
    • embeddings
  10. 9
    tensorflow/hub3,528 · ⑂ 1,639

    A library for transfer learning by reusing parts of TensorFlow models.

    • tensorflow
    • machine-learning
    • transfer-learning
    • embeddings
    • image-classification
    • python
  11. 10
    towhee-io/towhee3,451 · ⑂ 257

    Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.

    • machine-learning
    • convolutional-networks
    • embedding-vectors
    • embeddings
    • computer-vision
    • image-processing
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    embeddings-benchmark/mteb3,427 · ⑂ 704

    MTEB: State-of-the-art evaluation of embeddings across languages and modalities

    • benchmark
    • clustering
    • information-retrieval
    • sentence-transformers
    • sts
    • text-embedding
  14. 12
    superlinked/sie3,291 · ⑂ 311

    Open-source inference server and production cluster for all the models your agent needs.

    • embeddings
    • vector-search
    • data-pipeline
    • deep-learning
    • information-retrieval
    • llm
  15. 13
    qdrant/fastembed3,211 · ⑂ 249

    Fast, Accurate, Lightweight Python library to make State of the Art Embedding

    • embeddings
    • openai
    • rag
    • retrieval
    • retrieval-augmented-generation
    • vector-search
  16. 14
    hegelai/prompttools3,055 · ⑂ 256

    Open-source tools for prompt testing and experimentation, with support for both LLMs (e.g. OpenAI, LLaMA) and vector databases (e.g. Chroma, Weaviate, LanceDB).

    • deep-learning
    • large-language-models
    • machine-learning
    • prompt-engineering
    • python
    • embeddings
  17. 15
    zilliztech/memsearch2,626 · ⑂ 253

    A persistent, unified memory layer for all your AI agents (e.g. Claude Code, Codex, DSH), backed by Markdown and Milvus.

    • agent-memory
    • claude-code
    • claude-code-plugin
    • memory
    • openclaw
    • rag
  18. 16
    ailia-ai/ailia-models2,392 · ⑂ 365

    The collection of pre-trained, state-of-the-art AI models for ailia SDK

    • deep-learning
    • face-recognition
    • face-detection
    • object-detection
    • object-recognition
    • hand-detection
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    vearch/vearch2,326 · ⑂ 365

    Distributed vector search for AI-native applications

    • vectors
    • vector-search
    • cloud-native
    • document-retrieval
    • embeddings
    • vector-database
  21. 18
    PetrochukM/PyTorch-NLP2,220 · ⑂ 252

    Basic Utilities for PyTorch Natural Language Processing (NLP)

    • pytorch
    • nlp
    • natural-language-processing
    • pytorch-nlp
    • torchnlp
    • data-loader
  22. 19
    MinishLab/model2vec2,210 · ⑂ 127

    Fast State-of-the-Art Static Embeddings

    • embeddings
    • machine-learning
    • model2vec
    • nlp
    • python
    • sentence-transformers

Find Python engineers shipping Embeddings

The list above ranks the most-starred public Python repositories tagged with the Embeddings topic, drawn from the public GitHub graph. Across 325 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Embeddings experience.

That overlap is rare. Most Python engineers haven’t shipped Embeddings, and most Embeddings maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.

Refolk turns this list into a search. Ask for Python Embeddings maintainers hiring” or Python engineers shipping Embeddings in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.

How this list is built

Refolk searched GitHub for public Python repositories tagged with the Embeddings topic, ranked them by stargazer count, and kept those with at least 25 stars. The list refreshes once a day.

Last refreshed: Sat, 19 Sep 2026 20:44:50 GMT

Search this list

Need a more specific search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Related lists

See all repository lists.

Or zoom out

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Sat, 19 Sep 2026 20:44:50 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring