Refolk

Top Vector databases repositories on GitHub

Embedding stores and approximate nearest-neighbour search engines.

Ranked by stars across 361 repositories tagged vector-database. Refreshed daily.

  1. 1
    Mintplex-Labs/anything-llm66,227 · ⑂ 7,362

    Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience

    • rag
    • localai
    • vector-database
    • llm
    • ai-agents
    • multimodal
  2. 2
    meilisearch/meilisearch59,343 · ⑂ 2,712

    A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.

    • search-engine
    • typo-tolerance
    • site-search
    • database
    • enterprise-search
    • search
  3. 3
    pathwaycom/llm-app58,916 · ⑂ 1,498

    Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.

    • chatbot
    • hugging-face
    • llm
    • llm-local
    • llm-prompting
    • llm-security
  4. 4
    run-llama/llama_index52,236 · ⑂ 8,177

    LlamaIndex is the document processing platform for AI

    • agents
    • application
    • data
    • fine-tuning
    • framework
    • llamaindex
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    milvus-io/milvus46,161 · ⑂ 4,257

    Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search

    • anns
    • nearest-neighbor-search
    • faiss
    • vector-search
    • image-search
    • hnsw
  7. 6
    VectifyAI/PageIndex35,757 · ⑂ 3,152

    📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG

    • agentic-ai
    • agents
    • ai
    • ai-agents
    • context-engineering
    • llm
  8. 7
    qdrant/qdrant34,690 · ⑂ 2,686

    Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/

    • neural-network
    • search-engine
    • knn-algorithm
    • hnsw
    • vector-search
    • nearest-neighbor-search
  9. 8
    topoteretes/cognee30,844 · ⑂ 3,065

    Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine.

    • ai
    • cognitive-architecture
    • vector-database
    • ai-agents
    • graph-database
    • ai-memory
  10. 9
    NirDiamant/RAG_Techniques29,549 · ⑂ 3,612

    This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.

    • rag
    • tutorials
    • langchain
    • llama-index
    • llms
    • python
  11. 10
    weaviate/weaviate16,825 · ⑂ 1,404

    Weaviate is an open-source vector database that stores both objects and vectors, allowing for the combination of vector search with structured filtering with the fault tolerance and scalability of a cloud-native database​.

    • search-engine
    • semantic-search
    • semantic-search-engine
    • vector-search
    • vector-search-engine
    • vector-database
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    memvid/memvid16,549 · ⑂ 1,421

    Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.

    • ai
    • context
    • embedded
    • faiss
    • knowledge-base
    • knowledge-graph
  14. 12
    alibaba/zvec15,969 · ⑂ 998

    A lightweight, lightning-fast, in-process vector database

    • rag
    • agent-skills
    • embedded
    • faiss
    • hnsw
    • llm-memory
  15. 13
    langchain4j/langchain4j13,127 · ⑂ 2,561

    LangChain4j is an idiomatic, open-source Java library for building LLM-powered applications on the JVM. It offers a unified API over popular LLM providers and vector stores, and makes implementing tool calling (including MCP support), agents and RAG easy. It integrates seamlessly with enterprise Java frameworks like Quarkus and Spring Boot.

    • huggingface
    • java
    • langchain
    • openai
    • chatgpt
    • gpt
  16. 14
    neuml/txtai12,964 · ⑂ 891

    💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows

    • python
    • search
    • nlp
    • semantic-search
    • vector-search
    • txtai
  17. 15
    StarTrail-org/LEANN12,946 · ⑂ 1,169

    [MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.

    • ai
    • faiss
    • langchain
    • llama-index
    • llm
    • localstorage
  18. 16
    zilliztech/claude-context12,544 · ⑂ 926

    Code search MCP for Claude Code. Make entire codebase the context for any coding agent.

    • agent
    • agentic-rag
    • ai-coding
    • code-search
    • cursor
    • embedding
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    lancedb/lancedb11,473 · ⑂ 1,059

    Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.

    • approximate-nearest-neighbor-search
    • image-search
    • nearest-neighbor-search
    • recommender-system
    • search-engine
    • semantic-search
  21. 18
    oramasearch/orama10,557 · ⑂ 402

    🌌 A complete search engine and RAG pipeline in your browser, server or edge network with support for full-text, vector, and hybrid search in less than 2kb.

    • data-structures
    • full-text
    • search
    • typo-tolerance
    • algiorithm
    • search-engine
  22. 19
    oceanbase/oceanbase10,283 · ⑂ 1,919

    OceanBase is the unified distributed database for the AI era — open-source, multi-model, one engine for your most demanding workloads.

    • oceanbase
    • paxos
    • htap
    • cloud-native
    • mysql-compatibility
    • oltp

Find engineers shipping Vector databases

The list above ranks the most-starred public repositories tagged with the Vector databases topic, drawn from the public GitHub graph. Across 361 repositories tagged this way, the maintainers and top contributors are a tight cluster of the people actually building Vector databases.

Looking for engineers who’ve worked on Vector databases for real, not just listed it on LinkedIn? The fastest path is the contributor list of these repos. Their commits, issues, and READMEs are public proof of depth.

Refolk turns this list into a search. Ask for “maintainers of top Vector databases repos who are hiring”, Vector databases engineers in San Francisco”, or “founders shipping Vector databases” and Refolk returns a ranked shortlist with sources.

How this list is built

Refolk searched GitHub for public repositories tagged with the Vector databases topic, ranked them by stargazer count, and kept those with at least 50 stars. The list refreshes once a day.

Last refreshed: Sun, 20 Sep 2026 00:23:33 GMT

Search this list

Need a list like this for any search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Browse other topics

See all repository lists.

Vector databases by language

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Sun, 20 Sep 2026 00:23:33 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring