Refolk

Top RAG repositories on GitHub

Retrieval-augmented generation pipelines, embeddings, and grounding tooling.

Ranked by stars across 1,530 repositories tagged rag. Refreshed daily.

  1. 1
    open-webui/open-webui152,627 · ⑂ 22,336

    User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

    • ollama
    • ollama-webui
    • llm
    • webui
    • self-hosted
    • llm-ui
  2. 2
    langchain-ai/langchain146,746 · ⑂ 24,541

    The agent engineering platform.

    • ai
    • anthropic
    • gemini
    • langchain
    • llm
    • openai
  3. 3
    Shubhamsaboo/awesome-llm-apps139,148 · ⑂ 20,454

    100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.

    • llms
    • rag
    • python
    • agents
  4. 4
    Graphify-Labs/graphify119,860 · ⑂ 11,590

    Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.

    • claude-code
    • graphrag
    • knowledge-graph
    • codex
    • openclaw
    • skills
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    thedotmack/claude-mem94,334 · ⑂ 8,330

    Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

    • ai
    • ai-agents
    • ai-memory
    • anthropic
    • artificial-intelligence
    • claude
  7. 6
    infiniflow/ragflow91,065 · ⑂ 10,791

    RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs

    • ai
    • ai-agents
    • context-engine
    • rag
    • retrieval-augmented-generation
    • agentic-ai
  8. 7
    PaddlePaddle/PaddleOCR89,885 · ⑂ 11,373

    Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

    • ocr
    • chineseocr
    • pdf2markdown
    • pp-ocr
    • pp-structure
    • document-parsing
  9. 8
    datawhalechina/hello-agents80,086 · ⑂ 9,969

    📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程

    • agent
    • tutorial
    • llm
    • rag
  10. 9
    dair-ai/Prompt-Engineering-Guide78,503 · ⑂ 8,630

    🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.

    • deep-learning
    • prompt-engineering
    • openai
    • chatgpt
    • language-model
    • generative-ai
  11. 10
    headroomlabs-ai/headroom73,245 · ⑂ 5,636

    Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

    • agent
    • ai
    • anthropic
    • compression
    • context-engineering
    • context-window
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    Mintplex-Labs/anything-llm66,258 · ⑂ 7,370

    Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience

    • rag
    • localai
    • vector-database
    • llm
    • ai-agents
    • multimodal
  14. 12
    mem0ai/mem065,716 · ⑂ 7,717

    The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.

    • ai
    • chatgpt
    • llm
    • python
    • rag
    • long-term-memory
  15. 13
    pathwaycom/llm-app58,911 · ⑂ 1,499

    Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.

    • chatbot
    • hugging-face
    • llm
    • llm-local
    • llm-prompting
    • llm-security
  16. 14
    FlowiseAI/Flowise55,468 · ⑂ 25,033

    Build AI Agents, Visually

    • artificial-intelligence
    • chatgpt
    • large-language-models
    • low-code
    • no-code
    • javascript
  17. 15
    run-llama/llama_index52,249 · ⑂ 8,180

    LlamaIndex is the document processing platform for AI

    • agents
    • application
    • data
    • fine-tuning
    • framework
    • llamaindex
  18. 16
    bojieli/ai-agent-book48,971 · ⑂ 5,498

    《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码

    • agent
    • agent-memory
    • ai-agent
    • book
    • coding-agent
    • context-engineering
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    jeecgboot/JeecgBoot47,911 · ⑂ 16,195

    【低代码v2.0,一句话即可生成整个系统】企业级AI低代码平台,一键生成前后端代码甚至整个系统。 AI Skills 一句话画流程、设计表单、生成报表、大屏。内置 AI应用平台涵盖:AI聊天、知识库、流程编排、MCP插件等,兼容主流大模型。引领AI低代码「Skills 生成 → 在线配置 → 代码生成 → 手工合并->AI修改」开发模式,解决 Java 项目 90% 重复工作,提高效率又不失灵活。

    • antd
    • activiti
    • codegenerator
    • springcloud
    • springboot
    • low-code
  21. 18
    milvus-io/milvus46,176 · ⑂ 4,257

    Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search

    • anns
    • nearest-neighbor-search
    • faiss
    • vector-search
    • image-search
    • hnsw
  22. 19
    langchain-ai/langgraph42,025 · ⑂ 7,093

    Build resilient agents.

    • agents
    • ai
    • ai-agents
    • chatgpt
    • deepagents
    • enterprise

Find engineers shipping RAG

The list above ranks the most-starred public repositories tagged with the RAG topic, drawn from the public GitHub graph. Across 1,530 repositories tagged this way, the maintainers and top contributors are a tight cluster of the people actually building RAG.

Looking for engineers who’ve worked on RAG for real, not just listed it on LinkedIn? The fastest path is the contributor list of these repos. Their commits, issues, and READMEs are public proof of depth.

Refolk turns this list into a search. Ask for “maintainers of top RAG repos who are hiring”, RAG engineers in San Francisco”, or “founders shipping RAG” and Refolk returns a ranked shortlist with sources.

How this list is built

Refolk searched GitHub for public repositories tagged with the RAG topic, ranked them by stargazer count, and kept those with at least 50 stars. The list refreshes once a day.

Last refreshed: Sun, 20 Sep 2026 20:23:41 GMT

Search this list

Need a list like this for any search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Browse other topics

See all repository lists.

RAG by language

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Sun, 20 Sep 2026 20:23:41 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring