Refolk

Top Speech recognition repositories on GitHub

ASR models and pipelines for converting audio to text.

Ranked by stars across 702 repositories tagged speech-recognition. Refreshed daily.

  1. 1
    huggingface/transformers166,453 · ⑂ 34,645

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    • nlp
    • natural-language-processing
    • pytorch
    • pytorch-transformers
    • transformer
    • model-hub
  2. 2
    ggml-org/whisper.cpp53,811 · ⑂ 6,183

    Port of OpenAI's Whisper model in C/C++

    • openai
    • speech-to-text
    • transformer
    • whisper
    • inference
    • speech-recognition
  3. 3
    mozilla/DeepSpeech26,775 · ⑂ 4,075

    DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.

    • deep-learning
    • machine-learning
    • neural-networks
    • tensorflow
    • speech-recognition
    • speech-to-text
  4. 4
    SYSTRAN/faster-whisper25,483 · ⑂ 2,079

    Faster Whisper transcription with CTranslate2

    • deep-learning
    • inference
    • quantization
    • speech-recognition
    • speech-to-text
    • transformer
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    m-bain/whisperX24,151 · ⑂ 2,433

    WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

    • asr
    • speech
    • speech-recognition
    • speech-to-text
    • whisper
  7. 6
    modelscope/FunASR20,443 · ⑂ 2,040

    Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

    • pytorch
    • speech-recognition
    • paraformer
    • punctuation
    • speaker-diarization
    • voice-activity-detection
  8. 7
    leon-ai/leon17,531 · ⑂ 1,469

    🧠 Leon is your open-source personal assistant.

    • leon
    • personal-assistant
    • nodejs
    • python
    • ai
    • artificial-intelligence
  9. 8
    kaldi-asr/kaldi15,488 · ⑂ 5,359

    kaldi-asr/kaldi is the official location of the Kaldi project.

    • kaldi
    • c-plus-plus
    • cuda
    • shell
    • speech-recognition
    • speech-to-text
  10. 9
    alphacep/vosk-api15,141 · ⑂ 1,760

    Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

    • speech-recognition
    • asr
    • voice-recognition
    • speech-to-text
    • android
    • ios
  11. 10
    NVIDIA/DeepLearningExamples14,847 · ⑂ 3,407

    State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.

    • computer-vision
    • deep-learning
    • drug-discovery
    • forecasting
    • large-language-models
    • mxnet
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    kmario23/deep-learning-drizzle12,949 · ⑂ 2,983

    Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures!!

    • machine-learning
    • deep-learning
    • deep-neural-networks
    • pattern-recognition
    • computer-vision
    • optimization
  14. 12
    abus-aikorea/voice-pro12,851 · ⑂ 1,853

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    • faster-whisper
    • tts
    • whisper
    • gradio
    • subtitles
    • transcription
  15. 13
    PaddlePaddle/PaddleSpeech12,686 · ⑂ 1,957

    Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

    • transformer
    • conformer
    • speech-translation
    • streaming-asr
    • speech-alignment
    • punctuation-restoration
  16. 14
    speechbrain/speechbrain11,828 · ⑂ 1,725

    A PyTorch-based Speech Toolkit

    • speech-recognition
    • speech-toolkit
    • speaker-recognition
    • speech-to-text
    • speech-enhancement
    • speech-separation
  17. 15
    QuentinFuxa/WhisperLiveKit11,073 · ⑂ 1,137

    Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

    • python
    • real-time
    • speaker-diarization
    • speech-recognition
    • speech-to-text
    • streaming
  18. 16
    openvinotoolkit/openvino10,880 · ⑂ 3,388

    OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

    • inference
    • deep-learning
    • openvino
    • ai
    • computer-vision
    • diffusion-models
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    espnet/espnet9,963 · ⑂ 2,435

    End-to-End Speech Processing Toolkit

    • deep-learning
    • end-to-end
    • chainer
    • pytorch
    • kaldi
    • speech-recognition
  21. 18
    QwenAudio/SenseVoice9,334 · ⑂ 827

    Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

    • asr
    • speech-recognition
    • speech-to-text
    • cross-lingual
    • pytorch
    • speech-emotion-recognition
  22. 19
    Uberi/speech_recognition8,989 · ⑂ 2,416

    Speech recognition module for Python, supporting several engines and APIs, online and offline.

    • python
    • audio
    • speech-recognition
    • speech-to-text

Find engineers shipping Speech recognition

The list above ranks the most-starred public repositories tagged with the Speech recognition topic, drawn from the public GitHub graph. Across 702 repositories tagged this way, the maintainers and top contributors are a tight cluster of the people actually building Speech recognition.

Looking for engineers who’ve worked on Speech recognition for real, not just listed it on LinkedIn? The fastest path is the contributor list of these repos. Their commits, issues, and READMEs are public proof of depth.

Refolk turns this list into a search. Ask for “maintainers of top Speech recognition repos who are hiring”, Speech recognition engineers in San Francisco”, or “founders shipping Speech recognition” and Refolk returns a ranked shortlist with sources.

How this list is built

Refolk searched GitHub for public repositories tagged with the Speech recognition topic, ranked them by stargazer count, and kept those with at least 50 stars. The list refreshes once a day.

Last refreshed: Sun, 20 Sep 2026 23:21:29 GMT

Search this list

Need a list like this for any search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Browse other topics

See all repository lists.

Speech recognition by language

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Sun, 20 Sep 2026 23:21:29 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring