Top Speech recognition repositories on GitHub
ASR models and pipelines for converting audio to text.
Ranked by stars across 702 repositories tagged speech-recognition. Refreshed daily.
- 1huggingface/transformers★ 166,453 · ⑂ 34,645
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
- nlp
- natural-language-processing
- pytorch
- pytorch-transformers
- transformer
- model-hub
- 2ggml-org/whisper.cpp★ 53,811 · ⑂ 6,183
Port of OpenAI's Whisper model in C/C++
- openai
- speech-to-text
- transformer
- whisper
- inference
- speech-recognition
- 3mozilla/DeepSpeech★ 26,775 · ⑂ 4,075
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
- deep-learning
- machine-learning
- neural-networks
- tensorflow
- speech-recognition
- speech-to-text
- 4SYSTRAN/faster-whisper★ 25,483 · ⑂ 2,079
Faster Whisper transcription with CTranslate2
- deep-learning
- inference
- quantization
- speech-recognition
- speech-to-text
- transformer
- Live search
Find the people behind these repos
Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.
500 free credits on sign-up, no card needed.
- 5m-bain/whisperX★ 24,151 · ⑂ 2,433
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
- asr
- speech
- speech-recognition
- speech-to-text
- whisper
- 6modelscope/FunASR★ 20,443 · ⑂ 2,040
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
- pytorch
- speech-recognition
- paraformer
- punctuation
- speaker-diarization
- voice-activity-detection
- 7leon-ai/leon★ 17,531 · ⑂ 1,469
🧠 Leon is your open-source personal assistant.
- leon
- personal-assistant
- nodejs
- python
- ai
- artificial-intelligence
- 8kaldi-asr/kaldi★ 15,488 · ⑂ 5,359
kaldi-asr/kaldi is the official location of the Kaldi project.
- kaldi
- c-plus-plus
- cuda
- shell
- speech-recognition
- speech-to-text
- 9alphacep/vosk-api★ 15,141 · ⑂ 1,760
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
- speech-recognition
- asr
- voice-recognition
- speech-to-text
- android
- ios
- 10NVIDIA/DeepLearningExamples★ 14,847 · ⑂ 3,407
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
- computer-vision
- deep-learning
- drug-discovery
- forecasting
- large-language-models
- mxnet
- Live search
Who is hiring in this space?
I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.
500 free credits on sign-up, no card needed.
- 11kmario23/deep-learning-drizzle★ 12,949 · ⑂ 2,983
Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures!!
- machine-learning
- deep-learning
- deep-neural-networks
- pattern-recognition
- computer-vision
- optimization
- 12abus-aikorea/voice-pro★ 12,851 · ⑂ 1,853
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
- faster-whisper
- tts
- whisper
- gradio
- subtitles
- transcription
- 13PaddlePaddle/PaddleSpeech★ 12,686 · ⑂ 1,957
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
- transformer
- conformer
- speech-translation
- streaming-asr
- speech-alignment
- punctuation-restoration
- 14speechbrain/speechbrain★ 11,828 · ⑂ 1,725
A PyTorch-based Speech Toolkit
- speech-recognition
- speech-toolkit
- speaker-recognition
- speech-to-text
- speech-enhancement
- speech-separation
- 15QuentinFuxa/WhisperLiveKit★ 11,073 · ⑂ 1,137
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
- python
- real-time
- speaker-diarization
- speech-recognition
- speech-to-text
- streaming
- 16openvinotoolkit/openvino★ 10,880 · ⑂ 3,388
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
- inference
- deep-learning
- openvino
- ai
- computer-vision
- diffusion-models
- Live search
Turn any brief into a list like this
I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.
500 free credits on sign-up, no card needed.
- 17espnet/espnet★ 9,963 · ⑂ 2,435
End-to-End Speech Processing Toolkit
- deep-learning
- end-to-end
- chainer
- pytorch
- kaldi
- speech-recognition
- 18QwenAudio/SenseVoice★ 9,334 · ⑂ 827
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
- asr
- speech-recognition
- speech-to-text
- cross-lingual
- pytorch
- speech-emotion-recognition
- 19Uberi/speech_recognition★ 8,989 · ⑂ 2,416
Speech recognition module for Python, supporting several engines and APIs, online and offline.
- python
- audio
- speech-recognition
- speech-to-text
Find engineers shipping Speech recognition
The list above ranks the most-starred public repositories tagged with the Speech recognition topic, drawn from the public GitHub graph. Across 702 repositories tagged this way, the maintainers and top contributors are a tight cluster of the people actually building Speech recognition.
Looking for engineers who’ve worked on Speech recognition for real, not just listed it on LinkedIn? The fastest path is the contributor list of these repos. Their commits, issues, and READMEs are public proof of depth.
Refolk turns this list into a search. Ask for “maintainers of top Speech recognition repos who are hiring”, “Speech recognition engineers in San Francisco”, or “founders shipping Speech recognition” and Refolk returns a ranked shortlist with sources.
How this list is built
Last refreshed: Sun, 20 Sep 2026 23:21:29 GMT
Need a list like this for any search?
Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:
- Speech recognition maintainers hiringRun
- Speech recognition engineers in San FranciscoRun
- Companies shipping Speech recognitionRun
500 free credits on sign-up, no card needed.
Browse other topics
- Top Design systems repos
- Top Terminal UI repos
- Top Computer vision repos
- Top REST APIs repos
- Top Fine-tuning repos
- Top AI agents repos
- Top Game development repos
- Top Text-to-speech repos
See all repository lists.
Speech recognition by language
Common questions
How are these repositories ranked?
By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.
How fresh is the data?
The ranking re-renders at least daily. Last refreshed: Sun, 20 Sep 2026 23:21:29 GMT.
Can I find the maintainers and contributors behind these repos?
Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.
Can I use this list for hiring?
That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.