Refolk

Top Python Speech recognition repositories on GitHub

ASR models and pipelines for converting audio to text. Filtered to projects whose primary language is Python.

Ranked by stars across 486 Python repositories tagged speech-recognition. Refreshed daily.

  1. 1
    huggingface/transformers166,276 · ⑂ 34,620

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    • nlp
    • natural-language-processing
    • pytorch
    • pytorch-transformers
    • transformer
    • model-hub
  2. 2
    SYSTRAN/faster-whisper25,451 · ⑂ 2,075

    Faster Whisper transcription with CTranslate2

    • deep-learning
    • inference
    • quantization
    • speech-recognition
    • speech-to-text
    • transformer
  3. 3
    m-bain/whisperX24,092 · ⑂ 2,426

    WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

    • asr
    • speech
    • speech-recognition
    • speech-to-text
    • whisper
  4. 4
    modelscope/FunASR20,410 · ⑂ 2,035

    Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

    • pytorch
    • speech-recognition
    • paraformer
    • punctuation
    • speaker-diarization
    • voice-activity-detection
  5. Live search

    Find the people behind these repos

    Stars rank the projects. I can rank the engineers - maintainers, top contributors, and the people they work with. Fire one of these to see how it works.

    500 free credits on sign-up, no card needed.

  6. 5
    abus-aikorea/voice-pro12,842 · ⑂ 1,854

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    • faster-whisper
    • tts
    • whisper
    • gradio
    • subtitles
    • transcription
  7. 6
    PaddlePaddle/PaddleSpeech12,685 · ⑂ 1,957

    Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

    • transformer
    • conformer
    • speech-translation
    • streaming-asr
    • speech-alignment
    • punctuation-restoration
  8. 7
    speechbrain/speechbrain11,823 · ⑂ 1,727

    A PyTorch-based Speech Toolkit

    • speech-recognition
    • speech-toolkit
    • speaker-recognition
    • speech-to-text
    • speech-enhancement
    • speech-separation
  9. 8
    QuentinFuxa/WhisperLiveKit11,057 · ⑂ 1,137

    Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

    • python
    • real-time
    • speaker-diarization
    • speech-recognition
    • speech-to-text
    • streaming
  10. 9
    espnet/espnet9,965 · ⑂ 2,432

    End-to-End Speech Processing Toolkit

    • deep-learning
    • end-to-end
    • chainer
    • pytorch
    • kaldi
    • speech-recognition
  11. 10
    Uberi/speech_recognition8,988 · ⑂ 2,416

    Speech recognition module for Python, supporting several engines and APIs, online and offline.

    • python
    • audio
    • speech-recognition
    • speech-to-text
  12. Live search

    Who is hiring in this space?

    I read hiring signals across LinkedIn, GitHub, and the open web - so a topic list becomes a warm outreach list. Try one live.

    500 free credits on sign-up, no card needed.

  13. 11
    nl8590687/ASRT_SpeechRecognition8,390 · ⑂ 1,892

    A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统

    • tensorflow
    • cnn
    • ctc
    • python
    • keras
    • speech-recognition
  14. 12
    Blaizzy/mlx-audio7,908 · ⑂ 719

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

    • apple-silicon
    • audio-processing
    • mlx
    • multimodal
    • speech-recognition
    • speech-synthesis
  15. 13
    modelscope/FunClip6,327 · ⑂ 754

    FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

    • speech-recognition
    • video-subtitles
    • subtitles-generator
    • speech-to-text
    • gradio
    • llm
  16. 14
    PaddlePaddle/PaddleX6,265 · ⑂ 1,218

    All-in-One Development Tool based on PaddlePaddle

    • classification
    • segmentation
    • deployment
    • ocr
    • time-series
    • pp-chatocr
  17. 15
    wenet-e2e/wenet5,239 · ⑂ 1,188

    Production First and Production Ready End-to-End Speech Recognition Toolkit

    • e2e-models
    • pytorch
    • asr
    • transformer
    • conformer
    • production-ready
  18. 16
    Picovoice/porcupine4,939 · ⑂ 580

    On-device wake word detection powered by deep learning

    • wake-word-detection
    • hotword
    • keyword-spotting
    • keyword-spotter
    • wake-word
    • wake-word-engine
  19. Live search

    Turn any brief into a list like this

    I run natural-language searches across GitHub, LinkedIn, and the open web. Describe who you want and I'll build the shortlist.

    500 free credits on sign-up, no card needed.

  20. 17
    yanshengjia/ml-road4,937 · ⑂ 1,723

    Machine Learning and Agentic AI Resources, Practice and Research

    • machine-learning
    • deep-learning
    • nlp
    • computer-vision
    • speech-recognition
    • tensorflow
  21. 18
    jianchang512/stt4,796 · ⑂ 503

    Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式

    • speech
    • speech-recognition
    • speech-to-text
    • stt
  22. 19
    huggingface/distil-whisper4,117 · ⑂ 357

    Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.

    • audio
    • speech-recognition
    • whisper

Find Python engineers shipping Speech recognition

The list above ranks the most-starred public Python repositories tagged with the Speech recognition topic, drawn from the public GitHub graph. Across 486 matching repositories, the contributors are a tight cluster of engineers with both Python chops and real Speech recognition experience.

That overlap is rare. Most Python engineers haven’t shipped Speech recognition, and most Speech recognition maintainers don’t write Python. The people on this list’s contributor graph are the ones who do both.

Refolk turns this list into a search. Ask for Python Speech recognition maintainers hiring” or Python engineers shipping Speech recognition in 2025” and Refolk returns a ranked shortlist with the commits, profiles, and projects behind each name.

How this list is built

Refolk searched GitHub for public Python repositories tagged with the Speech recognition topic, ranked them by stargazer count, and kept those with at least 25 stars. The list refreshes once a day.

Last refreshed: Fri, 18 Sep 2026 08:00:02 GMT

Search this list

Need a more specific search?

Refolk runs natural-language searches across GitHub, LinkedIn, and the open web. Try one of these:

500 free credits on sign-up, no card needed.

Related lists

See all repository lists.

Or zoom out

Common questions

How are these repositories ranked?

By stars, with forks and recent activity as tiebreakers, read from the public GitHub API. The methodology section above has the details.

How fresh is the data?

The ranking re-renders at least daily. Last refreshed: Fri, 18 Sep 2026 08:00:02 GMT.

Can I find the maintainers and contributors behind these repos?

Yes. Stars rank the projects; I can rank the engineers - maintainers, top contributors, and the people they work with. You start with 500 free credits, no card required.

Can I use this list for hiring?

That's the point. I read hiring signals across GitHub, LinkedIn, and the open web, so a repo list turns into a shortlist of engineers worth talking to.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Keep exploring