Skip to content

Entry

SpeechBrain

Appears in 5 awesome lists

PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.

Open github.comspeechbrain/speechbrain

Found in these lists

Awesome Speaker Diarization

Section: Framework · SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.

FreshScore 84

awesome-huggingface

Section: Speech · A PyTorch-based speech toolkit.

ArchivedScore 49

Awesome Open Source AI

Section: 2. Model Codebases & Model Families · PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.

FreshScore 89

Awesome-Pytorch-list

Section: NLP & Speech Processing: · SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.

FreshScore 88

awesome-python

Section: Other · A PyTorch-based Speech Toolkit

FreshScore 81

Whisper

Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.

In 11 listsDetails

FunASR

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

VibeVoice

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

In 5 listsDetails

Kaldi

Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.

In 5 listsDetails

FluidAudio

SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.

In 4 listsDetails

gpt-oss

OpenAI open-weight model repository with inference examples, recipes, and deployment guidance.

In 4 listsDetails

sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…

In 4 listsDetails

MiniCPM-V

Compact vision-language model family with edge-focused deployment examples and strong OCR-oriented use cases.

In 4 listsDetails