Skip to content
84

Awesome Speaker Diarization

A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.

1.9k stars245 forks171 entriesLast push Sep 16, 2026 (13 days ago)License Apache-2.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Publications >Special topics

A Review of Speaker Diarization: Recent Advances with Deep Learning

, 2021

A review on speaker diarization systems and approaches

, 2012

Speaker diarization: A review of recent research

, 2010

DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

, 2024

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

, 2023

Lexical speaker error correction: Leveraging language models for speaker diarization error correction

, 2023

DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors

, 2023

TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization

, 2023

Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

, 2022

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

, 2021

Supervised online diarization with sample mean loss for multi-domain data

, 2019

Discriminative Neural Clustering for Speaker Diarisation

, 2019

End-to-End Neural Speaker Diarization with Permutation-Free Objectives

, 2019

End-to-End Neural Speaker Diarization with Self-attention

, 2019

Fully Supervised Speaker Diarization

, 2018

A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings

, 2022

Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection

, 2021

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed…

, 2021

Joint Speech Recognition and Speaker Diarization via Sequence Transduction

, 2019

Says who? Deep learning models for joint speech recognition, segmentation and diarization

, 2018

Speaker Diarization as a Fully Online Bandit Learning Problem in MiniVox

, 2021

Online Speaker Diarization with Relation Network

, 2020

VoiceID on the Fly: A Speaker Recognition System that Learns from Scratch

, 2020

M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

, 2022

The Hitachi-JHU DIHARD III system: Competitive end-to-end neural diarization and x-vector clustering systems combined…

Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge

, 2018

ODESSA at Albayzin Speaker Diarization Challenge 2018

, 2018

Joint Discriminative Embedding Learning, Speech Activity and Overlap Detection for the DIHARD Challenge

, 2018

AVA-AVD: Audio-Visual Speaker Diarization in the Wild

, 2022

DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization

, 2022

End-to-End Audio-Visual Neural Speaker Diarization

, 2022

MSDWild: Multi-modal Speaker Diarization Dataset in the Wild

, 2022

Publications >Other

Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation

End-to-end speaker segmentation for overlap-aware resegmentation

DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding

DOVER-Lap: A method for combining overlap-aware diarization outputs

Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: Theory, implementation and analysis on…

AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in…

, 2021

An End-to-End Speaker Diarization Service for improving Multimedia Content Access

Spot the conversation: speaker diarisation in the wild

Speaker Diarization with Region Proposal Network

Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Overlap-aware diarization: resegmentation using neural end-to-end overlapped speech detection

Speaker diarization using latent space clustering in generative adversarial network

A study of semi-supervised speaker diarization system using gan mixture model

Learning deep representations by multilayer bootstrap networks for speaker diarization

Enhancements for Audio-only Diarization Systems

LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization

Meeting Transcription Using Virtual Microphone Arrays

Speaker diarisation using 2D self-attentive combination of embeddings

Speaker Diarization with Lexical Information

Neural speech turn segmentation and affinity propagation for speaker diarization

Multimodal Speaker Segmentation and Diarization using Lexical and Acoustic Cues via Sequence to Sequence Neural Networks

Joint Speaker Diarization and Recognition Using Convolutional and Recurrent Neural Networks

Speaker Diarization with LSTM

Speaker diarization using deep neural network embeddings

Speaker diarization using convolutional neural network for statistics accumulation refinement

pyannote. metrics: a toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems

Speaker Change Detection in Broadcast TV using Bidirectional Long Short-Term Memory Networks

Speaker Diarization using Deep Recurrent Convolutional Neural Networks for Speaker Embeddings

A Speaker Diarization System for Studying Peer-Led Team Learning Groups

Diarization resegmentation in the factor analysis subspace

A study of the cosine distance-based mean shift for telephone speech diarization

Speaker diarization with PLDA i-vector scoring and unsupervised calibration

Artificial neural network features for speaker diarization

Unsupervised methods for speaker diarization: An integrated and iterative approach

PLDA-based Clustering for Speaker Diarization of Broadcast Streams

Speaker diarization of meetings based on speaker role n-gram models

Speaker Diarization for Meeting Room Audio

Stream-based speaker segmentation using speaker factors and eigenvoices

An overview of automatic speaker diarization systems

A spectral clustering approach to speaker diarization

Software >Framework

FunASR

Industrial-grade speech recognition toolkit with built-in speaker diarization (cam++), VAD, ASR (SenseVoice/Paraformer), and punctuation. 50+ languages, 170x realtime.

In 9 listsDetails

MiniVox

MiniVox is an open-source evaluation system for the online speaker diarization task.

SpeechBrain

SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.

In 5 listsDetails

SIDEKIT for diarization (s4d)

An open source package extension of SIDEKIT for Speaker diarization.

pyAudioAnalysis

Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications.

In 4 lists

AaltoASR

Speaker diarization scripts, based on AaltoASR.

LIUM SpkDiarization

LIUM_SpkDiarization is a software dedicated to speaker diarization (i.e. speaker segmentation and clustering). It is written in Java, and includes the most recent developments in the domain (as of 2013).

kaldi-asr

Example scripts for speaker diarization on a portion of CALLHOME used in the 2000 NIST speaker recognition evaluation.

In 5 listsDetails

kaldi-speaker-diarization

Icelandic speaker diarization scripts using kaldi.

Alize LIA_SpkSeg

ALIZÉ is an opensource platform for speaker recognition. LIA_SpkSeg is the tools for speaker diarization.

pyannote-audio

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, speaker embedding.

In 3 lists

pyBK

Speaker diarization using binary key speaker modelling. Computationally light solution that does not require external training data.

Speaker-Diarization

Speaker diarization using uis-rnn and GhostVLAD. An easier way to support openset speakers.

EEND

End-to-End Neural Diarization.

VBx

Variational Bayes HMM over x-vectors diarization. x-vector extractor recipe

RE-VERB

RE: VERB is speaker diarization system, it allows the user to send/record audio of a conversation and receive timestamps of who spoke when.

StreamingSpeakerDiarization

Streaming speaker diarization, extends pyannote.audio to online processing

simple_diarizer

Simplified diarization pipeline using some pretrained models. Made to be a simple as possible to go from an input audio file to diarized segments.

Picovoice Falcon

A lightweight, accurate, and fast speaker diarization engine written in C and available in Python, running on CPU with minimal overhead.

DiaPer

Pytorch implementation for DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors including models pre-trained on free and public data.

sherpa-onnx

Support speaker diarization, speech recognition, and text-to speech on various platforms with various language bindings.

In 4 listsDetails

FluidAudio

A native Swift speaker diarization library for Apple platforms, using CoreML for efficient, real-time audio processing with high accuracy.

In 4 listsDetails

Software >Evaluation

pyannote-metrics

A toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems.

SimpleDER

A lightweight library to compute Diarization Error Rate (DER).

DiarizationLM

Implements Word Error Rate (WER), Word Diarization Error Rate (WDER), and concatenated minimum-permutation Word Error Rate (cpWER).

dscore

Diarization scoring tools.

Sequence Match Accuracy

Match the accuracy of two sequences with Hungarian algorithm.

In 2 lists

spyder

Simple Python package for fast DER computation.

CDER

Conversational DER from The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines

Selective Collar

Selective collar for mitigating evaluation bias in speaker diarization from Beyond Uniform Forgiveness: Introducing the Selective Collar to Mitigate Evaluation Bias in Speaker Diarization

Software >Clustering

Sequence Match Accuracy

Match the accuracy of two sequences with Hungarian algorithm.

In 2 lists

uis-rnn-sml

A variant of UIS-RNN, for the paper Supervised Online Diarization with Sample Mean Loss for Multi-Domain Data.

DNC

Transformer-based Discriminative Neural Clustering (DNC) for Speaker Diarisation. Like UIS-RNN, it is supervised.

SpectralCluster

Spectral clustering with affinity matrix refinement operations, auto-tune, and speaker turn constraints.

sklearn.cluster

scikit-learn clustering algorithms.

In 2 lists

PLDA

Probabilistic Linear Discriminant Analysis & classification, written in Python.

PLDA

Open-source implementation of simplified PLDA (Probabilistic Linear Discriminant Analysis).

Auto-Tuning Spectral Clustering

Auto-tuning Spectral Clustering method that does not need development set or supervised tuning.

Software >Speaker embedding

resemble-ai/Resemblyzer

PyTorch implementation of generalized end-to-end loss for speaker verification, which can be used for voice cloning and diarization.

Speaker_Verification

Tensorflow implementation of generalized end-to-end loss for speaker verification.

PyTorch_Speaker_Verification

PyTorch implementation of "Generalized End-to-End Loss for Speaker Verification" by Wan, Li et al. With UIS-RNN integration.

Real-Time Voice Cloning

Implementation of "Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis" (SV2TTS) with a vocoder that works in real-time.

In 2 lists

conformer-speaker-encoder

Massively multilingual conformer-based speaker recognition models in TFLite format.

deep-speaker

Third party implementation of the Baidu paper Deep Speaker: an End-to-End Neural Speaker Embedding System.

x-vector-kaldi-tf

Tensorflow implementation of x-vector topology on top of Kaldi recipe.

kaldi-ivector

Extension to Kaldi implementing the standard i-vector hyperparameter estimation and i-vector extraction procedure.

voxceleb-ivector

Voxceleb1 i-vector based speaker recognition system.

pytorch_xvectors

PyTorch implementation of Voxceleb x-vectors. Additionaly, includes meta-learning architectures for embedding training. Evaluated with speaker diarization and speaker verification.

ASVtorch

ASVtorch is a toolkit for automatic speaker recognition.

asv-subtools

ASV-Subtools is developed based on Pytorch and Kaldi for the task of speaker recognition, language identification, etc. The 'sub' of 'subtools' means that there are many modular tools and the parts constitute the whole.

WeSpeaker

WeSpeaker is a research and production oriented speaker verification, recognition and diarization toolkit, which supports very strong recipes with on-the-fly data preparation, model training and evaluation, as well as runtime C++ codes.

ReDimNet

Neural network architecture presented in the paper Reshape Dimensions Network for Speaker Recognition

Software >Speaker change detection

change_detection

Code for Speaker Change Detection in Broadcast TV using Bidirectional Long Short-Term Memory Networks.

tidydiarize

Diarization inside OpenAI Whisper decoder

Software >Audio feature extraction

LibROSA

Python library for audio and music analysis. https://librosa.github.io/

In 5 listsDetails

python_speech_features

This library provides common speech features for ASR including MFCCs and filterbank energies. https://python-speech-features.readthedocs.io/en/latest/

In 3 lists

pyAudioAnalysis

Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications.

In 4 lists

Software >Audio data augmentation

pyroomacoustics

Pyroomacoustics is a package for audio signal processing for indoor applications. It was developed as a fast prototyping platform for beamforming algorithms in indoor scenarios. https://pyroomacoustics.readthedocs.io

In 2 lists

gpuRIR

Python library for Room Impulse Response (RIR) simulation with GPU acceleration

rir_simulator_python

Room impulse response simulator using python

WavAugment

WavAugment performs data augmentation on audio data. The audio data is represented as pytorch tensors

EEND_dataprep

Recipes for generating simulated conversations used to train end-to-end diarization models.

Software >Other software

VB Diarization

VB Diarization with Eigenvoice and HMM Priors.

DOVER-Lap

Python package for combining diarization system outputs

Diar-az

Data formatting tool to support the ruv-di dataset. Kaldi to Gecko to Kaldi and corpus and back

voxsolo

Verbatim per-speaker audio from diarization: keeps one speaker bit-exact, silences other speakers and overlapped speech; exports EDL/SRT/VTT/Audacity labels for editors.

Datasets >Diarization datasets

2000 NIST Speaker Recognition Evaluation

Evaluation Plan

2003 NIST Rich Transcription Evaluation Data

telephone speech, broadcast news

CALLHOME American English Speech

CH109 whitelist

The ICSI Meeting Corpus

License

The AMI Meeting Corpus

License

Fisher English Training Speech Part 1 Speech

Fisher English Training Speech Part 1 Transcripts

Fisher English Training Part 2, Speech

Fisher English Training Part 2, Transcripts

VoxConverse

VoxConverse is an audio-visual diarisation dataset consisting of over 50 hours of multispeaker clips of human speech, extracted from YouTube videos

MiniVox

MiniVox is an open-source evaluation system for the online speaker diarization task.

The AliMeeting Corpus

Together with audios

Datasets >Speaker embedding training sets

TIMIT

Published in 1993, the TIMIT corpus of read speech is one of the earliest speaker recognition datasets.

VCTK

Most were selected from a newspaper plus the Rainbow Passage and an elicitation paragraph intended to identify the speaker's accent.

LibriSpeech

Large-scale (1000 hours) corpus of read English speech.

In 2 lists

Multilingual LibriSpeech (MLS)

Multilingual LibriSpeech (MLS) dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.

LibriVox

Free public domain audiobooks. LibriSpeech is a processed subset of LibriVox. Each original unsegmented utterance could be very long.

In 2 lists

VoxCeleb 1&2

VoxCeleb is an audio-visual dataset consisting of short clips of human speech, extracted from interview videos uploaded to YouTube.

The Spoken Wikipedia Corpora

Volunteer readers reading Wikipedia articles.

CN-Celeb

A Free Chinese Speaker Recognition Corpus Released by CSLT@Tsinghua University.

BookTubeSpeech

Audio samples extracted from BookTube videos - videos where people share their opinions on books - from YouTube. The dataset can be downloaded using BookTubeSpeech-download.

DeepMine

A speech database in Persian and English designed to build and evaluate speaker verification, as well as Persian ASR systems.

NISP-Dataset

This dataset contains speech recordings along with speaker physical parameters (height, weight, ... ) as well as regional information and linguistic information.

VoxBlink2

Multilingual dataset from VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Datasets >Augmentation noise sources

AudioSet

A large-scale dataset of manually annotated audio events.

MUSAN

MUSAN is a corpus of music, speech, and noise recordings.

Other learning materials >Books

Voice Identity Techniques: From core algorithms to engineering practice (Chinese)

by Quan Wang, 2020

Other learning materials >Tech blogs

Literature Review For Speaker Change Detection

by Halil Erdoğan

Speaker Diarization: Separation of Multiple Speakers in an Audio File

by Jaspreet Singh

Speaker Diarization with Kaldi

by Yoav Ramon

Who spoke when! How to Build your own Speaker Diarization Module

by Rahul Saxena

Other learning materials >Video tutorials

pyannote audio: neural building blocks for speaker diarization

by Hervé Bredin

Google's Diarization System: Speaker Diarization with LSTM

by Google

Fully Supervised Speaker Diarization: Say Goodbye to clustering

by Google

Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection

by Google

Speaker Diarization: Optimal Clustering and Learning Speaker Embeddings

by Microsoft Research

Robust Speaker Diarization for Meetings: the ICSI system

by Microsoft Research

【机器之心&博文视点】入门声纹技术|第二讲:声纹分割聚类与其他应用

by Quan Wang

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago