Skip to content

Entry

VibeVoice

Appears in 5 awesome lists

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

Open github.commicrosoft/vibevoice

Found in these lists

AI Game DevTools (AI-GDT)

Section: Speech · VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

ActiveScore 77

Awesome LLM Resources

Section: 语音 Speech

FreshScore 87

Awesome Open Source AI

Section: 2. Model Codebases & Model Families · Open Frontier Voice AI toolkit spanning speech understanding, generation, and multilingual TTS workflows, with active research and deployment tooling.

FreshScore 89

Awesome Python

Section: AI and Agents · A family of open-source voice AI models from Microsoft for text-to-speech and long-form speech recognition.

FreshScore 94

awesome-python

Section: Other · Open-Source Frontier Voice AI

FreshScore 81

Whisper

Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.

In 11 listsDetails

FunASR

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

Bark

Bark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.

In 6 listsDetails

TorToiSe

A multi-voice TTS system trained with an emphasis on quality github | research paper | demo

In 5 listsDetails

SpeechBrain

PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.

In 5 listsDetails

kittentts

Kitten TTS is an open-source realistic text-to-speech model with just 15 million parameters, designed for lightweight deployment and high-quality voice synthesis.

In 4 listsDetails

TTS

🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

In 4 listsDetails

gpt-oss

OpenAI open-weight model repository with inference examples, recipes, and deployment guidance.

In 4 listsDetails