Repo Voice & Audio

coqui-ai/TTS

Coqui TTS is a Python deep learning toolkit for text-to-speech with pretrained models, voice cloning and conversion, training tools, a CLI and a Python API.

  • 46.1k GitHub stars
  • Python
  • ⚖️ MPL-2.0
  • 🎯 Intermediate
pip install TTS
coqui-ai/TTS preview image

What it is

🐸TTS is a library for advanced Text-to-Speech generation. It provides pretrained models in over 1100 languages, tools for training and fine-tuning models in any language, and utilities for dataset analysis and curation. It includes spectrogram models, end-to-end models such as ⓍTTS, VITS, YourTTS, Tortoise and Bark, vocoders, a speaker encoder, and voice conversion.

Who it's for

  • Developers who want to synthesize speech from released pretrained TTS models via Python or the command line
  • Researchers and engineers training or fine-tuning TTS models in any language
  • Users who need multi-speaker, multilingual or voice-cloning TTS
  • People curating Text2Speech datasets

Requirements

Requirements

  • Python >= 3.9, < 3.12
  • Tested on Ubuntu 18.04
  • make system-deps is intended for Ubuntu (Debian)

Setup

  1. Install from PyPI (synthesis only)

    The easiest option if you only want to synthesize speech with the released models.

    bash
    pip install TTS
  2. Install locally for coding or training

    Clone the repo and install it locally, selecting the relevant extras.

    bash
    git clone https://github.com/coqui-ai/TTS
    pip install -e .[all,dev,notebooks]  # Select the relevant extras
  3. Ubuntu/Debian install via make

    On Ubuntu (Debian) you can also use these commands.

    bash
    $ make system-deps  # intended to be used on Ubuntu (Debian). Let us know if you have a different OS.
    $ make install
  4. Try with Docker

    Run TTS without installing it by using the CPU docker image, then start a server.

    bash
    docker run --rm -it -p 5002:5002 --entrypoint /bin/bash ghcr.io/coqui-ai/tts-cpu
    python3 TTS/server/server.py --list_models #To get the list of available models
    python3 TTS/server/server.py --model_name tts_models/en/vctk/vits # To start a server

Examples

Multilingual voice cloning with ⓍTTS v2

python
python
import torch
from TTS.api import TTS

# Get device
device = "cuda" if torch.cuda.is_available() else "cpu"

# List available 🐸TTS models
print(TTS().list_models())

# Init TTS
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to(device)

wav = tts.tts(text="Hello world!", speaker_wav="my/cloning/audio.wav", language="en")
tts.tts_to_file(text="Hello world!", speaker_wav="my/cloning/audio.wav", language="en", file_path="output.wav")

What it does: Loads the multilingual ⓍTTS v2 model and synthesizes speech either as amplitude values or to a file, cloning the voice from a reference wav. Target speaker_wav and language must be set.

Voice conversion

python
python
tts = TTS(model_name="voice_conversion_models/multilingual/vctk/freevc24", progress_bar=False).to("cuda")
tts.voice_conversion_to_file(source_wav="my/source.wav", target_wav="my/target.wav", file_path="output.wav")

What it does: Converts the voice in source_wav to the voice of target_wav using the FreeVC model.

Voice cloning with any model via voice conversion

python
python
tts = TTS("tts_models/de/thorsten/tacotron2-DDC")
tts.tts_with_vc_to_file(
    "Wie sage ich auf Italienisch, dass ich dich liebe?",
    speaker_wav="target/speaker.wav",
    file_path="output.wav"
)

What it does: Combines a TTS model with voice conversion so any model in 🐸TTS can be used to clone a voice.

Command-line synthesis with a chosen model

bash
bash
$ tts --text "Text for TTS" --model_name "tts_models/en/ljspeech/glow-tts" --out_path output/path/speech.wav

What it does: Runs the glow-tts English LJSpeech model with its default vocoder and writes the audio to the given path.

List models from the CLI

bash
bash
$ tts --list_models

What it does: Lists the provided models so you can pick a name to use with --model_name.

Pros & cons

Pros

  • Pro:Wide model coverage: Tacotron, Glow-TTS, FastPitch, VITS, ⓍTTS, Tortoise, Bark, plus many vocoders
  • Pro:Supports multi-speaker, multilingual, voice cloning and voice conversion, with ~1100 Fairseq language models usable
  • Pro:Offers Python API, CLI, docker image and server, plus training and fine-tuning tools and dataset analysis utilities

Cons

  • Con:Officially tested only on Ubuntu 18.04 with Python >= 3.9 and < 3.12
  • Con:The README states make system-deps is intended for Ubuntu/Debian; Windows users are pointed to external instructions
  • Con:Some models in the performance chart (TTS*, Judy*) are internal and not released open-source

Images