Repo Voice & Audio

2noise/ChatTTS

ChatTTS is a text-to-speech model for dialogue scenarios such as LLM assistants, with English and Chinese support and control over laughter, pauses and interjections.

  • 39.9k GitHub stars
  • Python
  • ⚖️ AGPL-3.0
  • 🎯 Intermediate
pip install --upgrade -r requirements.txt
2noise/ChatTTS preview image

What it is

ChatTTS is a text-to-speech model designed specifically for dialogue scenarios such as LLM assistants. The repo contains the algorithm infrastructure and some simple examples. It supports multiple speakers and can predict and control fine-grained prosodic features including laughter, pauses, and interjections. The open-source model on HuggingFace is a 40,000-hour pre-trained model without SFT.

Who it's for

  • Researchers and developers exploring dialogue-oriented text-to-speech for academic or educational purposes
  • Developers building LLM assistant voice output who need English or Chinese speech synthesis
  • Users who want fine-grained prosody control (laughter, pauses, interjections) in synthesized speech

Requirements

Requirements

  • Python (the conda example uses python=3.11)
  • Dependencies from requirements.txt
  • At least 4GB of GPU memory for a 30-second audio clip
  • Optional vLLM install is Linux only
  • Model is for academic/research use only (CC BY-NC 4.0); code is AGPLv3+

Setup

  1. Clone the repo

    Clone the repository and enter the project directory.

    bash
    git clone https://github.com/2noise/ChatTTS
    cd ChatTTS
  2. Install requirements directly

    Install dependencies with pip.

    bash
    pip install --upgrade -r requirements.txt
  3. Or install from conda

    Create a conda environment and install the requirements.

    bash
    conda create -n chattts python=3.11
    conda activate chattts
    pip install -r requirements.txt
  4. Install from PyPI

    Install the stable version of the package.

    bash
    pip install ChatTTS

Examples

Launch the WebUI

bash
bash
python examples/web/webui.py

What it does: Quick start option; run from the project root directory.

Infer by command line

bash
bash
python examples/cmd/run.py "Your text 1." "Your text 2."

What it does: Saves audio to ./output_audio_n.mp3.

Basic Python usage

python
python
import ChatTTS
import torch
import torchaudio

chat = ChatTTS.Chat()
chat.load(compile=False) # Set to True for better performance

texts = ["PUT YOUR 1st TEXT HERE", "PUT YOUR 2nd TEXT HERE"]

wavs = chat.infer(texts)

for i in range(len(wavs)):
    try:
        torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]).unsqueeze(0), 24000)
    except:
        torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]), 24000)

What it does: Loads the model, synthesizes a list of texts, and saves each result as a 24000 Hz WAV file. The try/except handles torchaudio version differences.

Sample a speaker and set sentence-level control

python
python
rand_spk = chat.sample_random_speaker()
print(rand_spk) # save it for later timbre recovery

params_infer_code = ChatTTS.Chat.InferCodeParams(
    spk_emb = rand_spk, # add sampled speaker 
    temperature = .3,   # using custom temperature
    top_P = 0.7,        # top P decode
    top_K = 20,         # top K decode
)

params_refine_text = ChatTTS.Chat.RefineTextParams(
    prompt='[oral_2][laugh_0][break_6]',
)

wavs = chat.infer(
    texts,
    params_refine_text=params_refine_text,
    params_infer_code=params_infer_code,
)

What it does: Samples a random speaker, sets decoding parameters, and uses oral_(0-9), laugh_(0-2) and break_(0-7) special tokens for sentence-level control.

Word-level control

python
python
text = 'What is [uv_break]your favorite english food?[laugh][lbreak]'
wavs = chat.infer(text, skip_refine_text=True, params_refine_text=params_refine_text,  params_infer_code=params_infer_code)

What it does: Inserts tokens such as [uv_break], [laugh] and [lbreak] directly into the text and skips the text refinement step.

Pros & cons

Pros

  • Pro:Optimized for dialogue, with multi-speaker support
  • Pro:Fine-grained control over laughter, pauses and interjections at sentence and word level
  • Pro:Multiple ways to run it: WebUI, command line, Python API, or PyPI install
  • Pro:Streaming audio generation and zero-shot inferring code are listed as completed roadmap items

Cons

  • Con:Released model is for academic/research use only (CC BY-NC 4.0), not commercial use
  • Con:Model stability is limited: multiple speakers or poor audio quality can occur, and multiple samples may be needed
  • Con:Only English and Chinese are supported, and token-level control is limited to [laugh], [uv_break], and [lbreak]; multi-emotion control is not yet done

Images