What it is
ChatTTS is a text-to-speech model designed specifically for dialogue scenarios such as LLM assistants. The repo contains the algorithm infrastructure and some simple examples. It supports multiple speakers and can predict and control fine-grained prosodic features including laughter, pauses, and interjections. The open-source model on HuggingFace is a 40,000-hour pre-trained model without SFT.
Who it's for
- Researchers and developers exploring dialogue-oriented text-to-speech for academic or educational purposes
- Developers building LLM assistant voice output who need English or Chinese speech synthesis
- Users who want fine-grained prosody control (laughter, pauses, interjections) in synthesized speech
Requirements
Requirements
- Python (the conda example uses python=3.11)
- Dependencies from requirements.txt
- At least 4GB of GPU memory for a 30-second audio clip
- Optional vLLM install is Linux only
- Model is for academic/research use only (CC BY-NC 4.0); code is AGPLv3+
Setup
Clone the repo
Clone the repository and enter the project directory.
bashgit clone https://github.com/2noise/ChatTTS cd ChatTTSInstall requirements directly
Install dependencies with pip.
bashpip install --upgrade -r requirements.txtOr install from conda
Create a conda environment and install the requirements.
bashconda create -n chattts python=3.11 conda activate chattts pip install -r requirements.txtInstall from PyPI
Install the stable version of the package.
bashpip install ChatTTS
Examples
Launch the WebUI
bashpython examples/web/webui.pyWhat it does: Quick start option; run from the project root directory.
Infer by command line
bashpython examples/cmd/run.py "Your text 1." "Your text 2."What it does: Saves audio to ./output_audio_n.mp3.
Basic Python usage
pythonimport ChatTTS
import torch
import torchaudio
chat = ChatTTS.Chat()
chat.load(compile=False) # Set to True for better performance
texts = ["PUT YOUR 1st TEXT HERE", "PUT YOUR 2nd TEXT HERE"]
wavs = chat.infer(texts)
for i in range(len(wavs)):
try:
torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]).unsqueeze(0), 24000)
except:
torchaudio.save(f"basic_output{i}.wav", torch.from_numpy(wavs[i]), 24000)What it does: Loads the model, synthesizes a list of texts, and saves each result as a 24000 Hz WAV file. The try/except handles torchaudio version differences.
Sample a speaker and set sentence-level control
pythonrand_spk = chat.sample_random_speaker()
print(rand_spk) # save it for later timbre recovery
params_infer_code = ChatTTS.Chat.InferCodeParams(
spk_emb = rand_spk, # add sampled speaker
temperature = .3, # using custom temperature
top_P = 0.7, # top P decode
top_K = 20, # top K decode
)
params_refine_text = ChatTTS.Chat.RefineTextParams(
prompt='[oral_2][laugh_0][break_6]',
)
wavs = chat.infer(
texts,
params_refine_text=params_refine_text,
params_infer_code=params_infer_code,
)What it does: Samples a random speaker, sets decoding parameters, and uses oral_(0-9), laugh_(0-2) and break_(0-7) special tokens for sentence-level control.
Word-level control
pythontext = 'What is [uv_break]your favorite english food?[laugh][lbreak]'
wavs = chat.infer(text, skip_refine_text=True, params_refine_text=params_refine_text, params_infer_code=params_infer_code)What it does: Inserts tokens such as [uv_break], [laugh] and [lbreak] directly into the text and skips the text refinement step.
Pros & cons
Pros
- Pro:Optimized for dialogue, with multi-speaker support
- Pro:Fine-grained control over laughter, pauses and interjections at sentence and word level
- Pro:Multiple ways to run it: WebUI, command line, Python API, or PyPI install
- Pro:Streaming audio generation and zero-shot inferring code are listed as completed roadmap items
Cons
- Con:Released model is for academic/research use only (CC BY-NC 4.0), not commercial use
- Con:Model stability is limited: multiple speakers or poor audio quality can occur, and multiple samples may be needed
- Con:Only English and Chinese are supported, and token-level control is limited to [laugh], [uv_break], and [lbreak]; multi-emotion control is not yet done
Images
