What it is
IndexTTS is a zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest release, IndexTTS-2.5, supports Chinese, English, Japanese, Spanish and Arabic. It offers fine-grained emotion control, speaking speed control, pronunciation control (Pinyin / CMU phonemes / Japanese Kana), and faster inference than IndexTTS-2. It ships a WebUI and a Python API, and production deployment is supported via vLLM.
Who it's for
- Developers who need voice cloning from a single reference audio clip
- Teams building multilingual TTS (Chinese, English, Japanese, Spanish, Arabic)
- Users who want emotion, speed and pronunciation control over synthesized speech
- Engineers looking to deploy TTS in production via vLLM
Requirements
Requirements
- git
- uv (required for dependency management)
- NVIDIA CUDA Toolkit 12.8 or newer if a CUDA error appears during installation on Linux/Windows
- Model checkpoints downloaded from HuggingFace or ModelScope
- Network access to HuggingFace/ModelScope (a mirror can be set via HF_ENDPOINT)
Setup
Clone the repository
Make sure git is installed, then download the repository.
bashgit clone https://github.com/index-tts/index-tts.git && cd index-ttsInstall uv
uv is required to manage the project's dependency environment.
bashpip install -U uv # or see the link above for other install methodsInstall dependencies
Creates a .venv project directory and installs the correct Python and all required dependencies.
bashuv sync --all-extrasDownload IndexTTS-2.5 model via HuggingFace
Install the huggingface-hub tool and download the checkpoints.
bashuv tool install "huggingface-hub" # IndexTTS-2.5 hf download IndexTeam/IndexTTS-2.5 --local-dir=checkpointsCheck GPU acceleration
Diagnose your environment and see which GPUs are detected.
bashuv run tools/gpu_check.pyLaunch the WebUI
Then open http://127.0.0.1:7860 in your browser.
bash# IndexTTS-2.5 (default) uv run webui.py
Examples
Initialize IndexTTS-2.5
pythonfrom indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_bf16=True)What it does: Loads the IndexTTS-2.5 model with BF16 inference from the downloaded checkpoints.
Voice cloning from a single reference audio
pythontext = "Translate for me, what is a surprise!"
# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)What it does: Clones the voice in the reference clip and synthesizes the text in English.
Emotion control with a separate emotional reference and emo_alpha
pythontext = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)What it does: Uses a sad emotional reference audio; emo_alpha (0.0–1.0, default 1.0) sets how strongly it affects the output.
Emotion control with an emotion vector
pythontext = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)What it does: The 8-float vector is ordered [happy, angry, sad, afraid, disgusted, melancholic, surprised, calm]; here sad is 0.8.
Speaking speed control
pythontext = "大家好,欢迎来到IndexTTS的语速控制演示。"
# IndexTTS2.5
# Slow down (1.2x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_slow.wav", duration_factor=1.2, verbose=True)
# Speed up (0.8x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_fast.wav", duration_factor=0.8, verbose=True)What it does: duration_factor above 1.0 slows speech and below 1.0 speeds it up; valid range 0.5–2.0, default 1.0.
Pros & cons
Pros
- Pro:Clones a voice from a single reference audio clip
- Pro:Fine-grained control over emotion (audio, vector, or text-based), speaking speed and pronunciation
- Pro:IndexTTS-2.5 supports five languages and is documented as faster than IndexTTS-2 (RTF about 0.2065 vs 0.3257 overall on an RTX 4090)
- Pro:Supports FP16/BF16 inference for lower VRAM use, optional DeepSpeed, and vLLM for production deployment
Cons
- Con:uv is required for installation, and DeepSpeed may be difficult to install on Windows
- Con:Random sampling (use_random) reduces voice cloning fidelity
- Con:For IndexTTS-2.5, use_emo_text=True requires use_qwen_emo=True or it raises a RuntimeError
Images
