What it is
CosyVoice is a text-to-speech system based on large language models. Its latest release, Fun-CosyVoice 3.0, is designed for zero-shot multilingual speech synthesis in the wild. The repo provides pretrained models, inference scripts, training scripts under examples/libritts, a web demo, and deployment options (gRPC, FastAPI, Docker, TensorRT-LLM).
Who it's for
- Developers building multilingual or cross-lingual TTS and zero-shot voice cloning
- Teams needing low-latency streaming speech synthesis
- Engineers deploying TTS as a gRPC/FastAPI service or with TensorRT-LLM
- Researchers training or evaluating TTS models
Requirements
Requirements
- Conda with Python 3.10
- Dependencies from requirements.txt
- Git submodules (clone with --recursive)
- Pretrained model downloads via ModelScope or Hugging Face
- sox (and libsox-dev / sox-devel) if you hit sox compatibility issues
- vLLM 0.9.0 or 0.11.x+ for vLLM usage (older versions unsupported, e.g. 0.10.x untested)
- NVIDIA runtime for the documented Docker deployment commands
Setup
Clone the repo
Clone with submodules. If the submodule clone fails due to network problems, rerun the update command until it succeeds.
bashgit clone --recursive https://github.com/FunAudioLLM/CosyVoice.git # If you failed to clone the submodule due to network failures, please run the following command until success cd CosyVoice git submodule update --init --recursiveCreate the Conda environment
Create and activate a Python 3.10 env, then install requirements. Install sox packages if you hit compatibility issues.
bashconda create -n cosyvoice -y python=3.10 conda activate cosyvoice pip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com # If you encounter sox compatibility issues # ubuntu sudo apt-get install sox libsox-dev # centos sudo yum install sox sox-develDownload pretrained models
The README strongly recommends downloading the pretrained models. A ModelScope example is shown for Fun-CosyVoice3-0.5B; the README also lists other models and a Hugging Face alternative.
pythonfrom modelscope import snapshot_download snapshot_download('FunAudioLLM/Fun-CosyVoice3-0.5B-2512', local_dir='pretrained_models/Fun-CosyVoice3-0.5B')Optional: install ttsfrd
Optional step for better text normalization. If ttsfrd is not installed, wetext is used by default.
bashcd pretrained_models/CosyVoice-ttsfrd/ unzip resource.zip -d . pip install ttsfrd_dependency-0.1-py3-none-any.whl pip install ttsfrd-0.4.2-cp310-cp310-linux_x86_64.whl
Examples
Run basic usage
bashpython example.pyWhat it does: The README recommends Fun-CosyVoice3-0.5B and points to example.py for detailed usage of each model.
Run with vLLM
bashconda create -n cosyvoice_vllm --clone cosyvoice
conda activate cosyvoice_vllm
# for vllm>=0.11.0
pip install vllm==v0.11.0 transformers==4.57.1 numpy==1.26.4 -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
python vllm_example.pyWhat it does: Clones the env and installs vLLM 0.11.0 with the pinned transformers and numpy versions before running vllm_example.py.
Start the web demo
bashpython3 webui.py --port 50000 --model_dir pretrained_models/CosyVoice-300MWhat it does: Launches the web UI; the model_dir can be changed to the SFT or Instruct model for those modes.
Deploy via Docker with FastAPI
bashcd runtime/python
docker build -t cosyvoice:v1.0 .
docker run -d --runtime=nvidia -p 50000:50000 cosyvoice:v1.0 /bin/bash -c "cd /opt/CosyVoice/CosyVoice/runtime/python/fastapi && python3 server.py --port 50000 --model_dir iic/CosyVoice-300M && sleep infinity"
cd fastapi && python3 client.py --port 50000 --mode <sft|zero_shot|cross_lingual|instruct>What it does: Builds the image, starts the FastAPI server, and runs the client in one of the listed modes. A gRPC variant is also documented.
Pros & cons
Pros
- Pro:Covers 9 languages and 18+ Chinese dialects/accents, with multilingual and cross-lingual zero-shot voice cloning
- Pro:Bi-streaming (text-in and audio-out) with latency as low as 150ms
- Pro:Controllability via pronunciation inpainting (Pinyin, CMU phonemes) and instructions for emotion, speed, volume, dialect
- Pro:Full-stack: training scripts, inference, web demo, gRPC/FastAPI Docker deployment, vLLM and TensorRT-LLM support
Cons
- Con:Setup is involved: submodules, Conda env, multiple model downloads, and vLLM has specific version requirements
- Con:vLLM versions below 0.9.0 are unsupported and in-between versions such as 0.10.x are untested
- Con:Documented deployment commands assume an NVIDIA Docker runtime
Images
