Repo Voice & Audio

QwenAudio/CosyVoice

CosyVoice is an LLM-based multilingual zero-shot TTS system with streaming, voice cloning, and training, inference and deployment scripts.

  • 23.9k GitHub stars
  • Python
  • ⚖️ Apache-2.0
  • 🎯 Advanced
QwenAudio/CosyVoice preview image

What it is

CosyVoice is a text-to-speech system based on large language models. Its latest release, Fun-CosyVoice 3.0, is designed for zero-shot multilingual speech synthesis in the wild. The repo provides pretrained models, inference scripts, training scripts under examples/libritts, a web demo, and deployment options (gRPC, FastAPI, Docker, TensorRT-LLM).

Who it's for

  • Developers building multilingual or cross-lingual TTS and zero-shot voice cloning
  • Teams needing low-latency streaming speech synthesis
  • Engineers deploying TTS as a gRPC/FastAPI service or with TensorRT-LLM
  • Researchers training or evaluating TTS models

Requirements

Requirements

  • Conda with Python 3.10
  • Dependencies from requirements.txt
  • Git submodules (clone with --recursive)
  • Pretrained model downloads via ModelScope or Hugging Face
  • sox (and libsox-dev / sox-devel) if you hit sox compatibility issues
  • vLLM 0.9.0 or 0.11.x+ for vLLM usage (older versions unsupported, e.g. 0.10.x untested)
  • NVIDIA runtime for the documented Docker deployment commands

Setup

  1. Clone the repo

    Clone with submodules. If the submodule clone fails due to network problems, rerun the update command until it succeeds.

    bash
    git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git
    # If you failed to clone the submodule due to network failures, please run the following command until success
    cd CosyVoice
    git submodule update --init --recursive
  2. Create the Conda environment

    Create and activate a Python 3.10 env, then install requirements. Install sox packages if you hit compatibility issues.

    bash
    conda create -n cosyvoice -y python=3.10
    conda activate cosyvoice
    pip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
    
    # If you encounter sox compatibility issues
    # ubuntu
    sudo apt-get install sox libsox-dev
    # centos
    sudo yum install sox sox-devel
  3. Download pretrained models

    The README strongly recommends downloading the pretrained models. A ModelScope example is shown for Fun-CosyVoice3-0.5B; the README also lists other models and a Hugging Face alternative.

    python
    from modelscope import snapshot_download
    snapshot_download('FunAudioLLM/Fun-CosyVoice3-0.5B-2512', local_dir='pretrained_models/Fun-CosyVoice3-0.5B')
  4. Optional: install ttsfrd

    Optional step for better text normalization. If ttsfrd is not installed, wetext is used by default.

    bash
    cd pretrained_models/CosyVoice-ttsfrd/
    unzip resource.zip -d .
    pip install ttsfrd_dependency-0.1-py3-none-any.whl
    pip install ttsfrd-0.4.2-cp310-cp310-linux_x86_64.whl

Examples

Run basic usage

bash
bash
python example.py

What it does: The README recommends Fun-CosyVoice3-0.5B and points to example.py for detailed usage of each model.

Run with vLLM

bash
bash
conda create -n cosyvoice_vllm --clone cosyvoice
conda activate cosyvoice_vllm
# for vllm>=0.11.0
pip install vllm==v0.11.0 transformers==4.57.1 numpy==1.26.4 -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
python vllm_example.py

What it does: Clones the env and installs vLLM 0.11.0 with the pinned transformers and numpy versions before running vllm_example.py.

Start the web demo

bash
bash
python3 webui.py --port 50000 --model_dir pretrained_models/CosyVoice-300M

What it does: Launches the web UI; the model_dir can be changed to the SFT or Instruct model for those modes.

Deploy via Docker with FastAPI

bash
bash
cd runtime/python
docker build -t cosyvoice:v1.0 .
docker run -d --runtime=nvidia -p 50000:50000 cosyvoice:v1.0 /bin/bash -c "cd /opt/CosyVoice/CosyVoice/runtime/python/fastapi && python3 server.py --port 50000 --model_dir iic/CosyVoice-300M && sleep infinity"
cd fastapi && python3 client.py --port 50000 --mode <sft|zero_shot|cross_lingual|instruct>

What it does: Builds the image, starts the FastAPI server, and runs the client in one of the listed modes. A gRPC variant is also documented.

Pros & cons

Pros

  • Pro:Covers 9 languages and 18+ Chinese dialects/accents, with multilingual and cross-lingual zero-shot voice cloning
  • Pro:Bi-streaming (text-in and audio-out) with latency as low as 150ms
  • Pro:Controllability via pronunciation inpainting (Pinyin, CMU phonemes) and instructions for emotion, speed, volume, dialect
  • Pro:Full-stack: training scripts, inference, web demo, gRPC/FastAPI Docker deployment, vLLM and TensorRT-LLM support

Cons

  • Con:Setup is involved: submodules, Conda env, multiple model downloads, and vLLM has specific version requirements
  • Con:vLLM versions below 0.9.0 are unsupported and in-between versions such as 0.10.x are untested
  • Con:Documented deployment commands assume an NVIDIA Docker runtime

Images