Repo Voice & Audio

RVC-Boss/GPT-SoVITS

Few-shot voice cloning and TTS WebUI: zero-shot from a 5s sample, fine-tune with 1 minute of data, with built-in dataset prep tools and multilingual inference.

  • 62.6k GitHub stars
  • Python
  • ⚖️ MIT
  • 🎯 Intermediate
docker compose run --service-ports <GPT-SoVITS-CU126-Lite|GPT-SoVITS-CU128-Lite|GPT-SoVITS-CU126|GPT-SoVITS-CU128>
RVC-Boss/GPT-SoVITS — repo preview

What it is

GPT-SoVITS-WebUI is a few-shot voice conversion and text-to-speech WebUI. It supports zero-shot TTS from a 5-second vocal sample and few-shot fine-tuning with about 1 minute of training data. It includes integrated tools for vocal/accompaniment separation, automatic audio slicing, ASR and text labeling to help build training datasets and GPT/SoVITS models.

Who it's for

  • Beginners who want an integrated WebUI workflow for building voice datasets and training GPT/SoVITS models
  • Developers wanting to clone a voice from very little data (5s zero-shot or 1 minute fine-tuning)
  • Users needing cross-lingual TTS in English, Japanese, Korean, Cantonese and Chinese

Requirements

Requirements

  • Python 3.10-3.12 (README badge); tested environments include Python 3.9-3.12 with PyTorch 2.2.2-2.8.0dev
  • Device: CUDA 11.8/12.4/12.8, Apple silicon, or CPU per the tested environments table
  • FFmpeg (plus libsox-dev on Ubuntu/Debian; ffmpeg.exe and ffprobe.exe and Visual Studio 2017 on Windows)
  • Pretrained models placed in GPT_SoVITS/pretrained_models (skippable if install.sh runs successfully)
  • Training data in .list format: vocal_path|speaker_name|language|text

Setup

  1. Linux install

    Create a conda environment and run the install script, choosing device and model source.

    bash
    conda create -n GPTSoVits python=3.10
    conda activate GPTSoVits
    bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope> [--download-uvr5]
  2. Windows install

    Alternatively, Windows users can download the integrated package and double-click go-webui.bat. Otherwise run the install script in PowerShell.

    pwsh
    conda create -n GPTSoVits python=3.10
    conda activate GPTSoVits
    pwsh -F install.ps1 --Device <CU126|CU128|CPU> --Source <HF|HF-Mirror|ModelScope> [--DownloadUVR5]
  3. macOS install

    The README notes that models trained with GPUs on Macs have significantly lower quality, so CPUs are temporarily used instead.

    bash
    conda create -n GPTSoVits python=3.10
    conda activate GPTSoVits
    bash install.sh --device <MPS|CPU> --source <HF|HF-Mirror|ModelScope> [--download-uvr5]
  4. Manual dependency install

    Install Python dependencies manually, then install FFmpeg separately.

    bash
    conda create -n GPTSoVits python=3.10
    conda activate GPTSoVits
    
    pip install -r extra-req.txt --no-deps
    pip install -r requirements.txt
  5. Docker

    Run a chosen service from docker-compose.yaml. Docker Compose mounts all files in the current directory, so switch to the project root and pull the latest code first.

    bash
    docker compose run --service-ports <GPT-SoVITS-CU126-Lite|GPT-SoVITS-CU128-Lite|GPT-SoVITS-CU126|GPT-SoVITS-CU128>

Examples

Dataset annotation format

text
text
vocal_path|speaker_name|language|text

D:\GPT-SoVITS\xxx/xxx.wav|xxx|en|I like playing Genshin.

What it does: The .list format for TTS annotation. Language codes are zh, ja, en, ko and yue.

Open the WebUI

bash
bash
python webui.py <language(optional)>

What it does: Launches the main WebUI for dataset preparation and fine-tuning. Use python webui.py v1 <language(optional)> to switch to V1.

Open the inference WebUI

bash
bash
python GPT_SoVITS/inference_webui.py <language(optional)>

What it does: Starts the inference WebUI directly. Alternatively run python webui.py and open 1-GPT-SoVITS-TTS/1C-inference.

Command-line ASR (non-Chinese)

bash
bash
python ./tools/asr/fasterwhisper_asr.py -i <input> -o <output> -l <language> -p <precision>

What it does: Runs Faster Whisper ASR on a dataset for languages other than Chinese, without progress bars.

Command-line audio slicing

bash
bash
python audio_slicer.py \
    --input_path "<path_to_original_audio_file_or_directory>" \
    --output_root "<directory_where_subdivided_audio_clips_will_be_saved>" \
    --threshold <volume_threshold> \
    --min_length <minimum_duration_of_each_subclip> \
    --min_interval <shortest_time_gap_between_adjacent_subclips>
    --hop_size <step_size_for_computing_volume_curve>

What it does: Splits original audio into smaller clips for building a training dataset.

Pros & cons

Pros

  • Pro:Zero-shot from a 5-second sample and few-shot fine-tuning with 1 minute of data
  • Pro:Integrated WebUI tools for vocal separation, slicing, ASR and labeling
  • Pro:Cross-lingual inference across English, Japanese, Korean, Cantonese and Chinese
  • Pro:Multiple install paths: integrated Windows package, install scripts, manual, Docker, and Colab

Cons

  • Con:Models trained with GPUs on Macs have significantly lower quality, so CPU is used instead
  • Con:Multiple model versions (v1 to v5) each require separate pretrained model downloads and upgrade steps
  • Con:Docker images release more slowly than the codebase, so users must check Docker Hub for current tags