What it is
GPT-SoVITS-WebUI is a few-shot voice conversion and text-to-speech WebUI. It supports zero-shot TTS from a 5-second vocal sample and few-shot fine-tuning with about 1 minute of training data. It includes integrated tools for vocal/accompaniment separation, automatic audio slicing, ASR and text labeling to help build training datasets and GPT/SoVITS models.
Who it's for
- Beginners who want an integrated WebUI workflow for building voice datasets and training GPT/SoVITS models
- Developers wanting to clone a voice from very little data (5s zero-shot or 1 minute fine-tuning)
- Users needing cross-lingual TTS in English, Japanese, Korean, Cantonese and Chinese
Requirements
Requirements
- Python 3.10-3.12 (README badge); tested environments include Python 3.9-3.12 with PyTorch 2.2.2-2.8.0dev
- Device: CUDA 11.8/12.4/12.8, Apple silicon, or CPU per the tested environments table
- FFmpeg (plus libsox-dev on Ubuntu/Debian; ffmpeg.exe and ffprobe.exe and Visual Studio 2017 on Windows)
- Pretrained models placed in GPT_SoVITS/pretrained_models (skippable if install.sh runs successfully)
- Training data in .list format: vocal_path|speaker_name|language|text
Setup
Linux install
Create a conda environment and run the install script, choosing device and model source.
bashconda create -n GPTSoVits python=3.10 conda activate GPTSoVits bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope> [--download-uvr5]Windows install
Alternatively, Windows users can download the integrated package and double-click go-webui.bat. Otherwise run the install script in PowerShell.
pwshconda create -n GPTSoVits python=3.10 conda activate GPTSoVits pwsh -F install.ps1 --Device <CU126|CU128|CPU> --Source <HF|HF-Mirror|ModelScope> [--DownloadUVR5]macOS install
The README notes that models trained with GPUs on Macs have significantly lower quality, so CPUs are temporarily used instead.
bashconda create -n GPTSoVits python=3.10 conda activate GPTSoVits bash install.sh --device <MPS|CPU> --source <HF|HF-Mirror|ModelScope> [--download-uvr5]Manual dependency install
Install Python dependencies manually, then install FFmpeg separately.
bashconda create -n GPTSoVits python=3.10 conda activate GPTSoVits pip install -r extra-req.txt --no-deps pip install -r requirements.txtDocker
Run a chosen service from docker-compose.yaml. Docker Compose mounts all files in the current directory, so switch to the project root and pull the latest code first.
bashdocker compose run --service-ports <GPT-SoVITS-CU126-Lite|GPT-SoVITS-CU128-Lite|GPT-SoVITS-CU126|GPT-SoVITS-CU128>
Examples
Dataset annotation format
textvocal_path|speaker_name|language|text
D:\GPT-SoVITS\xxx/xxx.wav|xxx|en|I like playing Genshin.What it does: The .list format for TTS annotation. Language codes are zh, ja, en, ko and yue.
Open the WebUI
bashpython webui.py <language(optional)>What it does: Launches the main WebUI for dataset preparation and fine-tuning. Use python webui.py v1 <language(optional)> to switch to V1.
Open the inference WebUI
bashpython GPT_SoVITS/inference_webui.py <language(optional)>What it does: Starts the inference WebUI directly. Alternatively run python webui.py and open 1-GPT-SoVITS-TTS/1C-inference.
Command-line ASR (non-Chinese)
bashpython ./tools/asr/fasterwhisper_asr.py -i <input> -o <output> -l <language> -p <precision>What it does: Runs Faster Whisper ASR on a dataset for languages other than Chinese, without progress bars.
Command-line audio slicing
bashpython audio_slicer.py \
--input_path "<path_to_original_audio_file_or_directory>" \
--output_root "<directory_where_subdivided_audio_clips_will_be_saved>" \
--threshold <volume_threshold> \
--min_length <minimum_duration_of_each_subclip> \
--min_interval <shortest_time_gap_between_adjacent_subclips>
--hop_size <step_size_for_computing_volume_curve>What it does: Splits original audio into smaller clips for building a training dataset.
Pros & cons
Pros
- Pro:Zero-shot from a 5-second sample and few-shot fine-tuning with 1 minute of data
- Pro:Integrated WebUI tools for vocal separation, slicing, ASR and labeling
- Pro:Cross-lingual inference across English, Japanese, Korean, Cantonese and Chinese
- Pro:Multiple install paths: integrated Windows package, install scripts, manual, Docker, and Colab
Cons
- Con:Models trained with GPUs on Macs have significantly lower quality, so CPU is used instead
- Con:Multiple model versions (v1 to v5) each require separate pretrained model downloads and upgrade steps
- Con:Docker images release more slowly than the codebase, so users must check Docker Hub for current tags