Voice & AudioIntermediate

RVC-Boss/GPT-SoVITS

62.6kPython+63/daypushed 1d ago

Use it for: Clone a voice and train a text-to-speech model from about one minute of voice data using a web UI.

Quick start
  1. Clone the repo and follow the README install section (Python 3.10-3.12)
  2. Alternatively, try the Colab training notebook or Docker image
  3. Launch the WebUI
  4. Provide a short voice sample and train or run text-to-speech
Voice & AudioBeginner

debpalash/VoiceStudio

56.2kPython+307/daypushed 0d ago

Use it for: Clone or design voices, dub videos, dictate, transcribe and create audiobooks locally in 646 languages, as a free ElevenLabs alternative.

Quick start
  1. On macOS or Linux run: curl -fsSL https://voicestudio.sh/install | sh
  2. On Windows run in PowerShell: irm https://voicestudio.sh/install | iex
  3. Launch the desktop app
  4. Choose a workspace such as voice cloning, dubbing or voice design
Voice & AudioIntermediate

coqui-ai/TTS

46.1kPython+20/daypushed 784d ago

Use it for: Generate speech from text with pretrained models in many languages, or train and fine-tune your own text-to-speech models.

Quick start
  1. Install the TTS package from PyPI (pip install TTS).
  2. Pick a pretrained model, such as ⓍTTS, from the README or docs.
  3. Generate speech from text with the library or its command-line tools.
  4. To customize a voice, follow the example fine-tuning recipes in the repo.
Voice & AudioIntermediate

2noise/ChatTTS

39.9kPython+46/daypushed 182d ago

Use it for: Turn text into natural-sounding conversational speech in English or Chinese, suited to LLM assistants and dialogue.

Quick start
  1. Clone the repo or install the ChatTTS package from PyPI.
  2. Download the model weights from Hugging Face.
  3. Run the example scripts or the Colab notebook to synthesize dialogue speech.
  4. Try community end-user projects listed in Awesome-ChatTTS for ready-made apps.
Voice & AudioIntermediate

OpenBMB/VoxCPM

38.5kPython+99/daypushed 2d ago

Use it for: Generate multilingual speech in 30 languages, design new voices from text descriptions, and clone voices from short reference clips.

Quick start
  1. Clone the repo and follow the README install section.
  2. Download the VoxCPM2 model weights as described in the README.
  3. Enter text, optionally with a voice description or reference audio, to generate 48kHz speech.
  4. Provide reference audio plus its transcript if you want the most faithful cloning.
Voice & AudioIntermediate

myshell-ai/OpenVoice

37.8kPython+36/daypushed 538d ago

Use it for: Clone a voice's tone color from a short reference clip and generate speech in multiple languages with controllable style.

Quick start
  1. Clone the repo and follow the README install section.
  2. Download the OpenVoice V2 checkpoints as directed in the README.
  3. Supply a reference voice clip and the text you want spoken.
  4. Adjust style settings such as emotion, accent, rhythm, and pauses.
Voice & AudioIntermediate

babysor/MockingBird

36.9kPython+20/daypushed 220d ago

Use it for: Clone a voice from a few seconds of audio and generate arbitrary speech, with strong Mandarin support.

Quick start
  1. Install Python 3.7+, PyTorch, and ffmpeg.
  2. Run pip install -r requirements.txt to install the remaining dependencies.
  3. Prepare the pretrained models as described in the README.
  4. Launch the toolbox or web server to clone a voice and synthesize speech.
Voice & AudioIntermediate

index-tts/index-tts

24.4kPython+40/daypushed 10d ago

Use it for: Clone a voice from a single audio clip and synthesize speech in Chinese, English, Japanese, Spanish or Arabic with emotion and speed control.

Quick start
  1. Clone the repo and follow the README install section.
  2. Download the IndexTTS-2.5 model from HuggingFace or ModelScope.
  3. Provide one reference audio clip and your text.
  4. Tune emotion, speaking speed, or pronunciation controls as needed.
Voice & AudioAdvanced

QwenAudio/CosyVoice

23.9kPython+29/daypushed 137d ago

Use it for: Synthesize multilingual speech with zero-shot voice cloning, covering inference, training, and deployment, including many Chinese dialects.

Quick start
  1. Clone the repo and follow the README install section.
  2. Download a pretrained model such as Fun-CosyVoice 3.0 from ModelScope or HuggingFace.
  3. Run inference with a text prompt and optional reference voice.
  4. Use the training and deployment docs to fine-tune or serve the model.
Voice & AudioIntermediate

nari-labs/dia

19.4kPython+36/daypushed 324d ago

Use it for: Generate realistic English dialogue from a transcript in one pass, including nonverbal sounds like laughter and coughing.

Quick start
  1. Clone the repo and follow the README install section, or use Hugging Face Transformers.
  2. Download the Dia-1.6B weights from Hugging Face.
  3. Write a transcript of moderate length, using non-verbal tags sparingly.
  4. Optionally condition on reference audio for emotion and tone, then generate the audio.