Repo Voice & Audio

OpenBMB/VoxCPM

Apache-2.0, tokenizer-free 2B TTS model supporting 30 languages, voice design, controllable and ultimate voice cloning, and 48kHz output.

  • 38.5k GitHub stars
  • Python
  • ⚖️ Apache-2.0
  • 🎯 Intermediate
pip install voxcpm
OpenBMB/VoxCPM preview image

What it is

VoxCPM is a tokenizer-free text-to-speech system that generates continuous speech representations through an end-to-end diffusion autoregressive architecture. VoxCPM2 is the latest release: a 2B-parameter model built on a MiniCPM-4 backbone and trained on over 2 million hours of multilingual speech. It supports 30 languages, Voice Design, Controllable Voice Cloning and 48kHz audio output. Weights and code are released under Apache-2.0.

Who it's for

  • Developers who need multilingual speech synthesis across 30 languages and several Chinese dialects
  • Teams that want to create new voices from text descriptions or clone voices from short reference clips
  • Engineers deploying TTS at scale with Nano-vLLM or vLLM-Omni, or on-device via llama.cpp-omni
  • Projects needing a commercially usable open-source TTS model

Requirements

Requirements

  • Python ≥ 3.10 (<3.13)
  • PyTorch ≥ 2.5.0
  • CUDA ≥ 12.0
  • VRAM of about 8 GB for VoxCPM2 (per the models comparison table)

Setup

  1. Install the package

    Install VoxCPM from PyPI.

    bash
    pip install voxcpm
  2. Optional: download weights from ModelScope

    Install modelscope if you prefer downloading the model from ModelScope first.

    bash
    pip install modelscope
  3. Launch the web demo

    Run the demo app, then open it in a browser at http://localhost:8808.

    bash
    python app.py --port 8808  # then open in browser: http://localhost:8808

Examples

Basic text-to-speech

python
python
from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained(
  "openbmb/VoxCPM2",
  load_denoiser=False,
)

wav = model.generate(
    text="VoxCPM2 is the current recommended release for realistic multilingual speech synthesis.",
    cfg_value=2.0,
    inference_timesteps=10,
    seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)
print("saved: demo.wav")

What it does: Loads VoxCPM2 from Hugging Face, synthesizes a sentence and writes it to a WAV file at the model's sample rate.

Voice Design from a description

python
python
wav = model.generate(
    text="(A young woman, gentle and sweet voice)Hello, welcome to VoxCPM2!",
    cfg_value=2.0,
    inference_timesteps=10,
    seed=42,
)
sf.write("voice_design.wav", wav, model.tts_model.sample_rate)

What it does: Puts a natural-language voice description in parentheses at the start of the text, so no reference audio is needed.

Controllable voice cloning with style control

python
python
wav = model.generate(
    text="(slightly faster, cheerful tone)This is a cloned voice with style control.",
    reference_wav_path="path/to/voice.wav",
    cfg_value=2.0,
    inference_timesteps=10,
    seed=42,
)
sf.write("controllable_clone.wav", wav, model.tts_model.sample_rate)

What it does: Clones the timbre from a reference clip while a parenthesized instruction adjusts speed and emotion.

Streaming generation

python
python
import numpy as np

chunks = []
for chunk in model.generate_streaming(
    text="Streaming text to speech is easy with VoxCPM!",
):
    chunks.append(chunk)
wav = np.concatenate(chunks)
sf.write("streaming.wav", wav, model.tts_model.sample_rate)

What it does: Consumes audio chunks from generate_streaming and concatenates them into one waveform.

CLI voice cloning

bash
bash
voxcpm clone \
  --text "This is a voice cloning demo." \
  --reference-audio path/to/voice.wav \
  --output out.wav

What it does: Clones a voice from a reference audio file using the command-line interface.

Pros & cons

Pros

  • Pro:Supports 30 languages with no language tag needed, plus several Chinese dialects
  • Pro:Offers Voice Design, controllable cloning and ultimate cloning (reference audio plus transcript) in one model
  • Pro:Outputs 48kHz audio directly from 16kHz reference audio, with no external upsampler needed
  • Pro:Apache-2.0 license with multiple deployment paths: Python API, CLI, web demo, Nano-vLLM, vLLM-Omni and llama.cpp-omni

Cons

  • Con:Requires CUDA ≥ 12.0 and roughly 8 GB VRAM, more than VoxCPM1.5 (~6 GB) and VoxCPM-0.5B (~5 GB)
  • Con:Standard PyTorch RTF is ~0.30 on an RTX 4090, higher than VoxCPM1.5 (~0.15) and VoxCPM-0.5B (~0.17), so faster serving needs Nano-vLLM or vLLM-Omni
  • Con:The on-device llama.cpp-omni path is slower, at RTF ~1.76 (Q8_0) on Apple M4 Pro / Metal

Images