What it is
VoxCPM is a tokenizer-free text-to-speech system that generates continuous speech representations through an end-to-end diffusion autoregressive architecture. VoxCPM2 is the latest release: a 2B-parameter model built on a MiniCPM-4 backbone and trained on over 2 million hours of multilingual speech. It supports 30 languages, Voice Design, Controllable Voice Cloning and 48kHz audio output. Weights and code are released under Apache-2.0.
Who it's for
- Developers who need multilingual speech synthesis across 30 languages and several Chinese dialects
- Teams that want to create new voices from text descriptions or clone voices from short reference clips
- Engineers deploying TTS at scale with Nano-vLLM or vLLM-Omni, or on-device via llama.cpp-omni
- Projects needing a commercially usable open-source TTS model
Requirements
Requirements
- Python ≥ 3.10 (<3.13)
- PyTorch ≥ 2.5.0
- CUDA ≥ 12.0
- VRAM of about 8 GB for VoxCPM2 (per the models comparison table)
Setup
Install the package
Install VoxCPM from PyPI.
bashpip install voxcpmOptional: download weights from ModelScope
Install modelscope if you prefer downloading the model from ModelScope first.
bashpip install modelscopeLaunch the web demo
Run the demo app, then open it in a browser at http://localhost:8808.
bashpython app.py --port 8808 # then open in browser: http://localhost:8808
Examples
Basic text-to-speech
pythonfrom voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained(
"openbmb/VoxCPM2",
load_denoiser=False,
)
wav = model.generate(
text="VoxCPM2 is the current recommended release for realistic multilingual speech synthesis.",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)
print("saved: demo.wav")What it does: Loads VoxCPM2 from Hugging Face, synthesizes a sentence and writes it to a WAV file at the model's sample rate.
Voice Design from a description
pythonwav = model.generate(
text="(A young woman, gentle and sweet voice)Hello, welcome to VoxCPM2!",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("voice_design.wav", wav, model.tts_model.sample_rate)What it does: Puts a natural-language voice description in parentheses at the start of the text, so no reference audio is needed.
Controllable voice cloning with style control
pythonwav = model.generate(
text="(slightly faster, cheerful tone)This is a cloned voice with style control.",
reference_wav_path="path/to/voice.wav",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("controllable_clone.wav", wav, model.tts_model.sample_rate)What it does: Clones the timbre from a reference clip while a parenthesized instruction adjusts speed and emotion.
Streaming generation
pythonimport numpy as np
chunks = []
for chunk in model.generate_streaming(
text="Streaming text to speech is easy with VoxCPM!",
):
chunks.append(chunk)
wav = np.concatenate(chunks)
sf.write("streaming.wav", wav, model.tts_model.sample_rate)What it does: Consumes audio chunks from generate_streaming and concatenates them into one waveform.
CLI voice cloning
bashvoxcpm clone \
--text "This is a voice cloning demo." \
--reference-audio path/to/voice.wav \
--output out.wavWhat it does: Clones a voice from a reference audio file using the command-line interface.
Pros & cons
Pros
- Pro:Supports 30 languages with no language tag needed, plus several Chinese dialects
- Pro:Offers Voice Design, controllable cloning and ultimate cloning (reference audio plus transcript) in one model
- Pro:Outputs 48kHz audio directly from 16kHz reference audio, with no external upsampler needed
- Pro:Apache-2.0 license with multiple deployment paths: Python API, CLI, web demo, Nano-vLLM, vLLM-Omni and llama.cpp-omni
Cons
- Con:Requires CUDA ≥ 12.0 and roughly 8 GB VRAM, more than VoxCPM1.5 (~6 GB) and VoxCPM-0.5B (~5 GB)
- Con:Standard PyTorch RTF is ~0.30 on an RTX 4090, higher than VoxCPM1.5 (~0.15) and VoxCPM-0.5B (~0.17), so faster serving needs Nano-vLLM or vLLM-Omni
- Con:The on-device llama.cpp-omni path is slower, at RTF ~1.76 (Q8_0) on Apple M4 Pro / Metal
Images
