Audio#Text to speech#Voice cloning

VoxCPM2: open-source voice cloning from a short sample

VoxCPM2 by OpenBMB is a 2B-parameter open TTS: 30 languages, 48kHz output, Apache-2.0 licensed, so voiceovers and podcasts can run on your own machine.

Project facts

GitHub Ecosystem
Repositorygithub.com/OpenBMB/VoxCPM
License
Apache-2.0
Language
Python
Stars
38,243
Data checked
2026-10-01

Snapshot figures reflect the check date and may change over time.

Narrating a video or a podcast usually means paying a voice actor or living with obviously synthetic audio. VoxCPM2 from OpenBMB is a tokenizer-free, open-source TTS system: a 2B-parameter model trained on more than 2 million hours of speech. Feed it a reference clip of about ten seconds, and the cloned voice keeps the hesitations and delivery habits of the original speaker. A September video from the WeChat channel AIFeed claimed “you can’t tell the clone from the real recording” — we can’t verify that ranking claim, but the predecessor VoxCPM1.5 does carry a #1 GitHub Trending badge on the README from December 2025.

Core features

  • Controllable cloning: clone a voice from a reference clip, then steer emotion, pace, and expression with a text prompt while keeping the original timbre.
  • Ultimate cloning: give it both the reference audio and its transcript, and the model continues seamlessly from the reference, preserving rhythm and style — same capability VoxCPM1.5 shipped.
  • Voice design: describe a voice in one natural-language sentence (gender, age, tone, pace) and generate a brand-new one with no reference audio at all.
  • 30 languages: text goes in without language tags, and Chinese alone covers 9 dialects including Cantonese and Sichuanese.
  • 48kHz output: 16kHz reference audio goes through AudioVAE V2 straight to studio-quality 48kHz, with built-in super-resolution instead of an external upsampler.
  • Real-time streaming: per the README, RTF as low as ~0.3 on an RTX 4090, or ~0.13 accelerated by Nano-vLLM or vLLM-Omni, which exposes an OpenAI-compatible /v1/audio/speech API.

Typical use cases

  • Video and podcast narration: batch-generate voiceovers in one cloned voice, re-generating per script edit instead of re-booking studio time.
  • Product narration and serialized audio: design one voice up front so an entire series sounds consistent.
  • Localization: ship the same script in any of 30 languages without hiring per-language voice talent.

Quick start

pip install voxcpm

You need Python 3.10-3.12, PyTorch 2.5.0+, and CUDA 12.0+. First clip:

from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained("openbmb/VoxCPM2", load_denoiser=False)
wav = model.generate(
    text="The first line synthesized with VoxCPM2.",
    cfg_value=2.0,
    inference_timesteps=10,
    seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)

Without an NVIDIA GPU, the OpenBMB/VoxCPM-Demo playground on Hugging Face is enough for auditioning voices.

Summary

VoxCPM2 fits creators and developers who ship spoken audio in volume and want to keep one voice; if you narrate occasionally and don’t care about a signature voice, the hosted playground is simpler. Weights and code are both Apache-2.0, free for commercial use (about 38.2k stars as of 2026-10-01). The catch: local inference needs a CUDA environment, model downloads are large, and low-spec machines should serve it through vLLM instead. For a ready-made desktop app, YoVoice bundles VoxCPM2 with six other local TTS models behind one interface.