LiveTalking: a streaming digital human that talks back live
LiveTalking by lipku is a real-time streaming digital human: synced audio-video, interruptible speech, multi-session, low-latency WebRTC, Apache-2.0.
Published entries across all sections carrying the “Voice cloning” tag, newest first by publication date on this site.
6 entries
LiveTalking by lipku is a real-time streaming digital human: synced audio-video, interruptible speech, multi-session, low-latency WebRTC, Apache-2.0.
VoxCPM2 by OpenBMB is a 2B-parameter open TTS: 30 languages, 48kHz output, Apache-2.0 licensed, so voiceovers and podcasts can run on your own machine.
ElevenLabs announced a $300M employee tender at a $22B valuation, double its February price; Wellington and T. Rowe Price co-led.
ElevenLabs released v4 and v4 Turbo on September 28: 90+ languages, 10-second voice cloning and ~100ms Turbo latency for voice agents, available now via API.
The usual way to narrate a video is pasting a script into a cloud TTS: per-character pricing, generic voices, and your unpublished text on someone else's server. yovoice runs 7 local TTS models through audio.cpp on your own hardware, clones a voice from a 1-60 second clip, and ships a full record-trim-export workflow plus a CLI and an Agent Skill. Apache-2.0, for macOS and Windows.
An open-source tool that converts ebooks into audiobooks entirely on your own machine, with multiple built-in TTS engines, voice cloning, and support for 1,158 languages, producing an m4b with chapters and metadata — no cloud service needed, and it runs on as little as 2GB of RAM.