LiveTalking
a streaming digital human that talks back live
LiveTalking by lipku is a real-time streaming digital human: synced audio-video, interruptible speech, multi-session, low-latency WebRTC, Apache-2.0.
Published entries filed under “Audio” in GitHub, newest first by publication date on this site.
12 entries
a streaming digital human that talks back live
LiveTalking by lipku is a real-time streaming digital human: synced audio-video, interruptible speech, multi-session, low-latency WebRTC, Apache-2.0.
talking-head video from one photo
Meituan's LongCat team open-sourced a model that turns one photo and one audio track into a talking video — Whisper lip sync, 8-step distillation, MIT.
open-source voice cloning from a short sample
VoxCPM2 by OpenBMB is a 2B-parameter open TTS: 30 languages, 48kHz output, Apache-2.0 licensed, so voiceovers and podcasts can run on your own machine.
turn your AI coding assistant into a video studio
OpenMontage turns your AI coding assistant into a video studio: 12 pipelines, a storyboard approval gate with per-asset cost, free without any API key.
turn long videos into post-ready clips, locally
AutoClip is an open-source video clipper for long-form material: feed it an interview, podcast, lecture or gaming VOD and it mines the subtitles for an outline, scores segments for highlight-worthiness and cuts clips you approve before anything renders. Unlike cloud clipping services, cutting and rendering stay on your machine with your own model keys — fully offline via Ollama. It ships as a desktop app, a Docker stack and a CLI/MCP pipeline, with export presets for Douyin, Shorts and Bilibili.
turn any topic into a narrated explainer video
html-explainer is an open-source Agent Skill for Claude Code, Codex and other coding agents: give it a topic and it runs research, narration, voiceover, subtitles, frame-by-frame rendering and covers end to end, producing a hard-subtitled MP4. Scenes are written in HTML/CSS/GSAP and rendered with a deterministic seek renderer, so audio and animation stay locked in sync. The default edge-tts voiceover is free and everything runs locally.
an open-source voice studio that never leaves your machine
The usual way to narrate a video is pasting a script into a cloud TTS: per-character pricing, generic voices, and your unpublished text on someone else's server. yovoice runs 7 local TTS models through audio.cpp on your own hardware, clones a voice from a 1-60 second clip, and ships a full record-trim-export workflow plus a CLI and an Agent Skill. Apache-2.0, for macOS and Windows.
one-click AI-generated story videos
The open-source tool Story-Flicks takes nothing but a story theme and automatically produces a complete short video with AI illustrations, narration, and subtitles; the text model can connect to DeepSeek, Alibaba Cloud, OpenAI, and other providers, with one-command Docker deployment and multi-language output.
Turn Audio and Video into Structured Notes, Locally
An open-source audio/video note-taking tool that runs speech recognition locally with FunASR and organizes the transcript into structured Markdown using a local Ollama model, keeping data on your machine — well suited to meetings, interviews, and lecture recordings.
an open-source cross-platform player with pluggable sources
An open-source music client built with Flutter that covers desktop and mobile; music sources, playlists, and metadata are all plugged in, playback runs locally with no telemetry, and tracks can be downloaded with tags for offline listening.
A Go-based Cross-platform Command-line Video Downloader
Lux is a command-line video downloader written in Go that resolves real stream URLs across 45 sites including Bilibili, Douyin, and YouTube, with quality selection, subtitle downloads, batch playlist fetching, and resumable downloads.
A Local Ebook-to-Audiobook Converter
An open-source tool that converts ebooks into audiobooks entirely on your own machine, with multiple built-in TTS engines, voice cloning, and support for 1,158 languages, producing an m4b with chapters and metadata — no cloud service needed, and it runs on as little as 2GB of RAM.