Audio#Video generation

LongCat-Video-Avatar 1.5: talking-head video from one photo

Meituan's LongCat team open-sourced a model that turns one photo and one audio track into a talking video — Whisper lip sync, 8-step distillation, MIT.

Project facts

GitHub Ecosystem
Repositorygithub.com/meituan-longcat/LongCat-Video
License
MIT
Language
Python
Stars
8,446
Data checked
2026-10-01

Snapshot figures reflect the check date and may change over time.

The hard part of a talking-head video is never the script; it’s appearing on camera — lighting, shooting, editing, half a day gone. LongCat-Video-Avatar from Meituan’s LongCat team compresses all of that into two inputs: one photo, one audio track, out comes a lip-synced talking video. Version 1.5 shipped on May 21, 2026 with code and weights fully open-sourced. A Chinese short-video channel billed it as “digital humans at commodity prices,” and its official evaluation scenarios — news broadcasting, knowledge courses, commercial promotion — match exactly what those channels demonstrate.

Core features

  • Photo plus audio to video: no avatar training or fine-tuning; native tasks include Audio-Text-to-Video and Audio-Text-Image-to-Video.
  • Whisper-based lip sync: v1.5 swaps the Wav2Vec2 audio encoder for Whisper-Large, which the team credits for smoother lip dynamics.
  • Long-video stability: official claims include full-body temporal stability and identity consistency on long generations, a classic failure mode of avatar models.
  • Stylized domains: handles anime characters, animals, multi-person interactions, and object handling — not just realistic headshots.
  • Single- and multi-stream audio: drives one speaker or several characters from multi-track audio, for dialogue content.
  • 8-step inference: DMD2-based step distillation cuts inference to 8 NFE, pulling cost per clip down with it.

Typical use cases

  • News-style updates and courses: fix one presenter image, feed the day’s script, batch-produce talking videos.
  • E-commerce explainers: product imagery plus a voiceover track becomes a virtual host clip without a shoot.
  • Anime and virtual-IP content: stylized characters speak too, so a dubbing session turns straight into footage.

Quick start

Code lives in meituan-longcat/LongCat-Video, weights on Hugging Face:

git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video
pip install -r requirements.txt

The README also asks for torch 2.6.0 (CUDA 12.4) and flash_attn 2.7.4.post1. To skip the environment entirely, community ComfyUI nodes (such as rookiestar28/ComfyUI-LongCat-Avatar) and RunPod images already exist.

Summary

LongCat-Video-Avatar suits teams producing talking-head content in volume without appearing on camera, and developers studying audio-driven video; if you want “upload a photo, get a clip in minutes,” a hosted service like HeyGen stays simpler. The code is MIT-licensed, about 8.4k stars as of 2026-10-01. Caveats: local inference wants serious VRAM and a CUDA setup, and rendering is offline — it cannot do live streaming. For real-time interaction, see LiveTalking.