Tag
PD-GS introduces a phoneme-driven 3D Gaussian Splatting approach for audio-driven talking heads, using a Linguistic Fusion Module to improve lip articulation and reduce closure violations.
This tweet points out that open-source solutions can already turn a photo and an audio clip into a talking video, so there's no need to pay a monthly fee for AI avatar platforms. The repo link is in the reply.
Ultralight Digital Human is an open-source Python project that lets you train a person-specific, audio-driven talking head from a 3–5 minute video, with support for HuBERT/WeNet audio features, ONNX export, and streaming inference.
NVIDIA Audio2Face-3D generates high-fidelity 3D facial animations from audio, providing accurate lip-sync and emotional expression. It is released as an open-source SDK with pre-trained models and plugins for Maya and Unreal Engine 5.
LongCat-Video-Avatar 1.5 is an upgraded open-source framework for audio-driven human video generation with improved lip synchronization, production-ready stability, and efficient 8-step inference.