Tag
Ultralight Digital Human is an open-source Python project that lets you train a person-specific, audio-driven talking head from a 3–5 minute video, with support for HuBERT/WeNet audio features, ONNX export, and streaming inference.
Duix.Avatar is a local AI avatar toolkit that generates lip-synced avatar videos from video, voice, and script inputs without sending data to the cloud. It supports eight languages and is available under a community license on GitHub.
This paper presents PolyInterview, an LLM-based platform that conducts immersive mock interviews with a lip-synced digital human, providing comprehensive multimodal assessment of response content, vocal delivery, and non-verbal behavior. The system generates role-tailored questions and evaluates candidates using 13 behavioral features linked to KSA and STAR frameworks.
Fay is a fully open-source, commercially free digital human framework that supports offline operation, modular design, MCP tool invocation, and bionic memory. It features a built-in web management interface, enabling quick setup for virtual streamers, intelligent customer service, personal AI assistants, and more.
SpatialAvatar-0 introduces a multi-stage reconstruction method for high-quality 4D head avatars using a shared FLAME-mesh-bound Gaussian representation, achieving superior performance across benchmarks with reduced iterations.
Introducing the open-source project Pixelle-Video: a fully automated AI short video engine. Input a topic and it automatically generates a video with script, images, voiceover, and background music. Supports local and cloud models, modular design allows flexible replacement of each component model.
Meituan open-sourced the LongCat-Video-Avatar-1.5 model, which supports generating realistic talking videos from a single photo and voice, supports multiple languages and long videos, and outperforms commercial closed-source solutions.