Tag
UniSwap is a new framework for joint audio-visual identity swapping in talking videos, using a unified streaming audio-visual diffusion transformer to replace appearance and vocal timbre while preserving source content and dynamics.
This tweet points out that open-source solutions can already turn a photo and an audio clip into a talking video, so there's no need to pay a monthly fee for AI avatar platforms. The repo link is in the reply.