UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Hugging Face Daily Papers Papers

Summary

UniSwap is a new framework for joint audio-visual identity swapping in talking videos, using a unified streaming audio-visual diffusion transformer to replace appearance and vocal timbre while preserving source content and dynamics.

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.
Original Article
View Cached Full Text

Cached at: 08/14/26, 07:27 AM

Paper page - UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Source: https://huggingface.co/papers/2608.11752

Abstract

UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations.

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a singleaudio-visual diffusion transformerwhile preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce aswap-and-reconstruct pipelinethat removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model throughIn-context Pretrainingfor joint replacement,Conditional Streaming Adaptationforblock-causal KV-cached generation, andEfficient Self-forcing DMDfor mitigating exposure bias and reducing sampling from 30 to 3denoising stepsper block. EfficientMulti-LoRA Switchingenables the three DMD roles to share a single frozen backbone.Feature-RoPE Decompositionkeeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.

View arXiv pageView PDFProject pageGitHub4Add to collection

Get this paper in your agent:

hf papers read 2608\.11752

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.11752 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.11752 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.11752 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Hugging Face Daily Papers

AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.