UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Summary
UniSwap is a new framework for joint audio-visual identity swapping in talking videos, using a unified streaming audio-visual diffusion transformer to replace appearance and vocal timbre while preserving source content and dynamics.
View Cached Full Text
Cached at: 08/14/26, 07:27 AM
Paper page - UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Source: https://huggingface.co/papers/2608.11752
Abstract
UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations.
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a singleaudio-visual diffusion transformerwhile preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce aswap-and-reconstruct pipelinethat removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model throughIn-context Pretrainingfor joint replacement,Conditional Streaming Adaptationforblock-causal KV-cached generation, andEfficient Self-forcing DMDfor mitigating exposure bias and reducing sampling from 30 to 3denoising stepsper block. EfficientMulti-LoRA Switchingenables the three DMD roles to share a single frozen backbone.Feature-RoPE Decompositionkeeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2608\.11752
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.11752 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.11752 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.11752 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
The article discusses the UniVidX paper, which introduces a unified multimodal framework for video generation using diffusion priors and discusses its cross-modal coherence mechanisms.
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
UniSpace introduces a unified visual representation that unifies semantic understanding, high-fidelity reconstruction, and image generation in a single space using a reparameterized ViT, eliminating the need for a separate VAE.
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.
@multimodalart: UniSE: Unified Speech Enhancement high quality open source model for making an audio crisp & isolating speakers in mult…
UniSE is a unified, prompt-free, autoregressive speech enhancement model based on a decoder-only language model, supporting multiple tasks like speech restoration, target speaker extraction, and speech separation in a single model.
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
SwanSphere proposes a unified streaming framework for high-fidelity spatial audio generation from panoramic videos and text prompts using causal autoregressive diffusion transformers and multimodal learning strategies, achieving superior performance in both video-to-spatial and text-to-spatial audio tasks.