Tag
Presents a multimodal voice activity projection framework extending audio-only VAP to audio-visual inputs for turn-taking prediction in social robots, using pretrained backbones and low-rank adaptation. Achieves improvements on NoXi and Haru EDR corpora.
This paper introduces Representation Distribution Matching (RDM), a method for one-step image generation by matching feature distributions under pretrained encoders, achieving state-of-the-art results on ImageNet and enabling post-training of FLUX.2 into a one-step generator with improved performance.