Introducing SAM Audio: The First Unified Multimodal Model for Audio Separation
Summary
SAM Audio is introduced as the first unified multimodal model for audio separation, enabling users to isolate specific sounds from complex mixtures using text, visual, or temporal prompts.
Similar Articles
AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
AuralSAM2 integrates audio into SAM2 via an AuralFuser module that generates sparse and dense prompts from audio-visual features, enhancing cross-modal segmentation while maintaining interactive efficiency.
@multimodalart: UniSE: Unified Speech Enhancement high quality open source model for making an audio crisp & isolating speakers in mult…
UniSE is a unified, prompt-free, autoregressive speech enhancement model based on a decoder-only language model, supporting multiple tasks like speech restoration, target speaker extraction, and speech separation in a single model.
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
The paper proposes a multimodal solution for audio sentiment polarity classification that integrates audio and multilingual text transcripts via cross-modal transformers, and uses knowledge distillation to enhance an audio-only model without computational overhead during inference.
SAM 3: Segment Anything with Concepts
SAM 3 introduces a unified model for promptable concept segmentation and tracking, achieving state-of-the-art performance with a decoupled recognition and localization architecture and a scalable data engine.
Audio Interaction Model
This paper introduces Audio-Interaction, a unified streaming audio model that combines offline task execution with real-time audio instruction following via an end-to-end framework. It proposes SoundFlow for the perceive-decide-respond loop and evaluates competitive performance across benchmarks.