Tag
This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.
This paper presents an in silico simulation of the RAMPHO episodic buffer using phonetic entropy from wav2vec 2.0 to dissociate informational and energetic masking in multi-talker environments, revealing a cognitive-acoustic Pareto optimization problem.
OpenAI demonstrates the background robustness of its new voice model in noisy environments, accurately identifying conversation partners, understanding context, allowing users to interrupt naturally, and achieving smooth multi-person interaction.