Tag
This paper analyzes spurious onsets in full-duplex speech LLMs like Moshi and PersonaPlex and introduces an inference-time mitigation method based on causal analysis to suppress them without retraining.
This paper details the Eloquence team's approaches for Task 2 of the Interspeech 2026 MLC-SLM challenge, which involves multilingual multiple-choice question answering using speech LLMs with fine-tuning, in-context learning, and retrieval systems.
This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.
Introduces Latent-IM, a framework for recovering interaction management from frozen speech LLMs using activation-based selection and steering for conversational moves. It improves end-to-end move accuracy by 12.5 points over the unsteered backbone.
This paper analyzes synchronization and turn-taking dynamics in full-duplex speech dialogue models by simulating conversations between two instances of the Moshi model, measuring representational alignment via CKA and predicting turn boundaries with LSTM probes.