speech-llms

Tag

Cards List
#speech-llms

Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

arXiv cs.CL · 5d ago Cached

This paper analyzes spurious onsets in full-duplex speech LLMs like Moshi and PersonaPlex and introduces an inference-time mitigation method based on causal analysis to suppress them without retraining.

0 favorites 0 likes
#speech-llms

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

arXiv cs.CL · 2026-09-11 Cached

This paper details the Eloquence team's approaches for Task 2 of the Interspeech 2026 MLC-SLM challenge, which involves multilingual multiple-choice question answering using speech LLMs with fine-tuning, in-context learning, and retrieval systems.

0 favorites 0 likes
#speech-llms

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL · 2026-08-07 Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#speech-llms

Latent-IM: Latent Interaction Management for Speech LLMs

arXiv cs.CL · 2026-07-30 Cached

Introduces Latent-IM, a framework for recovering interaction management from frozen speech LLMs using activation-based selection and steering for conversational moves. It improves end-to-end move accuracy by 12.5 points over the unsteered backbone.

0 favorites 0 likes
#speech-llms

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

arXiv cs.CL · 2026-05-21 Cached

This paper analyzes synchronization and turn-taking dynamics in full-duplex speech dialogue models by simulating conversations between two instances of the Moshi model, measuring representational alignment via CKA and predicting turn boundaries with LSTM probes.

0 favorites 0 likes
← Back to home

Submit Feedback