speech-llm

Tag

Cards List
#speech-llm

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Hugging Face Daily Papers · yesterday Cached

X-AuT is a progressive framework for compressing audio-encoder layers in speech large language models, reducing inference cost while restoring accuracy via techniques like cross-scale distillation and LoRA adaptation.

0 favorites 0 likes
#speech-llm

Relative Time Intervals Representation for Word-level Timestamping with Masked Training

arXiv cs.AI · 2026-08-26 Cached

This paper introduces a method for word-level timestamping in Speech Large Language Models using relative time intervals and masked training to enhance prediction accuracy and robustness against noisy real-world annotations.

0 favorites 0 likes
#speech-llm

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

arXiv cs.CL · 2026-08-17 Cached

VoiceChat-TTS is a low-latency, continuous text-to-speech model designed for interactive agents, enabling real-time streaming and interruption handling without compromising speech quality.

0 favorites 0 likes
#speech-llm

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

arXiv cs.CL · 2026-07-29 Cached

This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems, evaluating on English and Italian tasks and achieving competitive word error rates with reduced communication costs.

0 favorites 0 likes
#speech-llm

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

arXiv cs.CL · 2026-07-21 Cached

This paper describes TalTech's systems for generating SOAP notes directly from doctor-patient conversation audio, using Voxtral models fine-tuned with supervised learning and DAPO reinforcement learning. Their submissions ranked first in both tracks of the BeTraC challenge, achieving high concept accuracy and low hallucination rates.

0 favorites 0 likes
#speech-llm

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

arXiv cs.CL · 2026-07-08 Cached

This paper introduces Lychee-FD, a native end-to-end full-duplex spoken language model that mitigates modality interference through a hierarchical parameter separation strategy, achieving significant improvements in speech intelligence and interaction fluidity.

0 favorites 0 likes
#speech-llm

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task

arXiv cs.CL · 2026-07-08 Cached

This paper presents an open-source re-implementation of the NAVER LABS instruction-following pipeline for IWSLT 2026, using SeamlessM4T-v2-large and Qwen3-4B-Instruct, with 100k synthetic examples.

0 favorites 0 likes
#speech-llm

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

arXiv cs.CL · 2026-07-03 Cached

This paper proposes Joint Speech-Text Interleaved Pretraining (JSTIP), a pretraining strategy that constructs word-level and segment-level interleaved speech-text sequences to improve ASR entity accuracy and reduce the modality gap between speech and text, showing competitive performance on domain adaptation and zero-shot speech question answering.

0 favorites 0 likes
#speech-llm

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

arXiv cs.CL · 2026-06-26 Cached

This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.

0 favorites 0 likes
#speech-llm

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention

arXiv cs.CL · 2026-06-04 Cached

This paper identifies a localized 'entity binding failure' in Speech Large Language Models (SLLMs) where logical reasoning involving entity tracking collapses to chance-level accuracy, and proposes Entity-Aware Chain-of-Thought (EA-CoT) prompting to resolve this, achieving up to 24.4% absolute accuracy improvement.

0 favorites 0 likes
#speech-llm

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hugging Face Daily Papers · 2026-06-03

SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.

0 favorites 0 likes
#speech-llm

LaSR: Context-Aware Speech Recognition via Latent Reasoning

arXiv cs.CL · 2026-06-02 Cached

LaSR proposes a latent reasoning training paradigm for context-aware speech recognition, aligning chain-of-thought supervision around acoustic features to improve terminology recognition without added latency, outperforming standard fine-tuning on Fun-Audio-Chat.

0 favorites 0 likes
#speech-llm

From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation

arXiv cs.CL · 2026-05-18 Cached

Proposes S2ST-Omni 2, a many-to-one compositional speech-to-speech translation framework that replaces flat language labels with structured typological priors to improve multilingual adaptation, achieving superior performance on CVSS-C.

0 favorites 0 likes
#speech-llm

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

arXiv cs.CL · 2026-05-08 Cached

Proposes TextPro-SLM, a speech large language model that minimizes the modality gap by processing spoken input to resemble prosody-aware text input, achieving strong paralinguistic understanding with low training data.

0 favorites 0 likes
#speech-llm

Liberating LLM Capabilities in Full-Duplex Speech Models

Hugging Face Daily Papers · 2026-05-04 Cached

Proposes Listen-Write-Speak (LWS), a text-first tri-channel paradigm that allows a single autoregressive LLM to continuously listen, write visible text, and speak in real-time, enabling full-duplex speech interaction without architectural modifications.

0 favorites 0 likes
← Back to home

Submit Feedback