speech-language-model

Tag

Cards List
#speech-language-model

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

arXiv cs.CL · 2026-07-07 Cached

Proposes speaker-disentangled syllabic tokenization using chunk-wise regression to learn linguistic content tokens from raw speech, achieving state-of-the-art in syllable boundary detection and clustering, and improving speech language model performance.

0 favorites 0 likes
#speech-language-model

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

arXiv cs.CL · 2026-07-03 Cached

This paper proposes SpeechCombine, an instruction-following speech language model trained without any instruction tuning, using only speech pre-training and a weight combination strategy that transfers text LLM capabilities to the speech domain.

0 favorites 0 likes
#speech-language-model

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

arXiv cs.CL · 2026-06-15 Cached

BayLing-Duplex is a native full-duplex speech language model that enables a single autoregressive LLM to manage turn-taking and interruptions without external VAD modules, achieving high success rates and improved response quality over prior models.

0 favorites 0 likes
#speech-language-model

Raon-Speech Technical Report

arXiv cs.CL · 2026-05-26 Cached

Raon-Speech is a 9B-parameter speech language model for English and Korean, supporting understanding, answering, and generation, with a full-duplex extension Raon-SpeechChat for natural real-time conversation. It achieves strong performance across 42 benchmarks and is fully open-sourced.

0 favorites 0 likes
← Back to home

Submit Feedback