spoken-language-model

Tag

Cards List
#spoken-language-model

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

arXiv cs.CL · 2026-07-08 Cached

This paper introduces Lychee-FD, a native end-to-end full-duplex spoken language model that mitigates modality interference through a hierarchical parameter separation strategy, achieving significant improvements in speech intelligence and interaction fluidity.

0 favorites 0 likes
#spoken-language-model

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Hugging Face Daily Papers · 2026-06-30 Cached

FlexiSLM introduces dynamic frame rate capabilities for speech input and output in spoken language models, outperforming fixed-frame-rate models and enabling controllable inference speed.

0 favorites 0 likes
#spoken-language-model

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

arXiv cs.CL · 2026-05-11 Cached

VITA-QinYu is an expressive end-to-end spoken language model capable of role-playing and singing, trained on 15.8K hours of data to outperform peers in expressiveness and conversational accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback