Tag
This paper analyzes spurious onsets in full-duplex speech LLMs like Moshi and PersonaPlex and introduces an inference-time mitigation method based on causal analysis to suppress them without retraining.
Nari Labs introduces Qwen3-TTS and Qwen3-ASR models, providing high accuracy, low latency, and cost-effectiveness in a free public beta, alongside optimized APIs and services for production deployment.
The Awesome-SpeechLM-Survey repository on GitHub systematically organizes the research lineage of speech language models, including classification frameworks, representative models, training datasets, and evaluation benchmarks. It serves as a knowledge map for understanding the field.