@wsl8297: GitHub 上有一份把「语音语言模型(SpeechLM)」研究脉络梳理得很清楚的资源库:Awesome-SpeechLM-Survey。 它把分类框架、代表模型、训练数据集到评测基准一站式整理成“知识地图”,查资料、补背景、找对标都很省…
摘要
GitHub 上的 Awesome-SpeechLM-Survey 仓库系统整理了语音语言模型的研究脉络,包括分类框架、代表模型、训练数据集和评测基准,是了解该领域的知识地图。
查看缓存全文
缓存时间: 2026/05/14 18:41
GitHub 上有一份把「语音语言模型(SpeechLM)」研究脉络梳理得很清楚的资源库:Awesome-SpeechLM-Survey。 它把分类框架、代表模型、训练数据集到评测基准一站式整理成“知识地图”,查资料、补背景、找对标都很省时间。 GitHub:https://github.com/dreamtheater123/Awesome-SpeechLM-Survey… 你能在里面看到: - 50+ 语音语言模型全景盘点:GPT-4o、Moshi、Mini-Omni 等都有条理化整理 - 语音 tokenizer 技术谱系:语义 / 声学 / 混合方案一网打尽 - 20+ 训练数据集汇总:从 LibriSpeech 到最新指令数据覆盖完整 - 10+ 评测基准对比:包含自研 VoxEval 等关键基准 想系统吃透语音 AI 的技术路线与研究进展,这份仓库很适合用来快速建立认知框架并深入追踪。
dreamtheater123/Awesome-SpeechLM-Survey
Source: https://github.com/dreamtheater123/Awesome-SpeechLM-Survey
Awesome-SpeechLM-Survey
🎉🎉🎉Our survey paper “Recent Advances in Speech Language Models: A Survey” has been accepted to ACL 2025 main conference!
This is the Github repository for paper: Recent Advances in Speech Language Models: A Survey. In this paper, we survey the field of Speech Language Models (SpeechLMs), which are capable of performing end-to-end speech interactions with humans and serve as autoregressive foundation models.
News
- We have released VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models on arXiv!🎉 Also check it out on Github.
Introduction
Why SpeechLMs? SpeechLMs are used for end-to-end speech-based interactions. Traditional ASR + LLM + TTS setups suffer from information loss and cumulative errors during conversion. SpeechLMs directly model speech data, capturing both semantic and paralinguistic information for richer interactions!

Taxonomy
We introduce a novel taxonomy for SpeechLMs, categorizing them based on their architecture and training recipes.

Existing SpeechLMs
| Model | Title | Url |
|---|---|---|
| OpenAI Advanced Voice Mode | OpenAI Advanced Voice Mode | Link |
| Claude Voice Mode | Claude Voice Mode | Link |
| MindGPT-4o-Audio | 理想同学MindGPT-4o-Audio实时语音对话大模型发布 | Link |
| VITA-Audio | VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model | Link |
| Voila | Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play | Link |
| Kimi-Audio | Kimi-Audio Technical Report | Link |
| Lyra | Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition | Link |
| Flow-Omni | Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners | Link |
| NTPP | NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction | Link |
| Qwen2.5-Omni | Qwen2.5-Omni Technical Report | Link |
| CSM | Conversational Speech Generation Model | Link |
| Minmo | MinMo: A Multimodal Large Language Model for Seamless Voice Interaction | Link |
| Slamming | Slamming: Training a Speech Language Model on One GPU in a Day | Link |
| VITA-1.5 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction | Link |
| Baichuan-Audio | Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction | Link |
| Step-Audio | Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction | Link |
| MiniCPM-o | A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your Phone | Link |
| SyncLLM | Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents | Link |
| OmniFlatten | OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation | Link |
| SLAM-Omni | SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training | Link |
| GLM-4-Voice | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot | Link |
| - | Scaling Speech-Text Pre-training with Synthetic Interleaved Data | Link |
| SALMONN-omni | SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation | Link |
| Mini-Omni2 | Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities | Link |
| Uniaudio | Uniaudio: An audio foundation model toward universal audio generation | Link |
| Parrot | Parrot: Autoregressive Spoken Dialogue Language Modeling with Decoder-only Transformers | Link |
| Moshi | Moshi: a speech-text foundation model for real-time dialogue | Link |
| Freeze-Omni | Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | Link |
| EMOVA | EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions | Link |
| IntrinsicVoice | IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities | Link |
| LSLM | Language Model Can Listen While Speaking | Link |
| SpiRit-LM | SpiRit-LM: Interleaved Spoken and Written Language Model | Link |
| SpeechGPT-Gen | SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation | Link |
| Spectron | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM | Link |
| SUTLM | Toward Joint Language Modeling for Speech Units and Text | Link |
| tGSLM | Generative Spoken Language Model based on continuous word-sized audio tokens | Link |
| LauraGPT | LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT | Link |
| VoxtLM | VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks | Link |
| VITA | VITA: Towards Open-Source Interactive Omni Multimodal LLM | Link |
| FunAudioLLM | FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs | Link |
| Voicebox | Voicebox: Text-guided multilingual universal speech generation at scale | Link |
| LLaMA-Omni | LLaMA-Omni: Seamless Speech Interaction with Large Language Models | Link |
| Mini-Omni | Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming | Link |
| TWIST | Textually pretrained speech language models | Link |
| GPST | Generative pre-trained speech language model with efficient hierarchical transformer | Link |
| AudioPaLM | AudioPaLM: A Large Language Model That Can Speak and Listen | Link |
| VioLA | VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation | Link |
| SpeechGPT | Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities | Link |
| dGSLM | Generative spoken dialogue language modeling | Link |
| pGSLM | Text-Free Prosody-Aware Generative Spoken Language Modeling | Link |
| GSLM | On generative spoken language modeling from raw audio | Link |
SpeechLM Tokenizers
Semantic Tokenizers
| Name | Title | Url |
|---|---|---|
| Whisper | Robust Speech Recognition via Large-Scale Weak Supervision | Link |
| CosyVoice | CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens | Link |
| Google USM | Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages | Link |
| WavLM | WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing | Link |
| HuBERT | HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units | Link |
| W2v-bert | W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training | Link |
| Wav2vec 2.0 | wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations | Link |
Acoustic Tokenizers
| Name | Title | Url |
|---|---|---|
| WavTokenizer | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling | Link |
| SNAC | SNAC: Multi-Scale Neural Audio Codec | Link |
| Encodec | High Fidelity Neural Audio Compression | Link |
| SoundStream | SoundStream: An End-to-End Neural Audio Codec | Link |
Mixed Tokenizers
| Name | Title | Url |
|---|---|---|
| SpeechTokenizer | SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models | Link |
| Mimi | Moshi: a speech-text foundation model for real-time dialogue | Link |
Popular Training Datasets
| Dataset | Type | Phase | Hours | Year |
|---|---|---|---|---|
| LibriSpeech | ASR | Pre-Training | 1k | 2015 |
| Multilingual LibriSpeech | ASR | Pre-Training | 50.5k | 2020 |
| LibriLight | ASR | Pre-Training | 60k | 2019 |
| People dataset | ASR | Pre-Training | 30k | 2021 |
| VoxPopuli | ASR | Pre-Training | 1.6k | 2021 |
| Gigaspeech | ASR | Pre-Training | 40k | 2021 |
| Common Voice | ASR | Pre-Training | 2.5k | 2019 |
| VCTK | ASR | Pre-Training | 0.3k | 2017 |
| WenetSpeech | ASR | Pre-Training | 22k | 2022 |
| LibriTTS | TTS | Pre-Training | 0.6k | 2019 |
| CoVoST2 | S2TT | Pre-Training | 2.8k | 2020 |
| CVSS | S2ST | Pre-Training | 1.9k | 2022 |
| VoxCeleb | Speaker Identification | Pre-Training | 0.4k | 2017 |
| VoxCeleb2 | Speaker Identification | Pre-Training | 2.4k | 2018 |
| Spotify Podcasts | Podcast | Pre-Training | 47k | 2020 |
| Fisher | Telephone conversation | Pre-Training | 2k | 2004 |
| SpeechInstruct | Instruction-following | Instruction-Tuning | - | 2023 |
| InstructS2S-200K | Instruction-following | Instruction-Tuning | - | 2024 |
| VoiceAssistant-400K | Instruction-following | Instruction-Tuning | - | 2024 |
Evaluation Benchmarks
| Name | Eval Type | # Tasks | Audio Type | I/O |
|---|---|---|---|---|
| ABX | Representation | 1 | Speech | A \rightarrow - |
| sWUGGY | Linguistic | 1 | Speech | A \rightarrow - |
| sBLIMP | Linguistic | 1 | Speech | A \rightarrow - |
| sStoryCloze | Linguistic | 1 | Speech | A/T \rightarrow - |
| STSP | Paralinguistic | 1 | Speech | A/T \rightarrow A/T |
| MMAU | Downstream | 27 | Speech, Sound, Music | A \rightarrow T |
| Audiobench | Downstream | 8 | Speech, Sound | A \rightarrow T |
| AIR-Bench | Downstream | 20 | Speech, Sound, Music | A \rightarrow T |
| SD-Eval | Downstream | 4 | Speech | A \rightarrow T |
| SUPERB | Downstream | 10 | Speech | A \rightarrow T |
| Dynamic-SUPERB | Downstream | 180 | Speech, Sound, Music | A \rightarrow T |
| SALMON | Downstream | 8 | Speech | A \rightarrow - |
| VoiceBench | Downstream | 8 | Speech | A \rightarrow A |
| VoxEval | Downstream | 56 | Speech | A \rightarrow A |
Citation
@article{cui2024recent,
title={Recent advances in speech language models: A survey},
author={Cui, Wenqian and Yu, Dianzhi and Jiao, Xiaoqi and Meng, Ziqiao and Zhang, Guangyan and Wang, Qichao and Guo, Yiwen and King, Irwin},
journal={arXiv preprint arXiv:2410.03751},
year={2024}
}
相似文章
@GitHub_Daily: 想搞懂大语言模型底层原理,大部分资料只介绍理论知识,或者只给源码,看完还是一头雾水。 偶然看到 EveryonesLLM 这个开源教程,手把手带我们在 Google Colab 上从零搭建一个完整的大语言模型,全程动手写代码。 整套教程分…
EveryonesLLM 是一个开源教程,提供29个章节的Colab笔记本,手把手教用户从零在Google Colab上搭建完整的大语言模型,包括预训练和指令微调,并支持中文。
@GitHub_Daily: 大语言模型内部是如何工作的,为什么会产生幻觉,为什么有时答非所问,想深入了解这些。 可以看下 Awesome LLM Interpretability 这份资源合集,提供一整套拆解 AI 黑盒的系统路径。 涵盖从注意力可视化、神经元分析到…
介绍了一个名为 Awesome LLM Interpretability 的资源合集,汇集了多种可解释性工具、论文和社区资源,帮助理解大语言模型的内部工作机制。
@Jolyne_AI: 一本开源实战书:《Hands-On Large Language Models》(《动手学大模型》)。 全书 12 章,从语言模型基础到提示词工程、语义搜索、模型微调,再到多模态应用,循序渐进,覆盖大模型落地的关键路径。 GitHub:h…
一本开源实战书《Hands-On Large Language Models》(《动手学大模型》),全书12章,覆盖语言模型基础、提示词工程、语义搜索、模型微调及多模态应用,提供可运行代码示例,适合实战学习。
@GitHub_Daily: 想了解大语言模型到底是怎么工作的,找到的资料都太过于学术看不懂,或者说的太浅只讲概念,就没一个从头到尾讲清楚的内容。 无独有偶,看到 how-llms-work 这个项目,把大模型的完整流程做成了一个可视化交互网页,内容基于 Karpat…
An interactive visual guide, 'how-llms-work', breaks down the entire lifecycle of Large Language Models based on Andrej Karpathy's lectures, covering data collection to post-training.
@MaxForAI: 如果你在做语音Agent,你应该试一下这个项目 来自南洋理工、新国立和上海 AI Lab的团队发布了:Mega-ASR 这个完全开源的ASR基于 Qwen3-ASR构建,目的是打破长期困扰ASR的在嘈杂、混响或其他受损现实环境中表现的瓶颈…
南洋理工、新国立和上海 AI Lab 联合发布 Mega-ASR,一个基于 Qwen3-ASR 构建的完全开源 ASR 模型,通过 Voices-in-the-Wild-2M 数据集和渐进式声学到语义优化,在真实世界嘈杂环境中实现最高 30% 的相对词错误率下降,且仅 1.7B 参数可在消费级硬件高效推理。