@wsl8297: GitHub 上有一份把「语音语言模型(SpeechLM)」研究脉络梳理得很清楚的资源库:Awesome-SpeechLM-Survey。 它把分类框架、代表模型、训练数据集到评测基准一站式整理成“知识地图”,查资料、补背景、找对标都很省…

X AI KOLs Timeline 工具

摘要

GitHub 上的 Awesome-SpeechLM-Survey 仓库系统整理了语音语言模型的研究脉络,包括分类框架、代表模型、训练数据集和评测基准,是了解该领域的知识地图。

GitHub 上有一份把「语音语言模型(SpeechLM)」研究脉络梳理得很清楚的资源库:Awesome-SpeechLM-Survey。 它把分类框架、代表模型、训练数据集到评测基准一站式整理成“知识地图”,查资料、补背景、找对标都很省时间。 GitHub:https://github.com/dreamtheater123/Awesome-SpeechLM-Survey… 你能在里面看到: - 50+ 语音语言模型全景盘点:GPT-4o、Moshi、Mini-Omni 等都有条理化整理 - 语音 tokenizer 技术谱系:语义 / 声学 / 混合方案一网打尽 - 20+ 训练数据集汇总:从 LibriSpeech 到最新指令数据覆盖完整 - 10+ 评测基准对比:包含自研 VoxEval 等关键基准 想系统吃透语音 AI 的技术路线与研究进展,这份仓库很适合用来快速建立认知框架并深入追踪。
查看原文
查看缓存全文

缓存时间: 2026/05/14 18:41

GitHub 上有一份把「语音语言模型(SpeechLM)」研究脉络梳理得很清楚的资源库:Awesome-SpeechLM-Survey。 它把分类框架、代表模型、训练数据集到评测基准一站式整理成“知识地图”,查资料、补背景、找对标都很省时间。 GitHub:https://github.com/dreamtheater123/Awesome-SpeechLM-Survey… 你能在里面看到: - 50+ 语音语言模型全景盘点:GPT-4o、Moshi、Mini-Omni 等都有条理化整理 - 语音 tokenizer 技术谱系:语义 / 声学 / 混合方案一网打尽 - 20+ 训练数据集汇总:从 LibriSpeech 到最新指令数据覆盖完整 - 10+ 评测基准对比:包含自研 VoxEval 等关键基准 想系统吃透语音 AI 的技术路线与研究进展,这份仓库很适合用来快速建立认知框架并深入追踪。


dreamtheater123/Awesome-SpeechLM-Survey

Source: https://github.com/dreamtheater123/Awesome-SpeechLM-Survey

Awesome-SpeechLM-Survey

arXiv

🎉🎉🎉Our survey paper “Recent Advances in Speech Language Models: A Survey” has been accepted to ACL 2025 main conference!

This is the Github repository for paper: Recent Advances in Speech Language Models: A Survey. In this paper, we survey the field of Speech Language Models (SpeechLMs), which are capable of performing end-to-end speech interactions with humans and serve as autoregressive foundation models.

News

Introduction

Why SpeechLMs? SpeechLMs are used for end-to-end speech-based interactions. Traditional ASR + LLM + TTS setups suffer from information loss and cumulative errors during conversion. SpeechLMs directly model speech data, capturing both semantic and paralinguistic information for richer interactions!

Intro Figure

Taxonomy

We introduce a novel taxonomy for SpeechLMs, categorizing them based on their architecture and training recipes.

taxonomy

Existing SpeechLMs

ModelTitleUrl
OpenAI Advanced Voice ModeOpenAI Advanced Voice ModeLink
Claude Voice ModeClaude Voice ModeLink
MindGPT-4o-Audio理想同学MindGPT-4o-Audio实时语音对话大模型发布Link
VITA-AudioVITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language ModelLink
VoilaVoila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-PlayLink
Kimi-AudioKimi-Audio Technical ReportLink
LyraLyra: An Efficient and Speech-Centric Framework for Omni-CognitionLink
Flow-OmniContinuous Speech Tokens Makes LLMs Robust Multi-Modality LearnersLink
NTPPNTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionLink
Qwen2.5-OmniQwen2.5-Omni Technical ReportLink
CSMConversational Speech Generation ModelLink
MinmoMinMo: A Multimodal Large Language Model for Seamless Voice InteractionLink
SlammingSlamming: Training a Speech Language Model on One GPU in a DayLink
VITA-1.5VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionLink
Baichuan-AudioBaichuan-Audio: A Unified Framework for End-to-End Speech InteractionLink
Step-AudioStep-Audio: Unified Understanding and Generation in Intelligent Speech InteractionLink
MiniCPM-oA GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your PhoneLink
SyncLLMBeyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsLink
OmniFlattenOmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationLink
SLAM-OmniSLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage TrainingLink
GLM-4-VoiceGLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken ChatbotLink
-Scaling Speech-Text Pre-training with Synthetic Interleaved DataLink
SALMONN-omniSALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and GenerationLink
Mini-Omni2Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex CapabilitiesLink
UniaudioUniaudio: An audio foundation model toward universal audio generationLink
ParrotParrot: Autoregressive Spoken Dialogue Language Modeling with Decoder-only TransformersLink
MoshiMoshi: a speech-text foundation model for real-time dialogueLink
Freeze-OmniFreeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMLink
EMOVAEMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsLink
IntrinsicVoiceIntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction AbilitiesLink
LSLMLanguage Model Can Listen While SpeakingLink
SpiRit-LMSpiRit-LM: Interleaved Spoken and Written Language ModelLink
SpeechGPT-GenSpeechGPT-Gen: Scaling Chain-of-Information Speech GenerationLink
SpectronSpoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMLink
SUTLMToward Joint Language Modeling for Speech Units and TextLink
tGSLMGenerative Spoken Language Model based on continuous word-sized audio tokensLink
LauraGPTLauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPTLink
VoxtLMVoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation TasksLink
VITAVITA: Towards Open-Source Interactive Omni Multimodal LLMLink
FunAudioLLMFunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMsLink
VoiceboxVoicebox: Text-guided multilingual universal speech generation at scaleLink
LLaMA-OmniLLaMA-Omni: Seamless Speech Interaction with Large Language ModelsLink
Mini-OmniMini-Omni: Language Models Can Hear, Talk While Thinking in StreamingLink
TWISTTextually pretrained speech language modelsLink
GPSTGenerative pre-trained speech language model with efficient hierarchical transformerLink
AudioPaLMAudioPaLM: A Large Language Model That Can Speak and ListenLink
VioLAVioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and TranslationLink
SpeechGPTSpeechgpt: Empowering large language models with intrinsic cross-modal conversational abilitiesLink
dGSLMGenerative spoken dialogue language modelingLink
pGSLMText-Free Prosody-Aware Generative Spoken Language ModelingLink
GSLMOn generative spoken language modeling from raw audioLink

SpeechLM Tokenizers

Semantic Tokenizers

NameTitleUrl
WhisperRobust Speech Recognition via Large-Scale Weak SupervisionLink
CosyVoiceCosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensLink
Google USMGoogle USM: Scaling Automatic Speech Recognition Beyond 100 LanguagesLink
WavLMWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech ProcessingLink
HuBERTHuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden UnitsLink
W2v-bertW2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-TrainingLink
Wav2vec 2.0wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsLink

Acoustic Tokenizers

NameTitleUrl
WavTokenizerWavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingLink
SNACSNAC: Multi-Scale Neural Audio CodecLink
EncodecHigh Fidelity Neural Audio CompressionLink
SoundStreamSoundStream: An End-to-End Neural Audio CodecLink

Mixed Tokenizers

NameTitleUrl
SpeechTokenizerSpeechTokenizer: Unified Speech Tokenizer for Speech Large Language ModelsLink
MimiMoshi: a speech-text foundation model for real-time dialogueLink

Popular Training Datasets

DatasetTypePhaseHoursYear
LibriSpeechASRPre-Training1k2015
Multilingual LibriSpeechASRPre-Training50.5k2020
LibriLightASRPre-Training60k2019
People datasetASRPre-Training30k2021
VoxPopuliASRPre-Training1.6k2021
GigaspeechASRPre-Training40k2021
Common VoiceASRPre-Training2.5k2019
VCTKASRPre-Training0.3k2017
WenetSpeechASRPre-Training22k2022
LibriTTSTTSPre-Training0.6k2019
CoVoST2S2TTPre-Training2.8k2020
CVSSS2STPre-Training1.9k2022
VoxCelebSpeaker IdentificationPre-Training0.4k2017
VoxCeleb2Speaker IdentificationPre-Training2.4k2018
Spotify PodcastsPodcastPre-Training47k2020
FisherTelephone conversationPre-Training2k2004
SpeechInstructInstruction-followingInstruction-Tuning-2023
InstructS2S-200KInstruction-followingInstruction-Tuning-2024
VoiceAssistant-400KInstruction-followingInstruction-Tuning-2024

Evaluation Benchmarks

NameEval Type# TasksAudio TypeI/O
ABXRepresentation1SpeechA \rightarrow -
sWUGGYLinguistic1SpeechA \rightarrow -
sBLIMPLinguistic1SpeechA \rightarrow -
sStoryClozeLinguistic1SpeechA/T \rightarrow -
STSPParalinguistic1SpeechA/T \rightarrow A/T
MMAUDownstream27Speech, Sound, MusicA \rightarrow T
AudiobenchDownstream8Speech, SoundA \rightarrow T
AIR-BenchDownstream20Speech, Sound, MusicA \rightarrow T
SD-EvalDownstream4SpeechA \rightarrow T
SUPERBDownstream10SpeechA \rightarrow T
Dynamic-SUPERBDownstream180Speech, Sound, MusicA \rightarrow T
SALMONDownstream8SpeechA \rightarrow -
VoiceBenchDownstream8SpeechA \rightarrow A
VoxEvalDownstream56SpeechA \rightarrow A

Citation

@article{cui2024recent,
  title={Recent advances in speech language models: A survey},
  author={Cui, Wenqian and Yu, Dianzhi and Jiao, Xiaoqi and Meng, Ziqiao and Zhang, Guangyan and Wang, Qichao and Guo, Yiwen and King, Irwin},
  journal={arXiv preprint arXiv:2410.03751},
  year={2024}
}

相似文章

@GitHub_Daily: 想搞懂大语言模型底层原理,大部分资料只介绍理论知识,或者只给源码,看完还是一头雾水。 偶然看到 EveryonesLLM 这个开源教程,手把手带我们在 Google Colab 上从零搭建一个完整的大语言模型,全程动手写代码。 整套教程分…

X AI KOLs Timeline

EveryonesLLM 是一个开源教程,提供29个章节的Colab笔记本,手把手教用户从零在Google Colab上搭建完整的大语言模型,包括预训练和指令微调,并支持中文。

@GitHub_Daily: 大语言模型内部是如何工作的,为什么会产生幻觉,为什么有时答非所问,想深入了解这些。 可以看下 Awesome LLM Interpretability 这份资源合集,提供一整套拆解 AI 黑盒的系统路径。 涵盖从注意力可视化、神经元分析到…

X AI KOLs Timeline

介绍了一个名为 Awesome LLM Interpretability 的资源合集,汇集了多种可解释性工具、论文和社区资源,帮助理解大语言模型的内部工作机制。

@Jolyne_AI: 一本开源实战书:《Hands-On Large Language Models》(《动手学大模型》)。 全书 12 章,从语言模型基础到提示词工程、语义搜索、模型微调,再到多模态应用,循序渐进,覆盖大模型落地的关键路径。 GitHub:h…

X AI KOLs Timeline

一本开源实战书《Hands-On Large Language Models》(《动手学大模型》),全书12章,覆盖语言模型基础、提示词工程、语义搜索、模型微调及多模态应用,提供可运行代码示例,适合实战学习。

@GitHub_Daily: 想了解大语言模型到底是怎么工作的,找到的资料都太过于学术看不懂,或者说的太浅只讲概念,就没一个从头到尾讲清楚的内容。 无独有偶,看到 how-llms-work 这个项目,把大模型的完整流程做成了一个可视化交互网页,内容基于 Karpat…

X AI KOLs Timeline

An interactive visual guide, 'how-llms-work', breaks down the entire lifecycle of Large Language Models based on Andrej Karpathy's lectures, covering data collection to post-training.

@MaxForAI: 如果你在做语音Agent,你应该试一下这个项目 来自南洋理工、新国立和上海 AI Lab的团队发布了:Mega-ASR 这个完全开源的ASR基于 Qwen3-ASR构建,目的是打破长期困扰ASR的在嘈杂、混响或其他受损现实环境中表现的瓶颈…

X AI KOLs Timeline

南洋理工、新国立和上海 AI Lab 联合发布 Mega-ASR,一个基于 Qwen3-ASR 构建的完全开源 ASR 模型,通过 Voices-in-the-Wild-2M 数据集和渐进式声学到语义优化,在真实世界嘈杂环境中实现最高 30% 的相对词错误率下降,且仅 1.7B 参数可在消费级硬件高效推理。