@wsl8297: There is a repository on GitHub that clearly organizes the research lineage of Speech Language Models (SpeechLM): Awesome-SpeechLM-Survey. It comprehensively organizes classification frameworks, representative models, training datasets, and evaluation benchmarks into a single 'knowledge map,' making it time-efficient to look up materials, fill in background knowledge, and find benchmarks.

X AI KOLs Timeline Tools

Summary

The Awesome-SpeechLM-Survey repository on GitHub systematically organizes the research lineage of speech language models, including classification frameworks, representative models, training datasets, and evaluation benchmarks. It serves as a knowledge map for understanding the field.

There is a repository on GitHub that clearly organizes the research lineage of Speech Language Models (SpeechLM): Awesome-SpeechLM-Survey. It comprehensively organizes everything from classification frameworks, representative models, training datasets, to evaluation benchmarks into a single 'knowledge map,' saving time when researching, filling in background knowledge, and finding benchmarks. GitHub: https://github.com/dreamtheater123/Awesome-SpeechLM-Survey... You can find within it: - A comprehensive overview of 50+ speech language models: GPT-4o, Moshi, Mini-Omni, etc., all systematically organized - A genealogy of speech tokenizer technologies: covering semantic, acoustic, and hybrid approaches - A summary of 20+ training datasets: covering everything from LibriSpeech to the latest instruction data - A comparison of 10+ evaluation benchmarks: including key benchmarks like the self-developed VoxEval If you want to systematically understand the technical routes and research progress of speech AI, this repository is well-suited for quickly building a cognitive framework and conducting in-depth tracking.
Original Article
View Cached Full Text

Cached at: 05/14/26, 06:41 PM

On GitHub, there is a well-organized resource repository that clearly lays out the research landscape of Speech Language Models: Awesome-SpeechLM-Survey. It integrates classification frameworks, representative models, training datasets, and evaluation benchmarks into a one-stop “knowledge map,” saving time for searching materials, filling in background knowledge, and finding comparable works.
GitHub: https://github.com/dreamtheater123/Awesome-SpeechLM-Survey…
Inside, you can find:

  • A comprehensive overview of 50+ Speech Language Models: GPT-4o, Moshi, Mini-Omni, etc., all systematically organized.
  • The technical lineage of speech tokenizers: covering semantic, acoustic, and hybrid approaches.
  • A summary of 20+ training datasets: from LibriSpeech to the latest instruction data, with full coverage.
  • Comparison of 10+ evaluation benchmarks: including key benchmarks like VoxEval developed in-house.
    If you want to systematically understand the technical routes and research progress of speech AI, this repository is ideal for quickly building a cognitive framework and diving deeper.

dreamtheater123/Awesome-SpeechLM-Survey

Source: https://github.com/dreamtheater123/Awesome-SpeechLM-Survey

Awesome-SpeechLM-Survey

arXiv (https://arxiv.org/abs/2410.03751)

🎉🎉🎉Our survey paper “Recent Advances in Speech Language Models: A Survey” has been accepted to ACL 2025 main conference!

This is the Github repository for paper: Recent Advances in Speech Language Models: A Survey (https://arxiv.org/abs/2410.03751). In this paper, we survey the field of Speech Language Models (SpeechLMs), which are capable of performing end-to-end speech interactions with humans and serve as autoregressive foundation models.

News

  • We have released VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models (https://arxiv.org/abs/2501.04962) on arXiv!🎉 Also check it out on Github (https://github.com/dreamtheater123/VoxEval).

Introduction

Why SpeechLMs? SpeechLMs are used for end-to-end speech-based interactions. Traditional ASR + LLM + TTS setups suffer from information loss and cumulative errors during conversion. SpeechLMs directly model speech data, capturing both semantic and paralinguistic information for richer interactions!

Intro Figure

Taxonomy

We introduce a novel taxonomy for SpeechLMs, categorizing them based on their architecture and training recipes.

taxonomy

Existing SpeechLMs

ModelTitleUrl
OpenAI Advanced Voice ModeOpenAI Advanced Voice ModeLink (https://help.openai.com/en/articles/9617425-advanced-voice-mode-faq)
Claude Voice ModeClaude Voice ModeLink (https://support.anthropic.com/en/articles/11101966-using-voice-mode-on-claude-mobile-apps)
MindGPT-4o-Audio理想同学MindGPT-4o-Audio实时语音对话大模型发布Link (https://mp.weixin.qq.com/s?__biz=MzkyNzc3ODYzMQ==&mid=2247483808&idx=1&sn=15b2d0fc5c415066e9e85a0e17fa4094&chksm=c313b6c42e4bac7f551e3ce6b314897e6c09b2829d202ae09b088c39b6208d14545221a82785&mpshare=1&scene=1&srcid=06157RruwKQJDuZvxSmt0ALH&sharer_shareinfo=4f156f5dabba628552a2429a555bca65&sharer_shareinfo_first=4f156f5dabba628552a2429a555bca65#rd)
VITA-AudioVITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language ModelLink (https://arxiv.org/abs/2505.03739)
VoilaVoila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-PlayLink (https://arxiv.org/abs/2505.02707)
Kimi-AudioKimi-Audio Technical ReportLink (https://arxiv.org/abs/2504.18425)
LyraLyra: An Efficient and Speech-Centric Framework for Omni-CognitionLink (https://arxiv.org/abs/2412.09501)
Flow-OmniContinuous Speech Tokens Makes LLMs Robust Multi-Modality LearnersLink (https://arxiv.org/abs/2412.04917)
NTPPNTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair PredictionLink (https://arxiv.org/abs/2506.00975)
Qwen2.5-OmniQwen2.5-Omni Technical ReportLink (https://arxiv.org/abs/2503.20215)
CSMConversational Speech Generation ModelLink (https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice)
MinmoMinMo: A Multimodal Large Language Model for Seamless Voice InteractionLink (https://arxiv.org/abs/2501.06282)
SlammingSlamming: Training a Speech Language Model on One GPU in a DayLink (https://arxiv.org/abs/2502.15814)
VITA-1.5VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionLink (https://arxiv.org/abs/2501.01957)
Baichuan-AudioBaichuan-Audio: A Unified Framework for End-to-End Speech InteractionLink (https://arxiv.org/abs/2502.17239)
Step-AudioStep-Audio: Unified Understanding and Generation in Intelligent Speech InteractionLink (https://arxiv.org/abs/2502.11946)
MiniCPM-oA GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your PhoneLink (https://github.com/OpenBMB/MiniCPM-o)
SyncLLMBeyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsLink (https://arxiv.org/abs/2409.15594)
OmniFlattenOmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationLink (https://arxiv.org/abs/2410.17799)
SLAM-OmniSLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage TrainingLink (https://arxiv.org/abs/2412.15649)
GLM-4-VoiceGLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken ChatbotLink (https://arxiv.org/abs/2412.02612)
-Scaling Speech-Text Pre-training with Synthetic Interleaved DataLink (http://arxiv.org/abs/2411.17607)
SALMONN-omniSALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and GenerationLink (http://arxiv.org/abs/2411.18138)
Mini-Omni2Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex CapabilitiesLink (http://arxiv.org/abs/2410.11190)
UniaudioUniaudio: An audio foundation model toward universal audio generationLink (https://arxiv.org/abs/2310.00704)
ParrotParrot: Autoregressive Spoken Dialogue Language Modeling with Decoder-only TransformersLink (https://openreview.net/forum?id=Ttndg2Jl5F)
MoshiMoshi: a speech-text foundation model for real-time dialogueLink (https://kyutai.org/Moshi.pdf)
Freeze-OmniFreeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLMLink (http://arxiv.org/abs/2411.00774)
EMOVAEMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsLink (http://arxiv.org/abs/2409.18042)
IntrinsicVoiceIntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction AbilitiesLink (http://arxiv.org/abs/2410.08035)
LSLMLanguage Model Can Listen While SpeakingLink (http://arxiv.org/abs/2408.02622)
SpiRit-LMSpiRit-LM: Interleaved Spoken and Written Language ModelLink (http://arxiv.org/abs/2402.05755)
SpeechGPT-GenSpeechGPT-Gen: Scaling Chain-of-Information Speech GenerationLink (https://arxiv.org/abs/2401.13527v2)
SpectronSpoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMLink (https://openreview.net/forum?id=izrOLJov5y)
SUTLMToward Joint Language Modeling for Speech Units and TextLink (http://arxiv.org/abs/2310.08715)
tGSLMGenerative Spoken Language Model based on continuous word-sized audio tokensLink (https://arxiv.org/abs/2310.05224v1)
LauraGPTLauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPTLink (https://arxiv.org/abs/2310.04673v4)
VoxtLMVoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation TasksLink (https://ieeexplore.ieee.org/abstract/document/10447112/?casa_token=rNHOTa7BbZMAAAAA:3Dk4RlgUcRbvDIewE9uUk-wk5D_0f2zm1z4hGgG1DSMkiH-KZwk7AVs5Z8PVMetvCKxFdV1C9o0)
VITAVITA: Towards Open-Source Interactive Omni Multimodal LLMLink (https://arxiv.org/abs/2408.05211)
FunAudioLLMFunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMsLink (https://arxiv.org/abs/2407.04051)
VoiceboxVoicebox: Text-guided multilingual universal speech generation at scaleLink (https://proceedings.neurips.cc/paper_files/paper/2023/hash/2d8911db9ecedf866015091b28946e15-Abstract-Conference.html)
LLaMA-OmniLLaMA-Omni: Seamless Speech Interaction with Large Language ModelsLink (https://arxiv.org/abs/2409.06666)
Mini-OmniMini-Omni: Language Models Can Hear, Talk While Thinking in StreamingLink (https://arxiv.org/abs/2408.16725)
TWISTTextually pretrained speech language modelsLink (https://proceedings.neurips.cc/paper_files/paper/2023/hash/c859b99b5d717c9035e79d43dfd69435-Abstract-Conference.html)
GPSTGenerative pre-trained speech language model with efficient hierarchical transformerLink (https://aclanthology.org/2024.acl-long.97)
AudioPaLMAudioPaLM: A Large Language Model That Can Speak and ListenLink (http://arxiv.org/abs/2306.12925)
VioLAVioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and TranslationLink (http://arxiv.org/abs/2305.16107)
SpeechGPTSpeechgpt: Empowering large language models with intrinsic cross-modal conversational abilitiesLink (https://arxiv.org/abs/2305.11000)
dGSLMGenerative spoken dialogue language modelingLink (https://direct.mit.edu/tacl/article-abstract/doi/10.1162/tacl_a_00545/115240)
pGSLMText-Free Prosody-Aware Generative Spoken Language ModelingLink (http://arxiv.org/abs/2109.03264)
GSLMOn generative spoken language modeling from raw audioLink (https://direct.mit.edu/tacl/article-abstract/doi/10.1162/tacl_a_00430/108611)

SpeechLM Tokenizers

Semantic Tokenizers

NameTitleUrl
WhisperRobust Speech Recognition via Large-Scale Weak SupervisionLink (https://arxiv.org/abs/2212.04356)
CosyVoiceCosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensLink (https://arxiv.org/abs/2407.05407)
Google USMGoogle USM: Scaling Automatic Speech Recognition Beyond 100 LanguagesLink (https://arxiv.org/abs/2303.01037)
WavLMWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech ProcessingLink (https://arxiv.org/abs/2110.13900)
HuBERTHuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden UnitsLink (https://arxiv.org/abs/2106.07447)
W2v-bertW2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-TrainingLink (https://arxiv.org/abs/2108.06209)
Wav2vec 2.0wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsLink (https://arxiv.org/abs/2006.11477)

Acoustic Tokenizers

NameTitleUrl
WavTokenizerWavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingLink (https://arxiv.org/abs/2408.16532)
SNACSNAC: Multi-Scale Neural Audio CodecLink (https://arxiv.org/abs/2410.14411)
EncodecHigh Fidelity Neural Audio CompressionLink (https://arxiv.org/abs/2210.13438)
SoundStreamSoundStream: An End-to-End Neural Audio CodecLink (https://arxiv.org/abs/2107.03312)

Mixed Tokenizers

NameTitleUrl
SpeechTokenizerSpeechTokenizer: Unified Speech Tokenizer for Speech Large Language ModelsLink (https://arxiv.org/abs/2308.16692)
MimiMoshi: a speech-text foundation model for real-time dialogueLink (https://arxiv.org/abs/2410.00037)

Popular Training Datasets

DatasetTypePhaseHoursYear
LibriSpeech (https://www.openslr.org/12)ASRPre-Training1k2015
Multilingual LibriSpeech (https://www.openslr.org/94/)ASRPre-Training50.5k2020
LibriLight (https://github.com/facebookresearch/libri-light)ASRPre-Training60k2019
People dataset (https://github.com/mlcommons/peoples-speech)ASRPre-Training30k2021
VoxPopuli (https://github.com/facebookresearch/voxpopuli)ASRPre-Training1.6k2021
Gigaspeech (https://github.com/SpeechColab/GigaSpeech)ASRPre-Training40k2021
Common Voice (https://commonvoice.mozilla.org/zh-CN)ASRPre-Training2.5k2019
VCTK (https://paperswithcode.com/dataset/voice-bank-demand)ASRPre-Training0.3k2017
WenetSpeech (https://wenet.org.cn/WenetSpeech/)ASRPre-Training22k2022
LibriTTS (https://www.openslr.org/60/)TTSPre-Training0.6k2019
CoVoST2 (https://github.com/facebookresearch/covost)S2TTPre-Training2.8k2020
CVSS (https://github.com/google-research-datasets/cvss)S2STPre-Training1.9k2022
VoxCeleb (https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html)Speaker IdentificationPre-Training0.4k2017
VoxCeleb2 (https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html)Speaker IdentificationPre-Training2.4k2018
Spotify Podcasts (https://podcastsdataset.byspotify.com/)PodcastPre-Training47k2020
Fisher (https://catalog.ldc.upenn.edu/LDC2004T19)Telephone conversationPre-Training2k2004
SpeechInstruct (https://huggingface.co/datasets/fnlp/SpeechInstruct)Instruction-followingInstruction-Tuning-2023
InstructS2S-200KInstruction-followingInstruction-Tuning-2024
VoiceAssistant-400K (https://huggingface.co/datasets/gpt-omni/VoiceAssistant-400K)Instruction-followingInstruction-Tuning-2024

Evaluation Benchmarks

NameEval Type# TasksAudio TypeI/O
ABXRepresentation1SpeechA \rightarrow -
sWUGGYLinguistic1SpeechA \rightarrow -
sBLIMPLinguistic1SpeechA \rightarrow -
sStoryCloze (https://github.com/slp-rl/SpokenStoryCloze)Linguistic1SpeechA/T \rightarrow -
STSP (https://github.com/facebookresearch/spiritlm/blob/main/spiritlm/eval/README.md)Paralinguistic1SpeechA/T \rightarrow A/T
MMAU (https://github.com/apple/axlearn/tree/main/docs/research/mmau)Downstream27Speech, Sound, MusicA \rightarrow T
Audiobench (https://github.com/AudioLLMs/AudioBench)Downstream8Speech, SoundA \rightarrow T
AIR-Bench (https://github.com/OFA-Sys/AIR-Bench)Downstream2

Similar Articles

@GitHub_Daily: Want to understand the underlying principles of large language models? Most resources only cover theory or provide source code, leaving you still confused. Stumbled upon this open-source tutorial, EveryonesLLM, which guides us step by step to build a complete large language model from scratch on Google Colab, writing code throughout. The whole tutorial is divided into...

X AI KOLs Timeline

EveryonesLLM is an open-source tutorial that provides 29 chapters of Colab notebooks. It teaches users step by step to build a complete large language model from scratch on Google Colab, including pre-training and instruction fine-tuning, and supports Chinese.

@GitHub_Daily: How do large language models work internally, why do they hallucinate, and why do they sometimes give irrelevant answers? For a deeper understanding, check out the Awesome LLM Interpretability resource collection, which provides a systematic path to unpack the AI black box. It covers attention visualization, neuron analysis, and more.

X AI KOLs Timeline

Introduces the Awesome LLM Interpretability resource collection, which gathers various interpretability tools, papers, and community resources to help understand the internal workings of large language models.

@Jolyne_AI: An open-source hands-on book: "Hands-On Large Language Models". The book has 12 chapters, progressing from language model fundamentals to prompt engineering, semantic search, model fine-tuning, and multimodal applications, covering the key paths to deploying large models in practice. GitHub: h…

X AI KOLs Timeline

An open-source hands-on book "Hands-On Large Language Models", with 12 chapters covering language model fundamentals, prompt engineering, semantic search, model fine-tuning, and multimodal applications. It provides runnable code examples, ideal for practical learning.

@GitHub_Daily: Want to understand how Large Language Models actually work? Existing resources are either too academic and hard to digest, or too superficial, focusing only on concepts, with nothing that clearly explains the entire process from start to finish. Similarly, I came across the 'how-llms-work' project, which turns the complete workflow of LLMs into a visual interactive webpage, based on Andrej Karpathy’s...

X AI KOLs Timeline

An interactive visual guide, 'how-llms-work', breaks down the entire lifecycle of Large Language Models based on Andrej Karpathy's lectures, covering data collection to post-training.

@MaxForAI: If you are working on voice agents, you should try this project. A team from NTU, NUS, and Shanghai AI Lab released: Mega-ASR. This fully open-source ASR is built on Qwen3-ASR, aiming to break the long-standing bottleneck of ASR performance in noisy, reverberant, or other impaired real-world environments...

X AI KOLs Timeline

NTU, NUS, and Shanghai AI Lab jointly released Mega-ASR, a fully open-source ASR model built on Qwen3-ASR. Using the Voices-in-the-Wild-2M dataset and progressive acoustic-to-semantic optimization, it achieves up to 30% relative Word Error Rate (WER) reduction in real-world noisy environments. With only 1.7B parameters, it enables efficient inference on consumer-grade hardware.