Tag
Proposes speaker-disentangled syllabic tokenization using chunk-wise regression to learn linguistic content tokens from raw speech, achieving state-of-the-art in syllable boundary detection and clustering, and improving speech language model performance.
This paper proposes a speaker-disentangled syllabic tokenizer that uses regression of perturbed student representations toward clean teacher targets, achieving state-of-the-art syllable boundary detection and a 7% relative improvement in speech language model understanding over SpiRit-LM.