Tag
StreamAlign is a streaming text-aligned speech tokenization framework that enables real-time speech–text joint modeling, reducing latency and achieving state-of-the-art results on speech recognition and spoken language modeling tasks.
This paper introduces PINT, a method for invariant speech tokenization that fine-tunes an SSL encoder using alignment losses across parallel utterances to distill linguistic content, achieving significant reductions in speaker probe accuracy and LM perplexity.
Proposes speaker-disentangled syllabic tokenization using chunk-wise regression to learn linguistic content tokens from raw speech, achieving state-of-the-art in syllable boundary detection and clustering, and improving speech language model performance.