speech-tokenization

Tag

Cards List
#speech-tokenization

StreamAlign: Streaming Text-Aligned Speech Tokenization

arXiv cs.CL · yesterday Cached

StreamAlign is a streaming text-aligned speech tokenization framework that enables real-time speech–text joint modeling, reducing latency and achieving state-of-the-art results on speech recognition and spoken language modeling tasks.

0 favorites 0 likes
#speech-tokenization

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

arXiv cs.CL · 2026-07-22 Cached

This paper introduces PINT, a method for invariant speech tokenization that fine-tunes an SSL encoder using alignment losses across parallel utterances to distill linguistic content, achieving significant reductions in speaker probe accuracy and LM perplexity.

0 favorites 0 likes
#speech-tokenization

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

arXiv cs.CL · 2026-07-07 Cached

Proposes speaker-disentangled syllabic tokenization using chunk-wise regression to learn linguistic content tokens from raw speech, achieving state-of-the-art in syllable boundary detection and clustering, and improving speech language model performance.

0 favorites 0 likes
← Back to home

Submit Feedback