hubert

Tag

Cards List
#hubert

Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

arXiv cs.CL · 2026-08-24 Cached

This paper proposes using self-supervised speech representations from HuBERT to track the convergence of children's speech to adult patterns in deaf and hard-of-hearing infants, demonstrating a scalable, language-neutral assessment method for language development.

0 favorites 0 likes
#hubert

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

arXiv cs.CL · 2026-07-22 Cached

This paper introduces PINT, a method for invariant speech tokenization that fine-tunes an SSL encoder using alignment losses across parallel utterances to distill linguistic content, achieving significant reductions in speaker probe accuracy and LM perplexity.

0 favorites 0 likes
#hubert

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

arXiv cs.CL · 2026-07-07 Cached

Proposes speaker-disentangled syllabic tokenization using chunk-wise regression to learn linguistic content tokens from raw speech, achieving state-of-the-art in syllable boundary detection and clustering, and improving speech language model performance.

0 favorites 0 likes
#hubert

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

Hugging Face Daily Papers · 2026-07-05 Cached

This paper proposes a speaker-disentangled syllabic tokenizer that uses regression of perturbed student representations toward clean teacher targets, achieving state-of-the-art syllable boundary detection and a 7% relative improvement in speech language model understanding over SpiRit-LM.

0 favorites 0 likes
#hubert

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

arXiv cs.AI · 2026-06-20 Cached

This paper presents a systematic empirical study of fine-tuning pretrained Transformer models (Wav2Vec2.0, HuBERT, XLS-R) for Quranic Automatic Speech Recognition (ASR), achieving a WER of 0.08 on the EveryAyah subset and reducing training time from 140 to 40 hours, with Wav2Vec2-XLSR-53 providing the best representation.

0 favorites 0 likes
#hubert

hubert.cpp, a C++ implementation of distilHuBERT [P]

Reddit r/MachineLearning · 2026-06-12

A C++ implementation of distilHuBERT with no runtime dependencies, compiled-in weights, dynamic sizing, and on-par performance with ONNX Runtime, designed for easy integration into CMake projects.

0 favorites 0 likes
← Back to home

Submit Feedback