Phone Segmentation and Recognition through Phonological Activation Mapping
Summary
This paper introduces SPAM (S3M-based Phonological Activation Mapping), a method that leverages self-supervised speech models to perform both phone segmentation and recognition simultaneously using lightweight, gradient-descent-free prediction heads requiring minimal phonetic transcriptions.
View Cached Full Text
Cached at: 07/13/26, 07:49 AM
Paper page - Phone Segmentation and Recognition through Phonological Activation Mapping
Source: https://huggingface.co/papers/2607.09020 Authors:
,
,
,
,
,
,
,
,
,
Abstract
Phonesegmentationandrecognitionareinherentlyrelatedtasks,yetmodernapproachestypicallymodelthemseparately.Wearguethatphoneticstructureisalreadylatentintherepresentationsofself-supervisedspeechmodels(S3Ms),andoneonlyneedstosteerthemtosolvebothtasks.WeleverageS3M-basedPhonologicalActivationMapping(SPAM),whichmapseachS3Mrepresentationframetoavectorofphonologicalfeatureactivations,suchasvoicingandnasality.OntopofSPAM,weintroducetwosimplebuteffectivelightweight,gradient-descent-freepredictionheads:arecognitionheadandasegmentationhead.Ourmethodrequireslessthanaminuteofphonetictranscriptions,andgeneralizestounseenphonesduringtraining.Acrossadiverserangeofdatasets,ourapproachattainsstrongsegmentationandrecognitionperformance.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2607\.09020
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.09020 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.09020 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.09020 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
Presents an age-aware multi-task learning method for phoneme recognition from children's speech, enabling a lightweight 94M-parameter model to outperform larger models and run on edge devices like phones.
Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese
This paper presents SAMPA, a Whisper-based segmenter fine-tuned on Brazilian Portuguese speech data to automatically mark terminal prosodic boundaries, achieving competitive F1 scores on held-out and out-of-distribution datasets.
SAM 3: Segment Anything with Concepts
SAM 3 introduces a unified model for promptable concept segmentation and tracking, achieving state-of-the-art performance with a decoupled recognition and localization architecture and a scalable data engine.
Phonological Perception of Sign Language Models
This paper evaluates whether Sign Language Recognition models exhibit phonological sensitivity by probing them with minimal pairs of signs, revealing architectural trade-offs and emergent but limited phonological perception.