Phone Segmentation and Recognition through Phonological Activation Mapping
2026-07-10 • Artificial Intelligence
Artificial IntelligenceComputation and LanguageMachine LearningSound
AI summaryⓘ
The authors show that speech models trained without supervision already capture hidden details about speech sounds. They use a method called SPAM to link these hidden speech features to specific sound traits like voicing or nasality. Then, they add two simple components to recognize and segment speech sounds without needing lots of training data. Their approach works well on many speech datasets and can even handle new sounds not seen during training.
phone segmentationphone recognitionself-supervised speech modelsphonological featuresSPAMgradient descentphonetic transcriptionspeech representation
Authors
Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
Abstract
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.