What Do Self-Supervised Speech Models Know About Words?

Pasad, Ankita; Chien, Chung-Ming; Settle, Shane; Livescu, Karen

doi:10.1162/tacl_a_00656

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

Supporting

Mentioning

Contrasting

Year Published

2024

Publication Types

Select...

Article3

Other2

Relationship

Self Cite0

Independent5

Authors

Journals

Cited by 6 publications

References 81 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

Choi,

Jung,

Watanabe

2024

ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

Choi,

Jung,

Watanabe

2024

ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

Convexity Based Pruning of Speech Representation Models

Dorszewski,

Tetkova,

Hansen

2024

2024 IEEE 34th International Workshop on Machine Learning for Signal Processing (MLSP)

View full text Add to dashboard Cite

A perceptual similarity space for speech based on self-supervised speech representations

Chernyak,

Bradlow,

Keshet

et al. 2024

The Journal of the Acoustical Society of America

View full text Add to dashboard Cite

Speech recognition by both humans and machines frequently fails in non-optimal yet common situations. For example, word recognition error rates for second-language (L2) speech can be high, especially under conditions involving background noise. At the same time, both human and machine speech recognition sometimes shows remarkable robustness against signal- and noise-related degradation. Which acoustic features of speech explain this substantial variation in intelligibility? Current approaches align speech to text to extract a small set of pre-defined spectro-temporal properties from specific sounds in particular words. However, variation in these properties leaves much cross-talker variation in intelligibility unexplained. We examine an alternative approach utilizing a perceptual similarity space acquired using self-supervised learning. This approach encodes distinctions between speech samples without requiring pre-defined acoustic features or speech-to-text alignment. We show that L2 English speech samples are less tightly clustered in the space than L1 samples reflecting variability in English proficiency among L2 talkers. Critically, distances in this similarity space are perceptually meaningful: L1 English listeners have lower recognition accuracy for L2 speakers whose speech is more distant in the space from L1 speech. These results indicate that perceptual similarity may form the basis for an entirely new speech and language analysis approach.

show abstract

What Do Self-Supervised Speech Models Know About Words?

Cited by 6 publications

References 81 publications

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

Understanding Probe Behaviors Through Variational Bounds of Mutual Information

Convexity Based Pruning of Speech Representation Models

A perceptual similarity space for speech based on self-supervised speech representations

Contact Info

Product

Resources

About