Skip to main content
eScholarship
Open Access Publications from the University of California

Latent speech representations learned through self-supervised learning predict listeners' generalization of adaptation across talkers

Creative Commons 'BY' version 4.0 license
Abstract

Unfamiliar accents can pose a challenge to speech recognition. However, listeners often adapt quickly to novel accents, and even generalize this adaptation across talkers with the same accent. We investigate how such cross-talker generalization---critical to effective speech perception---is achieved. We take advantage of advances in automatic speech recognition to test whether comparatively simple similarity-based inferences can explain cross-talker generalization in human listeners. We use the latent perceptual space learned by the HuBERT model---shaped by the statistics of the speech signal and the objective to recognize speech---to meaningfully measure the similarity between talkers' pronunciation. We find that word-level similarity in this latent space predict listeners' ability to successfully generalize across talkers. We discuss consequences for theories of adaptive speech perception. In particular, our results explain why cross-talker variability is not a prerequisite for cross-talker generalization (contrary to influential accounts).