- Main
Cross-language speech computations in the human temporal cortex
- Bhaya-Grossman, Ilina
- Advisor(s): Chang, Edward F
Abstract
Understanding speech is a rapid and seemingly effortless feat that humans perform every day. Yet this ability relies on complex neural computations that integrate language-specific knowledge in order to transform external acoustic signals into internal linguistic representations. This dissertation investigates how neural computations in the human temporal cortex, particularly in the superior temporal gyrus (STG), support speech processing across languages. In the first chapter, we introduce the theoretical and empirical foundations for studying speech processing in the STG and describe how high-resolution intracranial brain recordings, such as electrocorticography (ECoG), have enabled its functional description. Based on prior work, we argue that the STG performs fundamentally nonlinear and dynamic speech computations, including categorization, normalization, and the extraction of temporal structure. Open questions remain as to the extent to which speech computations in the STG are shared across speakers of different languages or language-experience dependent. In the second chapter, we focus on vowel perception to understand how abstract language-specific speech categories emerge from continuous acoustic variation. Vowels are acoustically cued by formants, resonance frequencies determined by vocal tract configuration, and serve as a fundamental building block across spoken languages. Using ECoG recordings from Spanish speakers listening to continuous speech, we demonstrate that in the STG, the predominant encoding of vowels is acoustic. Specifically, we find that neural responses recorded from single electrode contacts are highly selective to restricted zones within the two dimensional formant space. Only by aggregating across these local neural responses can vowel categories be decoded. A second experiment in which participants listened to synthetic formant-based sounds revealed that STG neural responses exhibit general formant sensitivity that extends beyond the natural speech range. In the third chapter, we look beyond single sounds (e.g. vowels) to address how language experience affects speech computations in the human STG. We collected ECoG recordings from speakers of Spanish, English and Mandarin as they listened to speech in both their native and an unfamiliar foreign language. Across participants, the STG exhibited shared acoustic-phonetic encoding for basic speech sound features such as vowel formants and consonants. However, enhanced neural encoding for language-specific sequence and word-level and features was observed when listening to native speech versus foreign speech. In Spanish-English bilinguals, language-specific neural encoding was observed for both familiar languages within the same STG neural populations. Together, these findings demonstrate that the encoding of acoustic-phonetics is integrated with language-specific word-level information within the STG. In the fourth chapter, we synthesize prior work and the empirical findings in this dissertation into an updated model of speech processing in the human STG. We propose that the STG performs multi-scale, recurrent computations to link acoustic-phonetic features with language-specific structures, like words. This model emphasizes the STG’s role as a critical interface between auditory and language systems, ultimately enabling the remarkable human capacity for speech understanding.