- Main
ROBUST SPEAKER VERIFICATION UNDER DOMAIN AND DEVELOPMENTAL VARIABILITY: ADAPTATION AND MODELING IN LOW-RESOURCE CHILD SPEECH
- Shetty, Vishwas
- Advisor(s): Alwan, Abeer Dr
Abstract
Modern speaker verification systems achieve strong performance on adult speech, but their performance often degrades when applied to children’s speech. This degradation arises from three challenges: children’s speech differs substantially from adult speech, children’s voices continue to change as they grow, and the lack of large child speech databases. Developmental changes affect acoustic and articulatory properties of speech and create age-related variability that can obscure speaker identity. This dissertation addresses the problem of robust speaker verification for children by studying developmental variability in child speech and developing adaptation methods that improve child speaker verification while preserving performance on adult speech. The first contribution of this dissertation is a longitudinal analysis of acoustic and articulatory development in children. By analyzing vowel formants, subglottal resonances, and ultrasound-based tongue-shape measures from children recorded longitudinally across four years, this dissertation quantifies both systematic developmental trends and child-specific variability. These findings motivate the need for speaker verification systems that are robust to age-related changes in children’s voices. Building on this analysis, the dissertation develops methods for age-robust child speaker verification, including approaches that improve verification when enrollment and test speech are separated by one or more years. It then addresses the broader problem of efficiently adapting adult speech trained speaker verification models to child speech under limited child speech data. A central challenge in this progression is catastrophic forgetting, i.e., adaptation to child speech can improve child speaker verification while degrading adult speaker verification performance. This dissertation addresses this challenge first in settings where source-domain data is available, and then in the more restrictive source-free setting where the original adult training data is unavailable during adaptation. Overall, this dissertation provides a progression from quantifying developmental variability in child speech to building robust, efficient, and age-agnostic speaker verification systems that better serve both children and adults.