- Main
Perceptually Grounded Modeling and Modification of Speaker Identity
- Netzorg, Robin
- Advisor(s): Anumanchipalli, Gopala Krishna
Abstract
Speech technology operates on representations of speaker identity that are either too high-level to be actionable (e.g., “masculine” or “old”) or too high-dimensional to be interpretable (e.g., learned embeddings). How can we bridge this gap to enable interpretable analysis and modification of what makes a voice sound the way it does? I present research address- ing this question in three stages. First, I establish that perceptual voice qualities, which are standardized descriptors of voice from speech pathology, speech forensics, and gender- affirming voice training, are reliably perceivable by non-experts and exhibit disentangled correlations with acoustic and physical measurements of the vocal tract. Second, I present the PQ-Encoder, a model that predicts these qualities from speech, and demonstrate that a compact 9-dimensional perceptual vector captures meaningful information about speaker identity, gender, and age. Finally, I introduce StyleMod, a conditional flow matching model that modifies speaker identity along perceptual voice quality axes, enabling interpretable and continuous manipulation of perceived gender and age while preserving non-targeted speaker attributes. Together, this work demonstrates that grounding speech technology in human perception yields representations that are both scientifically meaningful and practically use- ful for controllable voice generation.