- Main
Learning Vowel Harmony from Speech: What Emerges, What Requires Additional Structure
Abstract
Vowel harmony is a non-local phonological process in which vowels within a word share a common feature value. We examine which aspects of this process are learnable from speech alone, using Assamese Advanced Tongue Root (ATR) harmony as a test case. We use a generative (fiwGAN) and a predictive model (wav2vec2) as comparative probes of what different speech-learning objectives make accessible. Both models learn ATR-related phonetic distinctions: fiwGAN latent variables control F1 differences, and wav2vec2 separates ATR categories with _81% accuracy. Wav2vec2 also learns global feature agreement (98% accuracy) and recognizes the blocker /A/ (27.9% agreement drop in blocked contexts). However, neither model shows strong evidence for regressive spreading (d = 0.12) or iterative harmony across longer words. These results suggest that raw speech learning exposes the acoustic and co-occurrence structures, while directional phonological generalization may require morphology, lexical alternations, or temporal and architectural biases.