Skip to main content
eScholarship
Open Access Publications from the University of California

Phonological Perception of Sign Language Models

Creative Commons 'BY' version 4.0 license
Abstract

Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates SLR model phonological perception by probing sensitivity using minimal pairs and measuring representational alignment with human behavioral data. Our results reveal emergent phonological sensitivity with clear architectural trade-offs: pose-based models are more sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlate with human perceptual similarity judgments (r –ï 0.49). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to overcome their architectural inductive biases.