Skip to main content
eScholarship
Open Access Publications from the University of California

Japanese/Korean Linguistics

Japanese/Korean Linguistics banner

Can a Transformer Model Learn Sublexical Groups in Korean Like Us?

Creative Commons 'BY-NC-ND' version 4.0 license
Abstract

This study examined whether the selectivity of Korean L-Tensification (LT) extends to nonce words. LT applies only to a subset of eligible words, and this selective application has conventionally been attributed to Sino-Korean (SK) etymology. As an alternative, this study tested whether phonotactic patterns condition LT application. Forty-three speakers and a transformer model trained only on phonemic input-output pairs were tested on the same 144 nonce words, constructed to reflect SK, non-SK, or neutral phonotactics. LT application was quantified using acoustic measures for speakers and confidence scores for the transformer model. Both speakers and the model applied LT selectively. For speakers, non-SK phonotactics suppressed LT in plosive targets, whereas SK phonotactics promoted LT in affricate targets under specific structural conditions. However, the transformer diverged from the speakers, with both phonotactic manipulations lowering LT confidence. These results support a phonotactic contribution to LT selectivity, while the speaker-model divergence motivates future research.