Skip to main content
eScholarship
Open Access Publications from the University of California

Markedness as Surprisal: Information Theory and Language Inefficiency

Creative Commons 'BY' version 4.0 license
Abstract

Information-theoretic approaches to language predict pressure toward efficient, low-surprisal encoding, yet languages systematically retain forms that violate these pressures. Such departures are traditionally described as marked. In this study, we model markedness as localized phonological inefficiency, quantified using phonemic bigram surprisal. Focusing on iconic words in English (e.g., vroom) and ideophones in Japanese (e.g., fuyafuya), we examine how surprisal is distributed within words relative to language-specific phonotactic baselines. Across both languages, iconic forms exhibit reliably higher surprisal than non-iconic controls, indicating increased phonological unpredictability. This inefficiency is not uniform: surprisal is concentrated at specific word-internal positions, particularly at word onsets. In Japanese, baseline asymmetries in CV structure strongly shape surprisal, but ideophones retain distinct information-structural profiles once these constraints are controlled. The results presented here are limited in that they are correlational and do not by themselves establish that phonological inefficiency is functionally deployed, but nonetheless demonstrate a systematic association between iconicity and elevated surprisal.