- Main
Computational approaches to semantic change: the case of Christian Latin
- Lunardi, Valentina
- Advisor(s): Goldstein, David M;
- Vine, Brent H
Abstract
The theory of semantic change remains comparatively underdeveloped within historical linguistics, partly because change is often driven by extra-linguistic forces that resist prediction, and partly because close reading, the traditional method for tracing it, cannot realistically cover the vocabulary of a whole language. This dissertation investigates what and under which conditions static and contextual word embeddings can contribute to the study of semantic change in historical, low-resource languages. Latin, and the effect of the spread of Christianity on its vocabulary, serves as the case study through which this question is pursued. The dissertation builds and validates the resources this requires: a corrected and extended version of the LatinISE corpus, and a continuous, text-level measure of Christian lexical signal, developed collaboratively and validated against independently assessed texts. It also tests whether continued pretraining the contextual model on the working corpus yields a clear benefit (it does not). These resources support five case studies – dominus, pāgānus, commūnicō, commūniō, and uirtūs – selected computationally from a philologically motivated candidate pool to sample a deliberately uneven range of conditions. The static and contextual embedding models recover the semantic developments already documented by philological scholarship to varying degrees. Their accuracy is not uniform: modeling change at the level of the whole lemma, for instance, can dilute a development that is concentrated in one part of its paradigm, and other conditions can distort the picture in different ways. Despite these limits, the models remain valuable. Even an imperfect result offers a glimpse of where change may be occurring, worth investigating further through close reading, and the models organize every occurrence of a lemma by distributional similarity into interactive projections, where occurrences using a lemma in a similar way tend to cluster together – letting a researcher compare with ease passages where the word under investigation has similar distributional environments, rather than encountering them scattered across a large corpus. The dissertation offers this combination as a template other researchers can adapt to other historical, low-resource corpora.