- Main
Language experience and prediction across the lifespan: evidence from diachronic fine-tuning of language models
Abstract
Humans predict upcoming language input from context, which depends on prior language experience. This suggests that older adults' predictions may differ from those of young adults, due to longer language exposure. Here we use sentence completion data from two age cohorts (YA = 18-35 y.o.; OA = 50-80 y.o.) and language models fine-tuned to particular decades of a diachronic corpus of American English to examine the relationship between changes in language statistics and differences in linguistic prediction across different age groups. We observed greater consistency in contextual probabilities within age groups compared to across age groups, indicating that YA and OA make subtly different predictions given identical context. Next-word prediction performance for the fine-tuned models decreased as the temporal distance between the fine- tuning and testing decade increased, indicating that language usage statistics changed over the span of a few decades. Further, GPT-2 surprisal values are more predictive of YA than OA contextual probabilities, suggesting that the language statistics, as captured by a model trained largely on internet text, aligns more with YA's internal model than OA's. However, both age groups' data are better fit by models fine-tuned on more recent corpus decades.