- Main
Concepts in terms of other Concepts: a method for the analysis of collective word use
Abstract
Sometimes how concepts are used in day-to-day talk differs from the expectations set in explicit, socially accepted, definitions. For example, many occupations which are inherently gender-neutral exhibit a gender bias in their usage across different corpora. We propose a method to quantify interactions between concepts in collections of data using autoregressive language models. Instead of prompting or extracting internal representations (i.e. embeddings), we measure differences in token probability distributions on minimal pairs of subject-verb-complements sentences, where the main noun in the subject varies across different 'concept-words'. We demonstrate the expressivity of this approach with a case study where we recover well-known social biases in language from contemporary internet data: ranking concepts in relation to men-women reveals gender biases in what would be expected to be gender neutral vocabulary. We also show how our method can be used to examine concepts across different domains –time periods, authors or ideologies– by selecting sentences from different corpora (e.g. news articles from the 19th century). We extend this to arbitrary conceptual 'dimensions', enabling the study of concepts in terms of other concepts with state-of-the-art language models.