- Main
Vision-Language Model through the lens of Linguistic Relativity
Abstract
Linguistic relativity proposes that the categorical structure of a language systematically biases how continuous perceptual domains are partitioned. While this effect is well studied in human color perception, it remains unclear whether vision–language models (VLMs) exhibit a language-conditioned categorical structure analogous to it. In this work, we treat VLM as an \emph{artificial subject} and adapt psycholinguistic paradigms to test for Whorfian effects in a multimodal neural network. Using a tightly controlled set of color-discrimination triads that vary only in OKLab lightness, we quantify model decisions across English and Russian lexical frames. For a contrastive dual-encoder VLM, we observe a robust shift in the internal decision boundary between light and dark blue when the language of the textual input changes, despite identical visual input. This boundary displacement disappears in vision-only embeddings, demonstrating that the effect is not perceptual in origin but emerges from vision–language alignment. In contrast, when languages share the same color lexicon (e.g., English and French), category boundaries largely overlap, and when generic labels are used, no stable boundary emerges. Together, these findings show that a computational analogue of linguistic relativity can arise in artificial systems, but only when visual representations are explicitly aligned with language. Our results establish a framework for testing cognitive theories in multimodal models.