- Main
Multimodal Pragmatic Inference in Vision-Language Transformers
Abstract
Contemporary transformer models have achieved human-like performance on many text-based tasks. However, real-world communication requires the integration of language with non-linguistic context (e.g., visual, social, etc.). Here, we study such information integration in three multimodal transformer models. We test these models' pragmatic capabilities regarding referring expressions: when an object set contains two exemplars from the same category that differ in size, unambiguously referring to one of them requires a size adjective (e.g., the big hammer); the adjective is unnecessary if only one exemplar from the category is present. We evaluate these inferences when models process text-image inputs (via their surprisal for infelicitous vs. felicitous adjective use) and when they generate open-ended descriptions of images given text prompts. We find evidence for pragmatic integration of visual and linguistic context in all models. However, these inferences remain sensitive to the in-context statistics of visual inputs, unlike pragmatic inference in humans.