Skip to main content
eScholarship
Open Access Publications from the University of California

Visuospatial Perspective Taking in Multimodal Language Models

Creative Commons 'BY' version 4.0 license
Abstract

As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks largely rely on text-based vignettes or static scene understanding, leaving visuospatial perspective-taking (VPT) underexplored. We procedurally generate large stimulus batteries for two evaluation tasks adapted from human studies: the Rotating Figure Task, probing perspective-taking across angular disparities, and the Director Task, assessing VPT in a referential communication paradigm. Across both tasks, MLMs show pronounced deficits in Level 2 VPT, with failure patterns indicating reliance on simple mirroring heuristics rather than genuine perspective transformations. These results expose critical limitations in current MLMs' ability to represent and reason about alternative perspectives, with implications for use in collaborative contexts.