Skip to main content
Download PDF
- Main
Why Multimodal Models Struggle with Spatial Reasoning: Insights from Human Cognition
© 2025 by the author(s). Learn more.
Abstract
Multimodal models excel in tasks requiring semantic integra- tion of language and vision but struggle with spatial cognition. Using a visual perspective-taking task inspired by cognitive science, we find these models fail when the image and ref- erence view differ, reflecting spatial cognition comparable to a two-year-old child. To explore these disparities further, we analyze internal representations using a human action fMRI dataset and voxelwise encoding models, revealing key differ- ences between AI and human spatial encoding. This work pro- vides new benchmarks and insights into bridging artificial and biological cognition.