- Main
Dissecting the Neuro-Cognitive Dynamics of Vision-Language Models
Abstract
Vision-Language Models (VLMs) have achieved remarkable success in multimodal reasoning, yet the extent to which their internal representations mirror human neural dynamics remains an open scientific question. In this study, we propose the Component-Specific Residual Disentanglement (CSRD) framework to systematically evaluate the alignment between 17 state-of-the-art VLMs and human fMRI activity. Unlike traditional layer-wise approaches, CSRD disentangles the contributions of Attention and Feed-Forward Networks (FFNs) by selectively retaining or stripping residual connections, thereby isolating information integration from functional transformation. Leveraging the Natural Scenes Dataset (NSD), our analysis reveals a hierarchical "ascent" in brain-model alignment, but with a critical functional divergence in deep layers: Attention modules act as "cognitive hubs" maintaining high alignment with the human visual cortex, whereas FFNs progressively decouple from biological signals to prioritize task-specific predictions. Furthermore, we find that instruction-tuned generative architectures significantly outperform contrastive baselines (e.g., CLIP) and scaling-law-based expectations. These findings suggest that the generative objective fosters more biologically plausible representational structures, providing new theoretical guidance for designing brain-inspired multimodal systems.