- Main
Function Vectors for Relational Reasoning in Multimodal Large Language Models
- Fu, Shuhao
- Advisor(s): Wu, Ying Nian
Abstract
Multimodal large language models (MLLMs) exhibit impressive relational reasoning abilities from limited examples, yet the internal mechanisms supporting such behavior remain opaque. This thesis proposes a causal framework for interpreting and controlling relational behavior in MLLMs by extracting and manipulating function vectors: task-specific representations computed from attention head activations. Extending previous work in language-only models, we demonstrate that function vectors can also be identified in the vision-language model OpenFlamingo-4B and used to induce relational behavior in zero-shot settings. Using a synthetic image dataset designed to isolate spatial relations, we apply causal mediation analysis to identify a small subset of attention heads with high influence on relational predictions. These heads define compact function vectors that, when injected into the model’s hidden states, significantly improve zero-shot accuracy. We further show that these vectors can be fine-tuned—while keeping model parameters fixed—to enhance generalization and outperform in-context learning baselines. Our results reveal that MLLMs encode relational functions within localized internal structures, which can be systematically interpreted and optimized, advancing our understanding of model modularity and control.