- Main
Quantifying Semantic Priors vs Visual Evidence in Visual Language Models via Psychometric Curve Analysis
Abstract
Visual illusions are a classic tool in cognitive science for probing perceptual inference. Recent studies suggest that when visual facts conflict with semantic prior knowledge, vision-language models (VLMs) often make systematic errors in which priors override visual evidence. To quantify this effect, we treat VLMs as artificial participants and introduce IlluQuant, an illusion-diagnostic dataset that operationalizes template similarity and visual evidence strength as continuous, controllable variables. By fitting psychometric functions, we measure semantic prior dominance in VLMs. Results show that stronger semantic cues substantially increase the probability of prior-driven errors, while visual evidence must be amplified several-fold to partially offset this bias. We further propose MdCoT, an intervention that mitigates the effect. IlluQuant provides a psychophysics-style framework for quantitatively characterizing the balance between semantic priors and visual evidence in multimodal models under controlled illusion manipulations. The IlluQuant dataset is available at haitoooo/IlluQuant.