- Main
Assessing Perceptual Metacognition in Vision-Language Models
Abstract
Vision-language models (VLMs) have demonstrated unprecedented capabilities in perception and reasoning, yet their perceptual metacognitive abilities remain unexamined, a gap with implications for reliable deployment and welfare considerations. We evaluated six open-source VLMs on perceptual decision-making tasks, comparing implicit (logit-based) and explicit (self-reported) metacognition. Models exhibited robust implicit calibration and sensitivity on high-performance tasks but unreliable explicit confidence reports with poor calibration and task-dependent sensitivity. Process modeling further revealed that both confidence measures converged on classic Signal Detection Theory, suggesting VLMs lack the metacognitive monitoring mechanisms that differentiate decision and confidence processes in humans. Together, we showed that current VLMs thus limited perceptual metacognition with notable dissociations between implicit and explicit confidence.