Skip to main content
eScholarship
Open Access Publications from the University of California

Semantic bias in image-text matching in humans versus vision-language pretraining AI models

Creative Commons 'BY' version 4.0 license
Abstract

Recent research has shown that vision-language pretrain-ing (VLP) models using contrastive learning (Contrastive Language Image Pretraining, CLIP) has a semantic bias to-wards using concrete words during image-text matching. Here we showed that as compared with CLIP, humans at-tended more to abstract words during image-text matching. This difference likely results from CLIP's difficulty in de-veloping grounded understanding of abstract concepts and capturing contextual dependencies between images and captions through contrastive learning. While CLIP's caption attention aligned more closely with humans for concrete than for abstract captions, their alignment in im-age attention did not differ between the caption condi-tions, as human image attention was driven primarily by individual differences in explorative or focused perceptual style. Our findings thus revealed important differences in information processing mechanisms between humans and the VLP models, with important implications for not only human and AI image-text matching, but also for potential ethical issues resulting from such misalignment.