Decoding Free Viewing: Using Vision-Language Models to Reveal the Optimality of Human Eye Movements for Scene Understanding
Skip to main content
eScholarship
Open Access Publications from the University of California

UC Santa Barbara

UC Santa Barbara Electronic Theses and Dissertations bannerUC Santa Barbara

Decoding Free Viewing: Using Vision-Language Models to Reveal the Optimality of Human Eye Movements for Scene Understanding

Abstract

Eye movements are an important part of human vision. We constantly make them, directing our foveal processing to different locations. Research has shown eye movements to optimize accuracy in various tasks, such as search and navigation, implying that human eye movements are an active process, constantly seeking information that could aid decision-making. But what happens when there is no task to guide eye movements (free viewing task)? Is there any purpose to those eye movements? A prevalent idea is that during free viewing, humans make eye movements to low-level visually salient regions. However, decades of research have shown that humans fixate on people in scenes, on objects, on gaze, and on text, and, more recently, on local regions judged to be meaningful. The goal of this thesis is to understand which perceptual tasks humans engage in during free viewing and to assess whether free-viewing eye movements reflect the goal of optimizing task accuracy (optimal eye movements). Based on the findings, I propose that an important, general human default task during free viewing is to comprehend scenes, and that humans plan near-optimal eye movements, actively seeking information that maximizes comprehension. The first part of this thesis focuses on experimental evidence from eye-movement measurements suggesting that scene comprehension is an important default task for humans during free viewing. Eye movements were measured under four task instructions: free viewing, scene description, object search, and counting objects. The stimuli were image pairs, called Winograd images, containing minor visual alterations that drastically change scene interpretation without altering low-level saliency, thereby isolating the semantic factors guiding eye movements. Results indicate that free-viewing fixations closely resemble those of observers explicitly instructed to describe a scene, differing significantly from fixations during object search or counting. Furthermore, free-viewing fixations are disproportionately directed toward people and objects most critical to understanding the scene (objects that maximally impact scene descriptions when removed), rather than solely toward low-level visual saliency or locally meaningful regions. I also show that human fixations on these critical elements improve their understanding of the scene, demonstrating a causal influence of these fixations in accurate scene comprehension. The second part of this thesis evaluates whether human fixations during free viewing approximate the optimal strategy for maximizing scene comprehension. The challenge in this objective is to create a model that estimates optimal fixations for real-world scenes and high-level goals such as scene comprehension. Classical methods based on Bayesian ideal searchers have a strong mathematical foundation but cannot be applied to real-world scenes. I implemented a model that simulates human foveated vision and sequential eye-movement exploration of scenes. Leveraging state-of-the-art vision-language models (VLMs) capable of human-level scene comprehension, the model generated descriptions of the actively explored foveated scenes. I subsequently trained a reinforcement learning (RL) agent that uses a convolutional neural network (CNN) to optimize visual exploration (Q-network), systematically executing eye movements to maximize the semantic accuracy of the VLM’s descriptions at each step. By measuring free-viewing eye movements on images carefully curated to depict complex social interactions, actions, or implied actions, I categorized the elements present in these images and counted the human fixation frequencies for these categories (people, objects relevant and irrelevant to the understanding of the scene, text, gazed and grasped objects, salient regions). The optimized RL agent matched the fixation patterns observed in human data without any prior training on human fixation data, significantly outperforming the same RL model optimized for search or image classification tasks, as well as saliency prediction models. Furthermore, I found that VLM descriptions generated by simulating foveation at human-fixated locations, recorded during the free-viewing task, achieved semantic accuracy comparable to that obtained via fixations from an RL agent specifically optimized for scene understanding. The agreement between the RL model and human fixation frequencies also decreased when the model was trained with very low or extreme foveation, suggesting that human free viewing eye movements are an emergent property of an interaction between the goal to optimize scene comprehension and the specific foveated properties of the human visual system. I assessed the generalization of the RL model's results to a recently published data set on eye movements from 6720 observers aged 5-72. Together, this research demonstrates that eye movements during free viewing are an actively optimized, task-driven behavior aimed at comprehending the visual world. Fixations to people, text, objects relevant to understanding, and gazed/grasped objects are an emergent property of an interaction between the goal of optimizing scene understanding and the foveated properties of human sensory processing.