Skip to main content
eScholarship
Open Access Publications from the University of California

Verb Semantic Reasoning: a Semantics-Guided Approach for Improving Action Understanding in Vision--Language Models

Creative Commons 'BY' version 4.0 license
Abstract

Verbs are pivotal in human language and cognition. Psycholinguistic research has developed explicit, structured accounts of verb semantics, yet these theories have rarely been leveraged to improve models' visual action understanding. Meanwhile, current vision–language models (VLMs) often struggle with action semantics and exhibit unstable, weakly grounded judgments. Therefore, we introduce Verb Semantic Reasoning (VSR), a two-stage pipeline in which a Large Language Model (LLM) converts candidate action descriptions into structured event-semantic representations and a chain of semantic components questions; a VLM answers these questions over videos to select the best-supported action description. Results indicate that human performance was near ceiling, whereas both VLM baselines were substantially lower; VSR consistently improved accuracy across models and narrowed the human–model gap. These results suggest that explicit verb-semantic reasoning can significantly improve the accuracy of VLMs' action judgments, underscoring compositional, symbolic representations as an essential intermediate step for extracting semantics from the linguistic knowledge encoded in LLMs and leveraging it to guide VLMs' visual understanding.