Skip to main content
eScholarship
Open Access Publications from the University of California

Exploring Causal and Compositional Reasoning in Large Language Models

Creative Commons 'BY' version 4.0 license
Abstract

Large Language Models (LLMs) have shown surprising capabilities in reasoning tasks despite lacking direct physical experience with the world. We examine LLMs' ability to reason about object affordances through a tool innovation task where one must select unconventional objects to replace typical tools. In a study comparing GPT-3.5-turbo and GPT-4o with human participants (N=100), we found that while GPT-3.5 performed significantly worse than humans (38.7% vs. 85.8%), GPT-4o with chain-of-thought prompting achieved human-level performance (85.0%). Qualitative analysis revealed that both models could identify causally relevant object properties, but GPT-4o was superior in flexibly applying these properties in novel contexts. We argue that this success relies on compositional reasoning—the ability to decompose objects into abstract properties and recombine them for novel uses. Our findings suggest that LLMs' ability to reason about object affordances has progressed substantially, highlighting the need for further mechanistic research to characterise LLMs' underlying abilities.