- Main
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
- Langenfeld, Cassidy Marie;
- Beger, Claas;
- Geng, Gloria;
- Piriyakulkij, Top;
- Hu, Keya;
- Pu, Yewen;
- Ellis, Kevin
Abstract
Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual rea- soning abilities in radically new situations – a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parame- ter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.