Skip to main content
eScholarship
Open Access Publications from the University of California

Measuring IA Cognitive Competences: A Gap in Current Benchmarks

Creative Commons 'BY' version 4.0 license
Abstract

To what extent do current LLM benchmarks pose the same reasoning challenges that arise in real-world settings? Instead of relying on anthropocentric notions, here we analyze the cognitive competences demanded by a benchmark in terms of (a) formal complexity (e.g. number of inference steps required to reach the desired answer) and (b) amount of contextual interpretation needed to reduce language ambiguity to a level where the required inference can be performed. Our analysis of a representative sample of benchmarks reveals that there are almost no benchmarks with both, high formal complexity and high contextual interpretation requirements. This limitation may help explain the discrepancy between the strong benchmark performance of language models and the shortcomings observed when these models are deployed in real world applications.