- Main
Backward Digit Span Benchmarks Working Memory in LLMs
Abstract
The maintenance and manipulation of information in working memory (WM) is fundamental to human and artificial intelligence. Although human WM is famously capacity-limited, Large Language Models (LLMs) preserve inputs in the network's context window, profoundly reducing capacity limits that depend on maintenance alone. However, work in cognitive science suggests WM limits can also arise from representational interference during manipulation, raising the possibility of strong limits even in LLMs with perfect maintenance. Here, we evaluate 15 frontier LLMs and find models perform near-perfectly on forward digit span (recalling sequences in order) but collapse on backward digit span (recalling sequences in reverse). Moreover, backward span performance is significantly correlated with two measures of fluid intelligence (Raven's Progressive Matrices and ARC-AGI-1). These findings suggest backward digit span provides a plausible benchmark for the "working" component of working memory in LLMs and point to potential shared principles of information processing across systems.