Skip to main content
eScholarship
Open Access Publications from the University of California

Identifying Concepts Used by Human-Like Neural Network Chess Engines

Creative Commons 'BY' version 4.0 license
Abstract

As neural networks reach expert-level performance in many domains, there is growing interest in understanding the internal representations that guide their decisions. Models trained to mimic human decisions, rather than strive for optimal performance, offer a particularly interesting testbed for interpretability. Identifying what information these networks encode can generate hypotheses about what information humans rely on when performing similar tasks. Here, we use concept-based interpretability methods from prior work on superhuman-level chess neural networks to study neural networks trained to emulate human play across differing skill levels. We find that human-interpretable chess concepts are decodable from their latent representations. For many concepts, decodability increases with human-emulated skill level, and also varies with network depth: simpler concepts peak early, whereas complex concepts emerge later. Together, these results show that interpretability tools can be applied to AI systems that mimic human behavior and suggest how task representations may shift with expertise.