Skip to main content
eScholarship
Open Access Publications from the University of California

Estimating Task Representations in a Multidimensional Bandit Task: A Method for Revealing Meta-level Learning

Creative Commons 'BY' version 4.0 license
Abstract

Humans must learn not only which options pay off, but also which features are worth learning about. Yet such meta-level task representations are difficult to observe directly. We propose an inverse inference approach for estimating an individual's hypothesis space over feature dimensions in a multidimensional bandit task. Building on a model previously supported when the hypothesis space is known, we instantiate the model for each candidate hypothesis space and identify the space that best predicts each participant's choice sequence using hierarchical Bayesian estimation and WAIC-based model comparison. In an online experiment, participants' choice accuracy improved both within games and across games, suggesting adaptation at multiple timescales. Moreover, the inferred spaces shifted on average toward the true task structure. These results indicate that changes in task representations are associated with improvements in structured reward learning, and provide a measurement framework for developing process-level models of meta-level learning.