- Main
Separable contributions of value and choice to policy learning
Abstract
Reinforcement learning (RL) models have been highly successful in describing human behavior. However, in most experiments, reward history is highly correlated with choice history. Consequently, behavior could be partially driven by a value-free habitual learning (HL) mechanism but be misinterpreted as RL. Do people select actions they have chosen frequently in the past (HL), or actions that have been rewarding in the past (RL)? To disentangle the influence of value-based RL and value-free HL, we designed two experiments that unconfound reward and choice histories. Behavioral and modeling results suggest that both value-based and value-free mechanisms contribute to test phase choice. Their relative contributions differ in the two experiments, suggesting that different environments may recruit the two processes to different degrees. Our results highlight the existence of partially redundant processes that are often difficult to disambiguate, and call for careful consideration of underlying processes when interpreting experimental results.