- Main
Individual Differences in Human Teaching of Reinforcement Learning Agents: Evidence from Bayesian Hypothesis Testing
Abstract
When humans teach reinforcement learning (RL) agents through real-time interventions, does their teaching strategy adapt to environmental context or reflect systematic individual variation? We present a novel web-based platform for studying human teaching via state interventions (i.e.: physical relocation of learning agents). In a study with 82 participants, we manipulated world dynamics and how agents interpreted interventions using six formally-defined interpretation types. Bayesian analyses provided moderate-to-strong evidence that neither factor affected teaching behavior (BF01 = 5.22 for world type; BF01 = 27.08 for interpretation type). However, participants showed systematic individual variation: high-frequency teachers (n = 39; M = 36 interventions/round) targeted actions with lower Q-values compared to low-frequency teachers (d = -0.57, BF10 = 3.69). While frequent interventions boosted immediate scores, they often impaired long-term policy learning. The reset interpretation, which ignores the intervention, was an exception: it was uniquely robust to suboptimal human guidance. These findings suggest that human teaching in this task is characterized more by persistent individual behavioral profiles than by context-adaptive strategies, with implications for designing resilient human-AI interaction systems.