Abstract
<p>Humans are remarkably adept at transferring knowledge between experiences, and computational models offer insight into the various strategies humans may employ to optimize behavior. Reinforcement learning models are usually constructed to maximize rewards for a given task but often fail to generalize when the optimal policy changes substantially across contexts. Here, we test empirical predictions of a recent computational strategy based on reward-predictive representations (RPR; Lehnert et. al., 2020). RPRs compress complex environments into simpler representations by discovering abstractions of analogous events that may be superficially distinct but are equivalent in their ability to predict rewarding events. In contrast to most RL models, RPRs afford transfer to new settings even when both the rewards and the transitions between states change. To test whether humans could learn representations akin to RPRs, participants played a novel task where they learned through trial and error how to unlock a safe's keypad to collect diamonds. Uninstructed to the participants, the combinations contained a hidden rule in which some objects were analogous in their predictive transitions to successive rewarding states. We then tested human's ability to quickly adapt to rapidly changing environments that preserve this hidden rule when rewards and transitions both change. In line with RPR predictions, participants showed evidence of learning analogous structure and reapplying the corresponding abstractions to adapt when goals abruptly changed. Furthermore, we found evidence that RPRs were more robustly deployed after a time delay, supporting a role for memory mechanisms in promoting policy-agnostic abstractions in RL tasks.</p>