Abstract
<title>Abstract</title> <p>Addiction has been widely viewed as an example of reinforcement learning gone awry, often attributed to impaired model-based (MB) control and a compensatory shift toward model-free (MF) strategies. However, such “MB-MF” dichotomy is largely grounded in single-task settings and may not capture how learning operates in environments that require generalization across multiple tasks. Here, we examine learning in individuals with substance use disorders (SUD) using a two-step, multi-task paradigm in which participants navigate a tree-structured environment to achieve goals with varying reward contingencies. Our results show that, across groups, behavior is not well explained by canonical MB or MF accounts. Instead, participants acquire successor features that serve as reusable components for generalization across tasks. Leveraging information theory, we develop a resource-rational model that balances expected reward and cognitive cost in multi-task learning. This model successfully reproduced several characteristic behaviors in both groups. Critically, individuals with SUD show increased sensitivity to cognitive cost and preferentially adopt simplified policies and reuse them across tasks, at the expense of reward maximization. Individual differences in cost sensitivity and inferred cognitive effort predict addiction severity. These findings challenge the dominant MB–MF dichotomy of addiction and instead point to a systematic shift in how cognitive resources are allocated during learning. More broadly, our findings highlight resource-rational generalization as a key principle for understanding atypical behavior in complex, multi-task environments.</p>