Abstract
<p>Theoretical accounts of decision-making distinguish between two reinforcement-learning strategies: a model-based system that supports goal-directed choices through internal representations of action–outcome contingencies, and a model-free system that updates action values based on recent rewards, often approximating a win–stay/lose–shift (WSLS) pattern. While previous studies have characterized these mechanisms, their developmental trajectory remains unclear. Prevailing accounts attribute adolescents’ poorer performance in probabilistic learning to increased exploration and overestimation of environmental volatility, based on Bayesian models assuming a shared model-based strategy across age groups but differing latent parameters. Such approaches may overlook model-free contributions and individual variability, particularly during adolescence. Here, 36 adolescents and 67 adults completed a probabilistic reinforcement-learning task with binary choices and asymmetric reward contingencies (70:30), presented in sequential blocks to dissociate learning strategies. Individual behavior was fit with competing models: (I) a random baseline, (II) a WSLS heuristic, and (III) a hierarchical reinforcement-learning model implementing Bayesian adaptive learning. The results showed that both strategies contributed to behavior in each group, but in different proportions: model-based strategy was evident in only a subset of adolescents and became predominant in adults. Critically, individuals classified as model-based showed comparable latent parameters (e.g., learning rate, model strength, inverse temperature, precision) across age groups, and WSLS behavior also did not differ between adolescents and adults. Adolescents exhibiting model-based behavior had higher IQ but did not differ in age from their peers. Our findings indicate that in stable probabilistic environments, developmental differences in decision-making between adolescents and adults reflect variability in the engagement of model-based versus model-free strategies rather than differences in underlying computational parameters, with both groups behaving similarly when using the same strategy.</p>