Back to Search View Original Cite This Article

Abstract

<jats:p>Reinforcement learning is a key means by which animals learn appropriate actions in a given context. A large body of work suggests that such learning depends on interactions between cortico-basal ganglia circuits and the midbrain dopaminergic system, yet the underlying circuit mechanisms and plasticity rules are not fully understood. Here we present a biologically plausible, multi-region neural circuit model of songbird vocal learning and map it onto the actor-critic framework of reinforcement learning. In this model, stochastic spiking activity in the cortico-basal ganglia pathway implements action selection and drives behavioral exploration, while the pathways driving midbrain dopaminergic signaling evaluate behavioral outcomes and support a reward prediction error based learning rule that approximates stochastic gradient ascent. The model achieves millisecond-scale precise learning that matches observed behavior. We further use the model to examine two fundamental constraints on biological reinforcement learning. First, dopaminergic reinforcement signals are temporally imprecise, which can cause interference between neurons controlling actions that occur close in time. Second, dopaminergic signals are spatially imprecise, which can cause interference between neurons controlling different aspects of behavior but receiving a common reinforcement signal. By jointly modeling the actor and critic components of the circuit, we show that fast updating of reward prediction is crucial for precise and efficient learning under both forms of interference, and the model predicts the experimentally observed timescale of reward prediction updating. These results suggest a circuit-level mechanism by which biological systems achieve reinforcement learning despite the temporal and spatial limitations of global neuromodulatory signals.</jats:p>

Show More

Keywords

learning reinforcement model which dopaminergic

Related Articles

PORE

About

Connect