<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neural Circuits</journal-id>
<journal-title>Frontiers in Neural Circuits</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neural Circuits</abbrev-journal-title>
<issn pub-type="epub">1662-5110</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncir.2018.00116</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Minimal Circuit Model of Reward Prediction Error Computations and Effects of Nicotinic Modulations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Deperrois</surname> <given-names>Nicolas</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/541451/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Moiseeva</surname> <given-names>Victoria</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Gutkin</surname> <given-names>Boris</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/304/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Group for Neural Theory, LNC2 INSERM U960, DEC, &#x000C9;cole Normale Sup&#x000E9;rieure PSL&#x0002A; University</institution>, <addr-line>Paris</addr-line>, <country>France</country></aff>
<aff id="aff2"><sup>2</sup><institution>Center for Cognition and Decision Making, Institute for Cognitive Neuroscience, National Research University Higher School of Economics</institution>, <addr-line>Moscow</addr-line>, <country>Russia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Anita Disney, Vanderbilt University, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Kenji Morita, The University of Tokyo, Japan; Kasia M. Bieszczad, Rutgers, The State University of New Jersey, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Boris Gutkin <email>boris.gutkin&#x00040;ens.fr</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>01</month>
<year>2019</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>12</volume>
<elocation-id>116</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>09</month>
<year>2018</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>12</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2019 Deperrois, Moiseeva and Gutkin.</copyright-statement>
<copyright-year>2019</copyright-year>
<copyright-holder>Deperrois, Moiseeva and Gutkin</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Dopamine (DA) neurons in the ventral tegmental area (VTA) are thought to encode reward prediction errors (RPE) by comparing actual and expected rewards. In recent years, much work has been done to identify how the brain uses and computes this signal. While several lines of evidence suggest the interplay of the DA and the inhibitory interneurons in the VTA implements the RPE computation, it still remains unclear how the DA neurons learn key quantities, for example the amplitude and the timing of primary rewards during conditioning tasks. Furthermore, endogenous acetylcholine and exogenous nicotine, also likely affect these computations by acting on both VTA DA and GABA (&#x003B3; -aminobutyric acid) neurons via nicotinic-acetylcholine receptors (nAChRs). To explore the potential circuit-level mechanisms for RPE computations during classical-conditioning tasks, we developed a minimal computational model of the VTA circuitry. The model was designed to account for several reward-related properties of VTA afferents and recent findings on VTA GABA neuron dynamics during conditioning. With our minimal model, we showed that the RPE can be learned by a two-speed process computing reward timing and magnitude. By including models of nAChR-mediated currents in the VTA DA-GABA circuit, we showed that nicotine should reduce the acetylcholine action on the VTA GABA neurons by receptor desensitization and potentially boost DA responses to reward-related signals in a non-trivial manner. Together, our results delineate the mechanisms by which RPE are computed in the brain, and suggest a hypothesis on nicotine-mediated effects on reward-related perception and decision-making.</p></abstract>
<kwd-group>
<kwd>dopamine</kwd>
<kwd>reward-prediction error</kwd>
<kwd>ventral tegmental area</kwd>
<kwd>acetylcholine</kwd>
<kwd>nicotine</kwd>
</kwd-group>
<contract-num rid="cn001">ANR-10-LABX-0087</contract-num>
<contract-sponsor id="cn001">Agence Nationale de la Recherche<named-content content-type="fundref-id">10.13039/501100001665</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="0"/>
<equation-count count="14"/>
<ref-count count="63"/>
<page-count count="17"/>
<word-count count="13199"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>To adapt to their environment, animals constantly compare their predictions with new environmental outcomes (rewards, punishments, etc.). The difference between prediction and outcome is the prediction error, which in turn can serve as a teaching signal to allow the animal to update its predictions and render previously neutral stimuli predictive of rewards into reinforcers of behavior. Particularly, the dopamine (DA) neuron activity in the Ventral Tegmental Area (VTA) have been shown to encode the reward prediction error (RPE), or the difference between the actual reward the animal receives and the expected reward (Schultz et al., <xref ref-type="bibr" rid="B50">1997</xref>; Schultz, <xref ref-type="bibr" rid="B49">1998</xref>; Bayer and Glimcher, <xref ref-type="bibr" rid="B2">2005</xref>; Day and Carelli, <xref ref-type="bibr" rid="B8">2007</xref>; Matsumoto and Hikosaka, <xref ref-type="bibr" rid="B32">2009</xref>; Enomoto et al., <xref ref-type="bibr" rid="B12">2011</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>). During, for example, classical conditioning with appetitive rewards, unexpected rewards elicit strong transient increases in VTA DA neuron activity, but as a cue fully predicts the reward, the same reward produces little or no DA neurons response. Finally, after learning, if the reward is omitted, DA neurons pause their firing at the moment reward is expected (Schultz et al., <xref ref-type="bibr" rid="B50">1997</xref>; Schultz, <xref ref-type="bibr" rid="B49">1998</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>; Watabe-Uchida et al., <xref ref-type="bibr" rid="B57">2017</xref>). Thus DA neurons should either receive or compute the RPE. While several lines of evidence have pointed toward the RPE being computed by the VTA local circuitry, exactly how this is done vis-a-vis the inputs and how this computation is modulated by the endogenous acetylcholine and the endogenous substances that affect the VTA, e.g., nicotine, remains to be defined. Here we proceed to address these questions using a minimal computational modeling methodology.</p>
<p>In order to compute the RPE, the VTA should receive the relevant information from its inputs. Intuitively, distinct biological inputs to the VTA must differentially encode actual and expected rewards that are finally subtracted by a downstream target, the VTA DA neurons. For the last two decades, a great amount of experimental studies depicted which brain areas send this information to the VTA. Notably, a subpopulation of pedunculopontine tegmental nucleus (PPTg) has been found to send the actual reward signal to dopamine neurons (Kobayashi and Okada, <xref ref-type="bibr" rid="B25">2007</xref>; Okada et al., <xref ref-type="bibr" rid="B37">2009</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>), while other studies showed that the prefrontal cortex (PFC) and the nucleus accumbens (NAc) respond to the predictive cue (Funahashi, <xref ref-type="bibr" rid="B18">2006</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>; Oyama et al., <xref ref-type="bibr" rid="B39">2015</xref>; Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>; Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>), highly depending on VTA DA feedback projections in the PFC (Puig et al., <xref ref-type="bibr" rid="B45">2014</xref>; Popescu et al., <xref ref-type="bibr" rid="B44">2016</xref>) and the NAc (Yagishita et al., <xref ref-type="bibr" rid="B61">2014</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>; Fisher et al., <xref ref-type="bibr" rid="B17">2017</xref>). However, how each of these signals are integrated by VTA DA neurons during classical-conditioning remains elusive.</p>
<p>Recently, VTA GABA neurons were shown to encode reward expectation with a persistent cue response proportional to the expected reward (Cohen et al., <xref ref-type="bibr" rid="B4">2012</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>; Tian et al., <xref ref-type="bibr" rid="B53">2016</xref>). Additionally, selectively exciting and inhibiting VTA GABA neurons during a classical-conditioning task, Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) revealed that these neurons are likely source of the substraction operation, contributing to the inhibitory expectation signal in the RPE computation by DA neurons.</p>
<p>Furthermore, the presence of nicotinic acetylcholine receptors (nAChRs) in the VTA (Pontieri et al., <xref ref-type="bibr" rid="B42">1996</xref>; Maskos et al., <xref ref-type="bibr" rid="B31">2005</xref>; Changeux, <xref ref-type="bibr" rid="B3">2010</xref>; Faure et al., <xref ref-type="bibr" rid="B15">2014</xref>) provides a potential common route for acetylcholine (ACh) and nicotine (Nic) in modulating dopamine activity during a Pavlovian-conditioning task.</p>
<p>Particularly, the high-affinity &#x003B1;4&#x003B2;2 subunit-containing nAChRs desensitizing relatively slowly (&#x02243; sec) and located post-synaptically on VTA DA and GABA neurons have been shown to have the most prominent role in nicotine-induced DAergic bursting activity and self-administration, as suggested by mouse knock-out experiments (Maskos et al., <xref ref-type="bibr" rid="B31">2005</xref>; Changeux, <xref ref-type="bibr" rid="B3">2010</xref>; Faure et al., <xref ref-type="bibr" rid="B15">2014</xref>) and recent direct optogenetic modulation of these somatic receptors (Durand-de Cuttoli et al., <xref ref-type="bibr" rid="B10">2018</xref>).</p>
<p>We have previously developed and validated a population level circuit dynamics model (Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>; Tolu et al., <xref ref-type="bibr" rid="B55">2013</xref>; Maex et al., <xref ref-type="bibr" rid="B29">2014</xref>; Dumont et al., <xref ref-type="bibr" rid="B9">2018</xref>) of the influence nicotine and Ach interplay may have on the VTA dopamine cell activity. Using this model we showed that Nic action on &#x003B1;4&#x003B2;2 could result in either direct stimulation or disinhibition of DA neurons. The latter scenario suggests that relatively low nicotine concentrations (&#x0007E;500 nM) during and after smoking preferentially desensitize &#x003B1;4&#x003B2;2 nAChRs on GABA neurons (Fiorillo et al., <xref ref-type="bibr" rid="B16">2008</xref>). The endogenous cholinergic drive to GABA neurons would then decrease, resulting in decreased GABA neurons activity, and finally a disinhibition of DA neurons as confirmed <italic>in vitro</italic> (Mansvelder et al., <xref ref-type="bibr" rid="B30">2002</xref>) and suggested by Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>), Tolu et al. (<xref ref-type="bibr" rid="B55">2013</xref>), Maex et al. (<xref ref-type="bibr" rid="B29">2014</xref>), and Dumont et al. (<xref ref-type="bibr" rid="B9">2018</xref>) modeling work. Interestingly, this scenario requires that the high affinity nAChRs are in a pre-activated state, so that nicotine can desensitize them, which in turn implies a sufficiently high ambient cholinergic tone in the VTA. However, when the ACh tone is not sufficient, in this GABA-nAChR scenario, nicotine would lead to a significant inhibition of the DA neurons. Furthermore, a recent study showed that optogenetic inhibition of PPTg cholinergic fibers inhibit only the VTA non-DA neurons (Yau et al., <xref ref-type="bibr" rid="B62">2016</xref>), suggesting that ACh acts preferentially on VTA GABA neurons. However, the effects of Nic and ACh on dopamine responses to rewards via &#x003B1;4&#x003B2;2-nAChRs desensitization during classical-conditioning have remained elusive.</p>
<p>In addition to the above issues, a non-trivial issue arises from the timing structure of the conditioning tasks. Typically, the reward to be consumed is delivered after a temporal delay past the conditioning cue, which begs important related questions: how is the reward information transferred from the reward-delivery time to the earlier reward-predictive stimulus and how does the brain compute the precise timing of reward? In other words, how is the relative co-timing of the reward and the reinforcer learned in the brain? These issues generate further lines of enquiry on how this learning process may be altered by nicotine. In order to start clarifying the possible neural mechanisms underlying the observed RPE-like activity in DA neurons, we propose here a simple neuro-computational model inspired from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>), incorporating the mean dynamics of four neuron populations: the prefrontal cortex (PFC), the pedunculopontine tegmental nucleus (PPTg), the VTA dopamine and GABA neurons.</p>
<p>Note that we explicitly choose to base our model on the desensitization scenario from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>), where the nicotinic receptors are relatively efficient in controlling the GABA neuron populations activity. In this case, the positive dopamine response to nicotine is due to &#x003B1;4&#x003B2;2-nAChRs desensitization and requires a relatively high endogenous cholinergic tone-for low acetylcholine tone, nicotine is predicted to depress DA output in this scheme. Since the animal is performing experimental tasks in a state of cognitive effort, the disinhibition scenario we surmise could be relevant as it implies a high cholinergic tone impinging onto the VTA (Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>, <xref ref-type="bibr" rid="B41">2012</xref>).</p>
<p>Taking into account recent neurobiological data, particularly showing the activity of VTA GABA neurons during classical-conditioning (Cohen et al., <xref ref-type="bibr" rid="B4">2012</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>), we qualitatively and quantitively reproduce several aspects of a Pavlovian-conditioning task&#x02014;which we take as a paradigmatic example of reward-based conditioning&#x02014;such as the phasic components of dopaminergic activation with respect to reward magnitude, omission and timing, the working-memory activity in the PFC, the response of the PPTg to primary rewards, and the dopamine-induced plasticity in cortical and corticostriatal synapses.</p>
<p>Having built the minimal model that incorporates the influence of nAChRs on the computations of reward-related learning signals in the VTA circuit, we are poised to use the model to examine how acute nicotine may affect this computation. Notably, we qualitatively assessed the potential effects of nicotine-induced desensitization of &#x003B1;4&#x003B2;2-nAChRs on GABA neurons, leading to a disinhibition of DA burst-response to rewarding events. As we will show below, this effect would lead to pathological changes in evaluation of rewards and stimuli associated with nicotine and lead to a bias in boosting strong vs. weak rewards as observed recently experimentally. These last simulations imply an important role for nicotine in not only provoking a positive over-valuation of acute nicotine itself, but also in having an impact on the general rewarding quality of nicotine-associated environments. Additionally, our simulations also imply a heightened reward sensitivity in animals exposed to nicotine. We further analyze the potential behavioral and motivational implications of these predicted effects in the Discussion section.</p>
</sec>
<sec id="s2">
<title>2. Methods: Computational Model and Simulated Behavioral Tasks</title>
<p>In order to examine the VTA circuit level mechanisms of reward prediction error computation and effects of nicotine on this activity during classical-conditioning, we built a neural population model of the VTA and its afferent inputs inspired from the mean-field approach of Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>). This model incorporates the DA and GABA neuronal populations in the VTA and their glutamatergic and cholinergic afferents from the PFC and the PPTg (Figure <xref ref-type="fig" rid="F1">1</xref>). Based on recent neurobiological data, we propose a model for the activity of the PFC and PPTg inputs during classical-conditioning contributing to the observed VTA GABA and DA activity. Additionally, the activation and desensitization dynamics of the nAChR-mediated currents in response to Nic and ACh were described by a 4-state model taken from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Illustration of the VTA circuit and neural dynamics of each area during learning of a pavlovian-conditioning task. Afferents inputs and circuitry of the ventral tegmental area (VTA). The GABA neuron population (red) inhibits locally the DA neuron population (green). This local circuit receives excitatory glutamatergic input (blue axons) from the corticostriatal pathway and the pedunculopontine tegmental nucleus (PPTg). The PPTg furthermore furnishes cholinergic projections (purple axon) to the VTA neurons (&#x003B1;4&#x003B2;2 nAChRs). <italic>r</italic> is the parameter to change continuously the dominant site of &#x003B1;4&#x003B2;2 nAChR action. Dopaminergic efferents (green axon) project, amongst others, to the nucleus accumbens (NAc) and the prefrontal cortex (PFC) and modulates cortico-striatal projections <italic>w</italic><sub>PFC-D</sub> and <italic>w</italic><sub>PFC-G</sub> and PFC recurrent excitation <italic>J</italic><sub>PFC</sub> weights. The PFC integrates CS (tone) information, while the PPTg respond phasically to the water reward itself (US). Dopamine and acetylcholine outflows are represented by green and purple shaded areas, respectively. All parameters and description are summarized in Supplementary Table <xref ref-type="supplementary-material" rid="SM1">1</xref>.</p></caption>
<graphic xlink:href="fncir-12-00116-g0001.tif"/>
</fig>
<sec>
<title>2.1. Mean-Field Description of VTA Neurons and Their Afferents</title>
<p>First, the model from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) describing the dynamics of VTA neuron populations and the effects of Nic and ACh on nAChRs was re-implemented with several quantitative modifications according to experimental data.</p>
<p>The temporal dynamics of the average activities of DA and GABA neurons in the VTA taken from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) are described by the following equations:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mtext>D</mml:mtext></mml:msub><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mtext>D</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mtext>D</mml:mtext></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mtext>D</mml:mtext></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>G-D</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>Glu-D</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>r</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mn>4</mml:mn><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mtext>G</mml:mtext></mml:msub><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mtext>G</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mtext>G</mml:mtext></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x003A6;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mtext>G</mml:mtext></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>Glu-G</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mn>4</mml:mn><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BD;<sub>D</sub> and &#x003BD;<sub>G</sub> are the mean firing rates of the DA and GABAergic neuron populations, respectively. &#x003C4;<sub>D</sub> &#x0003D; 30 ms and &#x003C4;<sub>G</sub> &#x0003D; 30 ms are the membrane time constants of both neuron populations specifying how quickly the neurons integrate input changes. <italic>I</italic><sub>Glu</sub> characterize the excitatory inputs from PFC and PPTg mediated by glutamate receptors. <italic>I</italic><sub>&#x003B1;4&#x003B2;2</sub> represent the excitatory input mediated by &#x003B1;4&#x003B2;2-containing nAChRs, activated by PPTg ACh input and Nic. <italic>I</italic><sub>G-D</sub> is the local feed-forward inhibitory input to DA neurons emanating from VTA GABA neurons. <italic>B</italic><sub>D</sub> &#x0003D; 18 and <italic>B</italic><sub>G</sub> &#x0003D; 14 are the baseline firing rates of each neuron population in the absence of external inputs, according to Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) experimental data - with external inputs, the baseline activity of DA neurons is around 5 Hz.</p>
<p>The parameter <italic>r</italic> sets the balance of &#x003B1;4&#x003B2;2 nAChR action through GABA or DA neurons in the VTA. For <italic>r</italic> &#x0003D; 0, they act through GABA neurons only, whereas for <italic>r</italic> &#x0003D; 1 they influence DA neurons only. &#x003A6;(.) is the linear rectifier function, which only keeps the positive part of the operand and outputs 0 when it is negative. <italic>F</italic>(.) is a non-linear sigmoid transfer function for the dopaminergic neurons enabling to describe the high firing rates in the bursting mode and the low frequency activity in the tonic (pacemaker) mode, and their slow variation below their baseline activity with external inputs (&#x02243; 5 Hz):</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mi>F</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi>&#x003C9;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>exp</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C9; &#x0003D; 30 represent the maximum firing rate, &#x003B3; &#x0003D; 8 is the inflection point and &#x003B2; &#x0003D; 0.3 is the slope. These parameters were chosen in order to account for bursting activity of DA neurons starting from a certain threshold (&#x003B3;) of input and their maximal activity observed <italic>in vivo</italic> (Hyland et al., <xref ref-type="bibr" rid="B22">2002</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>). Indeed, physiologically, high firing rates (&#x0003E;8 Hz) are only attained during DA bursting activity and not tonic activity (&#x02243; 5 Hz).</p>
<p>The input currents in Equation (1) are given by:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>G-D</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>G-D</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mtext>G</mml:mtext></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>Glu-D</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-D</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PPT-D</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PPT</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>Glu-G</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-G</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PPT-G</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PPT</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mn>4</mml:mn><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mn>4</mml:mn><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mn>4</mml:mn><mml:mi>&#x003B2;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>w</italic><sub>x</sub>&#x00027;s (with x = G-D, PFC-D, PFC-G, PPT-D, PPT-G, &#x003B1;4&#x003B2;2) specify the total strength of the respective input (Figure <xref ref-type="fig" rid="F1">1</xref> and Supplementary Table <xref ref-type="supplementary-material" rid="SM1">1</xref>). For instance, <italic>w</italic><sub>PPT-D</sub> specifies the strength of the connection from the PPTg to the DA population.</p>
<p>The weight of &#x003B1;4&#x003B2;2-nAChRs, <italic>w</italic><sub>&#x003B1;4&#x003B2;2</sub> &#x0003D; 15 was chosen in order to account for the increase of baseline firing rates compared to Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) where <italic>w</italic><sub>&#x003B1;4&#x003B2;2</sub> &#x0003D; 1, <italic>B</italic><sub>D</sub> &#x0003D; 0.1 and <italic>B</italic><sub>G</sub> &#x0003D; 0. We also assumed that the PFC-DA and PFC-GABA connections were equal, which leads to the following important equality: <italic>w</italic><sub>PFC-D</sub>(<italic>n</italic>) &#x0003D; <italic>w</italic><sub>PFC-G</sub>(<italic>n</italic>) for any trial <italic>n</italic>.</p>
<p>In summary, inhibitory input to DA cells, <italic>I</italic><sub>G-D</sub>, depends on GABA neuron population activity, &#x003BD;<sub>G</sub> (Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>). Excitatory input to DA and GABA cells depends on PFC-NAc (Ishikawa et al., <xref ref-type="bibr" rid="B23">2008</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>) and PPTg (Lokwan et al., <xref ref-type="bibr" rid="B27">1999</xref>; Yoo et al., <xref ref-type="bibr" rid="B63">2017</xref>) glutamatergic inputs activities, &#x003BD;<sub>PFC</sub> and &#x003BD;<sub>PPT</sub> respectively (see next section). The activation of &#x003B1;4&#x003B2;2 nAChRs, &#x003BD;<sub>&#x003B1;4&#x003B2;2</sub>, determines the level of direct excitatory input <italic>I</italic><sub>&#x003B1;4&#x003B2;2</sub> evoked by nicotine or acetylcholine (see last section).</p>
</sec>
<sec>
<title>2.2. Neuronal Activities During Classical-Conditioning</title>
<p>As described above, previous studies identified signals from distinct brain areas that could be responsible for VTA DA neuron activity during classical conditioning. We thus consider a simple model that particularly accounts for Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) experimental data on VTA GABA neurons activity. In this approach, we propose that the sustained activity reflecting reward expectation in GABA neurons comes from the PFC (Schoenbaum et al., <xref ref-type="bibr" rid="B48">1998</xref>; Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>), that sends projections on both VTA DA and GABA neurons through the NAc (Morita et al., <xref ref-type="bibr" rid="B34">2013</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>). The PFC-NAc pathway thus drives feed-forward inhibition onto DA neurons by exciting VTA GABA neurons that in turn inhibit DA neurons (Figure <xref ref-type="fig" rid="F1">1</xref>). Second, we consider that a subpopulation of the PPTg provides the reward signal to the dopamine neurons at the US (Kobayashi and Okada, <xref ref-type="bibr" rid="B25">2007</xref>; Okada et al., <xref ref-type="bibr" rid="B37">2009</xref>).</p>
<sec>
<title>2.2.1. Classical-Conditioning Task and the Associated Signals</title>
<p>We modeled a VTA neural circuit (Figure <xref ref-type="fig" rid="F1">1</xref>) while mice are classically conditioned with a tone stimulus that predicts an appetitive outcome as in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>), but with 100% probability. Each simulated behavioral trial begins with a conditioned stimulus (CS; a tone, 0.5 s), followed by an unconditioned stimulus (US; the outcome, 0.5 s) separated by an interval of 1.5 s. (Figure <xref ref-type="fig" rid="F2">2A</xref>). This type of task, implying a delay between the CS offset and the US onset (here, 1 s), is then a trace-conditioning task, that differs from a delay-conditioning task where the CS and US overlap (Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Schematic of a classical-conditioning task. <bold>(A)</bold> Simulated thirsty mice receive a water reward ranging from 1 to 20 &#x003BC;L. Tone (CS) and reward (US) onsets are separated by 1.5 s. <bold>(B)</bold> Firing rates [mean &#x000B1; standard-error (s.e.)] of optogenetically identified dopamine neurons in response to different sizes of unexpected reward. Adapted from Eshel et al. (<xref ref-type="bibr" rid="B14">2016</xref>). <bold>(C)</bold> Temporal profile of the phasic function <italic>G</italic><sub>&#x003C4;</sub>(<italic>x</italic>(<italic>t</italic>)) (Equation 4) in response to a square input <italic>x</italic>(<italic>t</italic>).</p></caption>
<graphic xlink:href="fncir-12-00116-g0002.tif"/>
</fig>
<p>As the animal learns that a fixed reward predictably follows a predictive tone at a specific timing, our model proposes possible underlying biological mechanisms of Pavlovian-conditioning in PPTg, PFC, VTA DA, and GABA neurons (Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
<p>As represented in previous models (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B38">2007</xref>; Vitay and Hamker, <xref ref-type="bibr" rid="B56">2014</xref>), the CS signal is modeled by a square function (&#x003BD;<sub>CS</sub>(<italic>t</italic>)) equal to 1 during the CS presentation (0.5 s) and to 0 otherwise (Figure <xref ref-type="fig" rid="F2">2A</xref>). The US signal is modeled by a similar square function (&#x003BD;<sub>US</sub>(<italic>t</italic>)) as the CS but is equal to the reward size during the US presentation (0.5 s) and 0 otherwise (Figure <xref ref-type="fig" rid="F2">2A</xref>).</p>
</sec>
<sec>
<title>2.2.2. Neural Representation of the US Signal in the PPTg</title>
<p>Dopamine neurons in the VTA exhibit a relatively low tonic activity (around 5 Hz), but respond phasically with a short-latency (&#x0003C; 100 ms), short-duration (&#x0003C; 200 ms) burst of activity in response to unpredicted rewards (Schultz, <xref ref-type="bibr" rid="B49">1998</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>). These phasic bursts of activity are dependent on glutamatergic activation by a subpopulation of PPTg (Okada et al., <xref ref-type="bibr" rid="B37">2009</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>; Yoo et al., <xref ref-type="bibr" rid="B63">2017</xref>) found to discharge phasically at reward delivery, with the levels of activity associated with the actual reward and not affected by reward expectation.</p>
<p>To integrate the US input into a short-term phasic component we use the function <italic>G</italic><sub>&#x003C4;</sub>(<italic>x</italic>(<italic>t</italic>)) (Vitay and Hamker, <xref ref-type="bibr" rid="B56">2014</xref>) defined as follows:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>&#x003C4;</mml:mi><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x002D9;</mml:mo></mml:mover><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x003C4;</mml:mi><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x002D9;</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>G</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x003A6;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here when <italic>x</italic>(<italic>t</italic>) switches from 0 to 1 at time <italic>t</italic> &#x0003D; 0, <italic>G</italic><sub>&#x003C4;</sub>(<italic>x</italic>(<italic>t</italic>)) will display a localized bump of activation with a maximum at <italic>t</italic> &#x0003D; &#x003C4;. This function is thus convenient to integrate the square signal &#x003BD;<sub>US</sub>(<italic>t</italic>) into a short-latency response (Figure <xref ref-type="fig" rid="F2">2C</xref>).</p>
<p>Furthermore, dopamine response amplitudes to unexpected rewards follow a simple saturating function (fitted by a Hill function in Figure <xref ref-type="fig" rid="F2">2B</xref>) (Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>, <xref ref-type="bibr" rid="B14">2016</xref>). We thus consider that PPTg neurons respond to the reward delivery signal (US) in a same manner as DA neurons i.e., with a saturating dose-response function:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PPTg</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>PPTg</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>[</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>US</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>0.5</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>0.5</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mn>0.5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BD;<sub>PPTg</sub> is the mean activity of the PPTg neurons population, &#x003C4;<sub>PPTg</sub> &#x0003D; 100 ms (the short-latency response), and <italic>f</italic>(<italic>x</italic>) is a Hill function with two parameters: <italic>f</italic><sub>max</sub>, the saturating firing rate; and <italic>h</italic>, the reward size that elicits half-maximum firing rate. Here, we chose <italic>f</italic><sub>max</sub> &#x0003D; 70 and <italic>h</italic> &#x0003D; 20 in order to obtain a similar dose-response curve once PPTg activity is transferred to DA neurons as in Eshel et al. (<xref ref-type="bibr" rid="B14">2016</xref>) (Figure <xref ref-type="fig" rid="F2">2B</xref>).</p>
</sec>
<sec>
<title>2.2.3. Neural Representation of CS Signal in the PFC</title>
<p>In addition to their response to unpredicted rewards, learning drives the DA neurons to respond to reward-predictive cues and to reduce their response at the US (Schultz et al., <xref ref-type="bibr" rid="B50">1997</xref>; Schultz, <xref ref-type="bibr" rid="B49">1998</xref>; Matsumoto and Hikosaka, <xref ref-type="bibr" rid="B32">2009</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>). Neurons in the PFC respond to these cues through a sustained activation starting at the CS onset and ending at the reward-delivery (Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>; Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>). Furthermore, this activity has been shown to increase in the early stage of a classical-conditioning learning task (Schoenbaum et al., <xref ref-type="bibr" rid="B48">1998</xref>; Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>). Especially, the PFC participates in the association of temporally separated events in trace-conditioning task through working-memory mechanisms (Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>), maintaining a representation of the CS accross the CS-US interval, and this timing-association is dependent on dopamine modulation in the PFC (Puig et al., <xref ref-type="bibr" rid="B45">2014</xref>; Popescu et al., <xref ref-type="bibr" rid="B44">2016</xref>).</p>
<p>We thus assume that the PFC integrates the CS signal and learns to maintain its activity until the reward delivery. Consistently with previous neural-circuit working-memory models (Durstewitz et al., <xref ref-type="bibr" rid="B11">2000</xref>), we minimally described this mechanism by a neural population with recurrent excitation and slower adaptation dynamics blue (e.g., increase in calcium-dependent potassium hyperpolarizing currents <italic>I</italic><sub>KCa</sub>) inspired from Gerstner et al. (<xref ref-type="bibr" rid="B19">2014</xref>):</p>
<disp-formula id="E6"><mml:math id="M6"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>CS</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>CS</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>&#x0221E;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C4;<sub>PFC</sub> &#x0003D; 100 ms (short-latency response), <italic>a</italic>(<italic>t</italic>) describes the amount of adaptation that neurons have accumulated, <italic>a</italic><sub>&#x0221E;</sub> &#x0003D; <italic>c</italic>&#x000B7;&#x003BD;<sub>PFC</sub> is the asymptotic level of adaptation that is attained by a slow time constant &#x003C4;<sub><italic>a</italic></sub> &#x0003D; 1, 000 ms (Gerstner et al., <xref ref-type="bibr" rid="B19">2014</xref>) if the population continuously fires at a contant rate &#x003BD;<sub>PFC</sub>, <italic>J</italic><sub>PFC</sub>(<italic>n</italic>) represents the strength of the recurrent excitation exerted by the PFC depending on the learning trial <italic>n</italic> (initially <italic>J</italic>(1) &#x0003D; 0.2), <italic>w</italic><sub>CS</sub> the strength of the CS input. <italic>F</italic>(<italic>x</italic>) is the non-linear sigmoid transfer function defined in Equation (2) allowing the emergence of bistability network. We chose &#x003C9; &#x0003D; 30, &#x003B3; &#x0003D; 8 and &#x003B2; &#x0003D; 0.5 in order to account for the PFC activity changes in working-memory tasks (Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>).</p>
</sec>
<sec>
<title>2.2.4. Learning of the US Timing in the PFC</title>
<p>The dynamical system described above typically switches between two stables states: quasi absence of activity or maximal activity in the PFC. The latter stable state particularly appears as <italic>J</italic><sub>PFC</sub>(<italic>n</italic>) increases with learning:</p>
<disp-formula id="E7"><label>(6)</label><mml:math id="M7"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mtext>PFC</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x00394;</mml:mi><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>DA</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B1;<sub><italic>T</italic></sub> &#x0003D; 0.2 is the timing learning rate, &#x00394;<italic>t</italic><sub>DA</sub> &#x0003D; <italic>t</italic><sub>2</sub> &#x02212; <italic>t</italic><sub>1</sub> measures the difference between the time at which PFC activity declines (<italic>t</italic><sub>1</sub> such as &#x003BD;<sub>PFC</sub>(<italic>t</italic><sub>1</sub>) &#x02243; &#x003B3; after CS onset) and the time of DA maximal activity at the US, <italic>t</italic><sub>2</sub>. This learning mechanism of reward timing, simplified from Luzardo et al. (<xref ref-type="bibr" rid="B28">2013</xref>), triggers the increase of the recurrent connections (<italic>J</italic><sub>PFC</sub>) through dopamine-mediated modulation in the PFC (Puig et al., <xref ref-type="bibr" rid="B45">2014</xref>; Popescu et al., <xref ref-type="bibr" rid="B44">2016</xref>) such as &#x003BD;<sub>PFC</sub> collapses at the time of reward delivery. This learning process occurs in the early stage of the task (Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>) and is therefore much faster than the learning of reward expectation.</p>
</sec>
<sec>
<title>2.2.5. Learning of Reward Expectation in Cortico-Striatal Connections</title>
<p>According to studies showing a DA-dependent cortico-striatal plasticity (Reynolds et al., <xref ref-type="bibr" rid="B47">2001</xref>; Yagishita et al., <xref ref-type="bibr" rid="B61">2014</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>), we assumed that the reward value predicted from the tone (CS) is stored in the strength of cortico-striatal connections [<italic>w</italic><sub>PFC-D</sub>(<italic>n</italic>) and <italic>w</italic><sub>PFC-G</sub>(<italic>n</italic>)], i.e., between the PFC and the NAc, and is updated through plasticity mechanisms depending on phasic dopamine response after reward delivery as in the following equation proposed by Morita et al. (<xref ref-type="bibr" rid="B34">2013</xref>):</p>
<disp-formula id="E8"><label>(7)</label><mml:math id="M8"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-D</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-D</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mtext>V</mml:mtext></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-G</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>PFC-G</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mtext>V</mml:mtext></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B1;<sub>V</sub> is the cortico-striatal plasticity learning rate related to reward magnitude, &#x003B4;(<italic>n</italic>) is a deviation from the DA baseline firing rate, computed by the area under curve of &#x003BD;<sub>D</sub> in a 200 ms time-window following US onset, above a baseline defined by the value of &#x003BD;<sub>D</sub> at the time of US onset. &#x003B4;(<italic>n</italic>) is thus the reward-prediction error signal that updates the reward-expectation signal stored in the strength of the PFC input <italic>w</italic><sub>PFC-D</sub>(<italic>n</italic>) until the value of the reward is learned (Rescorla and Wagner, <xref ref-type="bibr" rid="B46">1972</xref>).</p>
<p>This assumption was taken from Morita et al. (<xref ref-type="bibr" rid="B34">2013</xref>) modeling work and various hypotheses on dopamine-mediated plasticity in associative-learning (Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>) and recent experimental data (Yagishita et al., <xref ref-type="bibr" rid="B61">2014</xref>; Fisher et al., <xref ref-type="bibr" rid="B17">2017</xref>). It implies that the excitatory signal from the PFC first activates the nucleus accumbens (NAc) and is then transferred via the direct disinhibitory pathway to the VTA. Here, we then considered that <italic>w</italic><sub>PFC-D</sub> and <italic>w</italic><sub>PFC-G</sub> are provided by the PFC-NAc pathway but we did not explicitly represent the NAc population (Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
</sec>
<sec>
<title>2.2.6. Cholinergic Input Activity</title>
<p>Our model also reflects the cholinergic (ACh) afferents to the DA and GABA cells in the VTA (Dautan et al., <xref ref-type="bibr" rid="B7">2016</xref>; Yau et al., <xref ref-type="bibr" rid="B62">2016</xref>). The &#x003B1;4&#x003B2;2 nAChRs are placed somatically on both the DA and the GABA neurons and their activity depends on ACh and Nic concentration within the VTA (see last section). As PPTg was found to be the main source of cholinergic input to the VTA, we assume that ACh concentration directly depends on PPTg activity, as modeled by the following equation:</p>
<disp-formula id="E9"><label>(8)</label><mml:math id="M9"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mtext>ACh</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>PPTg</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>w</italic><sub>ACh</sub> &#x0003D; 1 &#x003BC;M is the amplitude of the cholinergic connection that tunes concentration of acetylcholine <italic>ACh</italic> (in &#x003BC;M) at a physiologically relevant concentration (Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>).</p>
</sec>
</sec>
<sec>
<title>2.3. Modeling the Activation and Desensitization of nAChRs</title>
<p>We implemented nAChR activation and desensitization from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) as transitions of two independent state variables: an activation gate and a desensitization gate. The nAChR receptors can then be in four different states: deactivated/sensitized, activated/sensitized, activated/desensitized and deactivated/desensitized. The receptors are activated in response to both Nic and ACh, while desensitization is driven by Nic only (if &#x003B7; &#x0003D; 0). Once Nic or ACh is removed, the receptors can switch from activated to deactivated and from desensitized to sensitized.</p>
<p>The mean total activation level of nAChRs (&#x003BD;<sub>&#x003B1;4&#x003B2;2</sub>) is modeled as the product of the activation rate <italic>a</italic> (fraction of receptors in the activated state) and the sensitization rate <italic>s</italic> (fraction of receptors in the sensitized state). The total normalized nAChR activation is therefore: &#x003BD;<sub>&#x003B1;4&#x003B2;2</sub> &#x0003D; <italic>a</italic>&#x000B7;<italic>s</italic>. The time course of the activation and the sensitization variables is given by:</p>
<disp-formula id="E10"><label>(9)</label><mml:math id="M10"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>&#x0221E;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C4;<sub><italic>y</italic></sub>(<italic>Nic, ACh</italic>) refers to the Nic/ACh concentration-dependent time constant at which the steady-state <italic>y</italic><sub>&#x0221E;</sub>(<italic>Nic, ACh</italic>) is achieved. The maximal achievable activation or sensitization, for a given Nic/ACh concentration, <italic>a</italic><sub>&#x0221E;</sub>(<italic>Nic, ACh</italic>) and <italic>s</italic><sub>&#x0221E;</sub>(<italic>Nic, ACh</italic>) are given by Hill equations of the form:</p>
<disp-formula id="E11"><label>(10)</label><mml:math id="M11"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>a</mml:mi><mml:mi>&#x0221E;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>a</mml:mtext></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>a</mml:mtext></mml:msub></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>a</mml:mtext></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>s</mml:mi><mml:mi>&#x0221E;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>s</mml:mtext></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>s</mml:mtext></mml:msub></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>s</mml:mtext></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>EC</italic><sub>50</sub> and <italic>IC</italic><sub>50</sub> are the half-maximal concentrations of nAChR activation and sensitization, respectively. The factor &#x003B1;&#x0003E;1 accounts for the higher potency of Nic to evoke a response as compared to ACh: &#x003B1;<sub>&#x003B1;4&#x003B2;2</sub> &#x0003D; 3. <italic>n</italic><sub>a</sub> and <italic>n</italic><sub>s</sub> are the Hill coefficients of activation and sensitization. &#x003B7; varies between 0 and 1 and controls the fraction of the ACh concentration driving receptor desensitization. Here, as we only consider Nic-induced desensitization, we set &#x003B7; &#x0003D; 0.</p>
<p>As the transition from the deactivated to the activated state is fast (&#x0007E;&#x003BC;s), the activation time constant &#x003C4;<sub><italic>a</italic></sub> was simplified to be independent on ACh and Nic concentration: &#x003C4;<sub><italic>a</italic></sub>(<italic>Nic, ACh</italic>) &#x0003D; &#x003C4;<sub><italic>a</italic></sub> &#x0003D; <italic>const</italic>. The time course of Nic-driven desensitization is characterized by a concentration-dependent time constant</p>
<disp-formula id="E12"><label>(11)</label><mml:math id="M12"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub><mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub><mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>&#x003C4;</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C4;<sub>max</sub> refers to the recovery time constant from desensitization in the absence of ligands, &#x003C4;<sub>0</sub> is the fastest time constant at which the receptor is driven into the desensitized state at high ligand concentrations. <italic>K</italic><sub>&#x003C4;</sub> is the concentration at which the desensitization time constant attains half of its minimum. All model assumptions are further described in Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>).</p>
</sec>
<sec>
<title>2.4. Simulated Experiments</title>
<sec>
<title>2.4.1. Optogenetic Inhibition of VTA GABA Neurons</title>
<p>In order to qualitatively reproduce (Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>) experimental data, we simulated the photo-inhibition effect in a subpopulation of VTA GABA neurons with an exponential decrease between <italic>t</italic> &#x0003D; 1.5 s and <italic>t</italic> &#x0003D; 2.5s (&#x000B1;500 ms around reward-delivery). First, the light was modeled by a square signal &#x003BD;<sub>light</sub> equal to the laser intensity <italic>I</italic> &#x0003D; 4 for 1.5 &#x0003C; <italic>t</italic> &#x0003C; 2.5 and zero otherwise. Then, we subtracted this signal to VTA GABA neuron activity as follows:</p>
<disp-formula id="E13"><label>(12)</label><mml:math id="M13"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mtext>s</mml:mtext></mml:msub><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>light</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>G-opto</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003BD;</mml:mi><mml:mrow><mml:mtext>G-control</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>s</italic> is the subtracted signal that integrates the light signal &#x003BD;<sub>light</sub> with a time constant &#x003C4;<sub>s</sub> &#x0003D; 300 ms, &#x003BD;<sub>G-opto</sub> is the photo-inhibited GABA neurons activity, and &#x003BD;<sub>G-control</sub> is the normal GABA neurons activity with no opto-inhibition. All parameters (<italic>I</italic>, &#x003C4;<sub>s</sub>) were chosen in order to reproduce qualitatively the photo-inhibition effects revealed by Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) experiments. Furthermore, as the effects of GABA photo-inhibition onto DA neurons appear to be relatively weak in Figure 3 of Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>), we assumed that only a subpopulation of the total GABA neurons are photo-inhibited and we therefore applied (Equation 12) for only 20% of the VTA GABA population. This assumption was based on the partial expression of Archeorhodopsin (ArchT) in GABA neurons (Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>, Extended Data Figure <xref ref-type="fig" rid="F1">1</xref>) and the other possible optogenetic effects (recording distance, variability of the response among the population, laser intensity, etc.).</p>
</sec>
<sec>
<title>2.4.2. Nicotine Injection in the VTA</title>
<p>In order to model chronic nicotine injection in the VTA while mice perform classical-conditioning tasks with water reward, the above equations were simulated but after 5 min of 1 &#x003BC; M Nic injection in the model for each trial. This process allowed to focus only on the effects of &#x003B1;4&#x003B2;2-nAChRs desensitization (see next section) during conditioning trials.</p>
</sec>
<sec>
<title>2.4.3. Decision-Making Task</title>
<p>We simulated a protocol designed by Naud&#x000E9; et al. (<xref ref-type="bibr" rid="B36">2016</xref>) recording simultaneously the sequential choices of a mouse between three differently rewarding locations (associated with reward size) in a circular open-field (<bold>Figure 7A</bold>). These three locations form an equilateral triangle and provide respectively 2, 4, 8 &#x003BC; L water rewards. Each time the mouse reaches one of the rewarding locations, the reward is delivered. However, the mouse receives the reward only when it alternates between rewarding locations.</p>
<p>Before the simulated task, we considered that the mouse has already learned the value of each location (pre-training) and thus knows the expected associated reward. Each value was computed taking the maximal activity of DA neurons within a time window following the CS onset (here, the view of the location) for the three different reward sizes after learning. We also considered that each time the mouse reaches a new location, it enters in a new state <italic>i</italic>. Decision making-models inspired from Naud&#x000E9; et al. (<xref ref-type="bibr" rid="B36">2016</xref>) determine the probability <italic>P</italic><sub><italic>i</italic></sub> of choosing the next state <italic>i</italic> as a function of the expected value of this state. Because mice could not return to the same rewarding location, they had to choose between the two remaining locations. We thus modeled decisions between two alternatives. The probability <italic>P</italic><sub><italic>i</italic></sub> was computed according to the softmax choice rule:</p>
<disp-formula id="E14"><label>(13)</label><mml:math id="M14"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>exp</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>V</italic><sub><italic>i</italic></sub> and <italic>V</italic><sub><italic>j</italic></sub> are the values of the states <italic>i</italic> and <italic>j</italic> (the other option), respectively, <italic>b</italic> is an inverse temperature parameter reflecting the sensitivity of choice to the difference between both values. We chose <italic>b</italic> &#x0003D; 0.4 which corresponds to a reasonable exploration-exploitation ratio.</p>
<p>We simulated the task over 10,000 simulations and computed the number of times the mouse chose each location. We thus obtained the average repartition of the mouse over the three locations. A similar task was simulated for mice after 5 min Nic ingestion (see below).</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<p>We used the model developed above to understand the learning dynamics within the PFC-VTA circuitry and the mechanisms by which the RPE in the VTA is constructed. Our minimal circuit dynamics model of the VTA was inspired from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) and modified according to recent neurobiological studies (see Methods) in order to reproduce RPE computations in the VTA. This model reflects the glutamatergic (from PFC and PPTg) and cholinergic (from PPTg) afferents to VTA DA and GABA neurons, as well as local inhibition of DA neurons by GABA neurons. We also included the activation and desensitization dynamics of &#x003B1;4&#x003B2;2 nAChRs from Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>), placed somatically on both DA and GABA neurons, depending on a fraction parameter <italic>r</italic>.</p>
<p>We note that we explicitly set <italic>r</italic> so the majority of nAChRs are located on the inhibitory GABA interneurons, hence following the &#x0201C;disinhibition&#x0201D; scheme as per (Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>).</p>
<p>We simulated the proposed PFC and PPTg activity during the task, where corticostriatal connections between the PFC and the VTA and recurrent connections among the PFC were gradually modified by dopamine in the NAc. Finally, we studied the potential influence of nicotine exposure on DA responses to rewarding events.</p>
<p>We should note that most experiments we simulated herein concern the learning task of a CS-US association (Figure <xref ref-type="fig" rid="F2">2</xref>). The learning procedure consists of a conditioning phase where a tone (CS) and a constant water-reward (US) are presented together for 50 trials. Within each 3 s-trial, the CS is presented at <italic>t</italic> &#x0003D; 0.5 s (Figures <xref ref-type="fig" rid="F3">3</xref>, <bold>5</bold>, <bold>6</bold>, dashed gray line) followed by the US at <italic>t</italic> &#x0003D; 2 s (Figures <xref ref-type="fig" rid="F3">3</xref>, <bold>5</bold>, <bold>6</bold>, dashed cyan line).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Activity of VTA neurons and their afferents during a pavlovian-conditioning task. Simulated mean activity (Hz) of each neuron population during a pavlovian-conditioning task, where a tone is presented systematically 1.5 s before a water reward (4 &#x003BC;L). Three different trials are represented: the initial conditioning trial (<italic>n</italic> &#x0003D; 1, light colors), an intermediate trial (<italic>n</italic> &#x0003D; 6, medium colors) and the final trial (<italic>n</italic> &#x0003D; 50, dark colors) and when reward is omitted after learning (dotted lines). Vertical dashed gray and cyan lines represent CS and US onsets, respectively. <bold>(A)</bold> PFC neurons learn the timing of the task by maintaining their activity until US. <bold>(C)</bold> PPTg neurons activity responds to the US signal at all trials. <bold>(B)</bold> VTA GABA persistent activity increases with learning, <bold>(D)</bold> VTA DA activity increase at the CS and decrease at the US.</p></caption>
<graphic xlink:href="fncir-12-00116-g0003.tif"/>
</fig>
<sec>
<title>3.1. Pavlovian-Conditioning Task and VTA Activity</title>
<p>DA activity during a classical-conditioning task was first recorded by Schultz (<xref ref-type="bibr" rid="B49">1998</xref>) and tested in further several studies. Additionally, Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) also recorded the activity of their putative neighboring neurons, the VTA GABA neuron population. Our goal was first to qualitatively reproduce VTA GABA and DA activity during associative learning of a pavlovian-conditioning task.</p>
<p>In order to understand how different brain areas interact during the conditioning and also during reward omission, we examined the simulated time course of activity of four populations (PFC, PPTg, VTA DA and GABA), Figure <xref ref-type="fig" rid="F3">3</xref>, at the initial conditioning trial (<italic>n</italic> &#x0003D; 1, light color curves), an intermediary trial (<italic>n</italic> &#x0003D; 6, medium color curves) and at the final trial (<italic>n</italic> &#x0003D; 50, dark color curves). In line with experiments, the reward delivery (Figure <xref ref-type="fig" rid="F3">3</xref>, dashed cyan lines) activates the PPTg nucleus (Figure <xref ref-type="fig" rid="F3">3C</xref>) at each conditioning trial. These neurons activate in turn VTA DA and GABA neurons through glutamatergic connections, causing a phasic burst in DA neurons at the US when the reward is unexpected (Figure <xref ref-type="fig" rid="F3">3D</xref>, <italic>n</italic> &#x0003D; 1), and a small excitation in GABA neurons (Figure <xref ref-type="fig" rid="F3">3B</xref>, <italic>n</italic> &#x0003D; 1). PPTg fibers also stimulate VTA neurons through ACh-mediated &#x003B1;4&#x003B2;2 nAChRs activation, with a larger influence on GABA neurons (<italic>r</italic> &#x0003D; 0.2 in Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
<p>Early in the conditioning task, simulated PFC neurons respond to the tone (Figure <xref ref-type="fig" rid="F3">3A</xref>, <italic>n</italic> &#x0003D; 1), and this activity builds up until being maintained during the whole CS-US interval (Figure <xref ref-type="fig" rid="F3">3A</xref>, <italic>n</italic> &#x0003D; 6, <italic>n</italic> &#x0003D; 50). Thus, PFC neurons show a working-memory like activity now tuned to decay at the reward delivery time. Concurrently, the phasic activity of DA neurons at the US acts as prediction-error signal on corticostriatal synapses, increasing the glutamatergic input from the NAc onto VTA DA and GABA neurons (Figures <xref ref-type="fig" rid="F3">3B,D</xref>, <xref ref-type="fig" rid="F4">4B</xref>). Note that the NAc was not modeled explicitly, but we modeled the net effect of the PFC-NAc plasticity with the variables <italic>w</italic><sub>PFC-D</sub> and <italic>w</italic><sub>PFC-G</sub> (see next section).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Learning of reward timing and magnitude during classical-conditioning. <bold>(A)</bold> The maximal activity of the VTA DA neurons at the CS onset (blue line) and at the reward delivery (orange line) is plotted for each trial of the conditioning task. These values are computed by taking the maximum value of the firing rate of the DA neurons in a small time window (200 ms) after the CS and the US onsets. <bold>(B)</bold> PFC weights showing two phases of learning: learning of the US timing by PFC recurrent connections weight (<italic>J</italic><sub>PFC</sub>, orange line) and learning of the reward value by the weights of PFC neurons onto VTA neurons (<italic>w</italic><sub>PFC-D</sub> and <italic>w</italic><sub>PFC-G</sub>, magenta line). <bold>(C,D)</bold> Phase analysis of PFC neuron activity from Equation (8) before learning <bold>(C)</bold> and after learning <bold>(D)</bold>. Different times of the task are represented: <italic>t</italic> &#x0003C; 0.5 s (before CS onset, light blue) and 1 s &#x0003C; <italic>t</italic> &#x0003C; 2 s (between CS offset and US onset, light blue), 0.5 s &#x0003C; <italic>t</italic> &#x0003C; 1 s (during CS presentation, medium blue) and <italic>t</italic> &#x0003E; 2 s (after US onset, dark blue). Fixed points are represented by green (stable) or red (unstable) dots. Dashed arrows: trajectories of the system from <italic>t</italic> &#x0003D; 0 to <italic>t</italic> &#x0003D; 3 s.</p></caption>
<graphic xlink:href="fncir-12-00116-g0004.tif"/>
</fig>
<p>Consequently, with learning, VTA GABA neurons show a sustained activation during the CS-US interval (Figure <xref ref-type="fig" rid="F3">3B</xref>, <italic>n</italic> &#x0003D; 6, <italic>n</italic> &#x0003D; 50) as found in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) experiments and in turn inhibit their neighboring dopamine neurons. Thus, in DA neurons, the GABA neurons-induced inhibition occurs with a slight delay after the PFC-induced excitation, resulting in a phasic excitation at the CS and a phasic inhibition at the US (Figure <xref ref-type="fig" rid="F3">3D</xref>, <italic>n</italic> &#x0003D; 50).</p>
<p>The latter inhibition progressively cancels the reward-evoked excitation by the PPTg glutamatergic fibers in DA neurons. It also accounts for the pause in DA firing when reward is omitted after learning (Figures <xref ref-type="fig" rid="F3">3B,D</xref>, <italic>n</italic> &#x0003D; 50, dashed lines). In order to test whether this cancellation mode is robust to changes in GABA and PPTg time constants, we represented VTA GABA and DA neurons activity by varying &#x003C4;<sub>PPTg</sub> and &#x003C4;<sub>G</sub> (Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>). It results in slight variations of GABA and DA amplitudes, but their dynamics remain qualitatively robust. Together, these results propose a simple mechanism for RPE computation the VTA and its afferents.</p>
<p>Let us now take a closer look at the evolution of the phasic activity of DA neurons and their PFC-NAc afferents during the conditioning task. Figure <xref ref-type="fig" rid="F4">4A</xref> shows the evolution of CS- and US-mediated DA peaks over the 50 conditioning trials. Firstly, the US-related bursts (Figure <xref ref-type="fig" rid="F4">4A</xref>, red line) remain constant in the early trials until the timing is learnt by the PFC recurrent connections <italic>J</italic><sub>PFC</sub> (Figure <xref ref-type="fig" rid="F4">4B</xref>, orange line) following Equation (6). Secondly, US and CS (Figure <xref ref-type="fig" rid="F4">4A</xref>, blue line) responses respectively decrease and increase over all trials, following a slower learning process from cortico-striatal connections (Figure <xref ref-type="fig" rid="F4">4B</xref>, magenta line) described by Equation (7). This two-speed learning process enables to qualitatively reproduce the DA dynamics found experimentally, with almost no effect outside the CS and US time-windows (Figure <xref ref-type="fig" rid="F4">4D</xref>).</p>
<p>Particularly, the graphical analysis of the PFC system enables us to understand the timing learning mechanism. From Equation (6), we can see where the two functions &#x003BD;<sub>PFC</sub> &#x02192; &#x003BD;<sub>PFC</sub> and &#x003BD;<sub>PFC</sub> &#x02192; <italic>F</italic>[<italic>w</italic><sub>CS</sub> &#x000B7; &#x003BD;<sub>CS</sub>(<italic>t</italic>) &#x0002B; <italic>J</italic>(<italic>n</italic>) &#x000B7; &#x003BD;<sub>PFC</sub>(<italic>t</italic>) &#x02212; <italic>a</italic>(<italic>t</italic>)] intersect each other (fixed points analysis) at four different timings during the simulation: before and after the CS presentation (&#x003BD;<sub>CS</sub> &#x0003D; 0, <italic>a</italic> &#x0003D; 0), during CS presentation (&#x003BD;<sub>CS</sub> &#x0003D; 1, <italic>a</italic> &#x0003D; 0) and after the reward is delivered (&#x003BD;<sub>CS</sub> &#x0003D; 0, <italic>a</italic> &#x0003D; <italic>a</italic><sub>&#x0221E;</sub>). Before learning, as <italic>J</italic><sub>PFC</sub> is weak (Figure <xref ref-type="fig" rid="F4">4C</xref>), the system starts at one fixed point (&#x003BD;<sub>PFC</sub> &#x0003D; 0), then jumps to another stable point during CS presentation (&#x003BD;<sub>PFC</sub>&#x02243;30) and immediately goes back to the initial point (&#x003BD;<sub>PFC</sub> &#x0003D; 0) after CS presentation (<italic>t</italic> &#x0003D; 1 s) as shown in Figure <xref ref-type="fig" rid="F3">3A</xref>. After learning (Figure <xref ref-type="fig" rid="F4">4D</xref>), the system initially shows the same dynamics but when the CS is removed, the system is maintained at the second fixed point (30 Hz) until reward delivery (Figure <xref ref-type="fig" rid="F3">3A</xref>, <italic>n</italic> &#x0003D; 50) due to its bistability after CS presentation (cyan curve). Finally, with the adaptation dynamics, the PFC activity decays right after reward delivery (Figure <xref ref-type="fig" rid="F4">4D</xref>, dark blue). Indeed, through this timing learning mechanism, the strength of the recurrent connections maintains the Up state activity of the PFC exactly until the US timing (Equation 6). Together, these simulations show a two-speed learning process that enables VTA dopamine neurons to predict the value and the timing of the water reward from PFC plasticity mechanisms.</p>
</sec>
<sec>
<title>3.2. Photo-Inhibition of VTA GABA Neurons Modulates Prediction Errors</title>
<p>We next focus specifically on the local VTA neurons interactions at the end of the conditioning task. Particularly, we model the effects of VTA GABA optogenetic inhibition (Figure <xref ref-type="fig" rid="F5">5</xref>) revealed by one of Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) experiments. First, we pick the activity of VTA GABA and DA neurons at the last learning trial (<italic>n</italic> &#x0003D; 50), where DA neurons are excited by the cue (CS) rather than by the actual reward (US). Note that in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>), DA neurons were still activated at the US timing, which we suppose to be related to their experimental procedure consisting of delivering rewards stochastically (with 90% probability in this experiment). Second, as in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>), we simulated GABA photo-inhibition in a time-window (&#x000B1;500 ms) around the reward delivery time (Figure <xref ref-type="fig" rid="F5">5A</xref>, green shaded area). Considering that ArchT virus expression was partial in GABA neurons and that optogenetic effects do not account quantitatively for physiological effects, the photo-inhibition was simulated for only 20% of our GABA population. This simulated inhibition resulted in a disinhibition of DA neurons activity during laser stimulation (Figure <xref ref-type="fig" rid="F5">5B</xref>). If the inhibition was 100% efficient on GABA neurons, we assume that experimentally, DA neurons would then burst at high frequencies during the whole period of stimulation.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Photo-inhibition of VTA GABA neurons. <bold>(A)</bold> Activity of a subpopulation of GABA neurons (20%) in control (black) and with photo-inhibition (green) simulated by an exponential-like decrease of activity in a &#x000B1;500 ms time-window around the US (green shaded area) after learning (<italic>n</italic> &#x0003D; 50). <bold>(B)</bold> DA activity resulting from GABA neurons activity in control condition (black) and when GABA is photo-inhibited (green) after learning (<italic>n</italic> &#x0003D; 50).</p></caption>
<graphic xlink:href="fncir-12-00116-g0005.tif"/>
</fig>
<p>Inhibiting VTA GABA neurons partially reversed the expectation-dependent reduction of DA response at the US. As proposed by Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>), our model accounts for the burst-canceling expectation signal provided by VTA GABA neurons.</p>
</sec>
<sec>
<title>3.3. Effects of Nicotine on RPE Computations in the VTA</title>
<p>We next asked whether we can identify the effects of nicotine action in the VTA during the classical-conditioning task described in Figure <xref ref-type="fig" rid="F3">3</xref>. We compared the activity of DA neurons at different conditioning trials to their activity after 5 min of 1 &#x003BC; M nicotine injection, corresponding to physiologically relevant concentrations of Nic in the blood after cigarette-smoking (Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>; Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>). For our qualitative investigations, we assume that &#x003B1;4&#x003B2;2-nAChRs are mainly expressed on VTA GABA neurons (<italic>r</italic> &#x0003D; 0.2) and we study the effects of nicotine-induced desensitization on these receptors.</p>
<p>Nic-induced desensitization may potentially lead to several effects. First, under nicotine (Figure <xref ref-type="fig" rid="F6">6B</xref>), DA baseline activity slightly increases. Second, simulated exposure also raises DA responses to reward-delivery when the animal is naive (Figures <xref ref-type="fig" rid="F6">6A,B</xref>, <italic>n</italic> &#x0003D; 1), and therefore to reward-predictive cues when the animal has learnt the task (Figures <xref ref-type="fig" rid="F6">6A,B</xref>, <italic>n</italic> &#x0003D; 50). As expected, these effects derive from the reduction of the ACh-induced GABA activation provided by the PPTg nucleus (Figure <xref ref-type="fig" rid="F3">3C</xref>). Thus, our simulations predict that nicotine would up-regulate DA bursting activity at rewarding events.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Effects of nicotine on DA activity during classical-conditioning. <bold>(A)</bold> Activity of DA neurons during the pavlovian-conditioning (tone &#x0002B; 4 &#x003BC;L reward) task in three different trials as in Figure <xref ref-type="fig" rid="F3">3</xref>. <bold>(B)</bold> Same as <bold>(A)</bold> but after 5 min of 1 &#x003BC;M nicotine injection during all conditioning trials. <bold>(C)</bold> DA activity after learning under nicotine (magenta) or in the same condition but when nicotine is removed (dark red). <bold>(D)</bold> Dose-response curves of CS-related burst in DA neurons after learning in control condition (green) or under nicotine (magenta).</p></caption>
<graphic xlink:href="fncir-12-00116-g0006.tif"/>
</fig>
<p>What would happen if the animal, after having learned in the presence of nicotine, is not exposed to it anymore (nicotine withdrawal)? To answer this question, we investigate the effects of nicotine withdrawal on DA activity after the animal has learnt the CS-US association under nicotine (Figure <xref ref-type="fig" rid="F6">6C</xref>), with the same amount of reward (4 &#x003BC;L). In addition to a slight decrease in DA baseline activity, the DA response to the simulated water reward is reduced even below baseline (Figure <xref ref-type="fig" rid="F6">6C</xref>, dark red). DA neurons would then signal a negative reward-prediction error, consequently encoding a possible perceived insufficiency of the actual reward it usually receives. From these simulations, we could predict the effect of nicotine injection on the dose-response curve of DA neurons to rewarding events (Figure <xref ref-type="fig" rid="F6">6D</xref>).</p>
<p>Here, instead of plotting DA neuron response to different sizes of unexpected rewards as in Figure <xref ref-type="fig" rid="F2">2B</xref>, we plot DA response to the CS after the animal has learnt different sizes of rewards (Figure <xref ref-type="fig" rid="F6">6D</xref>), taking the maximum activity in a 200 ms time-window following the CS onset (Figures <xref ref-type="fig" rid="F6">6A,B</xref>, dark colors). Thus, when the animal learns under nicotine, the dose-response curve is elevated, assigning an amplification effect of nicotine on dopamine reward-prediction computations. Notably, the nicotine-induced increase in CS-related bursts grows with the increase of reward size for rewards ranging from 0 to 8 &#x003BC;L. Associating CS amplitude to the predicted value (Rescorla and Wagner, <xref ref-type="bibr" rid="B46">1972</xref>; Schultz, <xref ref-type="bibr" rid="B49">1998</xref>), this suggests that nicotine could increase the value of the cues predicting large rewards, therefore increasing the probability of choosing the associated states compared to control conditions.</p>
</sec>
<sec>
<title>3.4. Model-Based Analysis of Mouse Decision-Making Under Nicotine</title>
<p>In order to evaluate the effects of nicotine on the choice preferences among reward sizes, we simulated a decision-making task where a mouse chose between three locations providing different reward sizes (2, 4, 8 &#x003BC;L) in a circular open-field (Figure <xref ref-type="fig" rid="F7">7A</xref>) inspired by Naud&#x000E9; et al. (<xref ref-type="bibr" rid="B36">2016</xref>) experimental paradigm.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Effects of nicotine on mouse decision-making among reward sizes. <bold>(A)</bold> Illustration of the modeling of the task. Three explicit locations are placed in an open field. Mice receive a reward each time they reach one of the locations. Simulated mice, who could not receive two consecutive rewards at the same location, alternate between rewarding locations. The probability of transition from one state to another depends on the two available options. <bold>(B)</bold> Proportion of choices of the three rewarding locations as a function of reward value (2, 4, 8 &#x003BC;L) over 10,000 simulations in control mice (blue) or nicotine-ingested mice (red).</p></caption>
<graphic xlink:href="fncir-12-00116-g0007.tif"/>
</fig>
<p>Following reinforcement-learning theory (Rescorla and Wagner, <xref ref-type="bibr" rid="B46">1972</xref>; Sutton and Barto, <xref ref-type="bibr" rid="B52">1998</xref>), CS response to each reward size (computed from Figure <xref ref-type="fig" rid="F6">6D</xref>) was attributed to the expected value of each location. We then computed the repartition of the mouse between the three locations over 10,000 simulations in control conditions or after 5 min nicotine ingestion.</p>
<p>In control conditions, the simulated mice chose according to the location&#x00027;s estimated value (Figure <xref ref-type="fig" rid="F7">7B</xref>); the mice chose preferentially the locations that provide the greater amount of reward. Interestingly, under Nic-induced nAChRs desensitization, the simulations show a bias of mice choices toward large reward sizes; the proportion of choices for the small reward (2 &#x003BC;L) diminished by about 4%. Thus, these simulations suggested a differential amplifying effect of nicotine for large water rewards.</p>
<p>We can explain these simulation results in Figure <xref ref-type="fig" rid="F6">6D</xref>, by the fact that nicotine has a multiplicative effect on DA responses at the CS in the interval [0,8] &#x003BC;L compared to control condition. This then leads to a proportionally larger nicotine influence on the larger vs. the smaller rewards. We then expect that such bias would not appear for a set of larger rewards, as the nicotine effect is additive after 8 &#x003BC;L. This is a prediction of this model.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>The overarching aim of this study was to determine how dopamine neurons compute key quantities such as reward-prediction errors, and how these computations are affected by nicotine. In order to do so, we have developed a computational modeling approach extending the population activity of the VTA and its main afferents during a simple task of Pavlovian-conditioning. Including both theoretical and phenomenological conceptions, this model qualitatively reproduces several observations on the VTA activity during the task: phasic DA activity at the US and the CS and persistent activity of VTA GABA neurons. It particularly proposes a two-speed learning process of the reward timing and size mediated by the PFC working memory, coupled with the signaling of reward occurrence in the PPTg. Finally, using acetylcholine dynamics coupled with the desensitization kinetics of &#x003B1;4&#x003B2;2-nAChRs in the VTA, we revealed a potential effect of nicotine action on reward perception through up-regulation of DA phasic activity.</p>
<sec>
<title>4.1. Modeling Choices</title>
<p>Multiple studies have proposed a dual-pathway mechanism for RPE computation in the brain (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B38">2007</xref>; Vitay and Hamker, <xref ref-type="bibr" rid="B56">2014</xref>) through phenomenological bottom-up approaches. Although they propose different possible mechanisms, they mainly gather several components: regions that encode reward-expectation at the CS, regions that encode actual reward, regions that inhibit dopamine activity at the US, and final subtraction of these inputs at the VTA level. These models usually manage to reproduce the key properties of dopamine-related reward activity: progressive appearance of DA bursts at the CS onset, progressive decrease of DA bursts at the US onset, phasic inhibition when reward is omitted and early delivery of reward.</p>
<p>Additionally, a top-down theoretical approach as the temporal difference (TD) learning model assumes that the cue and reward cancellation signal both emerge from the same inputs (Sutton and Barto, <xref ref-type="bibr" rid="B52">1998</xref>; Morita et al., <xref ref-type="bibr" rid="B34">2013</xref>). After the task is learned, two sustained expectation signals <italic>V</italic>(<italic>t</italic>) and <italic>V</italic>(<italic>t</italic>&#x0002B;1) subtract each other (Figure <xref ref-type="fig" rid="F8">8</xref>), leading to the TD error: &#x003B4; &#x0003D; <italic>r</italic>&#x0002B;<italic>V</italic>(<italic>t</italic>&#x0002B;1)&#x02212;<italic>V</italic>(<italic>t</italic>). Notably, the temporary shift between both signals induce a phasic excitation at CS and an inhibition at the US.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>TD learning model (Watabe-Uchida et al., <xref ref-type="bibr" rid="B57">2017</xref>).</p></caption>
<graphic xlink:href="fncir-12-00116-g0008.tif"/>
</fig>
<p>TD models are reliable to describe many features of dopamine phasic activity and establish a link between reinforcement learning theory and dopamine activity. However, the biological evidence for such specific signals is still unclear.</p>
<p>In our study, we combine these two phenomenological and theoretical approaches to describe the VTA DA activity. Firstly, our simple model relies on neurobiological mechanisms such as PFC working memory activity (Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>; Le Merre et al., <xref ref-type="bibr" rid="B26">2018</xref>), PPTg activity (Kobayashi and Okada, <xref ref-type="bibr" rid="B25">2007</xref>; Okada et al., <xref ref-type="bibr" rid="B37">2009</xref>) and mostly VTA GABA neurons activity (Cohen et al., <xref ref-type="bibr" rid="B4">2012</xref>; Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>) and describe how these inputs could converge to VTA DA neurons. Secondly, at least at the end of learning, we also proposed a similar integration of inputs as in TD models, with two sustained signals that are temporally delayed. We note that in our model, like in the algorithmic TDRL models, late delivery of the reward would lead to a dip in the DA activity at the previously expected reward-time and same for early reward (simulations not shown). Arguably, the late reward response matches experimentally observed phasic DA activity, early reward remains a challenge for the model.</p>
<p>Indeed, the reward expectation signal comes from the same input (PFC): based on recent data on local circuitry in the VTA (Eshel et al., <xref ref-type="bibr" rid="B13">2015</xref>), we assumed that the PFC sends the <italic>V</italic>(<italic>t</italic> &#x0002B; 1) sustained signal to both VTA GABA and DA neurons. Only, via a feed-forward inhibition mechanism, this signal is shifted by VTA GABA neurons membrane time constant &#x003C4;<sub>G</sub>. Thus, in addition to the direct <italic>V</italic>(<italic>t</italic> &#x0002B; 1) excitatory signal from the PFC, VTA GABA neurons would send the <italic>V</italic>(<italic>t</italic>) inhibitory signal to VTA DA neurons (Figure <xref ref-type="fig" rid="F8">8</xref>). Adding the reward signal <italic>r</italic>(<italic>t</italic>) provided by the PPTg, our model integrates the TD error &#x003B4; into DA neurons. However, in our model, and as shown in several studies, CS- and US-related bursts gradually increase and decrease with learning, respectively, whereas TD learning predicts a progressive backward shift of the US-related burst during learning, what is not experimentally observed.</p>
<p>Although we make strong assumptions on VTA reward information integration that may be questioned at the level of detailed biology, it proposes a way to explain how the sustained activity in GABA neurons cancel the US-related dopamine burst without affecting the preceding tonic activity of DA neurons during the CS-US interval. Furthermore, this assumption can be strengthened by our simulation of optogenetic experiment (Figure <xref ref-type="fig" rid="F5">5</xref>) qualitatively reproducing DA increase in both baseline and phasic activity as found in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>).</p>
</sec>
<sec>
<title>4.2. Reliability of the VTA Afferents</title>
<p>As described above, our model includes two glutamatergic and one GABAergic input to the dopamine neurons, without considering the influence of all other brain areas.</p>
<p>Although the NAc disinhibitory input and the PPTg excitatory input were found to be important de-facto excitatory afferents to the VTA, it remains elusive whether these signals: (1) respectively encode reward expectation and actual reward and (2) are the only excitatory inputs to the VTA during a classical-conditioning task. As well, it is still unclear whether VTA GABA fully inhibit their dopamine neighbors. Here, we assumed that the activity of DA neurons with no GABAergic input was relatively high (<italic>B</italic><sub>D</sub> &#x0003D; 18 Hz) in order to compensate the observed high baseline activity of GABA neurons (<italic>B</italic><sub>G</sub> &#x0003D; 14 Hz) and get the observed DA tonic firing rate (&#x02243; 5 Hz). This brings up two issues: do these GABA neurons only partially inhibit their dopamine neighbors, for example, just when activated above their baseline? And also, is the inhibitory reward expectation signal mediated by other brain structures as the LHb (Watabe-Uchida et al., <xref ref-type="bibr" rid="B58">2012</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>; Tian and Uchida, <xref ref-type="bibr" rid="B54">2015</xref>)?</p>
<p>In an attempt to answer this question, Tian et al. (<xref ref-type="bibr" rid="B53">2016</xref>) recorded extracellular activity of monosynaptic inputs to dopamine neurons in seven input areas including the PPTg. Showing that many VTA inputs were affected by both CS and US signals, they proposed that DA neurons receive a mix of redundant information and compute a pure RPE signal. However, this does not elucidate which of these inputs effectively affect DA neurons activity during a classical-conditioning task.</p>
<p>While other areas might be implied in RPE computations in the VTA, within our minimal model, we used functional relevant inputs to the VTA that were shown to be strongly affected by reward information based on diverse recurrent studies in the last decades: the working-memory activity in the PFC integrating the timing of reward occurrence (Durstewitz et al., <xref ref-type="bibr" rid="B11">2000</xref>; Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>), the dopamine-mediated plasticity in the NAc via dopamine receptors (Morita et al., <xref ref-type="bibr" rid="B34">2013</xref>; Yagishita et al., <xref ref-type="bibr" rid="B61">2014</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>), the PPTg activation at the reward delivery (Okada et al., <xref ref-type="bibr" rid="B37">2009</xref>; Keiflin and Janak, <xref ref-type="bibr" rid="B24">2015</xref>). Notably, in most of our assumptions, we rely on experimental data that studied neuronal activity of mice performing a simple classical-conditioning task (reward delivery following conditioning cue with no instrumental actions required). In line with this modeling approach, further optogenetic manipulations implying photo-inhibition as in Eshel et al. (<xref ref-type="bibr" rid="B13">2015</xref>) would then be required to study the exact functional impact of the PFC, the NAc and the PPTg on dopamine RPE computations during a simple classical conditioning task.</p>
</sec>
<sec>
<title>4.3. Learning of Reward Expectation in the Corticostriatal Pathway</title>
<p>Our model proposes a specific scenario for PFC-NAc pathway integration of both reward timing and expectation, its biological plausibility is a significant discussion point.</p>
<p>The reward timing learning mechanism exposed in Equation (6) was inspired from Luzardo et al. (<xref ref-type="bibr" rid="B28">2013</xref>), who proposed that reward delivery timing can be learnt by adapting the drift rate of a neural accumulator whose firing rate is expected to reach a specific value at the reward delivery timing. If the reward occurs earlier than expected, the slope of this accumulator is increased. However, if the accumulator reaches its value before US timing, its slope is decreased. Therefore, the rule uses an error signal that is based on time discrepancy between the neural activity reaching a threshold and the reward. Here, we used the same error signal &#x00394;<italic>t</italic>, but the affected parameter is the recurrent excitation strength <italic>J</italic><sub>PFC</sub> and the neural activity dynamics is not an accumulator but an attractor.</p>
<p>We further assumed that this update mechanism could be linked with a potential dopamine-mediated modulation in the PFC (Puig et al., <xref ref-type="bibr" rid="B45">2014</xref>; Popescu et al., <xref ref-type="bibr" rid="B44">2016</xref>) such that &#x003BD;<sub>PFC</sub> rapidly decreases (transition from the active to the rest attractor) at the US timing. Although this dopamine-mediated timing representation hypothesis remains to be directly investigated experimentally, several lines of experimental evidence could support it. First, it is widely accepted that the PFC activity does represent timing information relevant to cognitive tasks through sustained firing activity (Curtis and D&#x00027;Esposito, <xref ref-type="bibr" rid="B6">2003</xref>; Morita et al., <xref ref-type="bibr" rid="B33">2012</xref>; Xu et al., <xref ref-type="bibr" rid="B59">2014</xref>; Connor and Gould, <xref ref-type="bibr" rid="B5">2016</xref>). Second, it has been shown that dopamine enables the induction of spike-timing dependent long-term potentiation (LTP) in layer V PFC pyramidal neurons by acting on D1-receptors (D1R) on excitatory synapses and D2-receptors on local PFC GABAergic interneurons to suppress inhibitory transmission (Xu and Yao, <xref ref-type="bibr" rid="B60">2010</xref>). Moreover, administration of D1 and D2-receptors antagonists in the PFC during learning has been found to impair discrimination of behaviorally relevant events (Popescu et al., <xref ref-type="bibr" rid="B44">2016</xref>).</p>
<p>Additionally, several DA-RPE models proposed a role for the PFC in providing an eligibility trace required in TD-learning algorithms (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B38">2007</xref>; Morita et al., <xref ref-type="bibr" rid="B33">2012</xref>, <xref ref-type="bibr" rid="B34">2013</xref>), considering working-memory representation as crucial in trace conditioning paradigms. Particularly, a specific PFC neuron population, called corticopontine/pyramidal tract (CPn/PT) cells, was assumed by Morita et al. (<xref ref-type="bibr" rid="B33">2012</xref>) and Morita et al. (<xref ref-type="bibr" rid="B34">2013</xref>) to represent the previous state <italic>s</italic>(<italic>t</italic>) or action <italic>a</italic>(<italic>t</italic>) as sustained activity due to the strong recurrent excitatory connections. Note however, that in their model, this signal was supposed to be inhibitory on DA neurons, as it was designed to go through the indirect cortico-striato-VTA pathway, which were assumed to represent <italic>V</italic>(<italic>t</italic>) (Figure <xref ref-type="fig" rid="F8">8</xref>). Here, we consider the sustained PFC signal to be excitatory by acting through the direct cortico-striato-VTA pathway, and that the inhibitory component was held by local VTA GABA neurons. In sum, these studies suggested us to consider the PFC as the main timing integrative component of dopaminergic RPE computations through DA-mediated plasticity.</p>
<p>It would be interesting to consider how CS-related sensory inputs (<italic>w</italic><sub>CS</sub> in the model) can be amplified with learning by sensory neuroplasticity, in addition to the dopamine-mediated effect on cortical recurrent connections (Equation 6). This possibility was tested in our model: by updating <italic>w</italic><sub>CS</sub> in addition to <italic>J</italic><sub>PFC</sub> (PFC recurrent connection strength), PFC neuron activity reaches the Up state earlier. It would then accelerate learning in the PFC but end up with the same maximal activity (obtained at <italic>n</italic> &#x0003D; 6 in Figure <xref ref-type="fig" rid="F3">3A</xref>). Thus, we see that considering sensory representation plasticity is relevant in our context, however it would add another variable to our model without changing the qualitative activity of our neuronal populations. We thus chose not to include these considerations in our minimal model explicitly.</p>
<p>Finally, it is still unclear how DA-mediated plasticity in the striatum could enable the learning of value by striatal neurons. In support of this assumption, it has been suggested that D1R signaling favors synaptic potentiation whereas D2R signaling has the opposite effect (Shen et al., <xref ref-type="bibr" rid="B51">2008</xref>). Moreover, it has been found that in absence of behaviorally important stimuli, DA neurons fire tonically to maintain striatal DA concentrations at levels sufficient to activate D2R, but not low affinity-D1R (Gonon, <xref ref-type="bibr" rid="B20">1997</xref>). We thus considered that dopamine-mediated corticostrial plasticity depended on DA phasic signaling on D1R containing-Medium spiny neurons (MSNs) leading to the activation of the direct excitatory (disinhibitory) pathway to the VTA. Future studies following Morita et al. (<xref ref-type="bibr" rid="B34">2013</xref>) modeling work could focus on the respective implication of D1 and D2R MSNs in corticostriatal plasticity during learning.</p>
</sec>
<sec>
<title>4.4. Nicotine-Induced Effects on nAChRs During Learning</title>
<p>As mentioned above, our local VTA circuit model including nAChRs-mediated current dynamics takes its cue from the minimal model introduced in Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>). This model was later used to explain effects of pharmacological manipulations on nicotinic receptors (Maex et al., <xref ref-type="bibr" rid="B29">2014</xref>), phasic DA response to nicotine injections (Tolu et al., <xref ref-type="bibr" rid="B55">2013</xref>) and the potential impact of receptor up-regulation following prolonged exposure to nicotine (Dumont et al., <xref ref-type="bibr" rid="B9">2018</xref>). In the original work, Graupner et al. (<xref ref-type="bibr" rid="B21">2013</xref>) examined, using computational models, under what conditions (e.g., endogenous cholinergic tone and inputs) one could explain the nicotine-evoked increases in dopamine cell activity and dopamine outflow. To do so, the relative expression of the receptors was parameterised between the DA neurons and the VTA GABA interneurons. In the former case, nicotine would act directly to excite the DA neurons by activating the receptors; in the latter, nicotine would disinhibit the dopamine neurons to increase their firing rate by receptor desensitization. In short, they concluded that both schemes are possible, yet under different endogenous ACh conditions. The direct excitation scheme requires a low Ach tone, while the disinhibition case would yield a robust DA increase under a high ACh tone. We followed the disinhibition scheme since we reasoned that it would be more relevant to behavioral situations where ACh tone is high - notably during motivation-guided behavior and reward seeking (Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>, <xref ref-type="bibr" rid="B41">2012</xref>). Had we considered the direct excitation scheme, certainly the outcomes of our model would be different. Notably, we reason that nicotine would lead to an immediate boost of RPE upon delivery, and then depress the RPE for subsequent CS-US pairings. Whether this is compatible with experimentally observed effects and behavior remains to be explored in subsequent studies.</p>
<p>Desensitization of &#x003B1;4&#x003B2;2-nAChRs on VTA GABA neurons following nicotine exposure results in increased activity of VTA DA neurons (Mansvelder et al., <xref ref-type="bibr" rid="B30">2002</xref>; Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>; Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>). Through the associative-learning mechanism suggested by our model, nicotine exposure would therefore up-regulate DA-response to rewarding events by decreasing the impact of endogenous acetylcholine on VTA GABA neurons provided by the PPTg nucleus activation (Figure <xref ref-type="fig" rid="F6">6</xref>). Together, our results propose that nicotine-mediated nAChRs desensitization potentially enhances the DA response to environmental cues encountered by a smoker (Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>).</p>
<p>Indeed, here, we considered that the rewarding effects of nicotine could be purely contextual: nicotine ingestion does not induce a short rewarding stimulus (US), but an internal state (here, after 5 min of ingestion) that would up-regulate smoker perception of environmental rewards (the taste of coffee) and consequently, when learned, the associated predictive cues (the view of a cup of coffee). While nicotine self-administration experiments considered nAChRs activation as the main rewarding effect of nicotine (Picciotto et al., <xref ref-type="bibr" rid="B40">2008</xref>; Changeux, <xref ref-type="bibr" rid="B3">2010</xref>; Faure et al., <xref ref-type="bibr" rid="B15">2014</xref>), our model focuses on the long-term (min to hours) effects of nicotine that a smoker usually seeks, that interestingly correlates with desensitization kinetics of &#x003B1;4&#x003B2;2-nAChRs (Changeux, <xref ref-type="bibr" rid="B3">2010</xref>).</p>
<p>However, the disinhibition hypothesis on nicotine effects in the VTA remains debated. Although demonstrated <italic>in vitro</italic> (Mansvelder et al., <xref ref-type="bibr" rid="B30">2002</xref>) and <italic>in silico</italic> (Graupner et al., <xref ref-type="bibr" rid="B21">2013</xref>), it is still not clear whether nicotine-induced nAChRs desensitization preferentially acts on GABA neurons within the VTA <italic>in vivo</italic>. This would depend on the ratio of &#x003B1;4&#x003B2;2-nAChRs expression levels <italic>r</italic> but also on the preferential VTA targets of cholinergic axons from the PPTg. While we gathered both components into the parameter <italic>r</italic>, recent studies found that PPTg-to-VTA cholinergic inputs preferentially target either DA neurons (Dautan et al., <xref ref-type="bibr" rid="B7">2016</xref>) or GABA neurons (Yau et al., <xref ref-type="bibr" rid="B62">2016</xref>). Notably, accounting for the relevance of Yau et al. (<xref ref-type="bibr" rid="B62">2016</xref>) experimental conditions&#x02014;photo-inhibition of PPTg-to-VTA cholinergic input during a Pavlovian-conditioning task&#x02014;we chose to preferentially express &#x003B1;4&#x003B2;2-nAChRs on GABA neurons (<italic>r</italic> &#x0003D; 0.2).</p>
<p>It is worth considering that the nicotinic receptors implied in this model are widely expressed throughout the brain. Notably, these are expressed in the PFC on both interneurons and pyramidal neurons, and direct effects of nicotine on the PFC activity has been shown (Picciotto et al., <xref ref-type="bibr" rid="B41">2012</xref>; Poorthuis et al., <xref ref-type="bibr" rid="B43">2013</xref>), together with an impact on VTA DA neurons. Nevertheless, previous work suggests that &#x003B2; 2-containing nAChRs in the VTA are crucial for the animals ability to require stable nicotine self-administration and control the firing patterns of the VTA dopamine neurons (Maskos et al., <xref ref-type="bibr" rid="B31">2005</xref>; Changeux, <xref ref-type="bibr" rid="B3">2010</xref>; Faure et al., <xref ref-type="bibr" rid="B15">2014</xref>). Clearly, our model does not give a full picture of how nicotine may affect learning of motivated behaviors as it does not yet explore the effect of nicotine on cortical dynamics. While we believe this to be a fruitful future direction of study, we would claim that our model gives a minimal sufficient description for the experimental observation that nicotine appears to preferentially boost large vs. small rewards choices through affecting specifically the RPE calculations in the VTA.</p>
</sec>
<sec>
<title>4.5. Predicted Potential Consequences of Nicotine Exposure on Human Decision-Making</title>
<p>In our behavioral simulations of a decision-making task (Figure <xref ref-type="fig" rid="F7">7</xref>), we report that nicotine exposure could potentially bias mice choices toward big rewards. Recent recordings from Faure and colleagues (unpublished data) showed a similar effect of chronic nicotine exposure, with mice showing increasing choices for locations with 100% and 50% reward probabilities at the expense of the location with 25% probability. In this line, future studies could investigate the effects of chronic nicotine on VTA activity during a classical conditioning task as presented here (Figure <xref ref-type="fig" rid="F6">6</xref>) but also on behavioral choices according to reward size (Figure <xref ref-type="fig" rid="F7">7</xref>).</p>
<p>In sum, our minimal model has shown that nicotine would have a double effect on the dopamine signaling of RPE. First, it reopens the window on previously learned rewarding stimuli, where positive error signals are again apparent after the animal has learnt the CS-US association under control conditions (Figure <xref ref-type="fig" rid="F6">6</xref>). Second, when we examine the effects of nicotine on reward-size choices, we see that the new nicotine-released phasic DA signals are disproportionally boosted for large rewards. Hence, we may speculate that nicotine could result in a pathologically increased reward sensitivity to large vs small rewards in decision making and behavior. Such reward sensitivity can lead to an apparent prevalence of exploitative behavior. In other words, if the nicotine-exposed animal overestimate the value of choices disproportionally to others, and base its choices on these values, it would essentially focus on its choices on the over-biased large reward choice at the expense of the under-biased small reward choice. Furthermore, some data indicate that in smokers, delay discounting is abnormal, but not for small immediate and very large delayed rewards (Addicott et al., <xref ref-type="bibr" rid="B1">2013</xref>). Here again, one may associate reward sensitivity as a vehicle, and the mechanisms we suggest playing a role. Nicotine abnormally boosts the value (utility) of the very large reward, relatively depressing the small reward and hence biasing the choice toward the delayed (large) reward, which would appear to resist discounting.</p>
<p>Speculatively, in an environment with high reward volatility, such nicotine-induced exploitation would look like an apparent behavioral rigidity. Several human studies have indeed suggested increased reward sensitivity in smokers (Naud&#x000E9; et al., <xref ref-type="bibr" rid="B35">2015</xref>) and an increase in exploitation vs exploration in smokers versus controls (Addicott et al., <xref ref-type="bibr" rid="B1">2013</xref>). Our model would predict that such behavior would arise from the boosted dopaminergic learning signals due to nicotine action on the VTA circuitry. This is of course with the caveat that in our model we did not discuss the multiple brain decision systems that intervene in real life, but focused exclusively on VTA computations.</p>
<p>The idea that dopamine neurons signal reward-prediction errors has revolutionized the neuronal interpretation of cognitive functions such as reward processing and decision-making. While our qualitative investigations are based on a minimal neuronal circuit dynamics model, our results suggest areas for future theoretical and experimental work that could potentially forge stronger links between dopamine, nicotine, learning, and drug-addiction.</p>
</sec>
</sec>
<sec id="s5">
<title>Author Contributions</title>
<p>ND designed research, performed research, wrote the manuscript. BG designed research, advised ND, obtained funding, wrote the manuscript. VM obtained funding, wrote the manuscript.</p>
<sec>
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="supplementary-material" id="s7">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fncir.2018.00116/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fncir.2018.00116/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Addicott</surname> <given-names>M. A.</given-names></name> <name><surname>Pearson</surname> <given-names>J. M.</given-names></name> <name><surname>Wilson</surname> <given-names>J.</given-names></name> <name><surname>Platt</surname> <given-names>M. L.</given-names></name> <name><surname>McClernon</surname> <given-names>F. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Smoking and the bandit: a preliminary study of smoker and nonsmoker differences in exploratory behavior measured with a multiarmed bandit task</article-title>. <source>Exp. Clin. Psychopharmacol.</source> <volume>21</volume>, <fpage>66</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1037/a0030843</pub-id><pub-id pub-id-type="pmid">23245198</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bayer</surname> <given-names>H. M.</given-names></name> <name><surname>Glimcher</surname> <given-names>P. W.</given-names></name></person-group> (<year>2005</year>). <article-title>Midbrain dopamine neurons encode a quantitative reward prediction error signal</article-title>. <source>Neuron</source> <volume>47</volume>, <fpage>129</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2005.05.020</pub-id><pub-id pub-id-type="pmid">15996553</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Changeux</surname> <given-names>J. P.</given-names></name></person-group> (<year>2010</year>). <article-title>Nicotine addiction and nicotinic receptors: lessons from genetically modified mice</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>11</volume>, <fpage>389</fpage>&#x02013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2849</pub-id><pub-id pub-id-type="pmid">20485364</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J. Y.</given-names></name> <name><surname>Haesler</surname> <given-names>S.</given-names></name> <name><surname>Vong</surname> <given-names>L.</given-names></name> <name><surname>Lowell</surname> <given-names>B. B.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2012</year>). <article-title>Neuron-type-specific signals for reward and punishment in the ventral tegmental area</article-title>. <source>Nature</source> <volume>482</volume>, <fpage>85</fpage>&#x02013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1038/nature10754</pub-id><pub-id pub-id-type="pmid">22258508</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Connor</surname> <given-names>D. A.</given-names></name> <name><surname>Gould</surname> <given-names>T. J.</given-names></name></person-group> (<year>2016</year>). <article-title>The role of working memory and declarative memory in trace conditioning</article-title>. <source>Neurobiol. Learn. Memory</source> <volume>134</volume>, <fpage>193</fpage>&#x02013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.1016/j.nlm.2016.07.009</pub-id><pub-id pub-id-type="pmid">27422017</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Curtis</surname> <given-names>C. E.</given-names></name> <name><surname>D&#x00027;Esposito</surname> <given-names>M.</given-names></name></person-group> (<year>2003</year>). <article-title>Persistent activity in the prefrontal cortex during working memory</article-title>. <source>Trends Cogn. Sci.</source> <volume>7</volume>, <fpage>415</fpage>&#x02013;<lpage>423</lpage>. <pub-id pub-id-type="doi">10.1016/S1364-6613(03)00197-9</pub-id><pub-id pub-id-type="pmid">12963473</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dautan</surname> <given-names>D.</given-names></name> <name><surname>Souza</surname> <given-names>A. S.</given-names></name> <name><surname>Huerta-Ocampo</surname> <given-names>I.</given-names></name> <name><surname>Valencia</surname> <given-names>M.</given-names></name> <name><surname>Assous</surname> <given-names>M.</given-names></name> <name><surname>Witten</surname> <given-names>I. B.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Segregated cholinergic transmission modulates dopamine neurons integrated in distinct functional circuits</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>1025</fpage>&#x02013;<lpage>1033</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4335</pub-id><pub-id pub-id-type="pmid">27348215</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Day</surname> <given-names>J. J.</given-names></name> <name><surname>Carelli</surname> <given-names>R. M.</given-names></name></person-group> (<year>2007</year>). <article-title>The nucleus accumbens and pavlovian reward learning</article-title>. <source>Neuroscientist</source> <volume>13</volume>, <fpage>148</fpage>&#x02013;<lpage>159</lpage>. <pub-id pub-id-type="doi">10.1177/1073858406295854</pub-id><pub-id pub-id-type="pmid">17404375</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dumont</surname> <given-names>G.</given-names></name> <name><surname>Maex</surname> <given-names>R.</given-names></name> <name><surname>Gutkin</surname> <given-names>B.</given-names></name></person-group> (<year>2018</year>). <article-title>Chapter 3-Dopaminergic neurons in the ventral tegmental area and their dysregulation in nicotine addiction</article-title>, in <source>Computational Psychiatry</source>, eds <person-group person-group-type="editor"><name><surname>Anticevic</surname> <given-names>A.</given-names></name> <name><surname>Murray</surname> <given-names>J. D.</given-names></name></person-group> (<publisher-loc>Cambridge</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>47</fpage>&#x02013;<lpage>84</lpage>.</citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Durand-de Cuttoli</surname> <given-names>R.</given-names></name> <name><surname>Mondoloni</surname> <given-names>S.</given-names></name> <name><surname>Marti</surname> <given-names>F.</given-names></name> <name><surname>Lemoine</surname> <given-names>D.</given-names></name> <name><surname>Nguyen</surname> <given-names>C.</given-names></name> <name><surname>Naud&#x000E9;</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Manipulating midbrain dopamine neurons and reward-related behaviors with light-controllable nicotinic acetylcholine receptors</article-title>. <source>eLife</source> <volume>7</volume>:<fpage>e37487</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.37487</pub-id><pub-id pub-id-type="pmid">30176987</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Durstewitz</surname> <given-names>D.</given-names></name> <name><surname>Seamans</surname> <given-names>J. K.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Neurocomputational models of working memory</article-title>. <source>Nat. Neurosci.</source> <volume>3</volume>:<fpage>1184</fpage>. <pub-id pub-id-type="doi">10.1038/81460</pub-id><pub-id pub-id-type="pmid">11127836</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Enomoto</surname> <given-names>K.</given-names></name> <name><surname>Matsumoto</surname> <given-names>N.</given-names></name> <name><surname>Nakai</surname> <given-names>S.</given-names></name> <name><surname>Satoh</surname> <given-names>T.</given-names></name> <name><surname>Sato</surname> <given-names>T. K.</given-names></name> <name><surname>Ueda</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Dopamine neurons learn to encode the long-term value of multiple future rewards</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>108</volume>, <fpage>15462</fpage>&#x02013;<lpage>15467</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1014457108</pub-id><pub-id pub-id-type="pmid">21896766</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eshel</surname> <given-names>N.</given-names></name> <name><surname>Bukwich</surname> <given-names>M.</given-names></name> <name><surname>Rao</surname> <given-names>V.</given-names></name> <name><surname>Hemmelder</surname> <given-names>V.</given-names></name> <name><surname>Tian</surname> <given-names>J.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Arithmetic and local circuitry underlying dopamine prediction errors</article-title>. <source>Nature</source> <volume>525</volume>:<fpage>243</fpage>. <pub-id pub-id-type="doi">10.1038/nature14855</pub-id><pub-id pub-id-type="pmid">26322583</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eshel</surname> <given-names>N.</given-names></name> <name><surname>Tian</surname> <given-names>J.</given-names></name> <name><surname>Bukwich</surname> <given-names>M.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>Dopamine neurons share common response function for reward prediction error</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>479</fpage>&#x02013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4239</pub-id><pub-id pub-id-type="pmid">26854803</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Faure</surname> <given-names>P.</given-names></name> <name><surname>Tolu</surname> <given-names>S.</given-names></name> <name><surname>Valverde</surname> <given-names>S.</given-names></name> <name><surname>Naud&#x000E9;</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Role of nicotinic acetylcholine receptors in regulating dopamine neuron activity</article-title>. <source>Neuroscience</source> <volume>282</volume>, <fpage>86</fpage>&#x02013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroscience.2014.05.040</pub-id><pub-id pub-id-type="pmid">24881574</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiorillo</surname> <given-names>C. D.</given-names></name> <name><surname>Newsome</surname> <given-names>W. T.</given-names></name> <name><surname>Schultz</surname> <given-names>W.</given-names></name></person-group> (<year>2008</year>). <article-title>The temporal precision of reward prediction in dopamine neurons</article-title>. <source>Nat. Neurosci.</source> <volume>11</volume>, <fpage>966</fpage>&#x02013;<lpage>973</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2159</pub-id><pub-id pub-id-type="pmid">18660807</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>S. D.</given-names></name> <name><surname>Robertson</surname> <given-names>P. B.</given-names></name> <name><surname>Black</surname> <given-names>M. J.</given-names></name> <name><surname>Redgrave</surname> <given-names>P.</given-names></name> <name><surname>Sagar</surname> <given-names>M. A.</given-names></name> <name><surname>Abraham</surname> <given-names>W. C.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Reinforcement determines the timing dependence of corticostriatal synaptic plasticity <italic>in vivo</italic></article-title>. <source>Nat. Commun.</source> <volume>8</volume>, <fpage>334</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-017-00394-x</pub-id><pub-id pub-id-type="pmid">28839128</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funahashi</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>Prefrontal cortex and working memory processes</article-title>. <source>Neuroscience</source> <volume>139</volume>, <fpage>251</fpage>&#x02013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroscience.2005.07.003</pub-id><pub-id pub-id-type="pmid">16325345</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Kistler</surname> <given-names>W. M.</given-names></name> <name><surname>Naud</surname> <given-names>R.</given-names></name> <name><surname>Paninski</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <source>Neuronal Dynamics: From Single Neurons to Networks and Models of Cognition</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9781107447615</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gonon</surname> <given-names>F.</given-names></name></person-group> (<year>1997</year>). <article-title>Prolonged and extrasynaptic excitatory action of dopamine mediated by D1 receptors in the rat striatum <italic>in vivo</italic></article-title>. <source>J. Neurosci.</source> <volume>17</volume>, <fpage>5972</fpage>&#x02013;<lpage>5978</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.17-15-05972.1997</pub-id><pub-id pub-id-type="pmid">9221793</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Graupner</surname> <given-names>M.</given-names></name> <name><surname>Maex</surname> <given-names>R.</given-names></name> <name><surname>Gutkin</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>Endogenous cholinergic inputs and local circuit mechanisms govern the phasic mesolimbic dopamine response to nicotine</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003183</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003183</pub-id><pub-id pub-id-type="pmid">23966848</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hyland</surname> <given-names>B.</given-names></name> <name><surname>Reynolds</surname> <given-names>J.</given-names></name> <name><surname>Hay</surname> <given-names>J.</given-names></name> <name><surname>Perk</surname> <given-names>C.</given-names></name> <name><surname>Miller</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>Firing modes of midbrain dopamine cells in the freely moving rat</article-title>. <source>Neuroscience</source> <volume>114</volume>, <fpage>475</fpage>&#x02013;<lpage>492</lpage>. <pub-id pub-id-type="doi">10.1016/S0306-4522(02)00267-1</pub-id><pub-id pub-id-type="pmid">12204216</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ishikawa</surname> <given-names>A.</given-names></name> <name><surname>Ambroggi</surname> <given-names>F.</given-names></name> <name><surname>Nicola</surname> <given-names>S. M.</given-names></name> <name><surname>Fields</surname> <given-names>H. L.</given-names></name></person-group> (<year>2008</year>). <article-title>Dorsomedial prefrontal cortex contribution to behavioral and nucleus accumbens neuronal responses to incentive cues</article-title>. <source>J. Neurosci.</source> <volume>28</volume>, <fpage>5088</fpage>&#x02013;<lpage>5098</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0253-08.2008</pub-id><pub-id pub-id-type="pmid">18463262</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keiflin</surname> <given-names>R.</given-names></name> <name><surname>Janak</surname> <given-names>P. H.</given-names></name></person-group> (<year>2015</year>). <article-title>Dopamine prediction errors in reward learning and addiction: from theory to neural circuitry</article-title>. <source>Neuron</source> <volume>88</volume>, <fpage>247</fpage>&#x02013;<lpage>263</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.08.037</pub-id><pub-id pub-id-type="pmid">26494275</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kobayashi</surname> <given-names>Y.</given-names></name> <name><surname>Okada</surname> <given-names>K. I.</given-names></name></person-group> (<year>2007</year>). <article-title>Reward prediction error computation in the pedunculopontine tegmental nucleus neurons</article-title>. <source>Ann. N.Y. Acad. Sci.</source> <volume>1104</volume>, <fpage>310</fpage>&#x02013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1196/annals.1390.003</pub-id><pub-id pub-id-type="pmid">17344541</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le Merre</surname> <given-names>P.</given-names></name> <name><surname>Esmaeili</surname> <given-names>V.</given-names></name> <name><surname>Charri&#x000E8;re</surname> <given-names>E.</given-names></name> <name><surname>Galan</surname> <given-names>K.</given-names></name> <name><surname>Salin</surname> <given-names>P.-A.</given-names></name> <name><surname>Petersen</surname> <given-names>C. C.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Reward-based learning drives rapid sensory signals in medial prefrontal cortex and dorsal hippocampus necessary for goal-directed behavior</article-title>. <source>Neuron</source> <volume>97</volume>, <fpage>83.e5</fpage>&#x02013;<lpage>91.e5</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2017.11.031</pub-id><pub-id pub-id-type="pmid">29249287</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lokwan</surname> <given-names>S. J. A.</given-names></name> <name><surname>Overton</surname> <given-names>P. G.</given-names></name> <name><surname>Berry</surname> <given-names>M. S.</given-names></name> <name><surname>Clark</surname> <given-names>D.</given-names></name></person-group> (<year>1999</year>). <article-title>Stimulation of the pedunculopontine tegmental nucleus in the rat produces burst firing in A9 dopaminergic neurons</article-title>. <source>Neuroscience</source> <volume>92</volume>, <fpage>245</fpage>&#x02013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1016/S0306-4522(98)00748-9</pub-id><pub-id pub-id-type="pmid">10392847</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luzardo</surname> <given-names>A.</given-names></name> <name><surname>Ludvig</surname> <given-names>E. A.</given-names></name> <name><surname>Rivest</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>An adaptive drift-diffusion model of interval timing dynamics</article-title>. <source>Behav. Proc.</source> <volume>95</volume>, <fpage>90</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.beproc.2013.02.003</pub-id><pub-id pub-id-type="pmid">23428705</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maex</surname> <given-names>R.</given-names></name> <name><surname>Grinevich</surname> <given-names>V. P.</given-names></name> <name><surname>Grinevich</surname> <given-names>V.</given-names></name> <name><surname>Budygin</surname> <given-names>E.</given-names></name> <name><surname>Bencherif</surname> <given-names>M.</given-names></name> <name><surname>Gutkin</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>Understanding the role &#x003B1;7 nicotinic receptors play in dopamine efflux in nucleus accumbens</article-title>. <source>ACS Chem. Neurosci.</source> <volume>5</volume>, <fpage>1032</fpage>&#x02013;<lpage>1040</lpage>. <pub-id pub-id-type="doi">10.1021/cn500126t</pub-id><pub-id pub-id-type="pmid">25147933</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mansvelder</surname> <given-names>H. D.</given-names></name> <name><surname>Keath</surname> <given-names>J.</given-names></name> <name><surname>McGehee</surname> <given-names>D. S.</given-names></name></person-group> (<year>2002</year>). <article-title>Synaptic mechanisms underlie nicotine-induced excitability of brain reward areas</article-title>. <source>Neuron</source> <volume>33</volume>, <fpage>905</fpage>&#x02013;<lpage>919</lpage>. <pub-id pub-id-type="doi">10.1016/S0896-6273(02)00625-6</pub-id><pub-id pub-id-type="pmid">11906697</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maskos</surname> <given-names>U.</given-names></name> <name><surname>Molles</surname> <given-names>B. E.</given-names></name> <name><surname>Pons</surname> <given-names>S.</given-names></name> <name><surname>Besson</surname> <given-names>M.</given-names></name> <name><surname>Guiard</surname> <given-names>B. P.</given-names></name> <name><surname>Guilloux</surname> <given-names>J.-P.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Nicotine reinforcement and cognition restored by targeted expression of nicotinic receptors</article-title>. <source>Nature</source> <volume>436</volume>, <fpage>103</fpage>&#x02013;<lpage>107</lpage>. <pub-id pub-id-type="doi">10.1038/nature03694</pub-id><pub-id pub-id-type="pmid">16001069</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matsumoto</surname> <given-names>M.</given-names></name> <name><surname>Hikosaka</surname> <given-names>O.</given-names></name></person-group> (<year>2009</year>). <article-title>Two types of dopamine neuron distinctly convey positive and negative motivational signals</article-title>. <source>Nature</source> <volume>459</volume>, <fpage>837</fpage>&#x02013;<lpage>841</lpage>. <pub-id pub-id-type="doi">10.1038/nature08028</pub-id><pub-id pub-id-type="pmid">19448610</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morita</surname> <given-names>K.</given-names></name> <name><surname>Morishima</surname> <given-names>M.</given-names></name> <name><surname>Sakai</surname> <given-names>K.</given-names></name> <name><surname>Kawaguchi</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Reinforcement learning: computing the temporal difference of values via distinct corticostriatal pathways</article-title>. <source>Trends Neurosci.</source> <volume>35</volume>, <fpage>457</fpage>&#x02013;<lpage>467</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2012.04.009</pub-id><pub-id pub-id-type="pmid">22658226</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morita</surname> <given-names>K.</given-names></name> <name><surname>Morishima</surname> <given-names>M.</given-names></name> <name><surname>Sakai</surname> <given-names>K.</given-names></name> <name><surname>Kawaguchi</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Dopaminergic control of motivation and reinforcement learning: a closed-circuit account for reward-oriented behavior</article-title>. <source>J. Neurosci.</source> <volume>33</volume>, <fpage>8866</fpage>&#x02013;<lpage>8890</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4614-12.2013</pub-id><pub-id pub-id-type="pmid">23678129</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naud&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Dongelmans</surname> <given-names>M.</given-names></name> <name><surname>Faure</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Nicotinic alteration of decision-making</article-title>. <source>Neuropharmacology</source> <volume>96</volume>, <fpage>244</fpage>&#x02013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuropharm.2014.11.021</pub-id><pub-id pub-id-type="pmid">25498234</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naud&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Tolu</surname> <given-names>S.</given-names></name> <name><surname>Dongelmans</surname> <given-names>M.</given-names></name> <name><surname>Torquet</surname> <given-names>N.</given-names></name> <name><surname>Valverde</surname> <given-names>S.</given-names></name> <name><surname>Rodriguez</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Nicotinic receptors in the ventral tegmental area promote uncertainty-seeking</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>471</fpage>&#x02013;<lpage>478</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4223</pub-id><pub-id pub-id-type="pmid">26780509</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Okada</surname> <given-names>K. I.</given-names></name> <name><surname>Toyama</surname> <given-names>K.</given-names></name> <name><surname>Inoue</surname> <given-names>Y.</given-names></name> <name><surname>Isa</surname> <given-names>T.</given-names></name> <name><surname>Kobayashi</surname> <given-names>Y.</given-names></name></person-group> (<year>2009</year>). <article-title>Different pedunculopontine tegmental neurons signal predicted and actual task rewards</article-title>. <source>J. Neurosci.</source> <volume>29</volume>, <fpage>4858</fpage>&#x02013;<lpage>4870</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4415-08.2009</pub-id><pub-id pub-id-type="pmid">19369554</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Frank</surname> <given-names>M. J.</given-names></name> <name><surname>Hazy</surname> <given-names>T. E.</given-names></name> <name><surname>Watz</surname> <given-names>B.</given-names></name></person-group> (<year>2007</year>). <article-title>PVLV: the primary value and learned value pavlovian learning algorithm</article-title>. <source>Behav. Neurosci.</source> <volume>121</volume>, <fpage>31</fpage>&#x02013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1037/0735-7044.121.1.31</pub-id><pub-id pub-id-type="pmid">17324049</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oyama</surname> <given-names>K.</given-names></name> <name><surname>Tateyama</surname> <given-names>Y.</given-names></name> <name><surname>Hern&#x000E1;di</surname> <given-names>I.</given-names></name> <name><surname>Tobler</surname> <given-names>P. N.</given-names></name> <name><surname>Iijima</surname> <given-names>T.</given-names></name> <name><surname>Tsutsui</surname> <given-names>K.-I.</given-names></name></person-group> (<year>2015</year>). <article-title>Discrete coding of stimulus value, reward expectation, and reward prediction error in the dorsal striatum</article-title>. <source>J. Neurophysiol.</source> <volume>114</volume>, <fpage>2600</fpage>&#x02013;<lpage>2615</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00097.2015</pub-id><pub-id pub-id-type="pmid">26378201</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Picciotto</surname> <given-names>M.</given-names></name> <name><surname>Addy</surname> <given-names>N.</given-names></name> <name><surname>Mineur</surname> <given-names>Y.</given-names></name> <name><surname>Brunzell</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>It is not &#x0201C;either/or&#x0201D;: Activation and desensitization of nicotinic acetylcholine receptors both contribute to behaviors related to nicotine addiction and mood</article-title>. <source>Prog. Neurobiol.</source> <volume>84</volume>, <fpage>329</fpage>&#x02013;<lpage>342</lpage>. <pub-id pub-id-type="doi">10.1016/j.pneurobio.2007.12.005</pub-id><pub-id pub-id-type="pmid">18242816</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Picciotto</surname> <given-names>M. R.</given-names></name> <name><surname>Higley</surname> <given-names>M. J.</given-names></name> <name><surname>Mineur</surname> <given-names>Y. S.</given-names></name></person-group> (<year>2012</year>). <article-title>Acetylcholine as a neuromodulator: cholinergic signaling shapes nervous system function and behavior</article-title>. <source>Neuron</source> <volume>76</volume>, <fpage>116</fpage>&#x02013;<lpage>129</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.08.036</pub-id><pub-id pub-id-type="pmid">23040810</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pontieri</surname> <given-names>F. E.</given-names></name> <name><surname>Tanda</surname> <given-names>G.</given-names></name> <name><surname>Orzi</surname> <given-names>F.</given-names></name> <name><surname>Chiara</surname> <given-names>G. D.</given-names></name></person-group> (<year>1996</year>). <article-title>Effects of nicotine on the nucleus accumbens and similarity to those of addictive drugs</article-title>. <source>Nature</source> <volume>382</volume>:<fpage>255</fpage>. <pub-id pub-id-type="doi">10.1038/382255a0</pub-id><pub-id pub-id-type="pmid">8717040</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poorthuis</surname> <given-names>R. B.</given-names></name> <name><surname>Bloem</surname> <given-names>B.</given-names></name> <name><surname>Schak</surname> <given-names>B.</given-names></name> <name><surname>Wester</surname> <given-names>J.</given-names></name> <name><surname>de Kock</surname> <given-names>C. P. J.</given-names></name> <name><surname>Mansvelder</surname> <given-names>H. D.</given-names></name></person-group> (<year>2013</year>). <article-title>Layer-specific modulation of the prefrontal cortex by nicotinic acetylcholine receptors</article-title>. <source>Cereb. Cortex</source> <volume>23</volume>, <fpage>148</fpage>&#x02013;<lpage>161</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhr390</pub-id><pub-id pub-id-type="pmid">22291029</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Popescu</surname> <given-names>A. T.</given-names></name> <name><surname>Zhou</surname> <given-names>M. R.</given-names></name> <name><surname>Poo</surname> <given-names>M. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Phasic dopamine release in the medial prefrontal cortex enhances stimulus discrimination</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>113</volume>, <fpage>E3169</fpage>&#x02013;<lpage>E3176</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1606098113</pub-id><pub-id pub-id-type="pmid">27185946</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Puig</surname> <given-names>M.</given-names></name> <name><surname>Antzoulatos</surname> <given-names>E.</given-names></name> <name><surname>Miller</surname> <given-names>E.</given-names></name></person-group> (<year>2014</year>). <article-title>Prefrontal dopamine in associative learning and memory</article-title>. <source>Neuroscience</source> <volume>282</volume>, <fpage>217</fpage>&#x02013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroscience.2014.09.026</pub-id><pub-id pub-id-type="pmid">25241063</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rescorla</surname> <given-names>R. A.</given-names></name> <name><surname>Wagner</surname> <given-names>A. W.</given-names></name></person-group> (<year>1972</year>). <article-title>A theory of Pavlovian conditioning: variations in the effectiveness of reinforcement and nonreinforcement</article-title>, in <source>Classical Conditioning II: Current Theory and Research</source>, <person-group person-group-type="editor"><name><surname>Black</surname> <given-names>A. H.</given-names></name> <name><surname>Prokazy</surname> <given-names>W. F.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Appleton-Century-Crofts</publisher-name>), <fpage>64</fpage>&#x02013;<lpage>99</lpage>.</citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reynolds</surname> <given-names>J. N. J.</given-names></name> <name><surname>Hyland</surname> <given-names>B. I.</given-names></name> <name><surname>Wickens</surname> <given-names>J. R.</given-names></name></person-group> (<year>2001</year>). <article-title>A cellular mechanism of reward-related learning</article-title>. <source>Nature</source> <volume>413</volume>:<fpage>67</fpage>. <pub-id pub-id-type="doi">10.1038/35092560</pub-id><pub-id pub-id-type="pmid">11544526</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schoenbaum</surname> <given-names>G.</given-names></name> <name><surname>Chiba</surname> <given-names>A. A.</given-names></name> <name><surname>Gallagher</surname> <given-names>M.</given-names></name></person-group> (<year>1998</year>). <article-title>Orbitofrontal cortex and basolateral amygdala encode expected outcomes during learning</article-title>. <source>Nat. Neurosci.</source> <volume>1</volume>:<fpage>155</fpage>. <pub-id pub-id-type="doi">10.1038/407</pub-id><pub-id pub-id-type="pmid">10195132</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schultz</surname> <given-names>W.</given-names></name></person-group> (<year>1998</year>). <article-title>Predictive reward signal of dopamine neurons</article-title>. <source>J. Neurophysiol.</source> <volume>80</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1152/jn.1998.80.1.1</pub-id><pub-id pub-id-type="pmid">9658025</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schultz</surname> <given-names>W.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Montague</surname> <given-names>R. P.</given-names></name></person-group> (<year>1997</year>). <article-title>A neural substrate of prediction and reward</article-title>. <source>Science</source> <volume>275</volume>, <fpage>1593</fpage>&#x02013;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1126/science.275.5306.1593</pub-id><pub-id pub-id-type="pmid">9054347</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>W.</given-names></name> <name><surname>Flajolet</surname> <given-names>M.</given-names></name> <name><surname>Greengard</surname> <given-names>P.</given-names></name> <name><surname>Surmeier</surname> <given-names>D. J.</given-names></name></person-group> (<year>2008</year>). <article-title>Dichotomous dopaminergic control of striatal synaptic plasticity</article-title>. <source>Science</source> <volume>321</volume>, <fpage>848</fpage>&#x02013;<lpage>851</lpage>. <pub-id pub-id-type="doi">10.1126/science.1160575</pub-id><pub-id pub-id-type="pmid">18687967</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source>Reinforcement Learning: An Introduction.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>R.</given-names></name> <name><surname>Cohen</surname> <given-names>J. Y.</given-names></name> <name><surname>Osakada</surname> <given-names>F.</given-names></name> <name><surname>Kobak</surname> <given-names>D.</given-names></name> <name><surname>Machens</surname> <given-names>C. K.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Distributed and mixed information in monosynaptic inputs to dopamine neurons</article-title>. <source>Neuron</source> <volume>91</volume>, <fpage>1374</fpage>&#x02013;<lpage>1389</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.08.018</pub-id><pub-id pub-id-type="pmid">27618675</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>J.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Habenula lesions reveal that multiple mechanisms underlie dopamine prediction errors</article-title>. <source>Neuron</source> <volume>87</volume>, <fpage>1304</fpage>&#x02013;<lpage>1316</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.08.028</pub-id><pub-id pub-id-type="pmid">26365765</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tolu</surname> <given-names>S.</given-names></name> <name><surname>Eddine</surname> <given-names>R.</given-names></name> <name><surname>Marti</surname> <given-names>F.</given-names></name> <name><surname>David</surname> <given-names>V.</given-names></name> <name><surname>Graupner</surname> <given-names>M.</given-names></name> <name><surname>Pons</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Co-activation of VTA DA and GABA neurons mediates nicotine reinforcement</article-title>. <source>Mol. Psychiatry</source> <volume>18</volume>, <fpage>382</fpage>&#x02013;<lpage>393</lpage>. <pub-id pub-id-type="doi">10.1038/mp.2012.83</pub-id><pub-id pub-id-type="pmid">22751493</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vitay</surname> <given-names>J.</given-names></name> <name><surname>Hamker</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>Timing and expectation of reward: a neuro-computational model of the afferents to the ventral tegmental area</article-title>. <source>Front. Neurorob.</source> <volume>8</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2014.00004</pub-id><pub-id pub-id-type="pmid">24550821</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watabe-Uchida</surname> <given-names>M.</given-names></name> <name><surname>Eshel</surname> <given-names>N.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>Neural circuitry of reward prediction error</article-title>. <source>Ann. Rev. Neurosci.</source> <volume>40</volume>, <fpage>373</fpage>&#x02013;<lpage>394</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-neuro-072116-031109</pub-id><pub-id pub-id-type="pmid">28441114</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watabe-Uchida</surname> <given-names>M.</given-names></name> <name><surname>Zhu</surname> <given-names>L.</given-names></name> <name><surname>Ogawa</surname> <given-names>S. K.</given-names></name> <name><surname>Vamanrao</surname> <given-names>A.</given-names></name> <name><surname>Uchida</surname> <given-names>N.</given-names></name></person-group> (<year>2012</year>). <article-title>Whole-brain mapping of direct inputs to midbrain dopamine neurons</article-title>. <source>Neuron</source> <volume>74</volume>, <fpage>858</fpage>&#x02013;<lpage>873</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.03.017</pub-id><pub-id pub-id-type="pmid">22681690</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>S.-Y.</given-names></name> <name><surname>Dan</surname> <given-names>Y.</given-names></name> <name><surname>Poo</surname> <given-names>M. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Representation of interval timing by temporally scalable firing patterns in rat prefrontal cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>480</fpage>&#x02013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1321314111</pub-id><pub-id pub-id-type="pmid">24367075</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>T.-X.</given-names></name> <name><surname>Yao</surname> <given-names>W.-D.</given-names></name></person-group> (<year>2010</year>). <article-title>D1 and D2 dopamine receptors in separate circuits cooperate to drive associative long-term potentiation in the prefrontal cortex</article-title>. <source>Proc. Natl. Acade. Sci. U.S.A.</source> <volume>107</volume>, <fpage>16366</fpage>&#x02013;<lpage>16371</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1004108107</pub-id><pub-id pub-id-type="pmid">20805489</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yagishita</surname> <given-names>S.</given-names></name> <name><surname>Hayashi-Takagi</surname> <given-names>A.</given-names></name> <name><surname>Ellis-Davies</surname> <given-names>G. C.</given-names></name> <name><surname>Urakubo</surname> <given-names>H.</given-names></name> <name><surname>Ishii</surname> <given-names>S.</given-names></name> <name><surname>Kasai</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>A critical time window for dopamine actions on the structural plasticity of dendritic spines</article-title>. <source>Science</source> <volume>345</volume>, <fpage>1616</fpage>&#x02013;<lpage>1620</lpage>. <pub-id pub-id-type="doi">10.1126/science.1255514</pub-id><pub-id pub-id-type="pmid">25258080</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yau</surname> <given-names>H.-J.</given-names></name> <name><surname>Wang</surname> <given-names>D. V.</given-names></name> <name><surname>Tsou</surname> <given-names>J.-H.</given-names></name> <name><surname>Chuang</surname> <given-names>Y.-F.</given-names></name> <name><surname>Chen</surname> <given-names>B. T.</given-names></name> <name><surname>Deisseroth</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Pontomesencephalic tegmental afferents to VTA non-dopamine neurons are necessary for appetitive pavlovian learning</article-title>. <source>Cell Rep.</source> <volume>16</volume>, <fpage>2699</fpage>&#x02013;<lpage>2710</lpage>. <pub-id pub-id-type="doi">10.1016/j.celrep.2016.08.007</pub-id><pub-id pub-id-type="pmid">27568569</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoo</surname> <given-names>J. H.</given-names></name> <name><surname>Zell</surname> <given-names>V.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Punta</surname> <given-names>C.</given-names></name> <name><surname>Ramajayam</surname> <given-names>N.</given-names></name> <name><surname>Shen</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Activation of pedunculopontine glutamate neurons is reinforcing</article-title>. <source>J. Neurosci.</source> <volume>37</volume>, <fpage>38</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3082-16.2016</pub-id><pub-id pub-id-type="pmid">28053028</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> ND acknowledges funding from the &#x000C9;cole Normale Sup&#x000E9;rieure and INSERM. BG acknowledges partial support from INSERM, CNRS, LABEX ANR-10-LABX-0087 IEC, and from IDEX ANR-10-IDEX-0001-02 PSL<sup>&#x0002A;</sup> as well as from HSE Basic Research Program and the Russian Academic Excellence Project &#x0201C;5-100.&#x0201D; VM received funding from HSE Basic Research Program and the Russian Academic Excellence Project &#x0201C;5-100.&#x0201D;</p></fn>
</fn-group>
</back>
</article>