<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Hum. Neurosci.</journal-id>
<journal-title>Frontiers in Human Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Hum. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5161</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnhum.2017.00592</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Functions of Learning Rate in Adaptive Reward Learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Wu</surname> <given-names>Xi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Ting</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Chang</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Wu</surname> <given-names>Tao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Jiang</surname> <given-names>Jiefeng</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhou</surname> <given-names>Dong</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/441320/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhou</surname> <given-names>Jiliu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/477263/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Computer Science, Chengdu University of Information Technology</institution>, <addr-line>Chengdu</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>College of Information Science and Engineering, Chengdu University</institution>, <addr-line>Chengdu</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Psychology, Stanford University</institution>, <addr-line>Stanford, CA</addr-line>, <country>United States</country></aff>
<aff id="aff4"><sup>4</sup><institution>Department of Neurology, West China Hospital, Sichuan University</institution>, <addr-line>Chengdu</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Carol Seger, Colorado State University, Fort Collins, United States</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Ling Wang, South China Normal University, China; Ekaterina Dobryakova, Kessler Foundation, United States</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Jiliu Zhou, <email>zhoujiliu@cuit.edu.cn</email></italic></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>12</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>592</elocation-id>
<history>
<date date-type="received">
<day>15</day>
<month>09</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>11</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Wu, Wang, Liu, Wu, Jiang, Zhou and Zhou.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Wu, Wang, Liu, Wu, Jiang, Zhou and Zhou</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>As a crucial cognitive function, learning applies prediction error (the discrepancy between the prediction from learning and the world state) to adjust predictions of the future. How much prediction error affects this adjustment also depends on the learning rate. Our understanding to the learning rate is still limited, in terms of (1) how it is modulated by other factors, and (2) the specific mechanisms of how learning rate interacts with prediction error to update learning. We applied computational modeling and functional magnetic resonance imaging to investigate these issues. We found that, when human participants performed a reward learning task, reward magnitude modulated learning rate. Modulation strength further predicted the difference in behavior following high vs. low reward across subjects. Imaging results further showed that this modulation was reflected in brain regions where the reward feedback is also encoded, such as the medial prefrontal cortex (MFC), precuneus, and posterior cingulate cortex. Furthermore, for the first time, we observed that the integration of the learning rate and the reward prediction error was represented in MFC activity. These findings extend our understanding of adaptive learning by demonstrating how it functions in a chain reaction of prediction updating.</p>
</abstract>
<kwd-group>
<kwd>adaptive learning</kwd>
<kwd>Bayesian modeling</kwd>
<kwd>fMRI</kwd>
<kwd>learning rate</kwd>
<kwd>reward</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="1"/>
<equation-count count="9"/>
<ref-count count="49"/>
<page-count count="15"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>The brain generalizes learned information to make predictions of the future. To improve the accuracy of these predictions, the learning process must incorporate new information to reflect the up-to-date states of the environment. The integration of new information involves two key factors: the prediction error that measures the discrepancy between current prediction and the observed environmental state, and the learning rate that determines to what degree the prediction error is applied to updating the prediction.</p>
<p>Prediction error has been thought to be calculated by the dopaminergic activity in the ventral tegmental area (VTA), then spreading to other brain regions via the afferent connections from the VTA (<xref ref-type="bibr" rid="B44">Schultz et al., 1997</xref>). Consistent with this theory, many functional magnetic resonance imaging (fMRI) studies have located brain structures in humans that reflect prediction error (for review, see <xref ref-type="bibr" rid="B16">Garrison et al., 2013</xref>). For example, <xref ref-type="bibr" rid="B14">D&#x2019;Ardenne et al. (2008)</xref> detected activity related to the reward prediction error in the VTA of human participants. Outside the VTA, several researchers have also documented activity related to reward prediction error in regions that are connected to the VTA. Examples include subcortical structures such as the striatum (<xref ref-type="bibr" rid="B34">O&#x2019;Doherty et al., 2003</xref>; <xref ref-type="bibr" rid="B45">Seymour et al., 2004</xref>; <xref ref-type="bibr" rid="B38">Pessiglione et al., 2006</xref>; <xref ref-type="bibr" rid="B19">Glascher et al., 2010</xref>; <xref ref-type="bibr" rid="B25">Jocham et al., 2011</xref>; <xref ref-type="bibr" rid="B49">Zhu et al., 2012</xref>; <xref ref-type="bibr" rid="B15">Eppinger et al., 2013</xref>) and nucleus accumbens (<xref ref-type="bibr" rid="B33">Niv et al., 2012</xref>). Cortical areas such as the medial prefrontal cortex (MFC) (<xref ref-type="bibr" rid="B25">Jocham et al., 2011</xref>; <xref ref-type="bibr" rid="B15">Eppinger et al., 2013</xref>) also appear to be involved.</p>
<p>A hallmark of learning is its flexibility, that is, the adaptive employment of prediction error in adjusting the prediction. This flexibility is attributed to the learning rate. To study this flexibility in adaptive learning, recent studies have employed computational models and algorithms to demonstrate how the learning rate shifts following changes in the environment (<xref ref-type="bibr" rid="B4">Behrens et al., 2007</xref>; <xref ref-type="bibr" rid="B32">Nassar et al., 2010</xref>; <xref ref-type="bibr" rid="B36">Payzan-LeNestour and Bossaerts, 2011</xref>; <xref ref-type="bibr" rid="B24">Jiang et al., 2014</xref>, <xref ref-type="bibr" rid="B23">2015</xref>; <xref ref-type="bibr" rid="B30">McGuire et al., 2014</xref>). A common finding is that learning rate should favor recent information more, if there is change in the environment, both to reflect the up-do-date world state and to dampen the influence of outdated information. By contrast, if the environment is stable, learning rate should depend more on information sampled over an extended period of time (as opposed to recent information), so that learning is more robust against noise. In the brain, this change in the learning rate is associated with the anterior cingulate cortex (ACC) (<xref ref-type="bibr" rid="B4">Behrens et al., 2007</xref>), the anterior insula and adjacent inferior frontal gyrus (IFG) (<xref ref-type="bibr" rid="B30">McGuire et al., 2014</xref>; <xref ref-type="bibr" rid="B23">Jiang et al., 2015</xref>), and the MFC (<xref ref-type="bibr" rid="B30">McGuire et al., 2014</xref>).</p>
<p>However, many questions regarding the mechanisms of the learning rate remain unanswered. One question is whether other factors (besides volatility or rate of change in the environment) mediate the learning rate. In reward learning tasks, a possible candidate factor is reward feedback, known from past research to affect the subsequent strategy of humans performing a gambling task (<xref ref-type="bibr" rid="B17">Gehring and Willoughby, 2002</xref>; <xref ref-type="bibr" rid="B48">Yeung and Sanfey, 2004</xref>). Similarly, humans seem to use asymmetric learning rates for positive and negative prediction errors (<xref ref-type="bibr" rid="B33">Niv et al., 2012</xref>; <xref ref-type="bibr" rid="B18">Gershman, 2015</xref>). Another important, yet unanswered, question is how the learning rate and the prediction error are integrated to drive learning. In the reinforcement learning model (<xref ref-type="bibr" rid="B41">Rescorla and Wagner, 1972</xref>), the updating of prediction at time i + 1 (denoted as &#x0394;p<sub>i+1</sub>) is the prediction error (PE<sub>i</sub>) multiplied by the learning rate (<italic>&#x03B1;<sub>i</sub></italic>) at time <italic>i</italic>; thus, &#x0394;p<sub>i+1</sub> = &#x03B1;<sub>i</sub> &#x00D7; PE<sub>i</sub>. Therefore, this joint effect of learning rate and prediction error on updating prediction can be tested as their interaction.</p>
<p>We hypothesized that: (1) the learning rate would be mediated by reward feedback; and (2) the integration of the learning rate and the prediction error would occur in the MFC, which is associated with both factors. To test these two hypotheses, we proposed a Bayesian model that provides computational mechanisms explaining how reward feedback influences the learning rate. Crucially, this model inferred the learning rate using the actual choices made by participants and the resulting reward, thus it accounts for individual differences and yield inference of the subjective learning rate. We further applied the model estimates to fMRI data, and provide new evidence of the neural substrates supporting these two hypotheses.</p>
</sec>
<sec id="s1" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec><title>Subjects</title>
<p>Twenty-nine college students participated in this study. All participants had normal or corrected-to-normal vision. Two participants did not finish the task and were excluded from analysis, so the sample for behavioral analysis consisted of 27 participants (14 females, 20&#x2013;24 years old, mean age = 22 years). In addition, four participants were excluded due to synchronization failure (the fMRI scanning and task did not start simultaneously) in at least one run, so the onsets of events in the behavioral task could not be mapped to the fMRI data. Data from two more participants were excluded due to normalization failure (i.e., SPM produced distorted normalized images after the normalization step; see below). Therefore, the final sample of fMRI analysis consisted of 21 participants (10 females, 20&#x2013;23 years old, mean age = 22 years). This study was approved by the institutional review board of Chengdu University of Information Technology.</p>
</sec>
<sec><title>Apparatus and Experimental Design</title>
<p>The task used in this study was programmed using Psychophysics Toolbox Version 3<sup><xref ref-type="fn" rid="fn01">1</xref></sup>. The stimuli (i.e., a red square and a green square) were displayed on a back projection screen. Participants viewed the display via a mirror attached to the head coil of the MR scanner and responded using two MR-compatible button boxes, one for each hand.</p>
<p><bold>Figure <xref ref-type="fig" rid="F1">1A</xref></bold> depicts the flow of events in the behavioral task trials. In each trial, participants chose between a red square and a green square to accumulate reward points, which determined the monetary reward they received after the task. At the beginning of each trial, a fixation cross appeared at the center of the screen, along with the two colored squares to the left and right of the cross for an exponentially jittered interval (4&#x2013;5.5 s, step size = 0.5 s), during which the participants chose one square by pressing the button on the same side as the selected square. Once a response was made, the unselected square disappeared. After this interval, the reward gained (either 1 point or 5 points, corresponding to a low reward or a high reward, respectively) was presented at the center of the screen for 1 s, followed by the fixation cross shown for another exponentially jittered inter-trial interval (4&#x2013;5.5 s, step size = 0.5 s), following which the stimulus display for the next trial appeared. If the participant did not respond, the trial was scored as <italic>no response</italic>. In case of no response (&#x003C;0.3% of all trials), no points were gained and a message &#x201C;+0&#x201D; was shown. The lack of response in these trials precludes the inference of the mental states, therefore, the no response trials were excluded from behavioral and imaging analyses. The total points gained in the current run were displayed at the bottom of the screen throughout the task.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Task, experimental design and the graphical representation of the flexible learning model. <bold>(A)</bold> An example trial. Two color squares were presented on the screen for an interval, during which the red square was chosen, causing the green square being removed from the screen. The choice of the red square resulted in a gain of five points, which was displayed on the screen after the interval. The trial ended with a presentation of a fixation cross. <bold>(B)</bold> The four different time courses of the probability of getting high reward by selecting the red square used in this task. <bold>(C)</bold> The graphical representation of the generative model. Each node represents the state of a model variable. The edges show the flow of the information.</p></caption>
<graphic xlink:href="fnhum-11-00592-g001.tif"/>
</fig>
<p>This task consisted of eight runs of 40 trials each. The participants were instructed to gain as many points as possible. Additionally, the participants were informed that at each trial, one color was more likely to lead to high reward than the other color. Further, we told participants that the more highly rewarded color would reset at the beginning of each run, and might change during the course of a run. Unbeknownst to the participants, at any trial, the sum of the two colors&#x2019; probabilities of getting a high reward was always 1. This constraint was used to keep the chance level of getting high reward at 50%. In order to create a wide range of high reward probabilities for the more rewarded color, we created two conditions (four runs for each condition, with the order of runs counterbalanced both within and across participants, <bold>Figure <xref ref-type="fig" rid="F1">1B</xref></bold>): In a <italic>hard</italic> run (i.e., probability shifted within a run thus making the more rewarded color difficult to track), the underlying probability of getting a high reward by selecting the red square, changed every four trials, either in the order of 0.2, 0.4, 0.6, 0.8, 0.6, 0.4, 0.2, 0.4, 0.6, 0.8, or its reverse. In an <italic>easy</italic> run, this probability of getting high reward by selecting the red square remained fixed at 0.2 or 0.8 (two runs for each probability) throughout the run. This design ensured that the mean probability of getting high reward by selecting either color constantly was 0.5 (because the two colors were equally likely to be the more rewarding one) across this task.</p>
</sec>
<sec><title>Procedure</title>
<p>All participants gave written informed consent before participating in the study. They then read the instructions for the task, performed a practice run to ensure that they understood the task, and underwent the scanning session. The scanning session consisted of one anatomical scan, eight runs of functional scans while the participants performed the task, one resting-state functional scan, and one diffuse tensor imaging (DTI) scan. The resting-state and DTI scans were not used in this study. After the scanning session, the participants received monetary compensation (a fixed amount for participation and a variable amount based on the score of a randomly selected run).</p>
</sec>
<sec><title>Dynamic Analysis</title>
<p>In order to assess how trial history modulated future choices of color, and whether/how this modulation changed as a function of the reward received at the most recent trial, we conducted a response dynamic analysis (<xref ref-type="bibr" rid="B27">Lau and Glimcher, 2005</xref>). We started by dividing all trials into two sets, depending on whether the reward received at the previous trial was high or low. Subsequently, for each set, we constructed a linear model in the following form:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mo>s</mml:mo><mml:mo>n</mml:mo></mml:msub><mml:msub><mml:mrow><mml:mo>&#x00A0;=&#x00A0;c</mml:mo></mml:mrow><mml:mrow><mml:mo>n-1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x00A0;s</mml:mo></mml:mrow><mml:mrow><mml:mo>n-1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x00A0;+&#x00A0;c</mml:mo></mml:mrow><mml:mrow><mml:mo>n-2</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mo>s</mml:mo><mml:mrow><mml:mo>n-2</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where s<sub>n</sub>, s<sub>n-1</sub>, and s<sub>n-2</sub> represented the color choice (red = 1, green = -1) at trial <italic>n, n</italic>-1, and <italic>n</italic>-2, respectively. This model only considered the two most recent trials, because of the fact that, in hard runs, the reward probability changed every four trials. For each set, we used this model and behavioral data to estimate c<sub>n-1</sub>, and c<sub>n-2</sub>, which were the dependence of the choice at trial <italic>n</italic> on trial <italic>n</italic>-1 and <italic>n</italic>-2, respectively. To estimate c<sub>n-1</sub>, and c<sub>n-2</sub>, a design matrix with two regressors was constructed. The two regressors represented normalized trial-wise color choice in the past trial and two trials ago. This design matrix was then regressed against the vector encoding normalized trial-wise color choice to obtain estimates of c<sub>n-1</sub>, and c<sub>n-2</sub>. In the end, we compared the c<sub>n-1</sub>, and c<sub>n-2</sub> estimates between the two sets (i.e., whether reward at trial <italic>n</italic>-1 was high or low).</p>
</sec>
<sec><title>The Flexible Learning Model</title>
<p>To simulate how individuals learn the color that leads to better a chance of receiving high reward, we adapted the flexible control model that <xref ref-type="bibr" rid="B24">Jiang et al. (2014</xref>, <xref ref-type="bibr" rid="B23">2015</xref>) have shown captures the flexible learning of control demand in a changing environment. This flexible learning model is also structurally similar to the model used by <xref ref-type="bibr" rid="B4">Behrens et al. (2007)</xref>. In that study, subjects repeatedly bet on one of two options, only one of which leads to reward at each trial. The probability of reward may stay constant (stable condition) or flip (volatile condition). To simulate the task, the <xref ref-type="bibr" rid="B4">Behrens et al. (2007)</xref> model tracks the belief of the volatility (rate of change in reward probability) and the belief of reward probability. After each trial, the winning option is revealed to the model. Using this information the model updates the beliefs based on Bayesian inference. Crucially, the participants&#x2019; choices are not used in the model. In practice, even given the same trial sequence, different participants are likely to produce different patterns of behavior due to their different mental states (e.g., the belief of volatility and which option is better). However, in this case, due to the fact that the models in <xref ref-type="bibr" rid="B4">Behrens et al. (2007)</xref> and <xref ref-type="bibr" rid="B24">Jiang et al. (2014)</xref> do not use subjects&#x2019; behavior to infer the mental states, these models are unable to account for individual differences in behavior and will yield identical model estimates of mental states for all participants. To account for individual differences, the present model includes the participants&#x2019; choices of colors, which reflect the participants&#x2019; specific beliefs about the states of the task. Therefore, when two participants underwent the same trial sequence but produced different choices, the present model would consider the differences in choices and produce different model estimates.</p>
<p>The flexible learning model represented in <bold>Figure <xref ref-type="fig" rid="F1">1C</xref></bold> has five variables, namely (1) the flexible learning rate, <italic>&#x03B1;</italic>, that quantifies the model&#x2019;s (or a participant&#x2019;s) belief concerning the relative weight of the most recent information (i.e., reward and choice of color observed) in learning; (2) the probability, <italic>p</italic>, of the red square leading to high reward; (3) observed selection of color, <italic>s</italic> (either 0 or 1, coded to correspond to green or red, respectively); (4) observed reward, <italic>r</italic> (either 0 or 1, coded to correspond to low or high reward, respectively), and (5) observed outcome, o, which is determined by <italic>s</italic> and <italic>r</italic> and encodes whether the selection resulted in the expected outcome (see below). The terms <italic>s</italic> and o are included to infer the hidden model belief states of <italic>&#x03B1;</italic> and <italic>p</italic>. The subscript <italic>i</italic> denotes the state of a variable at trial <italic>i</italic>. The dynamics of the distribution of learning rate across trials is defined such that the transition of the flexible learning rate is most likely to remain in its previous state; if it changes state, however, it is equally likely to jump to any other value, following a uniform probability distribution</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>k</mml:mi><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>k</italic> is the probability of the learning rate remaining the same as on the previous trial and 0 &#x003C; <italic>&#x03B1;<sub>i</sub>, k</italic> &#x003C; 1. &#x03B4;(&#x03B1;<sub>i+1</sub> -&#x03B1;<sub>i</sub>) equals to 1 if &#x03B1;<sub>i+1</sub> is the same as &#x03B1;<sub>i</sub> and equals to 0 otherwise.</p>
<p>Given the random sequencing of the task, it is not possible to make a precise prediction of the more rewarding color (e.g., predicting that the probability of the more rewarding color being red is 0.8 with 100% certainty, whereas this probability has 0 chance to be 0.799 or 0.801). Hence the prediction should be approximate, leading to a smooth distribution of <italic>p<sub>i</sub></italic>. For example, a high likelihood of <italic>p<sub>i</sub></italic> being 0.8 should also imply a high likelihood of <italic>p<sub>i</sub></italic> at values close to 0.8. Hence, we used a propagation process to smooth the distribution of <italic>p<sub>i</sub></italic> in the following manner:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>0.5</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>~</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where p<sub>i+0.5</sub> denotes the belief of the red color being the more rewarding color after the propagation process.</p>
<p>Up to this point, this model is identical to the flexible control model. The choices of equations and processing steps have been validated in <xref ref-type="bibr" rid="B23">Jiang et al. (2015)</xref>. The following steps were conceptually similar to the flexible control model and were tailored to suit the current task. Specifically, after propagation, p<sub>i+0.5</sub> is updated using a standard reinforcement learning rule with &#x03B1;<sub>i+1</sub> playing the role of learning rate:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>~</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>0.5</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>0.5</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where o<sub>i</sub> denotes the outcome at trial <italic>i</italic>. Thus, o<sub>i</sub> is 1 (i.e., supporting that the red color is associated with a better chance of obtaining high reward) when the chosen color is red and the reward is high, or when the chosen color is green and the reward is low. Otherwise, o<sub>i</sub> is 0, indicating that the outcome does not support the red color being the more rewarding color. At this point, the expectation of distributions p(p<sub>i+1</sub>) and p(&#x03B1;<sub>i+1</sub>) are used, respectively, as estimates of the probability of the red square being the more rewarding one and the learning rate for behavioral and fMRI analyses.</p>
<p>When the actual selection and outcome were observed, the model beliefs were updated in the following manner:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x200B;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>&#x221D;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x200B;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mo>p(si+1,&#x00A0;o</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x007C;p</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x200B;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x200B;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mo>=(1-&#x007C;o</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x00A0;-&#x00A0;p</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x007C;)(1&#x00A0;-&#x007C;s</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo>&#x00A0;-&#x00A0;sp</mml:mo></mml:mrow><mml:mrow><mml:mo>i+1</mml:mo></mml:mrow></mml:msub><mml:mo>&#x007C;)&#x00A0;</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>In Eq. (7), &#x007C;o<sub>i+1</sub> -p<sub>i+1</sub>&#x007C; quantified the discrepancy between estimated and actual outcomes; and &#x007C;s<sub>i+1</sub> - sp<sub>i+1</sub>&#x007C; quantified the discrepancy between estimated and actual human behavior (i.e., actual selection of color, please see Eq. (8) for the definition of <italic>sp</italic>), which was then used to fit the model to human behavior in order to better infer the participant&#x2019;s mental states and accounts for individual difference in reward learning. Therefore, p(S<sub>i+1</sub>, O<sub>i+1</sub>&#x007C;P<sub>i+1</sub>) integrated the prediction error of the outcome and the prediction error of behavior (i.e., which color was selected), in order to infer the participants&#x2019; internal states.</p>
<p>Using Eqs. (6 and 7), we updated the joint distribution of k, &#x03B1;<sub>i+1</sub>, p<sub>i+1</sub> with new observations s<sub>i+1</sub>, r<sub>i+1</sub> and o<sub>i+1</sub>. This updated joint distribution was then fed into the simulation of the next trial (i.e., trial <italic>i</italic>+2).</p>
<p>To apply this model to simulate a participant&#x2019;s behavior, the participant&#x2019;s trial-by-trial color selections and the received reward were submitted to this model to generate trial-by-trial estimates of <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic>. Specifically, the model maintained a joint distribution of p(k, &#x03B1;<sub>i</sub>, p<sub>i</sub>), which initialized as a uniform distribution at the beginning of each run. The distribution of a single variable, such as p(&#x03B1;<sub>i</sub>), could be calculated by marginalizing the joint distribution. At the beginning of trial <italic>i</italic>, we used Eqs. (2&#x2013;5) to update p(k, &#x03B1;<sub>i</sub>, p<sub>i</sub>) and produce estimates of <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic>. After the reward and the participant&#x2019;s choice of color were observed, Eqs. (6 and 7) were applied to update the joint distribution of p(k, &#x03B1;<sub>i</sub>, p<sub>i</sub>). The simulation then entered the next trial. Note that in the flexible learning model, <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic> represent random variables (i.e., probabilistic distributions). As mentioned above, in the following analyses, we only used the mean values of <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic>. So, from this point on, we refer <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic> to the mean of their corresponding random variables in order to keep the description of the methods and results concise.</p>
<p>After the trial-wise estimates of <italic>&#x03B1;<sub>i</sub></italic> and <italic>p<sub>i</sub></italic> were generated, the model&#x2019;s prediction of the probability of selecting the red color at trial <italic>i</italic>, or ps<sub>i</sub> was determined using the following softmax function:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M8"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x00A0;</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>Where &#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub> were hyper parameters that were estimated by fitting ps<sub>i</sub> to the actual choices of color made by each participant.</p>
<p>Similar to <xref ref-type="bibr" rid="B23">Jiang et al. (2015)</xref>, the flexible learning model and the softmax function were estimated iteratively using an EM algorithm. This algorithm started with p = ps and estimated trial-wise &#x03B1;<sub>i</sub> and <italic>p<sub>i</sub></italic> based on Eq. (2&#x2013;7) (E step). Then <italic>s<sub>i</sub></italic> and the estimates of &#x03B1;<sub>i</sub> and p<sub>i</sub> were used to estimate ps<sub>i</sub>, &#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub> (M step), which were used in the E step in the next iteration. These two steps continued until the estimates of the hyper parameters converged. In this study, we implemented these procedures using Matlab. The scripts are available on request.</p>
</sec>
<sec><title>Model Validation</title>
<p>To probe whether the flexible learning model accounted for behavioral data (i.e., actual color choices) better than typical, non-flexible reinforcement learning models, we performed model comparisons among four models: (a) the flexible learning model, (b) a reinforcement learning model with one fixed learning rate (RL_1 model), (c) a reinforcement learning model with one fixed learning rate for easy conditions and one for hard conditions (RL_2 model), and (d) a reinforcement learning model (PE-M) whose learning rate scales with the magnitude of prediction error (i.e., the learning rate is &#x03B1;&#x007C;r - p&#x007C;, where &#x03B1; is a base learning rate and &#x007C;r - p&#x007C; is the prediction error magnitude based modulation on &#x03B1;), which is similar to <xref ref-type="bibr" rid="B37">Pearce and Hall (1980)</xref>. For models (b&#x2013;d), the optimal learning rate(s) were obtained by an exhaustive search in the range of 0.01&#x2013;0.99 (step size = 0.01) for each participant. Each reinforcement learning model was also connected to a softmax function to predict behavior. The free parameters in these softmax functions were estimated similarly to the softmax function in the flexible learning model. Thus, the flexible learning model had two free parameters (&#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub>) for each participant; the RL_1, RL_2, and PE-M models had three (one learning rate plus &#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub>), four (two learning rates and &#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub>), and three (one base learning rate plus &#x03B2;<sub>1</sub> and &#x03B2;<sub>2</sub>) free parameters, respectively. Thus, compared to the flexible learning model, the RL_1 and RL_2 models had one and two more free parameters (i.e., the learning rates), respectively.</p>
<p>In order to control for over-fitting and to provide generalizable results, we conducted cross-validation, which is a common practice in assessing learner performance in machine learning and multi-voxel pattern recognition. Specifically, we divided the runs into two-fold. For the flexible learning model and RL_1 model, each fold consisted of data from two easy runs and two hard runs; for RL_2 model, easy and hard runs were processed separately, so each fold had two runs from the same difficulty condition. To assess model performance, one-fold served as the training set to estimate the free parameters that best accounted for the training set. These estimated free parameters were then applied to the other fold (test set) to produce a trial-by-trial sequence of simulated probability of choosing the red color. Because the training and test sets were independent, this cross validation effectively reduced over-fitting. We repeated this procedure after exchanging the training and test sets. In the end, each model had a simulated trial-by-trial probability of choosing the red color. The performances of these simulations were then compared across models: For each model and each participant, we calculated the Bayesian information criterion (BIC) in the following manner:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M9"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mo>BIC</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>nln</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x00A0;</mml:mo><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mo>e</mml:mo><mml:mo>2</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>n</italic> is the number of trials and <inline-formula><mml:math id="M10"><mml:msubsup><mml:mi mathvariant='normal' mathcolor='black'>&#x03C3;</mml:mi><mml:mi mathvariant='normal' mathcolor='black'>e</mml:mi><mml:mn mathvariant='normal' mathcolor='black'>2</mml:mn></mml:msubsup></mml:math></inline-formula> is the error variance (<xref ref-type="bibr" rid="B39">Priestley, 1981</xref>) between model simulations and human behavior. Importantly, in calculating the BIC, we omitted the penalty for having additional free parameters, so that we only compared how well each model accounts for the behavioral data. It should be noted that this omission did not give the flexible learning rate any advantage in the model comparison, because it indeed had fewer free parameters than the other models. According to the definition of BIC, smaller prediction errors translated into lower BIC values, so the model with the lowest BIC had best performance.</p>
</sec>
<sec><title>Statistical Analyses on Behavioral and Model Data</title>
<p>Repeated measure ANOVAs, two-tailed <italic>t</italic>-tests, and linear correlation analysis were conducted using SPSS or Matlab to analyze the behavioral and model data (see below for details).</p>
</sec>
<sec><title>Image Acquisition</title>
<p>Images were acquired on a GE MR750 3.0T scanner. The anatomical images were scanned using a T1-weighted axial sequence parallel to the anterior-commissure-posterior commissure line. Each anatomical scan had 156 axial slices (spatial resolution = 1 mm &#x00D7; 1 mm &#x00D7; 1 mm, field of view = 256 mm &#x00D7; 256 mm, time repetition [TR] = 8.124 ms). The functional images were scanned using a T2<sup>&#x2217;</sup>-weighted single-shot gradient EPI sequence with a TR of 2 s. Each functional volume contained 43 axial slices (spatial resolution = 3.75 mm &#x00D7; 3.75 mm &#x00D7; 3.3 mm, field of view = 240 mm &#x00D7; 240 mm, TE = 28 ms, flip angle = 90&#x00B0;). Each fMRI run lasted for 416 s (208 TRs). During image acquisition, software monitored head movement in real-time. When the head movement exceeded 3 mm or 3&#x00B0; within a run, the scanning for that run was re-started using a new trial sequence.</p>
</sec>
<sec><title>Image Analysis</title>
<p>The images were preprocessed using SPM12<sup><xref ref-type="fn" rid="fn02">2</xref></sup>. The first five volumes of each run were discarded before preprocessing. The remaining volumes were first realigned to the mean volume of the run, and went through slice-timing correction. The anatomical scan was co-registered to the mean volume, and then normalized to the Montreal Neurological Institute (MNI) template. The normalization parameters were applied to the slice-time corrected functional volumes, which were resampled to the spatial resolution of 3 mm &#x00D7; 3 mm &#x00D7; 3 mm). Finally, the resampled functional volumes were smoothed using a Gaussian kernel (FWHM = 8 mm).</p>
<p>We carried out a general linear model (GLM)-based analysis on the preprocessed fMRI data at each voxel for each individual (i.e., first-level analysis in SPM). This GLM consisted of up to 10 regressors, divided into three groups. The first group consisted of three regressors time-locked to the onset of the color squares at each trial: the stick function of each trial; <italic>&#x03B1;</italic> (the lack of a subscript indicates that this variable refers to the trial-by-trial time course of this variable); and the predicted reward probability of the chosen color (i.e., <italic>p</italic> or 1 -<italic>p</italic>, if the participant later chose red or green, respectively). We chose to time-lock <italic>&#x03B1;</italic> to the onset of the color squares to ensure that by that time the learning rate had been updated. The second group consisted of five regressors time-locked to the onset of the feedback at each trial: the stick function; the reward feedback, <italic>r</italic>; the signed reward prediction error, <italic>pe</italic> (<italic>r - p</italic> if red square was chosen, <italic>r -</italic> 1 + <italic>p</italic> if green square was chosen); the updating in learning, &#x03B1; &#x00D7; pe; and the &#x03B1; &#x00D7; r interaction that accounted for the behavioral pattern of post-high reward increase of learning rate (see below). The last group consisted of two regressors time-locked to the response at each trial: the stick function and the response (left or right). All regressors were concatenated across the eight runs of this experiment. The regressors were also normalized to remove the confounds from mean and magnitude. The time courses of model estimates (e.g., &#x03B1;, p) were obtained using the model parameters fit at the individual level. Therefore, the fact that the imaging analyses included fewer participants than the behavioral analyses did not change the model estimates for each participant.</p>
<p>To gauge the encoding strength of a variable (represented as a regressor), its regressor was first regressed against other regressors in the same group (i.e., sharing the same onset time) to remove shared variance so that the results could be uniquely attributed to the variable of interest. For example, to compute the encoding strength of <italic>&#x03B1;</italic>, the trial-by-trial time course of <italic>&#x03B1;</italic> was regressed against the other regressors in the same group. Then the post-regression <italic>&#x03B1;</italic> replaced the original <italic>&#x03B1;</italic> in the GLM.</p>
<p>This GLM was then convolved with SPM&#x2019;s hemodynamic function and appended with nuisance parameters, including six head movement parameters (translations and rotations relative to <italic>x, y</italic>, and <italic>z</italic> axes) and grand mean vectors for each run to remove run-specific baseline fMRI signal. The resulting GLM was subsequently fit to the preprocessed fMRI data to estimate the coefficient for the variable of interest&#x2019;s parametric modulator, which reflected the variable&#x2019;s encoding strength (one estimate at each voxel of each individual&#x2019;s data). To remove nuisance results at white matter and cerebrospinal fluid voxels, the statistical results were filtered using a gray matter mask obtained by segmenting the template and only keeping voxels with gray matter concentrations greater than 0.01. Finally, we conducted group-level <italic>t</italic>-tests on the estimates of encoding strength across participants to assess the group-level encoding strength (i.e., second-level analysis in SPM).</p>
</sec>
<sec><title>Control for Multiple Comparisons</title>
<p>Statistical results were corrected for multiple comparisons at <italic>P</italic> &#x003C; 0.05 for combined searchlight classification accuracy and cluster extent thresholds, using the AFNI ClustSim algorithm<sup><xref ref-type="fn" rid="fn03">3</xref></sup>. Specifically, 10,000 Monte Carlo simulations were conducted, each generating a random statistical map based on the smoothness of the map resulting from the group-level <italic>t</italic>-tests. For each randomly generated map, the algorithm searched for clusters using a voxel-wise <italic>P</italic>-value threshold of &#x003C;0.001. The identified clusters were then grouped to produce a null distribution of cluster size. As a result, the ClustSim algorithm determined that an uncorrected voxel-wise <italic>P</italic>-value threshold of &#x003C;0.001 in combination with a searchlight cluster size of 78&#x2013;84 voxels (depending on the specific contrast) ensured a false discovery rate of &#x003C;0.05.</p>
</sec>
</sec>
<sec><title>Results</title>
<sec><title>Behavioral Results</title>
<p>Twenty-seven participants performed the reward learning task (see Materials and Methods). At each trial, participants chose between a red square and a green square, and received either a high or a low reward based on the probability of high reward associated with the chosen color. Importantly, at each moment, one color had a better chance leading to a high reward than the other color (the sum of the two probabilities was always 1). In order to maximize reward, the participants must learn which color was more rewarding. To create a wide range of belief regarding the more rewarding color, the task contained two conditions. In the &#x201C;easy&#x201D; condition, the more rewarding color and its probability of yielding high reward remained constant at 80% throughout a run. Conversely, in the &#x201C;hard&#x201D; condition, the more rewarding color and its probability of obtaining high reward varied across time (from 20 to 80%, step size = 20%). Across the whole task, each color had a 50% chance of being the more rewarding color.</p>
<p>Participants chose the more rewarding color more frequently than chance-level (i.e., 50%) for both conditions: easy condition: 84.0 &#x00B1; 1.5%, <italic>t</italic>(26) = 23.05, <italic>P</italic> &#x003C; 0.001, one-sampled <italic>t</italic>-test; hard condition: 60.5 &#x00B1; 1.2%, <italic>t</italic>(26) = 9.12, <italic>P</italic> &#x003C; 0.001, one-sampled <italic>t</italic>-tests (<bold>Figure <xref ref-type="fig" rid="F2">2A</xref></bold>). Those outcomes indicate that participants followed the task instructions to learn the more rewarding color. Additionally, participants chose the more rewarding color more frequently in the easy condition than the hard condition, <italic>t</italic>(26) = 13.80, <italic>P</italic> &#x003C; 0.001, paired <italic>t</italic>-test (<bold>Figure <xref ref-type="fig" rid="F2">2A</xref></bold>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Behavioral results. <bold>(A)</bold> Group mean &#x00B1; standard error of the percentage of choosing the more rewarding color, plotted as a function of run difficulty. <bold>(B)</bold> Group mean &#x00B1; standard error of the percentage of repeating the previous trial&#x2019;s choice, plotted as a function of reward magnitude and run difficulty. <bold>(C)</bold> Group mean and standard error of the dependence of the choice at trial <italic>n</italic> on choice history, plotted as a function of whether a high reward was received at trial <italic>n</italic>-1 and trial <italic>n</italic>-2. <bold>(D)</bold> Group mean and standard error of the percentage of choosing the more rewarding color, plotted as a function of the temporal order of hard runs. <bold>(E)</bold> Group mean and standard error of the percentage of choosing the more rewarding color, plotted as a function of the three trials (No. 9, 21, and 33) that immediately follow the change of the more rewarding color in hard runs. <bold>(F)</bold> Group mean and standard error of the percentage of choosing the more rewarding color, plotted as a function of trials in easy and hard runs separately.</p></caption>
<graphic xlink:href="fnhum-11-00592-g002.tif"/>
</fig>
<p>To further probe how the participants adjusted their choice of color based on the reward received, we tested the frequency of the participants repeating their previous color choice. Because the more rewarding color was more likely to remain the same as on the previous trial than to switch to the other color, the participants should repeat their choices more frequently than chance level (50%). As expected, the overall frequency of choice repetition was significantly higher than chance [68.4 &#x00B1; 1.8%, <italic>t</italic>(26) = 38.10, <italic>P</italic> &#x003C; 0.001, one-sampled <italic>t</italic>-test]. To further test the differences in the frequency of choice repetition among experimental conditions, we conducted a repeated measure 2 (received reward: high, low) &#x00D7; 2 (difficulty: easy, hard) ANOVA (<bold>Figure <xref ref-type="fig" rid="F2">2B</xref></bold>). This ANOVA revealed a significant main effect of received reward, <italic>F</italic>(1,26) = 291.74, <italic>P</italic> &#x003C; 0.001, driven by a higher frequency of repeating color choice after receiving a high reward (89.3 &#x00B1; 1.8%) than after receiving a low reward (47.6 &#x00B1; 2.6%), suggesting that high reward enhanced the participants&#x2019; belief that the selected color was the more rewarding one. The main effect of difficulty was also significant, <italic>F</italic>(1,26) = 69.37, <italic>P</italic> &#x003C; 0.001, driven by higher likelihood of repeating color choice in easy condition (72.2 &#x00B1; 1.8%) than hard condition (64.7 &#x00B1; 2.0%). This difference possibly reflected the fact that the more rewarding color changed more frequently in hard than easy condition. The reward type &#x00D7; difficulty interaction was not significant [<italic>F</italic>(1,26) = 0.66].</p>
<p>We also conducted an additional dynamic analysis (<xref ref-type="bibr" rid="B27">Lau and Glimcher, 2005</xref>) that compared how the choice of color at trial <italic>n</italic> relied on the interaction between trial history and the reward received at trial <italic>n-</italic>1 (<xref ref-type="bibr" rid="B8">Browning et al., 2015</xref>). The results are shown in <bold>Figure <xref ref-type="fig" rid="F2">2C</xref></bold>: although the contribution of color choice at trial <italic>n-</italic>2 to color choice at trial <italic>n</italic> did not vary as a function of reward feedback at trial <italic>n</italic>-1, <italic>t</italic>(26) = 0.31, paired <italic>t</italic>-test, the contribution of color choice at trial <italic>n-</italic>1 differed significantly between reward feedback levels, <italic>t</italic>(26) = 17.16, <italic>P</italic> &#x003C; 0.001, paired <italic>t</italic>-test. Specifically, when receiving a high reward at trial <italic>n-</italic>1, color choice at trial <italic>n-</italic>1 had a positive influence (i.e., promoting choice repetition) on the choice at trial <italic>n</italic>, reliance: 0.72 &#x00B1; 0.04, <italic>t</italic>(26) = 20.30, <italic>P</italic> &#x003C; 0.001, one-sample <italic>t</italic>-test, whereas this influence became negative if the reward was negative (i.e., promoting choice change), reliance: -0.12 &#x00B1; 0.04, <italic>t</italic>(26) = -3.00, <italic>P</italic> = 0.006, one-sample <italic>t</italic>-test. This analysis again confirmed that the reward received at the current trial modulated how likely the choice would be repeated at the forthcoming trial.</p>
<p>In hard runs, it may be possible that the participants became aware of the change patterns of the rewarding probability and proactively altered their learning rate to adapt to the change. If this were true, task performance in hard runs should increase with time on task. We conducted a repeated-measures one-way ANOVA on the four hard runs&#x2019; probability of choosing the more rewarding color, and did not find a significant change of performance across runs, <italic>F</italic>(3,24) = 0.08 (<bold>Figure <xref ref-type="fig" rid="F2">2D</xref></bold>). Furthermore, we tested the effect of within-run learning of change patterns. To this end, we compared the probability of choosing the more rewarding colors among the three trials (No. 9, 21, and 33; <bold>Figure <xref ref-type="fig" rid="F1">1B</xref></bold>) that are immediately after the change of the more rewarding color (i.e., <italic>p</italic> changed from 0.4 to 0.6 or vice versa). If the participants learned the change patterns and proactively adjusted to the change, we expected an increase in the probability of choosing the more rewarded color across the three trials. However, a repeated-measures one way ANOVA did not find any effect, <italic>F</italic>(2,25) = 0.08 (<bold>Figure <xref ref-type="fig" rid="F2">2E</xref></bold>). Collectively, the participants appeared not to be able to apply the change patterns to boost their performance in this task.</p>
<p>A closer look at the time course of the participants&#x2019; choices (<bold>Figure <xref ref-type="fig" rid="F2">2F</xref></bold>) showed that, in easy runs, the probability of choosing the more rewarded color kept increasing, suggesting the participants readily learned the underlying reward probability. However, the probability of choosing the more rewarded color fluctuated in hard runs, due to the change of the underlying reward probability. A general trend in the hard condition was a sudden drop in the ability to identify the more rewarded color following the change of the more rewarded color (i.e., after trials #9, 21, and 33, similar to <bold>Figure <xref ref-type="fig" rid="F2">2E</xref></bold>), and a gradual recovery that suggests that the participants continued to learn the current more rewarded color.</p>
</sec>
<sec><title>Model Comparison</title>
<p>In order to model human behavior in this task and to infer latent learning related states, we employed a flexible learning model (see Materials and Methods) that self-adjusts the learning rate, &#x03B1;, and the belief that the red color was the more rewarded color, <italic>p</italic>, based on observed color choice and received reward. To ascertain that this model accounted for the behavior better than conventional reinforcement learning models, we conducted a model comparison analysis among the flexible learning model and reinforcement learners with one fixed learning rate (RL_1), reinforcement learners with two fixed learning rates (RL_2, one learning rate for each difficulty condition), and reinforcement learners whose learning rate scales with the magnitude of prediction error (PE-M). The flexible learning model accounted for trial-by-trial color choice best in the four models (i.e., the flexible learning model had the lowest BIC) for all participants (<bold>Figure <xref ref-type="fig" rid="F3">3A</xref></bold>). Therefore, we concluded that the flexible learning model explained variance in human behavior better than conventional reinforcement learners in this task. Consequently, we focused on the flexible learning model in the following analyses.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Model-based behavioral results. <bold>(A)</bold> Model comparison results. From left to right, each bar represents group mean Bayesian information criterion (BIC) and mean standard error of the flexible learning model (FLM), reinforcement learning model with one (RL_1) learning rate and two (RL_2) learning rates (one for each level of difficulty), and a reinforcement learning model whose learning rate scales with the magnitude of prediction error (PE-M). <bold>(B)</bold> Sample time courses of the learning rate and the expected probability that high reward for the red square. The locations of asterisks indicate high or low reward received at each trial. <bold>(C)</bold> Model belief that red was the more rewarding color, plotted as a function of difficulty and the chosen color. <bold>(D)</bold> Model simulation of the percentage of choice repetition, plotted as a function of reward magnitude and difficulty. <bold>(E)</bold> Model belief of the learning rate, plotted as a function of reward magnitude and difficulty. <bold>(F)</bold> Individual increments of the probability of choice repetition (following high reward minus following low reward), plotted against the individual increments (following high reward minus following low reward) of the model estimates of the learning rate. The dotted line depicts the trend line.</p></caption>
<graphic xlink:href="fnhum-11-00592-g003.tif"/>
</fig>
</sec>
<sec><title>Model-Based Behavioral Analysis</title>
<p>The flexible learning model generated trial-by-trial estimates of &#x03B1; and <italic>p</italic>, which drove the learning of the more rewarding color (<bold>Figure <xref ref-type="fig" rid="F3">3B</xref></bold>). To test whether these model estimates reflected behavioral patterns, we first conducted a repeated measures 2 (chosen color: red, green) &#x00D7; 2 (difficulty: easy, hard) ANOVA on the <italic>p</italic> estimates (<bold>Figure <xref ref-type="fig" rid="F3">3C</xref></bold>). That analysis revealed a significant main effect of color choice, <italic>F</italic>(1,26) = 507.48, <italic>P</italic> &#x003C; 0.001, whereby red square choices were associated with higher <italic>p</italic> estimates (i.e., stronger belief that red was the more rewarding color) than green square choices, with <italic>p</italic> estimates in red square choices: 0.64 &#x00B1; 0.01 and estimates in green square choices: 0.36 &#x00B1; 0.01. This result matches the model&#x2019;s specification that <italic>p</italic> was associated with the red square. In other words, this result showed that the participants tended to choose the more rewarded color predicted by the model. Additionally, we discovered a significant interaction between chosen color and difficulty, <italic>F</italic>(1,26) = 177.49, <italic>P</italic> &#x003C; 0.001, driven by a smaller effect of chosen color on <italic>p</italic> estimates in hard (0.35 &#x00B1; 0.01) than easy (0.21 &#x00B1; 0.01) conditions. This reduced effect may reflect the fact that the more rewarded color was more difficult to learn in the hard condition. The main effect of difficulty was not significant. Next, to confirm that <italic>p</italic> guided the prediction of chosen color in the right direction, we found that the responsible parameter &#x03B2;<sub>2</sub> (see Materials and Methods) was greater than 0 for all participants (i.e., higher <italic>p</italic> leads to higher probability of choosing the red square; range: 4.5&#x2013;25.0).</p>
<p>Recall that participants repeated the choice of color more frequently following a high reward than a low reward. The flexible learning model accounted for this result in two ways: First, the pattern of simulated probability of repeating the previous choice, calculated by applying trial-based <italic>p</italic> estimate to Eq. (8), then comparing the result to the previous choice, should resemble <bold>Figure <xref ref-type="fig" rid="F2">2B</xref></bold>. Second, the model accounts for this observation by increasing <italic>&#x03B1;</italic> after high reward to augment the current selection&#x2019;s influence on selecting (the same) color in the next trial. To test these model predictions, we conducted separate 2 (received reward: high, low) &#x00D7; 2 (difficulty: easy, hard) ANOVAs on the probability of repeating the previous choice (<bold>Figure <xref ref-type="fig" rid="F3">3D</xref></bold>) and <italic>&#x03B1;</italic> estimates (<bold>Figure <xref ref-type="fig" rid="F3">3E</xref></bold>). Similar to the behavioral results, the first ANOVA yielded a significant main effect of both reward type, <italic>F</italic>(1,26) = 384.22, <italic>P</italic> &#x003C; 0.001, driven by a higher likelihood of repeating a color choice after high reward (90.0 &#x00B1; 1.7%) than low reward (47.2 &#x00B1; 1.2%); and a main effect of difficulty, <italic>F</italic>(1,26) = 20.11, <italic>P</italic> &#x003C; 0.001, driven by a higher likelihood of repeating color choice after an easy condition (70.2 &#x00B1; 1.4%) than hard condition (67.1 &#x00B1; 1.5%). The interaction was not significant [<italic>F</italic>(1,26) = 0.19].</p>
<p>Also consistent with the model prediction, the second ANOVA revealed a main effect of reward type, <italic>F</italic>(1,26) = 265.37, <italic>P</italic> &#x003C; 0.001; learning rate estimates were higher after high reward (0.167 &#x00B1; 0.002) than low reward (0.150 &#x00B1; 0.001). No other effects were significant, <italic>F</italic>(1,26) &#x003C; 2.76, <italic>P</italic> = 0.11. As a control analysis, we also tested whether the prediction error of reward is a better predictor than the magnitude of reward for selecting the same response in the next trial, given their high correlation. To this end, we used a binary vector to represent whether the color selection was repeated in the next trial for each participant; and compared how much variance in this vector could be explained by trial-wise reward magnitude vs. trial-wise prediction error of reward. For all participants, reward magnitude was a better predictor than prediction error (additional variance explained by reward magnitude ranged from 2.5 to 13.5%). The results supported the notion that reward magnitude is a more likely driving factor for repeating selection than reward prediction error.</p>
<p>Finally, to test how <italic>&#x03B1;</italic> accounted for individual differences in repetition of color choice, we conducted the linear correlation analysis between the increase of <italic>&#x03B1;</italic> (high reward &#x2013; low reward) and the increase of the frequency of choice repetition (high reward &#x2013; low reward) across participants, and found a strong positive linear correlation (<italic>r</italic> = 0.86, <italic>P</italic> &#x003C; 0.001; <bold>Figure <xref ref-type="fig" rid="F3">3F</xref></bold>). Note that in Eq. (2), we did not constrain which way the learning rate should go conditioned on the type of reward (high or low) received (i.e., the model is not pre-defined to show the increase of learning rate following high reward). Furthermore, the same model was fit to individual behavior data. Therefore, the individual difference in the amount of learning rate increase following high reward is solely attributable to the participants&#x2019; behavior. In other words, the results in <bold>Figure <xref ref-type="fig" rid="F3">3F</xref></bold> indeed indicate that the change in the likelihood of choice repetition was captured by the change in learning rate estimates. Given that the flexible learning model takes the participants&#x2019; choices of colors as input in order to account for individual differences, this model revealed individual differences in raising the learning rate following a high reward in relation to a low reward. According to the reinforcement learning algorithm, the learning rate is also the weight of the current choice on determining the next choice. Therefore, participants&#x2019; increasing the learning rate more after a high reward than a low reward will let a choice that led to high reward have more weight on determining the next choice (i.e., more likely to repeat the choice), as compared to a choice that led to low reward, thus producing the correlation in <bold>Figure <xref ref-type="fig" rid="F3">3F</xref></bold>.</p>
<p>Taken together, these results suggested that the flexible learning model estimates of <italic>&#x03B1;</italic> and <italic>p</italic> captured the behavioral patterns. Therefore, the model provided meaningful learning-related information for the following imaging analyses.</p>
</sec>
<sec><title>Imaging Results</title>
<p>Using trial-by-trial model estimates derived from the flexible learning model, we examined the encoding of model variables and their interactions in the brain, based on data obtained from fMRI scans while the participants performed the task (<bold>Table <xref ref-type="table" rid="T1">1</xref></bold>). We found that the learning rate reliably co-varied with fMRI signals in the left IFG and adjacent anterior insula (<italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: -51, 17, 11, <bold>Figure <xref ref-type="fig" rid="F4">4A</xref></bold>). Although not statistically significant after controlling for multiple comparisons, three scattered ACC clusters (size: 7&#x2013;13 voxels) showed encoding of the learning rate at the <italic>P</italic> &#x003C; 0.001 (uncorrected) level (<bold>Figure <xref ref-type="fig" rid="F4">4A</xref></bold>). Given that the learning rate in the flexible learning model reflected the volatility in the task, these ACC clusters were consistent with earlier findings from <xref ref-type="bibr" rid="B4">Behrens et al. (2007)</xref>. No other clusters surpassed significance threshold.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of fMRI results.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Location</th>
<th valign="top" align="center">Peak MNI</th>
<th valign="top" align="center">Peak <italic>t</italic>-value</th>
<th valign="top" align="center">Cluster size (#voxels)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="4"><bold>High reward > low reward</bold></td></tr>
<tr>
<td valign="top" align="left">Middle Cingulate Gyrus, Posterior Cingulate Gyrus, Precuneus</td>
<td valign="top" align="center">(&#x2013;15, &#x2013;49, 11)</td>
<td valign="top" align="center">11.16</td>
<td valign="top" align="center">787</td>
</tr>
<tr>
<td valign="top" align="left">Medial Superior Frontal Gyrus</td>
<td valign="top" align="center">(0, 56, 17)</td>
<td valign="top" align="center">11.18</td>
<td valign="top" align="center">771</td>
</tr>
<tr>
<td valign="top" align="left">R. Precentral Gyrus, R. Rolandic Oper</td>
<td valign="top" align="center">(33, &#x2013;13, 38)</td>
<td valign="top" align="center">9.26</td>
<td valign="top" align="center">604</td>
</tr>
<tr>
<td valign="top" align="left">L. Hippocampus, L. Parahippocampal Gyrus, L. Fusiform Gyrus</td>
<td valign="top" align="center">(&#x2013;33, &#x2013;37, &#x2013;16)</td>
<td valign="top" align="center">9.36</td>
<td valign="top" align="center">555</td>
</tr>
<tr>
<td valign="top" align="left">R. Middle Temporal Gyrus, R. Superior Temporal Gyrus</td>
<td valign="top" align="center">(57, 2, &#x2013;10)</td>
<td valign="top" align="center">9.26</td>
<td valign="top" align="center">548</td>
</tr>
<tr>
<td valign="top" align="left">L. Middle Temporal Gyrus, L. Superior Temporal Gyrus</td>
<td valign="top" align="center">(&#x2013;51, &#x2013;7, &#x2013;7)</td>
<td valign="top" align="center">8.93</td>
<td valign="top" align="center">526</td>
</tr>
<tr>
<td valign="top" align="left">R. Hippocampus, R. Parahippocampal gyrus, R. Fusiform Gyrus</td>
<td valign="top" align="center">(36, &#x2013;22, &#x2013;7)</td>
<td valign="top" align="center">10.28</td>
<td valign="top" align="center">449</td>
</tr>
<tr>
<td valign="top" align="left">R. Precentral Gyrus, R. Postcentral Gyrus, R. Rolandic Oper</td>
<td valign="top" align="center">(&#x2013;45, &#x2013;10, 20)</td>
<td valign="top" align="center">9.15</td>
<td valign="top" align="center">360</td>
</tr>
<tr>
<td valign="top" align="left">L. Superior Occipital Gyrus, L. Middle Occipital Gyrus</td>
<td valign="top" align="center">(&#x2013;45, &#x2013;79, 17)</td>
<td valign="top" align="center">7.77</td>
<td valign="top" align="center">338</td>
</tr>
<tr>
<td valign="top" align="left">R. Superior Occipital Gyrus, R. Middle Occipital Gyrus</td>
<td valign="top" align="center">(33, &#x2013;91, 8)</td>
<td valign="top" align="center">7.50</td>
<td valign="top" align="center">218</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4"><bold>Negative encoding of reward feedback</bold></td></tr>
<tr>
<td valign="top" align="left">L. Inferior Parietal Gyrus</td>
<td valign="top" align="center">(&#x2013;45, &#x2013;49, 44)</td>
<td valign="top" align="center">&#x2013;7.21</td>
<td valign="top" align="center">306</td>
</tr>
<tr>
<td valign="top" align="left">L. Medial Superior Frontal Gyrus</td>
<td valign="top" align="center">(&#x2013;6, 26, 41)</td>
<td valign="top" align="center">&#x2013;8.7</td>
<td valign="top" align="center">291</td>
</tr>
<tr>
<td valign="top" align="left">R. Middle Frontal Gyrus</td>
<td valign="top" align="center">(48, 29, 35)</td>
<td valign="top" align="center">&#x2013;6.38</td>
<td valign="top" align="center">282</td>
</tr>
<tr>
<td valign="top" align="left">R. Supramarginal Gyrus</td>
<td valign="top" align="center">(45, &#x2013;40, 44)</td>
<td valign="top" align="center">&#x2013;6.97</td>
<td valign="top" align="center">247</td>
</tr>
<tr>
<td valign="top" align="left">L. Middle Frontal Gyrus</td>
<td valign="top" align="center">(&#x2013;48, 29, 32)</td>
<td valign="top" align="center">&#x2013;7.77</td>
<td valign="top" align="center">106</td>
</tr>
<tr>
<td valign="top" align="left">L. Insula</td>
<td valign="top" align="center">(&#x2013;30, 26, &#x2013;7)</td>
<td valign="top" align="center">&#x2013;8.43</td>
<td valign="top" align="center">92</td>
</tr>
<tr>
<td valign="top" align="left">R. Insula</td>
<td valign="top" align="center">(33, 23, &#x2013;4)</td>
<td valign="top" align="center">&#x2013;7.44</td>
<td valign="top" align="center">90</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4"><bold>Interaction between feedback and reward prediction error</bold></td></tr>
<tr>
<td valign="top" align="left">L. Orbital medial Frontal Gyrus, L. Medial Superior Frontal Gyrus</td>
<td valign="top" align="center">(&#x2013;3, 62, &#x2013;7)</td>
<td valign="top" align="center">6.02</td>
<td valign="top" align="center">914</td>
</tr>
<tr>
<td valign="top" align="left">R. Precuneus, R. Posterior Cingulate, R. Fusiform Gyrus</td>
<td valign="top" align="center">(30, &#x2013;37, &#x2013;16)</td>
<td valign="top" align="center">4.99</td>
<td valign="top" align="center">334</td>
</tr>
<tr>
<td valign="top" align="left">L. Fusiform Gyrus, L. Lingual Gyrus, L. Parahippocampal Gyrus</td>
<td valign="top" align="center">(&#x2013;27, &#x2013;61, &#x2013;4)</td>
<td valign="top" align="center">6.50</td>
<td valign="top" align="center">162</td>
</tr>
<tr>
<td valign="top" align="left">R. Superior Temporal Gyrus, R. Middle Temporal Gyrus</td>
<td valign="top" align="center">(51, &#x2013;55, 5)</td>
<td valign="top" align="center">4.98</td>
<td valign="top" align="center">150</td>
</tr>
<tr>
<td valign="top" align="left">L. Middle Temporal Gyrus, L. Inferior Temporal Gyrus</td>
<td valign="top" align="center">(&#x2013;48, &#x2013;13, &#x2013;16)</td>
<td valign="top" align="center">5.06</td>
<td valign="top" align="center">117</td>
</tr>
<tr>
<td valign="top" align="left">R. Rolandic Oper, R. Precentral Gyrus, R. Postcentral Gyrus</td>
<td valign="top" align="center">(69, &#x2013;10, 14)</td>
<td valign="top" align="center">5.09</td>
<td valign="top" align="center">98</td></tr>
</tbody></table>
<table-wrap-foot>
<attrib><italic>Thresholds were set at <italic>P</italic> &#x003C; 0.05, multiple comparisons corrected. MNI, Montreal Neurological Institute; L., left; R., right.</italic></attrib>
</table-wrap-foot>
</table-wrap>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Imaging results. Positive encoding and negative encoding are shown in red and green, respectively. The ch2bet template from mricron was used to be the reference brain to make figures from the statistical maps. <bold>(A)</bold> Encoding of the learning rate. The left panel shows the IFG/anterior insula region (<italic>P</italic> &#x003C; 0.05, corrected). The two other panels show the dACC clusters (<italic>P</italic> &#x003C; 0.001, uncorrected). <bold>(B)</bold> Encoding of the reward feedback (<italic>P</italic> &#x003C; 0.05, corrected). The lower left panel shows the overlap (in purple) of the clusters in <bold>(A)</bold> and brain regions showing significantly stronger encoding of reward feedback than reward prediction error (in blue). The lower right panel shows the overlap (in purple) of the clusters in <bold>(A)</bold> and brain regions significantly encoding of reward feedback when the reward prediction error regressor was also included in the GLM (in blue). <bold>(C)</bold> Brian regions showing significant positive reward feedback &#x00D7; learning rate interaction (<italic>P</italic> &#x003C; 0.05, corrected). <bold>(D)</bold> An MFC region showing significant learning rate &#x00D7; reward prediction error interaction (<italic>P</italic> &#x003C; 0.05, corrected).</p></caption>
<graphic xlink:href="fnhum-11-00592-g004.tif"/>
</fig>
<p>The reward feedback was represented both positively (i.e., higher brain activity if reward was higher than expected) and negatively (i.e., higher brain activity if reward was 1 point) in the brain. Specifically, positive encoding of the feedback was revealed in a widespread brain network, most significantly in the MFC (<italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: 0, 56, 17, <bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>) and the precuneus and posterior cingulate cortex (PCC, <italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: -15, -49, 11, <bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>), whereas negative encoding of the reward feedback was primarily found in the attentional control networks (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>; all regions reported had <italic>P</italic>s &#x003C; 0.05 after correction for multiple comparisons), including dorsal ACC (peak MNI coordinates: -6, 26, 41), bilateral insular and surrounding IFG (peak MNI coordinates: -30, 26, -7; 33, 23, -4), bilateral middle frontal gyri (peak MNI coordinates: -48, 29, 32; 48, 29, 35), and bilateral inferior parietal lobules (peak MNI coordinates: -45, -49, 44; 45, -40, 44).</p>
<p>An alternative explanation was that these regions encoded the reward prediction error, which was correlated with the reward feedback. This is because a high reward always produces positive prediction error and a low reward always produces negative prediction error. One way to tease apart the contribution of reward feedback from the contribution of prediction error would be to regress them against each other and compare the encoding strength of the residues. However, this regression may result in strong anti-correlation between the residues and confound the results. Therefore, as a control analysis, we constructed two GLMs based on the original GLM, with one excluding the reward prediction error regressor and the other excluding the reward feedback regressor, and compared the encoding strength (i.e., beta estimates) of the reward feedback with the encoding strength of the reward prediction error throughout the brain. By separating the regressors in different GLMs, we addressed the anti-correlation issue by not regressing reward feedback and prediction error against each other. The reward feedback regressor displayed stronger encoding strength than the reward prediction error regressor mostly in brain regions showing positive encoding of the reward feedback (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>, lower left panel, <italic>P</italic> &#x003C; 0.05, corrected). Moreover, when both reward feedback and reward prediction error were included in the same model (again without regressing them against each other) for direct comparison, the reward feedback still displayed significant encoding strength in these regions (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>, lower right panel, <italic>P</italic> &#x003C; 0.05, corrected), whereas the reward prediction error did not show significant encoding strength. On the other hand, there was no significant difference in encoding strength in the regions showing negative encoding of the reward feedback.</p>
<p>The key imaging analyses in this study concern how the flexible learning rate interacts with feedback and reward prediction error. The behavioral analysis revealed increased learning rates following high reward. Accordingly, we tested the interaction between the updated learning rate and the reward feedback. Given the positive interaction (i.e., high reward associated with high learning rate) found in the behavioral data, we also expected this interaction to be positive in the fMRI data. The results showed that the MFC (<italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: -3, 62, -7, <bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold>) and the precuneus and PCC (<italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: 30, -37, -16, <bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold>) reported above also showed positive learning rate &#x00D7; feedback interaction. Note that our analysis removed variance shared between model estimates and interactions prior to imaging analysis, so the overlapping results reported are unlikely to be attributed to similarity between model variables and their interaction terms. Finally, we performed an fMRI analysis seeking brain regions encoding the learning rate &#x00D7; prediction error interaction, and found that only a left MFC region negatively encoded this interaction term (<italic>P</italic> &#x003C; 0.05, corrected; peak MNI coordinates: -3, 47, 32, <bold>Figure <xref ref-type="fig" rid="F4">4D</xref></bold>).</p>
</sec>
</sec>
<sec><title>Discussion</title>
<p>The brain adaptively adjusts its learning rate to incorporate recent information in order to make more precise predictions of the future. To address the question of how the learning rate is adjusted, we applied a computational model based on the reinforcement learning model with a flexible learning rate to account for human behavior and brain activity in a reward learning task. Unlike other studies (<xref ref-type="bibr" rid="B33">Niv et al., 2012</xref>) that traced the reward probability for each option, the present study focused on learning which option is more rewarding. This difference of modeling strategy is because, regardless of the manipulation of rewarding probability for individual options, the optimal strategy for the task is always to select the most rewarding option. From this perspective, a novel finding in behavioral analysis was that, in the learning of the more rewarding color, the reward magnitude modulated the learning rate, which further predicted a greater likelihood that participants would repeat a previous choice after obtaining high reward than low reward. In subsequent fMRI analyses, this learning rate &#x00D7; outcome interaction was found in brain regions where the reward feedback was also encoded. Furthermore, to address how the learning rate mediates learning, we probed the representation of the learning rate &#x00D7; prediction error interaction, and found that the trial-by-trial fluctuation in this interaction correlated with the fMRI activity in the MFC.</p>
<p>We started by validating the proposed flexible learning model. Behavioral data was better explained by this model, as compared to reinforcement learning models assuming fixed learning rates (<bold>Figure <xref ref-type="fig" rid="F3">3A</xref></bold>). This suggested that participants indeed adjusted the learning rate during the task (<bold>Figure <xref ref-type="fig" rid="F3">3F</xref></bold>). The result that the flexible learning model outperformed a reinforcement learning model whose learning rate scales with the magnitude of prediction error also suggests that the change of learning rate is not simply mediated by prediction error alone. The estimates of learning rate can be influenced by response history and the interaction between the choice of color and reward (e.g., repeating previous choice after high reward). This integration also allows the model to account for individual differences (<bold>Figure <xref ref-type="fig" rid="F3">3F</xref></bold>).</p>
<p>The finding that the condition-mean learning rate did not differ significantly between the easy and hard conditions suggested that the learning rate adapted to changes at the trial level rather than in a more tonic way (i.e., the block level). This result is also consistent with the flexible learning model&#x2018;s design principle of feedback-driven, trial-by-trial level learning. Previous studies have shown higher estimate of learning rate (volatility) in faster changing environment (<xref ref-type="bibr" rid="B4">Behrens et al., 2007</xref>; <xref ref-type="bibr" rid="B23">Jiang et al., 2015</xref>). This study revealed additional contribution to the learning rate from reward magnitude. Because the subjects obtained higher reward in more frequently in easy condition than hard condition, learning rate may be boosted higher in easy condition. Taken together, the lack of difference in the learning rate between easy and hard conditions may be a result of the modulation of volatility and reward canceling out each other.</p>
<p>The fMRI analysis showed that the learning rate was encoded in the ACC and the IFG (<bold>Figure <xref ref-type="fig" rid="F4">4A</xref></bold>). The former finding replicated the <xref ref-type="bibr" rid="B4">Behrens et al. (2007)</xref> study, which documents that volatility (high volatility translated into high learning rate; see <xref ref-type="bibr" rid="B23">Jiang et al., 2015</xref> for details) is encoded in the ACC. The other finding, that the learning rate was encoded in the IFG, echoes previous studies demonstrating that the IFG activity reflects strategy changes in reading (<xref ref-type="bibr" rid="B31">Moss et al., 2011</xref>) and memory-encoding (<xref ref-type="bibr" rid="B12">Cohen et al., 2014</xref>). Additionally, the IFG is involved in affective switching (<xref ref-type="bibr" rid="B26">Kringelbach and Rolls, 2003</xref>; <xref ref-type="bibr" rid="B40">Remijnse et al., 2005</xref>), and task switching (<xref ref-type="bibr" rid="B13">Crone et al., 2006</xref>). The IFG finding in the present study is also in line with the involvement of the IFG in shifting learning rate/strategy.</p>
<p>A main finding in the behavioral results was that the participants tended to choose the same color more often after receiving a high reward than a low reward (<bold>Figure <xref ref-type="fig" rid="F2">2B</xref></bold>). Given the fact that the participants learned the more rewarding color and chose it more often than chance level (<bold>Figure <xref ref-type="fig" rid="F2">2A</xref></bold>), it is likely that this difference of choice repetition is (in part) due to the larger prediction error from unexpected low reward. Additionally, this change may also be attributed to the learning rate, which influences the updating of the prediction and hence the choice at the next trial. This hypothesis was tested using the flexible learning model, which treats both prediction error and the learning rate as variables that can change after each trial. Like the probability of choice repetition, the model estimate of the learning rate in the subsequent trial increased with reward magnitude. Moreover, the amount of increment in the learning rate also predicted the increase of choice repetition across participants, thus providing strong evidence that the learning rate served as a mechanism leading to the more frequent choice repetition after high reward. An alternative explanation is that the increased prediction error magnitude that signaled the importance of this trial led to the increase in the learning rate. However, this explanation was not supported in that the learning rate was indeed lower after low reward, when the prediction error magnitude should be high (because the participants successfully tracked the more rewarding color most of the time; <bold>Figure <xref ref-type="fig" rid="F2">2A</xref></bold>). According to the flexible learning model, increasing the learning rate after receiving high reward would increase the influence of the current high reward on future predictions of the more rewarding color. As a result, the current color that yielded high reward would be more likely to be selected than if a low reward were received at the current trial. This finding is also related to the literature of the exploration&#x2013;exploitation tradeoff in reward learning (<xref ref-type="bibr" rid="B11">Cohen et al., 2007</xref>), in that high reward is more likely to result in exploitation (i.e., repeatedly choosing the same color to accumulate high reward).</p>
<p>To locate the neural substrates that support feedback mediation of the learning rate, we first conducted an fMRI analysis to test the encoding of the reward feedback. Large-scale encoding of the feedback was observed (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>), including both positive (i.e., high reward > low reward) encoding in the MPFC, precuneus and the PCC, and negative (i.e., low reward > high reward) encoding in the control network. Activity in these regions is highly consistent with the findings reported in <xref ref-type="bibr" rid="B15">Eppinger et al. (2013)</xref>, who studied brain activity associated with positive learning and negative learning. Specifically, positive encoding of reward feedback may reflect reward valuation in the brain (<xref ref-type="bibr" rid="B20">Grabenhorst and Rolls, 2011</xref>), whereas negative encoding may suggest the engagement of the control network in error processing (e.g., obtaining a low reward while anticipating a high reward may be considered as an &#x201C;error&#x201D;) or performance monitoring (<xref ref-type="bibr" rid="B7">Botvinick et al., 2004</xref>), or both. Given the high correlation between the feedback and the reward prediction error, we performed an additional control analysis by comparing the encoding strength between the feedback regressor and the reward prediction error regressor throughout the brain, while keeping other regressors in the GLM. The feedback regressor displayed stronger encoding strength in all the aforementioned regions showing positive encoding of reward feedback, implying a higher likelihood of the feedback being encoded than the reward prediction error (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>). Moreover, when we included both regressors of reward feedback and reward prediction error in the same GLM, only the former showed significant encoding in the reported regions (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>). Therefore, the reported regions in <bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold> were more likely to represent reward feedback than reward prediction error.</p>
<p>Our reported regions did not include striatum, which has been shown by other research to encode prediction error (<xref ref-type="bibr" rid="B34">O&#x2019;Doherty et al., 2003</xref>; <xref ref-type="bibr" rid="B45">Seymour et al., 2004</xref>; <xref ref-type="bibr" rid="B38">Pessiglione et al., 2006</xref>; <xref ref-type="bibr" rid="B19">Glascher et al., 2010</xref>; <xref ref-type="bibr" rid="B25">Jocham et al., 2011</xref>; <xref ref-type="bibr" rid="B49">Zhu et al., 2012</xref>; <xref ref-type="bibr" rid="B15">Eppinger et al., 2013</xref>). We speculate that this lack of striatum finding is because, as we have shown, these regions were likely to encode reward feedback rather than reward prediction error. Although reward feedback and reward prediction error were highly correlated, the latter was defined as the difference between reward feedback and the prediction from the learning model. Therefore, reward feedback, as compared to reward prediction error, may be less involved in learning and prediction, and in turn less likely to be represented in the striatum that supports prediction from learning (<xref ref-type="bibr" rid="B23">Jiang et al., 2015</xref>).</p>
<p>We then performed an independent analysis seeking the brain regions encoding the interaction between the reward feedback and the learning rate that integrated this feedback (via prediction error), and found high degree of overlap between regions encoding this interaction and regions positively encoding the reward feedback (<bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold>). Moreover, these results shown in <bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold> were obtained after removing the shared variance with these two variables, thus excluding the confound of their correlation. Interestingly, the direction of the interaction in these regions was also in line with the positive encoding of reward feedback (<bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold>), which further suggested that the feedback played an important role in mediating the learning rate in these regions. Therefore, these results, when combined together, strongly supported the notion that the learning rate was updated as the feedback was processed.</p>
<p>In a final set of analyses, we examined the integration of the learning rate and the reward prediction error, and discovered an interaction between the learning rate and the reward prediction error in the MFC. Surprisingly, that interaction is negative: For example, when keeping the learning rate constant, fMRI activity decreased as the reward prediction error increased. Nevertheless, an explanation for this negative interaction was that prediction error could be encoded reversely by MFC neurons that signaled negative prediction error (e.g., receiving low reward while high reward was expected: <xref ref-type="bibr" rid="B2">Amiez et al., 2006</xref>; <xref ref-type="bibr" rid="B29">Matsumoto et al., 2007</xref>; <xref ref-type="bibr" rid="B35">Oliveira et al., 2007</xref>; <xref ref-type="bibr" rid="B22">Jessup et al., 2010</xref>). Consistent with this explanation, our fMRI results showed higher activity in the ACC (which is usually considered as part of the MFC) for low reward feedback (corresponding to negative prediction error, <bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>). Moreover, (<xref ref-type="bibr" rid="B22">Jessup et al., 2010</xref>) demonstrated that fMRI activity in the MFC is higher when participants lose in a trial in a gambling task. Interestingly, higher MFC activity when losing only occurs when participants were more likely to win than lose, which was exactly the case in the task used in the present research (<bold>Figure <xref ref-type="fig" rid="F2">2A</xref></bold>). Similar findings showing adaptive learning have also been reported in other domains of learning. For example, based on a multi-level Bayesian model that accounts for multiple forms of uncertainty (<xref ref-type="bibr" rid="B28">Mathys et al., 2011</xref>; <xref ref-type="bibr" rid="B21">Iglesias et al., 2013</xref>) showed that prediction error is modulated by perceptual precision in sensory learning. Moreover, <xref ref-type="bibr" rid="B21">Iglesias et al. (2013)</xref> also reported activity in an MFC region, similar to that illustrated in <bold>Figure <xref ref-type="fig" rid="F4">4D</xref></bold>, that encodes precision-modulated prediction error.</p>
<p>In addition to showing that the integration of learning rate and prediction error co-varies with neural signals in MFC, <xref ref-type="bibr" rid="B10">Chien et al. (2016)</xref> demonstrated that fMRI signals from two striatum regions that separately represent learning rate and reward prediction error jointly account for fMRI signals in the ventral MFC, in which the reward prediction was represented. This finding further links the integration of learning rate and prediction error to the updating of reward prediction. Published results and our own collectively suggest that MFC may serve the function of modulating prediction error-driven updating of future predictions across different domains of learning.</p>
<p>The MFC is involved in multiple cognitive functions, such as error and control monitoring (<xref ref-type="bibr" rid="B5">Botvinick et al., 1999</xref>, <xref ref-type="bibr" rid="B6">2001</xref>, <xref ref-type="bibr" rid="B7">2004</xref>), speed-accuracy tradeoff (<xref ref-type="bibr" rid="B47">Yeung and Nieuwenhuis, 2009</xref>), reward learning (<xref ref-type="bibr" rid="B43">Rushworth and Behrens, 2008</xref>; <xref ref-type="bibr" rid="B22">Jessup et al., 2010</xref>), decision making (<xref ref-type="bibr" rid="B42">Rushworth et al., 2004</xref>), and social cognition (<xref ref-type="bibr" rid="B3">Amodio and Frith, 2006</xref>). Recent modeling work that attempts to summarize the role of the MFC in those functions stresses the importance of learning in MFC functions (<xref ref-type="bibr" rid="B1">Alexander and Brown, 2011</xref>; <xref ref-type="bibr" rid="B46">Silvetti et al., 2011</xref>). In particular, prediction error is considered as the driving force of learning. Our findings further extend this notion by showing that (a) the MFC encoded the modulation of the reward feedback on the learning rate, and (b) the MFC is also involved in adaptive learning, such that the impact of the prediction error on learning can be flexibly adjusted based on factors such as the reward feedback and the volatility (i.e., rate of change in the environment). Therefore, the learning mechanism can determine how important and reliable (or both) the new piece of information is, and adjusts its weight in the new prediction accordingly.</p>
<p>The fMRI activity patterns showing representation of reward magnitude and its interaction with the learning rate in MFC and PCC overlap the default mode network. We speculated that this overlap may be related to a possible function of the default mode network of supporting internal simulations while ignoring external stimulation (<xref ref-type="bibr" rid="B9">Buckner et al., 2008</xref>). That is, the default mode network may be more engaged to support the simulations of performing future trials after high reward and when both reward and learning rate were high.</p>
<p>We employed two conditions with different degrees of difficulty in this task. The difficulty manipulation affected both behavioral (<bold>Figure <xref ref-type="fig" rid="F2">2B</xref></bold>) and simulation (<bold>Figures <xref ref-type="fig" rid="F3">3C,D</xref></bold>) data. Consequently, could difficulty be a confounding factor for the results reported? A closer look at the experimental task revealed that task difficulty consisted of two aspects: First, the hard condition included both a change of the more rewarding color and its rewarding probability. In the flexible learning model, this difference was accounted for by the flexible learning rate that adapted to the change in the environment. Second, the hard condition included additional rewarding probabilities of 0.4 and 0.6, which were, by definition, more ambiguous than probabilities of 0.2 and 0.8 in inferring which color was more rewarding. In other words, probabilities of 0.4 and 0.6 would generate larger reward prediction errors than probabilities of 0.2 and 0.8. Reward prediction error was also modeled by the flexible learning model. Therefore, both aspects of difficulty manipulation in this task have been accounted for. In other words, in this task, we attempted to use learning models to quantify and integrate the effects of abstract factors such as difficulty and change in the environment, so that behavior and neural activity patterns may be parsimoniously explained by concrete, quantifiable factors such as learning rate and reward prediction error.</p>
<p>One caveat of this study is that the task only used two levels of reward magnitude, which resulted in correlation between reward magnitude and reward prediction error. In addition, given that the goal of this study is to maximize reward, obtaining a low reward can be seen as the outcome of an &#x201C;incorrect&#x201D; response, which may further drive learning through prediction error. Even though we alleviated this confound by conducting additional behavioral and fMRI control analyses to show that reported results were better explained by reward magnitude than reward prediction error, an experimental design that varies reward magnitude on a trial-by-trial basis can de-correlate these two factors in the first place, and thus would be a better option than a design with only high vs. low reward.</p>
</sec>
<sec><title>Conclusion</title>
<p>Our findings provide novel evidence suggesting that, in order to achieve the goal of accumulate more reward, the reward feedback (high or low reward) mediated learning rate; and the learning rate further drove the reward prediction error to update the future decision. Importantly, the MFC seems to underlie both functions.</p>
</sec>
<sec><title>Ethics Statement</title>
<p>This study was carried out in accordance with the recommendations of Institutional Review Board of Chengdu University of Information Technology with written informed consent from all subjects. All subjects gave written informed consent in accordance with the Declaration of Helsinki. The protocol was approved by the Institutional Review Board of Chengdu University of Information Technology.</p>
</sec>
<sec><title>Author Contributions</title>
<p>XW, DZ, and JZ designed research. XW, TiW, CL, TaW, and JJ analyzed data. CL acquired the data. XW, TiW, JJ, and JZ wrote the paper. DZ, JJ, and JZ revised the draft.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alexander</surname> <given-names>W. H.</given-names></name> <name><surname>Brown</surname> <given-names>J. W.</given-names></name></person-group> (<year>2011</year>). <article-title>Medial prefrontal cortex as an action-outcome predictor.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>14</volume> <fpage>1338</fpage>&#x2013;<lpage>1344</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2921</pub-id> <pub-id pub-id-type="pmid">21926982</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amiez</surname> <given-names>C.</given-names></name> <name><surname>Joseph</surname> <given-names>J. P.</given-names></name> <name><surname>Procyk</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>Reward encoding in the monkey anterior cingulate cortex.</article-title> <source><italic>Cereb. Cortex</italic></source> <volume>16</volume> <fpage>1040</fpage>&#x2013;<lpage>1055</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhj046</pub-id> <pub-id pub-id-type="pmid">16207931</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amodio</surname> <given-names>D. M.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Meeting of minds: the medial frontal cortex and social cognition.</article-title> <source><italic>Nat. Rev. Neurosci.</italic></source> <volume>7</volume> <fpage>268</fpage>&#x2013;<lpage>277</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1884</pub-id> <pub-id pub-id-type="pmid">16552413</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Behrens</surname> <given-names>T. E. J.</given-names></name> <name><surname>Woolrich</surname> <given-names>M. W.</given-names></name> <name><surname>Walton</surname> <given-names>M. E.</given-names></name> <name><surname>Rushworth</surname> <given-names>M. F. S.</given-names></name></person-group> (<year>2007</year>). <article-title>Learning the value of information in an uncertain world.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>10</volume> <fpage>1214</fpage>&#x2013;<lpage>1221</lpage>. <pub-id pub-id-type="doi">10.1038/nn1954</pub-id> <pub-id pub-id-type="pmid">17676057</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M.</given-names></name> <name><surname>Nystrom</surname> <given-names>L. E.</given-names></name> <name><surname>Fissell</surname> <given-names>K.</given-names></name> <name><surname>Carter</surname> <given-names>C. S.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>1999</year>). <article-title>Conflict monitoring versus selection-for-action in anterior cingulate cortex.</article-title> <source><italic>Nature</italic></source> <volume>402</volume> <fpage>179</fpage>&#x2013;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.1038/46035</pub-id> <pub-id pub-id-type="pmid">10647008</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Braver</surname> <given-names>T. S.</given-names></name> <name><surname>Barch</surname> <given-names>D. M.</given-names></name> <name><surname>Carter</surname> <given-names>C. S.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>2001</year>). <article-title>Conflict monitoring and cognitive control.</article-title> <source><italic>Psychol. Rev.</italic></source> <volume>108</volume> <fpage>624</fpage>&#x2013;<lpage>652</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.108.3.624</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name> <name><surname>Carter</surname> <given-names>C. S.</given-names></name></person-group> (<year>2004</year>). <article-title>Conflict monitoring and anterior cingulate cortex: an update.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>8</volume> <fpage>539</fpage>&#x2013;<lpage>546</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2004.10.003</pub-id> <pub-id pub-id-type="pmid">15556023</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Browning</surname> <given-names>M.</given-names></name> <name><surname>Behrens</surname> <given-names>T. E.</given-names></name> <name><surname>Jocham</surname> <given-names>G.</given-names></name> <name><surname>O&#x2019;Reilly</surname> <given-names>J. X.</given-names></name> <name><surname>Bishop</surname> <given-names>S. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Anxious individuals have difficulty learning the causal statistics of aversive environments.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>18</volume> <fpage>590</fpage>&#x2013;<lpage>596</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3961</pub-id> <pub-id pub-id-type="pmid">25730669</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buckner</surname> <given-names>R. L.</given-names></name> <name><surname>Andrews-Hanna</surname> <given-names>J. R.</given-names></name> <name><surname>Schacter</surname> <given-names>D. L.</given-names></name></person-group> (<year>2008</year>). <article-title>The brain&#x2019;s default network: anatomy, function, and relevance to disease.</article-title> <source><italic>Ann. N. Y. Acad. Sci.</italic></source> <volume>1124</volume> <fpage>1</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1196/annals.1440.011</pub-id> <pub-id pub-id-type="pmid">18400922</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chien</surname> <given-names>S.</given-names></name> <name><surname>Wiehler</surname> <given-names>A.</given-names></name> <name><surname>Spezio</surname> <given-names>M.</given-names></name> <name><surname>Glascher</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Congruence of inherent and acquired values facilitates reward-based decision-making.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>36</volume> <fpage>5003</fpage>&#x2013;<lpage>5012</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3084-15.2016</pub-id> <pub-id pub-id-type="pmid">27147653</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J. D.</given-names></name> <name><surname>McClure</surname> <given-names>S. M.</given-names></name> <name><surname>Yu</surname> <given-names>A. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration.</article-title> <source><italic>Philos. Trans. R. Soc. B Biol. Sci.</italic></source> <volume>362</volume> <fpage>933</fpage>&#x2013;<lpage>942</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2007.2098</pub-id> <pub-id pub-id-type="pmid">17395573</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>M. S.</given-names></name> <name><surname>Rissman</surname> <given-names>J.</given-names></name> <name><surname>Suthana</surname> <given-names>N. A.</given-names></name> <name><surname>Castel</surname> <given-names>A. D.</given-names></name> <name><surname>Knowlton</surname> <given-names>B. J.</given-names></name></person-group> (<year>2014</year>). <article-title>Value-based modulation of memory encoding involves strategic engagement of fronto-temporal semantic processing regions.</article-title> <source><italic>Cogn. Affect. Behav. Neurosci.</italic></source> <volume>14</volume> <fpage>578</fpage>&#x2013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.3758/s13415-014-0275-x</pub-id> <pub-id pub-id-type="pmid">24683066</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crone</surname> <given-names>E. A.</given-names></name> <name><surname>Wendelken</surname> <given-names>C.</given-names></name> <name><surname>Donohue</surname> <given-names>S. E.</given-names></name> <name><surname>Bunge</surname> <given-names>S. A.</given-names></name></person-group> (<year>2006</year>). <article-title>Neural evidence for dissociable components of task-switching.</article-title> <source><italic>Cereb. Cortex</italic></source> <volume>16</volume> <fpage>475</fpage>&#x2013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhi127</pub-id> <pub-id pub-id-type="pmid">16000652</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>D&#x2019;Ardenne</surname> <given-names>K.</given-names></name> <name><surname>McClure</surname> <given-names>S. M.</given-names></name> <name><surname>Nystrom</surname> <given-names>L. E.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>2008</year>). <article-title>BOLD responses reflecting dopaminergic signals in the human ventral tegmental area.</article-title> <source><italic>Science</italic></source> <volume>319</volume> <fpage>1264</fpage>&#x2013;<lpage>1267</lpage>. <pub-id pub-id-type="doi">10.1126/science.1150605</pub-id> <pub-id pub-id-type="pmid">18309087</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eppinger</surname> <given-names>B.</given-names></name> <name><surname>Schuck</surname> <given-names>N. W.</given-names></name> <name><surname>Nystrom</surname> <given-names>L. E.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>2013</year>). <article-title>Reduced striatal responses to reward prediction errors in older compared with younger adults.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>33</volume> <fpage>9905</fpage>&#x2013;<lpage>9912</lpage>. <pub-id pub-id-type="doi">10.1523/Jneurosci.2942-12.2013</pub-id> <pub-id pub-id-type="pmid">23761885</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garrison</surname> <given-names>J.</given-names></name> <name><surname>Erdeniz</surname> <given-names>B.</given-names></name> <name><surname>Done</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Prediction error in reinforcement learning: a meta-analysis of neuroimaging studies.</article-title> <source><italic>Neurosci. Biobehav. Rev.</italic></source> <volume>37</volume> <fpage>1297</fpage>&#x2013;<lpage>1310</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2013.03.023</pub-id> <pub-id pub-id-type="pmid">23567522</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gehring</surname> <given-names>W. J.</given-names></name> <name><surname>Willoughby</surname> <given-names>A. R.</given-names></name></person-group> (<year>2002</year>). <article-title>The medial frontal cortex and the rapid processing of monetary gains and losses.</article-title> <source><italic>Science</italic></source> <volume>295</volume> <fpage>2279</fpage>&#x2013;<lpage>2282</lpage>. <pub-id pub-id-type="doi">10.1126/science.1066893</pub-id> <pub-id pub-id-type="pmid">11910116</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Do learning rates adapt to the distribution of rewards?</article-title> <source><italic>Psychon. Bull. Rev.</italic></source> <volume>22</volume> <fpage>1320</fpage>&#x2013;<lpage>1327</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-014-0790-3</pub-id> <pub-id pub-id-type="pmid">25582684</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glascher</surname> <given-names>J.</given-names></name> <name><surname>Daw</surname> <given-names>N.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>O&#x2019;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2010</year>). <article-title>States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning.</article-title> <source><italic>Neuron</italic></source> <volume>66</volume> <fpage>585</fpage>&#x2013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2010.04.016</pub-id> <pub-id pub-id-type="pmid">20510862</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grabenhorst</surname> <given-names>F.</given-names></name> <name><surname>Rolls</surname> <given-names>E. T.</given-names></name></person-group> (<year>2011</year>). <article-title>Value, pleasure and choice in the ventral prefrontal cortex.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>15</volume> <fpage>56</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.12.004.</pub-id> <pub-id pub-id-type="pmid">21216655</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iglesias</surname> <given-names>S.</given-names></name> <name><surname>Mathys</surname> <given-names>C.</given-names></name> <name><surname>Brodersen</surname> <given-names>K. H.</given-names></name> <name><surname>Kasper</surname> <given-names>L.</given-names></name> <name><surname>Piccirelli</surname> <given-names>M.</given-names></name> <name><surname>den Ouden</surname> <given-names>H. E. M.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Hierarchical prediction errors in midbrain and basal forebrain during sensory learning.</article-title> <source><italic>Neuron</italic></source> <volume>80</volume> <fpage>519</fpage>&#x2013;<lpage>530</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2013.09.009</pub-id> <pub-id pub-id-type="pmid">24139048</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jessup</surname> <given-names>R. K.</given-names></name> <name><surname>Busemeyer</surname> <given-names>J. R.</given-names></name> <name><surname>Brown</surname> <given-names>J. W.</given-names></name></person-group> (<year>2010</year>). <article-title>Error effects in anterior cingulate cortex reverse when error likelihood is high.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>30</volume> <fpage>3467</fpage>&#x2013;<lpage>3472</lpage>. <pub-id pub-id-type="doi">10.1523/Jneurosci.4130-09.2010</pub-id> <pub-id pub-id-type="pmid">20203206</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J. F.</given-names></name> <name><surname>Beck</surname> <given-names>J.</given-names></name> <name><surname>Heller</surname> <given-names>K.</given-names></name> <name><surname>Egner</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>An insula-frontostriatal network mediates flexible cognitive control by adaptively predicting changing control demands.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>6</volume>:<issue>8165</issue>. <pub-id pub-id-type="doi">10.1038/Ncomms9165</pub-id> <pub-id pub-id-type="pmid">26391305</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J. F.</given-names></name> <name><surname>Heller</surname> <given-names>K.</given-names></name> <name><surname>Egner</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Bayesian modeling of flexible cognitive control.</article-title> <source><italic>Neurosci. Biobehav. Rev.</italic></source> <volume>46</volume> <fpage>30</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2014.06.001</pub-id> <pub-id pub-id-type="pmid">24929218</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jocham</surname> <given-names>G.</given-names></name> <name><surname>Klein</surname> <given-names>T. A.</given-names></name> <name><surname>Ullsperger</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Dopamine-mediated reinforcement learning signals in the striatum and ventromedial prefrontal cortex underlie value-based choices.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>31</volume> <fpage>1606</fpage>&#x2013;<lpage>1613</lpage>. <pub-id pub-id-type="doi">10.1523/Jneurosci.3904-10.2011</pub-id> <pub-id pub-id-type="pmid">21289169</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kringelbach</surname> <given-names>M. L.</given-names></name> <name><surname>Rolls</surname> <given-names>E. T.</given-names></name></person-group> (<year>2003</year>). <article-title>Neural correlates of rapid reversal learning in a simple model of human social interaction.</article-title> <source><italic>Neuroimage</italic></source> <volume>20</volume> <fpage>1371</fpage>&#x2013;<lpage>1383</lpage>. <pub-id pub-id-type="doi">10.1016/S1053-8119(03)00393-398</pub-id> <pub-id pub-id-type="pmid">14568506</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lau</surname> <given-names>B.</given-names></name> <name><surname>Glimcher</surname> <given-names>P. W.</given-names></name></person-group> (<year>2005</year>). <article-title>Dynamic response-by-response models of matching behavior in rhesus monkeys.</article-title> <source><italic>J. Exp. Anal. Behav.</italic></source> <volume>84</volume> <fpage>555</fpage>&#x2013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1901/jeab.2005.110-04</pub-id> <pub-id pub-id-type="pmid">16596980</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mathys</surname> <given-names>C.</given-names></name> <name><surname>Daunizeau</surname> <given-names>J.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name></person-group> (<year>2011</year>). <article-title>A bayesian foundation for individual learning under uncertainty.</article-title> <source><italic>Front. Hum. Neurosci.</italic></source> <volume>5</volume>:<issue>39</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2011.00039</pub-id> <pub-id pub-id-type="pmid">21629826</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matsumoto</surname> <given-names>M.</given-names></name> <name><surname>Matsumoto</surname> <given-names>K.</given-names></name> <name><surname>Abe</surname> <given-names>H.</given-names></name> <name><surname>Tanaka</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <article-title>Medial prefrontal cell activity signaling prediction errors of action values.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>10</volume> <fpage>647</fpage>&#x2013;<lpage>656</lpage>. <pub-id pub-id-type="doi">10.1038/nn1890</pub-id> <pub-id pub-id-type="pmid">17450137</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGuire</surname> <given-names>J. T.</given-names></name> <name><surname>Nassar</surname> <given-names>M. R.</given-names></name> <name><surname>Gold</surname> <given-names>J. I.</given-names></name> <name><surname>Kable</surname> <given-names>J. W.</given-names></name></person-group> (<year>2014</year>). <article-title>Functionally dissociable influences on learning rate in a dynamic environment.</article-title> <source><italic>Neuron</italic></source> <volume>84</volume> <fpage>870</fpage>&#x2013;<lpage>881</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2014.10.013</pub-id> <pub-id pub-id-type="pmid">25459409</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moss</surname> <given-names>J.</given-names></name> <name><surname>Schunn</surname> <given-names>C. D.</given-names></name> <name><surname>Schneider</surname> <given-names>W.</given-names></name> <name><surname>McNamara</surname> <given-names>D. S.</given-names></name> <name><surname>Vanlehn</surname> <given-names>K.</given-names></name></person-group> (<year>2011</year>). <article-title>The neural correlates of strategic reading comprehension: cognitive control and discourse comprehension.</article-title> <source><italic>Neuroimage</italic></source> <volume>58</volume> <fpage>675</fpage>&#x2013;<lpage>686</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2011.06.034</pub-id> <pub-id pub-id-type="pmid">21741484</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nassar</surname> <given-names>M. R.</given-names></name> <name><surname>Wilson</surname> <given-names>R. C.</given-names></name> <name><surname>Heasly</surname> <given-names>B.</given-names></name> <name><surname>Gold</surname> <given-names>J. I.</given-names></name></person-group> (<year>2010</year>). <article-title>An approximately Bayesian delta-rule model explains the dynamics of belief updating in a changing environment.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>30</volume> <fpage>12366</fpage>&#x2013;<lpage>12378</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0822-10.2010</pub-id> <pub-id pub-id-type="pmid">20844132</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Niv</surname> <given-names>Y.</given-names></name> <name><surname>Edlund</surname> <given-names>J. A.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>O&#x2019;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2012</year>). <article-title>Neural prediction errors reveal a risk-sensitive reinforcement-learning process in the human brain.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>32</volume> <fpage>551</fpage>&#x2013;<lpage>562</lpage>. <pub-id pub-id-type="doi">10.1523/Jneurosci.5498-10.2012</pub-id> <pub-id pub-id-type="pmid">22238090</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x2019;Doherty</surname> <given-names>J. P.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Critchley</surname> <given-names>H.</given-names></name> <name><surname>Dolan</surname> <given-names>R. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Temporal difference models and reward-related learning in the human brain.</article-title> <source><italic>Neuron</italic></source> <volume>38</volume> <fpage>329</fpage>&#x2013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1016/S0896-6273(03)00169-7</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oliveira</surname> <given-names>F. T.</given-names></name> <name><surname>McDonald</surname> <given-names>J. J.</given-names></name> <name><surname>Goodman</surname> <given-names>D.</given-names></name></person-group> (<year>2007</year>). <article-title>Performance monitoring in the anterior cingulate is not all error related: expectancy deviation and the representation of action-outcome associations.</article-title> <source><italic>J. Cogn. Neurosci.</italic></source> <volume>19</volume> <fpage>1994</fpage>&#x2013;<lpage>2004</lpage>. <pub-id pub-id-type="doi">10.1162/jocn.2007.19.12.1994</pub-id> <pub-id pub-id-type="pmid">17892382</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Payzan-LeNestour</surname> <given-names>E.</given-names></name> <name><surname>Bossaerts</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>Risk, unexpected uncertainty, and estimation uncertainty: bayesian learning in unstable settings.</article-title> <source><italic>PLOS Comput. Biol.</italic></source> <volume>7</volume>:<issue>e1001048</issue>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1001048</pub-id> <pub-id pub-id-type="pmid">21283774</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pearce</surname> <given-names>J. M.</given-names></name> <name><surname>Hall</surname> <given-names>G.</given-names></name></person-group> (<year>1980</year>). <article-title>A model for Pavlovian learning: variations in the effectiveness of conditioned but not of unconditioned stimuli.</article-title> <source><italic>Psychol. Rev.</italic></source> <volume>87</volume> <fpage>532</fpage>&#x2013;<lpage>552</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.87.6.532</pub-id> <pub-id pub-id-type="pmid">7443916</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pessiglione</surname> <given-names>M.</given-names></name> <name><surname>Seymour</surname> <given-names>B.</given-names></name> <name><surname>Flandin</surname> <given-names>G.</given-names></name> <name><surname>Dolan</surname> <given-names>R. J.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Dopamine-dependent prediction errors underpin reward-seeking behaviour in humans.</article-title> <source><italic>Nature</italic></source> <volume>442</volume> <fpage>1042</fpage>&#x2013;<lpage>1045</lpage>. <pub-id pub-id-type="doi">10.1038/nature05051</pub-id> <pub-id pub-id-type="pmid">16929307</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Priestley</surname> <given-names>M. B.</given-names></name></person-group> (<year>1981</year>). <source><italic>Spectral Analysis and Time Series.</italic></source> <publisher-loc>London</publisher-loc>: <publisher-name>Academic Press</publisher-name>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Remijnse</surname> <given-names>P. L.</given-names></name> <name><surname>Nielen</surname> <given-names>M. M. A.</given-names></name> <name><surname>Uylings</surname> <given-names>H. B. M.</given-names></name> <name><surname>Veltman</surname> <given-names>D. J.</given-names></name></person-group> (<year>2005</year>). <article-title>Neural correlates of a reversal learning task with an affectively neutral baseline: an event-related fMRI study.</article-title> <source><italic>Neuroimage</italic></source> <volume>26</volume> <fpage>609</fpage>&#x2013;<lpage>618</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2005.02.009</pub-id> <pub-id pub-id-type="pmid">15907318</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rescorla</surname> <given-names>R. A.</given-names></name> <name><surname>Wagner</surname> <given-names>A. R.</given-names></name></person-group> (<year>1972</year>). <article-title>&#x201C;A theory of pavlovian conditioning: variations in the effectiveness of reinforcement and nonreinforcement,&#x201D; in</article-title> <source><italic>Classical Conditioning II. Appleton-Century-Crofts</italic></source> <role>eds</role> <person-group person-group-type="editor"><name><surname>Black</surname> <given-names>A. H.</given-names></name> <name><surname>Prokasy</surname> <given-names>W. F.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Appleton-Century-Crofts</publisher-name>) <fpage>64</fpage>&#x2013;<lpage>99</lpage>.</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rushworth</surname> <given-names>M. F.</given-names></name> <name><surname>Walton</surname> <given-names>M. E.</given-names></name> <name><surname>Kennerley</surname> <given-names>S. W.</given-names></name> <name><surname>Bannerman</surname> <given-names>D. M.</given-names></name></person-group> (<year>2004</year>). <article-title>Action sets and decisions in the medial frontal cortex.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>8</volume> <fpage>410</fpage>&#x2013;<lpage>417</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2004.07.009</pub-id> <pub-id pub-id-type="pmid">15350242</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rushworth</surname> <given-names>M. F. S.</given-names></name> <name><surname>Behrens</surname> <given-names>T. E. J.</given-names></name></person-group> (<year>2008</year>). <article-title>Choice, uncertainty and value in prefrontal and cingulate cortex.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>11</volume> <fpage>389</fpage>&#x2013;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1038/nn2066</pub-id> <pub-id pub-id-type="pmid">18368045</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schultz</surname> <given-names>W.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Montague</surname> <given-names>P. R.</given-names></name></person-group> (<year>1997</year>). <article-title>A neural substrate of prediction and reward.</article-title> <source><italic>Science</italic></source> <volume>275</volume> <fpage>1593</fpage>&#x2013;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1126/science.275.5306.1593</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seymour</surname> <given-names>B.</given-names></name> <name><surname>O&#x2019;Doherty</surname> <given-names>J. P.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Koltzenburg</surname> <given-names>M.</given-names></name> <name><surname>Jones</surname> <given-names>A. K.</given-names></name> <name><surname>Dolan</surname> <given-names>R. J.</given-names></name><etal/></person-group> (<year>2004</year>). <article-title>Temporal difference models describe higher-order learning in humans.</article-title> <source><italic>Nature</italic></source> <volume>429</volume> <fpage>664</fpage>&#x2013;<lpage>667</lpage>. <pub-id pub-id-type="doi">10.1038/nature02581</pub-id> <pub-id pub-id-type="pmid">15190354</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silvetti</surname> <given-names>M.</given-names></name> <name><surname>Seurinck</surname> <given-names>R.</given-names></name> <name><surname>Verguts</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>Value and prediction error in medial frontal cortex: integrating the single-unit and systems levels of analysis.</article-title> <source><italic>Front. Hum. Neurosci.</italic></source> <volume>5</volume>:<issue>75</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2011.00075</pub-id> <pub-id pub-id-type="pmid">21886616</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeung</surname> <given-names>N.</given-names></name> <name><surname>Nieuwenhuis</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Dissociating response conflict and error likelihood in anterior cingulate cortex.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>29</volume> <fpage>14506</fpage>&#x2013;<lpage>14510</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3615-09.2009</pub-id> <pub-id pub-id-type="pmid">19923284</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeung</surname> <given-names>N.</given-names></name> <name><surname>Sanfey</surname> <given-names>A. G.</given-names></name></person-group> (<year>2004</year>). <article-title>Independent coding of reward magnitude and valence in the human brain.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>24</volume> <fpage>6258</fpage>&#x2013;<lpage>6264</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4537-03.2004</pub-id> <pub-id pub-id-type="pmid">15254080</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>L.</given-names></name> <name><surname>Mathewson</surname> <given-names>K. E.</given-names></name> <name><surname>Hsu</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Dissociable neural representations of reinforcement and belief prediction errors underlie strategic learning.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>109</volume> <fpage>1419</fpage>&#x2013;<lpage>1424</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1116783109</pub-id> <pub-id pub-id-type="pmid">22307594</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="http://psychtoolbox.org/">http://psychtoolbox.org/</ext-link></p></fn>
<fn id="fn02"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="http://www.fil.ion.ucl.ac.uk/spm/">http://www.fil.ion.ucl.ac.uk/spm/</ext-link></p></fn>
<fn id="fn03"><label>3</label><p><ext-link ext-link-type="uri" xlink:href="http://afni.nimh.nih.gov/pub/dist/doc/program_help/3dClustSim.html">http://afni.nimh.nih.gov/pub/dist/doc/program_help/3dClustSim.html</ext-link></p></fn>
</fn-group>
</back>
</article>