<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Behav. Neurosci.</journal-id>
<journal-title>Frontiers in Behavioral Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Behav. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5153</issn>
<publisher>
<publisher-name>Frontiers Research Foundation</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbeh.2010.00184</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Reinforcement Learning Model of Precommitment in Decision Making</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Kurth-Nelson</surname> <given-names>Zeb</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Redish</surname> <given-names>A. David</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001">&#x0002A;</xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Neuroscience, University of Minnesota</institution> <country>Minneapolis, MN, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Daeyeol Lee, Yale University School of Medicine, USA</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Joseph W. Kable, University of Pennsylvania, USA; Veit Stuphorn, Johns Hopkins University, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: A. David Redish, Department of Neuroscience, University of Minnesota, 6-145 Jackson Hall, 321 Church Street, SE, Minneapolis, MN 55455, USA. e-mail: <email>redish&#x00040;umn.edu</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>12</month>
<year>2010</year>
</pub-date>
<pub-date pub-type="collection">
<year>2010</year>
</pub-date>
<volume>4</volume>
<elocation-id>184</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>09</month>
<year>2010</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>11</month>
<year>2010</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2010 Kurth-Nelson and Redish.</copyright-statement>
<copyright-year>2010</copyright-year>
<license license-type="open-access" xlink:href="http://www.frontiersin.org/licenseagreement"><p>This is an open-access article subject to an exclusive license agreement between the authors and the Frontiers Research Foundation, which permits unrestricted use, distribution, and reproduction in any medium, provided the original authors and source are credited.</p></license>
</permissions>
<abstract>
<p>Addiction and many other disorders are linked to impulsivity, where a suboptimal choice is preferred when it is immediately available. One solution to impulsivity is precommitment: constraining one&#x00027;s future to avoid being offered a suboptimal choice. A form of impulsivity can be measured experimentally by offering a choice between a smaller reward delivered sooner and a larger reward delivered later. Impulsive subjects are more likely to select the smaller-sooner choice; however, when offered an option to precommit, even impulsive subjects can precommit to the larger-later choice. To precommit or not is a decision between two conditions: (A) the original choice (smaller-sooner vs. larger-later), and (B) a new condition with only larger-later available. It has been observed that precommitment appears as a consequence of the preference reversal inherent in non-exponential delay-discounting. Here we show that most models of hyperbolic discounting cannot precommit, but a distributed model of hyperbolic discounting does precommit. Using this model, we find (1) faster discounters may be more or less likely than slow discounters to precommit, depending on the precommitment delay, (2) for a constant smaller-sooner vs. larger-later preference, a higher ratio of larger reward to smaller reward increases the probability of precommitment, and (3) precommitment is highly sensitive to the shape of the discount curve. These predictions imply that manipulations that alter the discount curve, such as diet or context, may qualitatively affect precommitment.</p>
</abstract>
<kwd-group>
<kwd>precommitment</kwd>
<kwd>delay discounting</kwd>
<kwd>reinforcement learning</kwd>
<kwd>hyperbolic discounting</kwd>
<kwd>decision making</kwd>
<kwd>impulsivity</kwd>
<kwd>addiction</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="1"/>
<equation-count count="15"/>
<ref-count count="58"/>
<page-count count="13"/>
<word-count count="10046"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction">
<title>Introduction</title>
<p>Precommitment is a general mechanism to control impulsive behavior (Ainslie, <xref ref-type="bibr" rid="B2">1975</xref>, <xref ref-type="bibr" rid="B4">2001</xref>; Dripps, <xref ref-type="bibr" rid="B18">1993</xref>). An alcoholic trying to quit may decide to avoid going to the bar, knowing that if he goes, he will drink. By avoiding the bar, he precommits to the decision of not drinking. Similarly, a heroin addict may take methadone even though it will preclude the euphoria of heroin. In general, precommitment is an action that alters the external environment (in the methadone example, the external environment includes the neuropharmacology of the individual) to foreclose the possibility of a future impulsive choice. It is important to note that precommitment is not synonymous with self-control: self-control entails an act of willpower to avoid an impulsive choice; precommitment strategies actually constrain the agent&#x00027;s future choices. The strategy of precommitment is ubiquitous in decision-making outside of addiction as well. Putting the ice cream out of sight, investing money in an inaccessible retirement fund, and pre-paid gym memberships can be precommitment devices.</p>
<p>Precommitment devices combat <italic>impulsivity</italic>, the overvaluation of immediate rewards relative to delayed rewards<xref ref-type="fn" rid="fn1"><sup>1</sup></xref>. To describe impulsivity quantitatively, we use the generalized notion of <italic>delay discounting</italic>, the decrease in subjective value associated with rewards that are more distant in the future (Koopmans, <xref ref-type="bibr" rid="B31">1960</xref>; Fishburn and Rubinstein, <xref ref-type="bibr" rid="B20">1982</xref>; Mazur, <xref ref-type="bibr" rid="B36">1987</xref>; Madden and Bickel, <xref ref-type="bibr" rid="B34">2010</xref>). A <italic>discounting function</italic> specifies how much less subjective value a reward has at any given time in the future. An individual&#x00027;s discounting function can be inferred from choices. For example, if a subject prefers &#x00024;20 in a week equally to &#x00024;10 today, then the subject discounts by 50% over that one week. Drug addicts discount faster than non-addicts (Madden et al., <xref ref-type="bibr" rid="B35">1997</xref>; Bickel et al., <xref ref-type="bibr" rid="B9">1999</xref>; Coffey et al., <xref ref-type="bibr" rid="B10">2003</xref>; Dom et al., <xref ref-type="bibr" rid="B16">2006</xref>). Thus one potential driver for addiction is that an addict may prefer a small immediate reward (drugs) over a large delayed reward (academic and career success, family, health, etc.). The ability to precommit is especially valuable when impulsive behavior is leading to serious problems such as drug abuse.</p>
<p>If each unit of time by which the reward is delayed causes the same attenuation of the reward&#x00027;s value, then discounting is <italic>exponential</italic>. In other words, the subjective value is attenuated by &#x003B3;<italic><sup>d</sup></italic>, where <italic>d</italic> is the delay to the reward, and &#x003B3; is the amount of attenuation incurred by each unit of time. Exponential discounting is theoretically optimal in certain situations (Samuelson, <xref ref-type="bibr" rid="B48">1937</xref>) and can be calculated recursively (Bellman, <xref ref-type="bibr" rid="B6">1958</xref>; Sutton and Barto, <xref ref-type="bibr" rid="B52">1998</xref>). Exponential discounting also has the property that two rewards separated by a given delay will maintain their relative values whether they are considered well in advance or they are near at hand. However, behavioral studies show that humans and animals do not discount exponentially; real discounting is usually better fit by a <italic>hyperbolic</italic> function (Ainslie, <xref ref-type="bibr" rid="B2">1975</xref>; Madden and Bickel, <xref ref-type="bibr" rid="B34">2010</xref>). In hyperbolic discounting, subjective value is attenuated by 1/(1&#x02009;&#x0002B;&#x02009;<italic>kd</italic>), where <italic>d</italic> is again the delay to reward, and <italic>k</italic> determines the steepness of the hyperbolic curve.</p>
<p>An agent using hyperbolic discounting will exhibit <italic>preference reversal</italic> (Strotz, <xref ref-type="bibr" rid="B51">1955</xref>; Ainslie, <xref ref-type="bibr" rid="B3">1992</xref>; Frederick et al., <xref ref-type="bibr" rid="B21">2002</xref>). As the time at which a choice is considered changes, the preference order reverses. For example, the agent may prefer to receive &#x00024;10 today over &#x00024;15 in a week, but prefer &#x00024;15 in 53 weeks over &#x00024;10 in 52 weeks. Humans and animals consistently display preference reversal (Madden and Bickel, <xref ref-type="bibr" rid="B34">2010</xref>). In fact, preference reversal is not exclusive to hyperbolic discounting. If discounting consists of a function that maps delay to an attenuation of subjective value, then exponential decay is the only discounting function in which preference reversal does not appear.</p>
<p>It has been suggested that preference reversal is the basis for precommitment (Ainslie, <xref ref-type="bibr" rid="B3">1992</xref>). Preference reversal entails a conflict between current (non-impulsive) and future (impulsive) preferences. This conflict leads the individual to commit to current preferences, to prevent the future self from undermining these preferences. For example, in 52 weeks, the agent will be given a choice between &#x00024;10 then or &#x00024;15 a week from then. If that choice is made freely, he will choose the &#x00024;10. Because he presently prefers the &#x00024;15, he may enter now into a contract that binds him to choosing the &#x00024;15 option when the choice becomes available.</p>
<p>To explicitly test whether precommitment can be learned in a controlled setting, Rachlin and Green (<xref ref-type="bibr" rid="B42">1972</xref>) and Ainslie (<xref ref-type="bibr" rid="B1">1974</xref>) trained pigeons on tasks that required choosing between smaller-sooner and larger-later options. After learning this paradigm, the pigeons were given an option preceding this choice, to inactivate the smaller-sooner option. Some pigeons that preferred smaller-sooner over larger-later would nonetheless elect to inactivate the smaller-sooner option &#x02013; thereby precommitting to the larger-later option. The pigeons&#x02019; willingness to precommit increased with the delay between precommitment and choice.</p>
<p>Here we develop a quantitative theory of how such precommitment may occur, based on reinforcement learning. Precommitment has not previously been implemented in a reinforcement learning model. Four models have been proposed to explain how a biological learning system could plausibly calculate hyperbolic discounting. First, in the average reward model (Tsitsiklis and Van Roy, <xref ref-type="bibr" rid="B57">1999</xref>; Daw and Touretzky, <xref ref-type="bibr" rid="B12">2000</xref>; Dezfouli et al., <xref ref-type="bibr" rid="B14">2009</xref>), discounting across states is linear (i.e., each additional unit of delay subtracts a constant from the subjective value), but the slope of this linear discounting is set, based on the reward magnitude, such that the total discounting over <italic>d</italic> delay is 1/(1&#x02009;&#x0002B;&#x02009;<italic>d</italic>). An average reward variable keeps track of the slope so that it is available for the linear discounting calculation at each state. Second, in a variant of the average reward model (Alexander and Brown, <xref ref-type="bibr" rid="B5">2010</xref>), if <italic>R</italic> is the average reward per trial and <italic>V</italic> is the value after <italic>d</italic> delay, then the value after <italic>d</italic>&#x02009;&#x0002B;&#x02009;1 delay is calculated as <italic>VR</italic>/(<italic>V</italic>&#x02009;&#x0002B;&#x02009;<italic>R</italic>). This produces hyperbolic discounting across states, given a linear state-space (i.e., a chain of states with no branches or choices) leading to a reward. The third method of calculating hyperbolic discounting is semi-Markov state representations, where a variable amount of time can elapse while the agent dwells within a single state (Daw, <xref ref-type="bibr" rid="B11">2003</xref>). The agent can simply compute the hyperbolic discount factor (1/(1&#x02009;&#x0002B;&#x02009;<italic>d</italic>)) over the entire duration, <italic>d</italic>, of the state. In the fourth model, exponential discounting is performed in parallel at various rates by a set of reinforcement learning &#x0201C;&#x003BC;Agents,&#x0201D; who collectively form the decision-making system of the overall agent (Kurth-Nelson and Redish, <xref ref-type="bibr" rid="B32">2009</xref>). In the &#x003BC;Agents model, choices of the overall agent are derived by taking the average value belief over the set of &#x003BC;Agents. Averaging across a set of different exponential discount curves yields a good approximation of hyperbolic discounting (e.g., Sozou, <xref ref-type="bibr" rid="B50">1998</xref>) that functions over multiple state transitions.</p>
<p>In this paper we first show that, of the four available hyperbolic discounting models, only the &#x003BC;Agents model can precommit. We then use the &#x003BC;Agents model to test specific predictions about the properties of precommitment behavior. Understanding the basis of precommitment may help to create situations where precommitment will be successful in the treatment of addiction as well as inform the study of decision-making in general.</p>
</sec>
<sec sec-type="materials|methods">
<title>Materials and Methods</title>
<p>We compare four models in this paper. Each model is an implementation of <italic>temporal difference reinforcement learning</italic> (TDRL) (Sutton and Barto, <xref ref-type="bibr" rid="B52">1998</xref>). Each model consisted of a simulated <italic>agent</italic> operating in an external <italic>world</italic>. The agent performed <italic>actions</italic> that influenced the state of the world, and in certain states the world supplied <italic>rewards</italic> to the agent. The available <italic>states</italic> of the world, together with the set of possible <italic>transitions</italic> between states, formed a <italic>state-space</italic>. Each state <italic>i</italic> was associated with a number <italic>R</italic>(<italic>i</italic>) (which may be 0) specifying how much reward the agent received upon leaving that state. In <italic>semi-Markov</italic> models, each state was also associated with a delay specifying the temporal extent of the state (how long the agent must wait before receiving the reward and/or transitioning to another state) (Daw, <xref ref-type="bibr" rid="B11">2003</xref>). In <italic>fully-Markov</italic> models, each state had a delay of one time unit, so variable delays were modeled by increasing or decreasing the number of states.</p>
<p>The agent learned, for each state, the total expected future reward from that state, discounted by the delay to reach that reward (or rewards). This discounted expected future reward is called <italic>value</italic>. The value of state <italic>i</italic> is called <italic>V</italic>(<italic>i</italic>). Learning these values allowed the agent, faced with a choice between two states, to choose the state that would lead to more total expected reward<xref ref-type="fn" rid="fn2"><sup>2</sup></xref>.</p>
<p>The parameter values used in the simulations are listed in Table <xref ref-type="table" rid="T1">1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Parameters used in the model (except where noted otherwise)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left">Parameter</th>
<th align="left">Description</th>
<th align="right">Default value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" colspan="3"><bold>COMMON</bold></td>
</tr>
<tr>
<td align="left"><italic>D</italic><sub>C</sub></td>
<td align="left">Delay between P and (C or N)</td>
<td align="right">100</td>
</tr>
<tr>
<td align="left"><italic>D</italic><sub>S</sub></td>
<td align="left">Delay between C and SS</td>
<td align="right">1</td>
</tr>
<tr>
<td align="left"><italic>R</italic><sub>S</sub></td>
<td align="left">Magnitude of smaller-sooner reward</td>
<td align="right">10</td>
</tr>
<tr>
<td align="left"><italic>D</italic><sub>L</sub></td>
<td align="left">Delay between C and LL</td>
<td align="right">50</td>
</tr>
<tr>
<td align="left"><italic>R</italic><sub>L</sub></td>
<td align="left">Magnitude of larger-later reward</td>
<td align="right">50</td>
</tr>
<tr>
<td align="left">&#x003B1;</td>
<td align="left">Learning rate</td>
<td align="right">0.1</td>
</tr>
<tr>
<td align="left" colspan="3"><bold>&#x003BC;AGENTS MODEL</bold></td>
</tr>
<tr>
<td align="left"><italic>N</italic><sub>&#x003BC;</sub></td>
<td align="left">Number of &#x003BC;Agents</td>
<td align="right">1000</td>
</tr>
<tr>
<td align="left"><italic>K</italic></td>
<td align="left">Hyperbolic discount rate</td>
<td align="right">1</td>
</tr>
<tr>
<td align="left" colspan="3"><bold>AVERAGE REWARD MODEL</bold></td>
</tr>
<tr>
<td align="left">&#x003C3;</td>
<td align="left">Average reward update rate</td>
<td align="right">0.002</td>
</tr>
<tr>
<td align="left" colspan="3"><bold>HDTD MODEL</bold></td>
</tr>
<tr>
<td align="left">&#x003C3;</td>
<td align="left">Average reward update rate</td>
<td align="right">0.01</td>
</tr>
<tr>
<td align="left" colspan="3"><bold>SEMI-MARKOV MODEL</bold></td>
</tr>
<tr>
<td align="left"><italic>k</italic></td>
<td align="left">Hyperbolic discount rate</td>
<td align="right">1</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec>
<title>&#x003BC;agents model</title>
<p>The &#x003BC;Agents model is described in detail in Kurth-Nelson and Redish (<xref ref-type="bibr" rid="B32">2009</xref>). The model produces identical behavior (including precommitment) in semi-Markov or fully-Markov state-spaces.</p>
<p>In the &#x003BC;Agents model, the agent (which we will sometimes call &#x0201C;macro-agent&#x0201D; for clarity) consisted of a set of <italic>&#x003BC;Agents</italic>, each performing TDRL independently. For a given state, different &#x003BC;Agents could learn different values. To select actions, the different values across &#x003BC;Agents were averaged. The only difference between &#x003BC;Agents was that each &#x003BC;Agent had a different discount rate.</p>
<p>Upon each state-transition from state <italic>x</italic> to state <italic>y</italic>, each &#x003BC;Agent <italic>i</italic> generated an <italic>error signal</italic>, &#x003B4;<italic><sub>i</sub></italic>, reflecting the discrepancy between (discounted) value observed and value predicted:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:msub><mml:mo>&#x003B4;</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>V<sub>i</sub></italic>(<italic>x</italic>) is the value of state <italic>x</italic> learned by &#x003BC;Agent <italic>i</italic>, &#x003B3;<italic><sub>i</sub></italic> is the discount rate of &#x003BC;Agent <italic>i</italic>, and <italic>d</italic> is the delay spent in state <italic>x</italic>. Note that the total benefit of moving from state <italic>x</italic> to state <italic>y</italic> is the reward received (<italic>R</italic>(<italic>x</italic>)) plus the reward expected in the future of the new state (<italic>V<sub>i</sub></italic>(<italic>y</italic>)). Because future value is attenuated by a constant multiple (&#x003B3;<italic><sub>i</sub></italic>) for each unit of delay, the discounting of each &#x003BC;Agent is exponential. The values of &#x003B3; were spread uniformly over the interval: [1/(<italic>N</italic><sub>&#x003BC;</sub>&#x02009;&#x0002B;&#x02009;1), 1&#x02009;&#x02212;&#x02009;1/(<italic>N</italic><sub>&#x003BC;</sub>&#x02009;&#x0002B;&#x02009;1)], where <italic>N</italic><sub>&#x003BC;</sub> is the number of &#x003BC;Agents. Thus if there were nine &#x003BC;Agents, they would have discount rates 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9.</p>
<p>To improve value estimates, each &#x003BC;Agent <italic>i</italic> used the error signal to update its <italic>V<sub>i</sub></italic>(<italic>x</italic>) after each state-transition:</p>
<disp-formula id="E2"><mml:math id="M2"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:msub><mml:mo>&#x003B4;</mml:mo><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>
<p>where &#x003B1; is a learning rate in (0,1) common to all &#x003BC;Agents. &#x003B1;&#x02009;&#x0003D;&#x02009;0.1 was used in all simulations.</p>
<p>From some states, actions were available to the macro-agent. Let <italic>A</italic> be the set of possible actions. Since each action in our simulations leads to a unique state, <italic>A</italic> is equivalently a set of states. The probability of selecting action <italic>a</italic>&#x02009;&#x02208;&#x02009;<italic>A</italic> was:</p>
<disp-formula id="E3"><mml:math id="M3"><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M4"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the average value of state <italic>a</italic> across &#x003BC;Agents. Note that the probabilities sum to one across the set of available actions; exactly one action from <italic>A</italic> was chosen.</p>
<p>Because each &#x003BC;Agent had an independent discount rate, this model is considered to perform <italic>distributed discounting</italic>. One consequence of distributed discounting is that although each individual &#x003BC;Agent performs exponential discounting, the overall discounting produced by the macro-agent approaches hyperbolic as the number of &#x003BC;Agents increases:</p>
<disp-formula id="E4"><mml:math id="M5"><mml:mrow><mml:munder><mml:mrow><mml:mi>lim</mml:mi><mml:mo>&#x02061;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mo>&#x003BC;</mml:mo></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mo>&#x0221E;</mml:mo></mml:mrow></mml:munder><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x02009;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mo>&#x003BC;</mml:mo></mml:msub></mml:mrow></mml:munderover><mml:mrow><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mi>x</mml:mi></mml:msubsup></mml:mrow></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:munderover><mml:mo>&#x0222B;</mml:mo><mml:mn>0</mml:mn><mml:mn>1</mml:mn></mml:munderover><mml:mrow><mml:msup><mml:mo>&#x003B3;</mml:mo><mml:mi>x</mml:mi></mml:msup></mml:mrow></mml:mrow></mml:mstyle><mml:mi>d</mml:mi><mml:mo>&#x003B3;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>when hyperbolic discounting is implemented as a sum of exponentials, it is hyperbolic across multiple state transitions. For more details on how distributed exponential discounting produces hyperbolic discounting, see Kurth-Nelson and Redish (<xref ref-type="bibr" rid="B32">2009</xref>). In this model we were also able to adjust the effective hyperbolic parameter <italic>k</italic> by biasing the distribution of &#x003BC;Agent exponential discount rates (&#x003B3;) (Kurth-Nelson and Redish, <xref ref-type="bibr" rid="B32">2009</xref>).</p>
</sec>
<sec>
<title>Average reward model</title>
<p>The average reward model (Tsitsiklis and Van Roy, <xref ref-type="bibr" rid="B57">1999</xref>; Daw and Touretzky, <xref ref-type="bibr" rid="B12">2000</xref>; Dezfouli et al., <xref ref-type="bibr" rid="B14">2009</xref>) uses a fully-Markov state-space. A variable <inline-formula><mml:math id="M6"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> tracked the average reward per timestep:</p>
<disp-formula id="E5"><mml:math id="M7"><mml:mrow><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>&#x02190;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003C3;</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mo>&#x003C3;</mml:mo><mml:mi>R</mml:mi></mml:mrow></mml:math></disp-formula>
<p>where <italic>R</italic> was the reward received on this timestep, and &#x003C3; controlled the rate at which <inline-formula><mml:math id="M8"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> changed. The reward prediction error &#x003B4;, upon transition from state <italic>x</italic> to state <italic>y</italic>, was calculated as:</p>
<disp-formula id="E6"><mml:math id="M9"><mml:mrow><mml:mo>&#x003B4;</mml:mo><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:mrow></mml:math></disp-formula>
<p>This produced linear discounting, because the value of <italic>x</italic> approached the value of y minus <inline-formula><mml:math id="M10"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> (which is effectively a constant because &#x003C3; is very small). In a linear state space with <italic>d</italic>&#x02032; delay leading to <italic>R</italic>&#x02032; reward, <inline-formula><mml:math id="M11"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> would approach <italic>R</italic>&#x02032;/(1&#x02009;&#x0002B;&#x02009;<italic>d</italic>&#x02032;) where (1&#x02009;&#x0002B;&#x02009;<italic>d</italic>&#x02032;) is the total length of a trial, including one time step to receive the reward), which is hyperbolic discounting as a function of total delay. In other words, the average reward model discounts linearly across a given state-space, but the total discounting across this state space is hyperbolic because the linear rate depends on the total delay of the state-space. The average reward model does not show hyperbolic discounting in a state-space with choices (branch points), because <inline-formula><mml:math id="M12"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> no longer approaches <italic>R</italic>&#x02032;/(1&#x02009;&#x0002B;&#x02009;<italic>d</italic>&#x02032;).</p>
</sec>
<sec>
<title>HDTD model</title>
<p>The HDTD model (Alexander and Brown, <xref ref-type="bibr" rid="B5">2010</xref>) also uses a fully-Markov state-space. In this model, average reward is tracked per trial rather than per timestep, but using the same update rule as the average reward model:</p>
<disp-formula id="E7"><mml:math id="M13"><mml:mrow><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>&#x02190;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003C3;</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mo>&#x003C3;</mml:mo><mml:mi>R</mml:mi></mml:mrow></mml:math></disp-formula>
<p>The reward prediction error was calculated as:</p>
<disp-formula id="E8"><mml:math id="M14"><mml:mrow><mml:mo>&#x003B4;</mml:mo><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula>
<p>In a linear state space with <italic>d</italic>&#x02032; delay leading to <italic>R</italic>&#x02032; reward, <inline-formula><mml:math id="M15"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> would approach <italic>R</italic>&#x00027;. Through algebra, <italic>V</italic>(<italic>x</italic>) would approach:</p>
<disp-formula id="E9"><mml:math id="M16"><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>+</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>which would produce hyperbolic discounting across states. In other words, in a linear chain of states, the value of a state is attenuated as a hyperbolic function of the temporal distance from that state to the reward. The HDTD model does not show hyperbolic discounting in a state-space with choices (branch points), because <inline-formula><mml:math id="M17"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> no longer approaches <italic>R</italic>&#x00027;.</p>
</sec>
<sec>
<title>Semi-markov model</title>
<p>The semi-Markov model is a standard temporal difference reinforcement learning model in a semi-Markov state-space (Daw, <xref ref-type="bibr" rid="B11">2003</xref>). This model was identical to &#x003BC;Agents, except there was only a single reinforcement learning entity, and it used the following rule instead of Eq <xref ref-type="disp-formula" rid="E1">1</xref>:</p>
<disp-formula id="E10"><mml:math id="M18"><mml:mrow><mml:mo>&#x003B4;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>R</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x02212;</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where again <italic>k</italic> is the discount rate (set to 1 for these simulations), and <italic>d</italic> is the delay spent in state <italic>x</italic> before transitioning to state <italic>y</italic>.</p>
</sec>
</sec>
<sec>
<title>Results</title>
<p>We ran the &#x003BC;Agents model on the precommitment state-space illustrated in Figure <xref ref-type="fig" rid="F1">1</xref>. The states and transitions inside the dashed box represent a simple choice between a smaller reward (<italic>R</italic><sub>S</sub>) available after a short delay (<italic>D</italic><sub>S</sub>) and a larger reward (<italic>R</italic><sub>L</sub>) available after a long delay (<italic>D</italic><sub>L</sub>). We will refer to the smaller-sooner choice as SS and the larger-later choice as LL. C was the state from which this choice is available. Although <italic>R</italic><sub>L</sub> was a larger reward than <italic>R</italic><sub>S</sub>, SS could be the preferred choice if <italic>D</italic><sub>S</sub> was sufficiently shorter than <italic>D</italic><sub>L</sub>, due to temporal discounting of future rewards.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>A reinforcement learning state-space with an option to precommit to a larger-later reward</bold>. This is the task on which the model was run. Time is schematically portrayed along the horizontal axis. The dashed box encloses the state-space for a simple two-alternative choice. From state C, the agent could choose a large reward <italic>R</italic><sub>L</sub> following a long delay <italic>D</italic><sub>L</sub>, or a small reward <italic>R</italic><sub>S</sub> following a short delay <italic>D</italic><sub>S</sub>. Preceding this simple choice was the option to precommit. From state P, the agent could either enter the simple choice by choosing C, or to precommit to the large delayed choice by choosing N. A precommitment delay <italic>D</italic><sub>C</sub> separated P from the subsequent state.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g001.tif"/>
</fig>
<p>Over the course of learning, each &#x003BC;Agent independently approached a steady-state estimate of the correct exponentially discounted value of each state. When values were fully learned, <inline-formula><mml:math id="M19"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mtext>S</mml:mtext></mml:msub><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>S</mml:mtext></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M20"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mtext>L</mml:mtext></mml:msub><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>L</mml:mtext></mml:msub></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> for a given &#x003BC;Agent <italic>i</italic>. Thus, the value estimates averaged across &#x003BC;Agents approximated the hyperbolically discounted values:</p>
<disp-formula id="E11"><label>(2)</label><mml:math id="M21"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mtext>S</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mtext>S</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E12"><label>(3)</label><mml:math id="M22"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mtext>L</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mtext>L</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>At state C, if the agent always selected SS, then the value of C would approach the discounted value of SS. Since the agent sometimes selected LL, the actual value of C was between the discounted values of SS and LL. The steady-state value of state C, for a given &#x003BC;Agent <italic>i</italic>, was:</p>
<disp-formula id="E13"><mml:math id="M23"><mml:mtable><mml:mtr><mml:mtd><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>C</mml:mtext></mml:msub></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mtext>S</mml:mtext></mml:msub><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>S</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>C</mml:mtext></mml:msub></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mtext>L</mml:mtext></mml:msub><mml:mo>&#x022C5;</mml:mo><mml:msubsup><mml:mo>&#x003B3;</mml:mo><mml:mi>i</mml:mi><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mtext>L</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>C</mml:mtext></mml:msub></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M24"><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability of selecting SS from C, and <inline-formula><mml:math id="M25"><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability of selecting LL from C. Thus, the average value of the C state approximated</p>
<disp-formula id="E14"><label>(4)</label><mml:math id="M26"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>SS</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>S</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>S</mml:mtext></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mtext>LL</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x022C5;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mtext>L</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>L</mml:mtext></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>Precommitment can be defined as limiting one&#x00027;s future options so that only the pre-selected option is available. This is represented in Figure <xref ref-type="fig" rid="F1">1</xref> by the state P which precedes the SS vs. LL choice. At state P, a choice was available to either enter the SS vs. LL choice, or to enter a situation (state N) from which only the LL option was available. In either case, there was a delay <italic>D</italic><sub>C</sub> following the choice made at state P.</p>
<p>The value of state N, averaged across &#x003BC;Agents, approximated</p>
<disp-formula id="E15"><label>(5)</label><mml:math id="M27"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mtext>L</mml:mtext></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>in the steady-state. By definition, the macro-agent preferred to precommit if and only if <inline-formula><mml:math id="M28"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0003E;</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<sec>
<title>&#x003BC;agents model precommits</title>
<p>We examined choice behavior in the model, using a specific set of parameters designed to simulate the choice between a small reward available immediately (<italic>R</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;10, <italic>D</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;1) and a large reward available later (<italic>R</italic><sub>L</sub>&#x02009;&#x0003D;&#x02009;50, <italic>D</italic><sub>L</sub>&#x02009;&#x0003D;&#x02009;50). The discounting rate <italic>k</italic> was set to 1. The number of &#x003BC;Agents, <italic>N</italic><sub>&#x003BC;</sub>, was set to 1000. The preference of the model was measured from the choices made at steady-state (after learning). Using these parameters, the model preferred SS over LL by a ratio of 5.2:1 (Figure <xref ref-type="fig" rid="F2">2</xref>A).</p>
<fig id="F2" position="float">
<label>Figure 2</label><caption><p><bold>The &#x003BC;Agents model exhibited precommitment behavior</bold>. <bold>(A)</bold> From P, the model chose N more often than C. However, when C was reached, the model chose SS more often than LL. This bar graph represents the same data as <bold>(B)</bold> at <italic>R</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;10. <bold>(B)</bold> plots the same information as <bold>(A)</bold>, over a range of values of <italic>R</italic><sub>S</sub>. For very small <italic>R</italic><sub>S</sub>, LL&#x02009;&#x0003E;&#x02009;SS and N&#x02009;&#x0003E;&#x02009;C. For very large <italic>R</italic><sub>S</sub>, SS&#x02009;&#x0003E;&#x02009;LL, and C&#x02009;&#x0003E;&#x02009;N. However, precommitment could occur because SS crossed LL at a different point than C crossed N, meaning that there was a range of <italic>R</italic><sub>S</sub> for which SS&#x02009;&#x0003E;&#x02009;LL and N&#x02009;&#x0003E;&#x02009;C. <bold>(C)</bold> illustrates the change in the average value of each option over time. The distributed value representation of &#x003BC;Agents allowed average values to cross. The value of C prior to discounting across the <italic>D</italic><sub>C</sub> interval was between the discounted values of SS and LL, and was therefore above N. However, during the <italic>D</italic><sub>C</sub> interval, the value of C crossed under N, such that N&#x02009;&#x0003E;&#x02009;C at the time of the C/N choice. Note that <bold>(C)</bold> uses a fully Markov state-space to show how values are discounted through each intermediate time point. Each circle shows the value of one state after values are fully learned.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g002.tif"/>
</fig>
<p>We also looked at the precommitment behavior of the model; that is, the preference of the model for the <italic>N</italic> state over the C state. The precommitment delay, <italic>D</italic><sub>C</sub>, was set to 100. Choices were again counted after the model had reached steady-state. Despite a strong preference for SS over LL, the model also exhibited a preference for N over C, by a ratio of 2.3:1 (Figure <xref ref-type="fig" rid="F2">2</xref>A). It is interesting to note that each &#x003BC;Agent that prefers SS over LL also prefers C over N (and vice versa), yet the system as a whole can prefer SS over LL while preferring N over C.</p>
<p>Thus, the model precommitted to LL even though SS was strongly preferred over LL. This shows that precommitment can occur for particular task parameters. We next varied the small reward magnitude, <italic>R</italic><sub>S</sub>, to see in what range precommitment was possible (Figure <xref ref-type="fig" rid="F2">2</xref>B). If <italic>R</italic><sub>S</sub> was very small, then SS was not valuable and LL was preferred over SS. If <italic>R</italic><sub>S</sub> was very large, the model would not precommit because the SS choice was too valuable. However, there was a range of <italic>R</italic><sub>S</sub> where SS was preferred over LL but N was preferred over C. This range was where the model would choose to precommit to avoid an impulsive choice. Note that in this graph we have manipulated only <italic>R</italic><sub>S</sub> for convenience; similar plots can be generated by manipulating any of these task parameters: <italic>R</italic><sub>S</sub>, <italic>D</italic><sub>S</sub>, <italic>R</italic><sub>L</sub>, and <italic>D</italic><sub>L</sub>.</p>
<p>To illustrate how precommitment arose in the &#x003BC;Agents model, we ran the model on a fully-Markov version of the precommitment state-space (Figure <xref ref-type="fig" rid="F2">2</xref>C). Here, each state corresponds to a single time-step, so a delay is represented by a chain of states. Each circle represents the average value of one state. First, note that the values of states in the LL and N chains overlapped because they were the same temporal distance from the same reward (<italic>R</italic><sub>L</sub>). The last state in the C chain necessarily had a value intermediate to the first states in the SS and LL chains<xref ref-type="fn" rid="fn3"><sup>3</sup></xref>. Thus, if SS was preferred to LL, then the last state in the C chain must have a value above the corresponding state in the N chain. This means that if there was no delay <italic>D</italic><sub>C</sub>, C would be preferred to N. However, during the delay <italic>D</italic><sub>C</sub>, the C chain crossed under the N chain. This was possible because the &#x003BC;Agents collectively maintain a distribution of values for each state (which is collapsed to a single average value for the sake of action selection). The same average value can come from different distributions. In particular, the last state in the C chain had more value concentrated in the fast-discounting &#x003BC;Agents (because of the contribution from SS), while the corresponding state in the N chain had more value concentrated in the slow-discounting &#x003BC;Agents. Thus more of the value in the C chain was attenuated by the delay <italic>D</italic><sub>C</sub>, allowing N to be preferred to C.</p>
<p>Note that in the limit as the number of &#x003BC;Agents goes to infinity and the learning rate &#x003B1; goes to 0, the choice behavior of the &#x003BC;Agents model (including precommitment) becomes analytically equivalent to &#x0201C;<italic>mathematical</italic>&#x0201D; hyperbolic discounting. For example, Figure <xref ref-type="fig" rid="F8">8</xref>A compares the precommitment behavior produced by either mathematical hyperbolic discounting or 1000 &#x003BC;Agents. In mathematical hyperbolic discounting, the value of each choice is calculated by the right-hand side of Eqs. <xref ref-type="disp-formula" rid="E11">2</xref>&#x02013;<xref ref-type="disp-formula" rid="E15">5</xref>. This produces &#x0201C;sophisticated&#x0201D; decision-making (O&#x00027;Donoghue and Rabin, <xref ref-type="bibr" rid="B40">1999</xref>), because the value of <italic>C</italic> is influenced by the relative probability of selecting SS vs. LL.</p>
</sec>
<sec>
<title>Other hyperbolic discounting models cannot precommit</title>
<p>We investigated the behavior of three other models of hyperbolic discounting on the precommitment task. None of these models produce precommitment behavior, because they do not correctly implement hyperbolic discounting across two choices. In general, it is impossible to precommit (where precommitment is defined, using the state-space given in this paper, as preferring N over C while also preferring SS over LL) with temporal difference learning if we make the assumption that discounting preserves order (i.e., if <italic>x</italic><sub>1</sub> is greater than <italic>x</italic><sub>2</sub>, then <italic>x</italic><sub>1</sub> discounted by <italic>d</italic> delay is greater than <italic>x</italic><sub>2</sub> discounted by the same <italic>d</italic> delay). Prior to discounting across the <italic>D</italic><sub>C</sub> interval, <italic>V</italic>(C) lies between <italic>V</italic>(SS) and <italic>V</italic>(LL), while <italic>V</italic>(N) is equal to <italic>V</italic>(LL). Thus the undiscounted <italic>V</italic>(C) is greater than the undiscounted <italic>V</italic>(N) if and only if <italic>V</italic>(SS)&#x02009;&#x0003E;&#x02009;<italic>V</italic>(LL). But both <italic>V</italic>(C) and <italic>V</italic>(N) are discounted by <italic>D</italic><sub>C</sub>. Assuming discounting preserves order, then <italic>V</italic>(C)&#x02009;&#x0003E;&#x02009;<italic>V</italic>(N)&#x02009;&#x021D4;&#x02009;<italic>V</italic>(SS)&#x02009;&#x0003E;&#x02009;<italic>V</italic>(LL).</p>
<p>Note that precommitment is possible in the &#x003BC;Agents model because the assumption that discounting preserves order is violated. Value is represented as a distribution across &#x003BC;Agents, so it is possible that the average undiscounted value of V(C) is greater than the average undiscounted value of V(N), but after discounting both distributions by the same delay, the average value of V(C) is less than the average value of V(N).</p>
<p>We implemented the average reward model of hyperbolic discounting (Tsitsiklis and Van Roy, <xref ref-type="bibr" rid="B57">1999</xref>; Daw and Touretzky, <xref ref-type="bibr" rid="B12">2000</xref>; Dezfouli et al., <xref ref-type="bibr" rid="B14">2009</xref>). In this model, when SS was preferred to LL, C was also preferred to N (Figures <xref ref-type="fig" rid="F3">3</xref>A,B). Although discounting as a function of total delay is hyperbolic in this model, the discounting from state to state is approximately linear (Figure <xref ref-type="fig" rid="F3">3</xref>C). The value of the last state in the C chain is between the discounted values of SS and LL, and this value is greater than the value of the corresponding state in the N chain (provided that SS is preferred to LL). Unlike in the &#x003BC;Agents model, the C and N chains never cross, so C is preferred to N at the time of the C/N choice (Figure <xref ref-type="fig" rid="F3">3</xref>C). The average reward model fundamentally fails to precommit because it only produces hyperbolic discounting across a linear state-space. When the state-space includes choices (branch points), discounting is no longer hyperbolic.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>The average reward model of hyperbolic discounting cannot produce precommitment behavior</bold>. <bold>(A)</bold> At the selected parameters, SS was preferred to LL and C was preferred to N. This bar graph represents the same data as (<bold>B)</bold> at <italic>R</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;44.5. <bold>(B)</bold> As <italic>R</italic><sub>S</sub> increased, the preference for C overtook N exactly at the same point as the preference for SS overtook LL, meaning that this model would not precommit to avoid an impulsive choice. <bold>(C)</bold> In the average reward model, the undiscounted value of C is again between the discounted values of SS and LL. However, the values of C and N never cross, so C is preferred to N whenever SS is preferred to LL. Note that discounting across states in this model is approximately linear.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g003.tif"/>
</fig>
<p>The HDTD model (Alexander and Brown, <xref ref-type="bibr" rid="B5">2010</xref>) is a variant of the average reward model which allows for hyperbolic discounting from state to state. Like the average reward model, HDTD preferred C to N whenever it preferred SS to LL (Figures <xref ref-type="fig" rid="F4">4</xref>A,B). Like the average reward model, the C and N chains never cross (Figure <xref ref-type="fig" rid="F4">4</xref>C), so if SS is preferred to LL, then C is preferred to N. Again, the reason the HDTD model cannot precommit is that it only produces hyperbolic discounting across a linear state-space. Both the average reward and HDTD models use an &#x0201C;average reward&#x0201D; variable to violate the Markov property, altering the discount rate based on the delay to reward. Because there is only a single &#x0201C;average reward&#x0201D; variable in each model, this mechanism works only when there is a single reward to track the delay to.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>The HDTD model of hyperbolic discounting cannot produce precommitment behavior</bold>. <bold>(A)</bold> At the selected parameters, SS was preferred to LL and C was preferred to N. This bar graph represents the same data as <bold>(B)</bold> at <italic>R</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;1.5. <bold>(B)</bold> As <italic>R</italic><sub>S</sub> increased, the preference for C overtook N exactly at the same point as the preference for SS overtook LL, meaning that this model would not precommit to avoid an impulsive choice. <bold>(C)</bold> In the HDTD model, the undiscounted value of C is between the discounted values of SS and LL. However, the values of C and N never cross, so C is preferred to N whenever SS is preferred to LL.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g004.tif"/>
</fig>
<p>We also tested a semi-Markov model without distributed discounting (Daw, <xref ref-type="bibr" rid="B11">2003</xref>). This model did not exhibit precommitment behavior (Figures <xref ref-type="fig" rid="F5">5</xref>A,B). Whenever SS was preferred to LL, C was preferred to N.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>The semi-Markov model of hyperbolic discounting (without distributed discounting) cannot produce precommitment behavior</bold>. <bold>(A)</bold> At the selected parameters, SS was preferred to LL and C was preferred to N. This bar graph represents the same data as <bold>(B)</bold> at <italic>R</italic><sub>S</sub>&#x02009;&#x0003D;&#x02009;3.46. <bold>(B)</bold> As <italic>R</italic><sub>S</sub> increased, the preference for C overtook N exactly at the same point as the preference for SS overtook LL, meaning that this model would not precommit to avoid an impulsive choice. Note that in this model, each state has a long temporal extent (and replacing long states with one-step states produces different discounting), so it is not possible to plot intermediate discounted values as in Figures <xref ref-type="fig" rid="F2">2</xref>C, <xref ref-type="fig" rid="F3">3</xref>C, and <xref ref-type="fig" rid="F4">4</xref>C.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g005.tif"/>
</fig>
</sec>
<sec>
<title>Precommitment depends on task parameters</title>
<p>Precommitment is behavior that avoids the opportunity to choose an impulsive option even though that option would be preferred given the choice. In order to understand what kind of situations are most favorable to precommitment behavior, we looked at how precommitment preference can be maximized for a given ratio of preference between SS and LL. We again looked at precommitment over a range of <italic>R</italic><sub>S</sub>, but this time adjusted <italic>D</italic><sub>S</sub> concurrently with <italic>R</italic><sub>S</sub> such that <italic>R</italic><sub>S</sub>&#x02032;/(1&#x02009;&#x0002B;&#x02009;<italic>kD</italic><sub>S</sub>) (the discounted value of the SS option) was held constant. We found that a smaller <italic>R<sub>S</sub></italic> (paired with a correspondingly shorter <italic>D</italic><sub>S</sub>) was always more favorable to precommitment (Figure <xref ref-type="fig" rid="F6">6</xref>A). In other words, the model was more likely to precommit when the impulsive choice was smaller and more immediate. Equivalently, a very large, very late LL choice always produced greater precommitment than an equivalently valued but modestly large and late LL choice. (Decreasing <italic>R</italic><sub>S</sub> is equivalent to increasing <italic>R</italic><sub>L</sub>, because choice is unaffected by equal scaling of the two reward magnitudes.)</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Precommitment depends on task parameters</bold>. <bold>(A)</bold> Precommitment was greater when the reward ratio <italic>R</italic><sub>L</sub>/<italic>R</italic><sub>S</sub> was larger. On the <italic>x</italic> axis, <italic>D</italic><sub>S</sub> and <italic>R</italic><sub>S</sub> were simultaneously manipulated so that the SS vs. LL preference was held constant. Despite the constant SS vs. LL preference, the preference for precommitment increased as the size of the small reward decreased. <bold>(B)</bold> Precommitment increases with the delay (<italic>D</italic><sub>C</sub>) between commitment and choice. At the selected parameters, LL was preferred 0.196 as much as SS. When there was no delay between commitment and choice, C was preferred over N by the same ratio. As <italic>D</italic><sub>C</sub> increased, the relative preference for N increased, despite no change in the relative preference for SS vs. LL. This figure was generated with mathematical hyperbolic discounting.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g006.tif"/>
</fig>
<p>We next investigated the effect of changing the delay <italic>D</italic><sub>C</sub> on precommitment behavior. Rachlin and Green (<xref ref-type="bibr" rid="B42">1972</xref>) observed that precommitment increases as the delay increases between the first and second choices. To look for a similar effect in the model, we plotted the relative preference for N (i.e., <inline-formula><mml:math id="M29"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula>) against a changing <italic>D</italic><sub>C</sub> (Figure <xref ref-type="fig" rid="F6">6</xref>B). As <italic>D</italic><sub>C</sub> increased, we observed an increase in preference for N, asymptotically approaching a constant as <italic>D</italic><sub>C</sub>&#x02009;&#x02192;&#x02009;&#x0221E;. This matches the result of Rachlin and Green (<xref ref-type="bibr" rid="B42">1972</xref>). We found a qualitatively similar effect of varying <italic>D</italic><sub>C</sub> for any given values of <italic>k</italic>, <italic>D</italic><sub>S</sub>, <italic>R</italic><sub>S</sub>, <italic>D</italic><sub>L</sub>, and <italic>R</italic><sub>L</sub> (data not shown).</p>
<p>The delay <italic>D</italic><sub>C</sub> occurred before the SS vs. LL choice and therefore did not affect the relative preference for these options. This means that although the model held a constant preference for SS over LL as <italic>D</italic><sub>C</sub> varied, the model switched from strongly preferring not to precommit to strongly preferring to precommit as <italic>D</italic><sub>C</sub> increased.</p>
</sec>
<sec>
<title>Precommitment depends on discount rate</title>
<p>In the model, hyperbolic discounting arises as the average of many exponential discount curves. If the discount rates of the individual exponential functions are spread uniformly over the interval (0,1), then the sum of these functions approaches (as <italic>N</italic><sub>&#x003BC;</sub>&#x02009;&#x02192;&#x02009;&#x0221E;) a hyperbolic function 1/(1&#x02009;&#x0002B;&#x02009;<italic>kd</italic>) with <italic>k</italic>&#x02009;&#x0003D;&#x02009;1. Altering the distribution of exponential discount rates to a non-uniform distribution alters the resulting average and can produce hyperbolic functions with any desired value of <italic>k</italic> (Kurth-Nelson and Redish, <xref ref-type="bibr" rid="B32">2009</xref>).</p>
<p>We took advantage of this to test the behavior of the model at different values of <italic>k</italic>. We found that the effect of varying <italic>k</italic> depended on the task parameters. Specifically, the effect of changing <italic>k</italic> was opposite for different values of <italic>D</italic><sub>C</sub> (Figure <xref ref-type="fig" rid="F7">7</xref>A). For small <italic>D</italic><sub>C</sub>, faster discounting (larger <italic>k</italic>) led to less precommitment (Figure <xref ref-type="fig" rid="F7">7</xref>B). But for large <italic>D</italic><sub>C</sub>, faster discounting led to more precommitment (Figure <xref ref-type="fig" rid="F7">7</xref>C). As <italic>k</italic> increased, the relative preference for smaller-sooner over larger-later also increased (not shown). The reason for these opposite results at different values of <italic>D</italic><sub>C</sub> is that increasing <italic>k</italic> has two effects. First, it boosts the preference for SS over LL. Second, it makes the <italic>D</italic><sub>C</sub> interval effectively longer (because <italic>k</italic> is simply a time dilation factor), so the difference in discounting between <italic>D</italic><sub>C</sub>&#x02009;&#x0002B;&#x02009;<italic>D</italic><sub>S</sub> and <italic>D</italic><sub>C</sub>&#x02009;&#x0002B;&#x02009;<italic>D</italic><sub>L</sub> is diminished. The first effect dominates when <italic>D</italic><sub>C</sub> is small, and the second effect dominates when <italic>D</italic><sub>C</sub> is large.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>The discounting rate has a complex interaction with precommitment behavior</bold>. <bold>(A)</bold> Precommitment preference plotted against <italic>k</italic> and <italic>D</italic><sub>C</sub>. The heavy black lines show the intersection of the planes <italic>D</italic><sub>C</sub>&#x02009;&#x0003D;&#x02009;5 and <italic>D</italic><sub>C</sub>&#x02009;&#x0003D;&#x02009;100 with this surface, representing the lines in <bold>(B)</bold> and <bold>(C)</bold>. <bold>(B)</bold> When <italic>D</italic><sub>C</sub> was small (<italic>D</italic><sub>C</sub>&#x02009;&#x0003D;&#x02009;5), faster discounting yielded less precommitment. The dashed line represents the value of <italic>k</italic> above which SS was preferred over LL. <bold>(C)</bold> When <italic>D</italic><sub>C</sub> was large (<italic>D</italic><sub>C</sub>&#x02009;&#x0003D;&#x02009;100), faster discounting yielded greater precommitment. The dashed line represents the value of k above which SS was preferred over LL. <bold>(D)</bold> As k increased, the amount of <italic>D</italic><sub>C</sub> required to prefer precommitment also increased. Faster discounting increased the amount of <italic>D</italic><sub>C</sub> required to produce precommitment. When <italic>k</italic> was less than 4/45, LL was preferred over SS, so N was preferred over C even when there was no precommitment delay. <bold>(E)</bold> The SS&#x02013;LL preference ratio was held constant by adjusting <italic>R</italic><sub>S</sub>. Under this condition, the amount of <italic>D</italic><sub>C</sub> required to precommit grew as <italic>k</italic> increased. In other words, for a given SS&#x02013;LL preference, agents with slower discounting will require a much longer <italic>D</italic><sub>C</sub> to achieve precommitment. This figure was generated with mathematical hyperbolic discounting.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g007.tif"/>
</fig>
<p>As <italic>k</italic> increased, the amount of delay <italic>D</italic><sub>C</sub> needed to produce precommitment also increased (Figure <xref ref-type="fig" rid="F7">7</xref>D). This is because faster discounters have a stronger preference for SS over LL, and more <italic>D</italic><sub>C</sub> was needed to overcome this preference. However, if <italic>k</italic> was manipulated while holding the SS vs. LL preference constant (by adjusting the magnitude of <italic>R</italic><sub>S</sub>), then the opposite effect was seen. For a given degree of preference for SS over LL, faster discounters required less <italic>D</italic><sub>C</sub> to achieve precommitment (Figure <xref ref-type="fig" rid="F7">7</xref>E).</p>
</sec>
<sec>
<title>Precommitment depends on shape of discount curve</title>
<p>As described above, the fidelity of the model&#x00027;s approximation to true hyperbolic discounting depended on the number of &#x003BC;Agents. Because precommitment in the model depended on non-exponential discounting, we investigated how changing the number of &#x003BC;Agents influenced precommitment behavior. With 1000 &#x003BC;Agents, the model produced precommitment that was very similar to the precommitment produced by true hyperbolic discounting. However, we found that reducing the number of &#x003BC;Agents to 100 eliminated precommitment preference under the selected parameters (Figure <xref ref-type="fig" rid="F8">8</xref>A). The difference between the discount curves with 100 vs. 1000 exponentials was slight (Figure <xref ref-type="fig" rid="F8">8</xref>B), indicating that precommitment behavior is highly sensitive to the precise shape of the discount curve.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Precommitment is highly dependent on the shape of the discount curve</bold>. (<bold>A)</bold> These graphs plot the average (across 200 trials) number of times each choice was selected at different values of <italic>D</italic><sub>C</sub>. The blue line shows choices made when <italic>N</italic><sub>&#x003BC;</sub>&#x02009;&#x0003D;&#x02009;1000, and the red line shows choices made when <italic>N</italic><sub>&#x003BC;</sub>&#x02009;&#x0003D;&#x02009;100. <bold>(B)</bold> The shape of the average discount function depends on the number of &#x003BC;Agents. Red shows the average discounting of 100 &#x003BC;Agents and green shows the average discounting of 1000 &#x003BC;Agents. If a perfect hyperbolic function is plotted on this graph, it overlies the green curve. For reference, exponential discount curves with &#x003B3; &#x0003D;&#x02009;0.1 and &#x003B3; &#x0003D;&#x02009;0.9 are plotted in blue. The vertical axis is log scale to make visible the separation of the red and green curves. In this figure, <italic>k</italic> &#x0003D;&#x02009;5.</p></caption>
<graphic xlink:href="fnbeh-04-00184-g008.tif"/>
</fig>
<p>To understand why precommitment occurred with 1000 but not with 100 &#x003BC;Agents, we tested the model with 100 &#x003BC;Agents but adjusted the slowest &#x003B3; from 0.99 to 0.999 (0.999 is the slowest &#x003B3; when there are 1000 &#x003BC;Agents). This adjustment was sufficient to recover precommitment behavior (data not shown), suggesting that the presence of this very slow discounting component was necessary for precommitment. The &#x003B3;&#x02009;&#x0003D;&#x02009;0.999 &#x003BC;Agent had little relative effect on values that are discounted over short delays, because those average values received a significant contribution from other &#x003BC;Agents. But on values discounted over long delays, the &#x003B3;&#x02009;&#x0003D;&#x02009;0.999 &#x003BC;Agent had a predominant effect, contributing far more than 1/100th of the average value. The &#x003B3;&#x02009;&#x0003D;&#x02009;0.999 &#x003BC;Agent effectively propped up the tail of the average discount curve without having a significant impact on the early part of the curve. This enhanced the degree of preference reversal inherent in the curve, which is the feature essential for precommitment.</p>
</sec>
</sec>
<sec sec-type="discussion">
<title>Discussion</title>
<p>In this paper, we have presented a reinforcement learning account of precommitment. The advance decision (precommitment) to avoid an impulsive choice is a natural consequence of hyperbolic discounting, because in hyperbolic discounting, preferences reverse as a choice is viewed from a distance. Our model performs hyperbolic discounting by using a set of independent &#x0201C;&#x003BC;Agents&#x0201D; performing exponential discounting in parallel, each at a different rate. This model demonstrates that a reinforcement learning system, implementing hyperbolic discounting, can exhibit precommitment. Precommitment also illustrates the more general problem of non-exponential discounting in complex state-spaces that include choices. To our knowledge, no other reinforcement learning models of hyperbolic discounting function correctly in such state-spaces. It is interesting to note that our model also matches Ainslie&#x00027;s (<xref ref-type="bibr" rid="B2">1975</xref>) prediction for <italic>bundled</italic> choices. Ainslie observed that even if SS is preferred when a choice is considered in isolation, hyperbolic discounting implies that if the present choice dictates the outcome of several future choices, LL may be preferred. Because our model implements hyperbolic discounting over complex state-spaces, it matches this prediction (data not shown).</p>
<p>In this paper, we have also made quantitative predictions about precommitment behavior, extending the work of Ainslie (<xref ref-type="bibr" rid="B2">1975</xref>, <xref ref-type="bibr" rid="B4">2001</xref>). Except for the prediction that subtle changes in the discount curve affect precommitment, all of the predictions here are general consequences of hyperbolic discounting, whether in a model-free or model-based system, and do not depend specifically on the &#x003BC;Agents model. To our knowledge, none of these predictions have been tested behaviorally. These predictions may inform the development of strategies to encourage precommitment behavior in patients with addiction.</p>
<sec>
<title>Predictions of the model and implications for treating addiction</title>
<p>Avoiding situations where a drug choice is immediately available may be the most important step in recovery from addiction. In this section, we outline the predictions of our model and discuss the implications for what factors may bias addicts&#x02019; decisions toward choosing to avoid such situations.</p>
<p>Pigeons show increased precommitment behavior when the delay (called <italic>D</italic><sub>C</sub> in this paper) between the first and second choice is increased (Rachlin and Green, <xref ref-type="bibr" rid="B42">1972</xref>). Our model reproduced this result (Figure <xref ref-type="fig" rid="F6">6</xref>B). This suggests that when designing precommitment interventions, long pre-choice delays are critical. An addict with a bag of heroin in his pocket may not choose to take methadone, but if the choice to take methadone could be presented further in advance, the addict may be willing to precommit. In general, if we can find out when addicts are going to have access to drug choices, we should offer precommitment devices as far in advance from these choices as possible.</p>
<p>Hyperbolic discounting also implies other properties of precommitment behavior. These properties would hold in any model that correctly implements hyperbolic discounting across multiple choices. First, hyperbolic discounting predicts that precommitment will be most differentially reinforced when the smaller-sooner reward is small in magnitude relative to the larger-later reward (Figure <xref ref-type="fig" rid="F6">6</xref>A). This is true even when the delays are modulated such that the two rewards themselves retain the same preference ratio. Thus a very large, very late reward should be more effective at producing precommitment than a modest but earlier delayed reward. Likewise, precommitment interventions are likely to be more successful when the impulsive choice is more immediate. This is promising for treatment because often the &#x0201C;larger-later&#x0201D; reward is a healthy, productive life, which is very large and very late. This prediction also suggests that precommitment devices may be more useful for drugs that deliver a small reward following a short latency, such as cigarettes.</p>
<p>Second, hyperbolic discounting predicts that precommitment behavior is influenced by discount rate. In situations where <italic>D</italic><sub>C</sub> is small, a faster discount rate makes the model less likely to precommit (Figure <xref ref-type="fig" rid="F7">7</xref>B). On the other hand, when <italic>D</italic><sub>C</sub> is large, a faster discount rate actually makes the model more likely to precommit (Figure <xref ref-type="fig" rid="F7">7</xref>C). Conversely, this means that individuals with a faster discount rate will see more benefit from lengthening <italic>D</italic><sub>C</sub>. This result runs counter to the intuition that impulsive, fast discounting individuals would be less likely to commit to long-range strategies. In fact, they are very likely to commit, because when the choice is viewed in advance, the SS and LL discount nearly identically, and the LL has a larger magnitude. In other words, for faster discounters, the hyperbolic curve flattens out faster. For a given degree of preference for SS over LL, faster discounters require less <italic>D</italic><sub>C</sub> to prefer precommitment (Figure <xref ref-type="fig" rid="F7">7</xref>E). This implies that measuring an individual&#x00027;s discount rate could help to determine what intervention strategies will be effective. If an individual is on the cusp of indecision, offering a precommitment device with a short delay may be sufficient for faster discounters, but ineffective for slow discounters. This also suggests the idea of two different addiction phenotypes, one for slow discounters and one for fast discounters. Slow discounters may not respond to the intervention strategies that work for fast discounters.</p>
<p>The &#x003BC;Agents model also makes a prediction that is not made by hyperbolic discounting alone. Subtle changes in the shape of the discount function, independent of overall discount rate, alter the preference for precommitment (Figure <xref ref-type="fig" rid="F8">8</xref>). This has two implications. First, it is likely that whatever the mechanism of discounting in humans and animals, it is not perfectly hyperbolic. Small fluctuations in the shape of the curve could determine whether an individual is willing to precommit. These differences could occur between individuals, in which case measuring precisely the shape of the discounting curve for an individual could help establish treatment patterns. The fluctuations could also occur within an individual over time (Mobini et al., <xref ref-type="bibr" rid="B37">2000</xref>; Schweighofer et al., <xref ref-type="bibr" rid="B49">2008</xref>), and may help to explain both spontaneous relapse and spontaneous recovery. Second, if humans and animals implement some form of distributed discounting, precommitment could be a very sensitive assay to determine exactly how discounting is being calculated.</p>
<p>Discount curves can be changed by context (Dixon et al., <xref ref-type="bibr" rid="B15">2006</xref>), diet (Schweighofer et al., <xref ref-type="bibr" rid="B49">2008</xref>), pharmacological state (de Wit et al., <xref ref-type="bibr" rid="B13">2002</xref>), availability of working memory capacity (Hinson et al., <xref ref-type="bibr" rid="B23">2003</xref>), and possibly cognitive training (Kendall and Wilcox, <xref ref-type="bibr" rid="B30">1980</xref>; Nelson and Behler, <xref ref-type="bibr" rid="B39">1989</xref>). For example, Schweighofer et al. (<xref ref-type="bibr" rid="B49">2008</xref>) showed that increasing dietary tryptophan slows discounting, and specifically increases task-related activation of parts of the striatum associated with slow discounting (Tanaka et al., <xref ref-type="bibr" rid="B54">2007</xref>). Such manipulations could potentially have a large impact on whether people decide to engage in precommitment strategies.</p>
</sec>
<sec>
<title>Neurobiology of precommitment</title>
<p>Reinforcement learning models have been used to describe the learning processes embodied in the basal ganglia (Doya, <xref ref-type="bibr" rid="B17">1999</xref>). During learning tasks, midbrain dopamine neurons fire in a pattern that closely matches the &#x003B4; signal of temporal difference reinforcement learning (Ljungberg et al., <xref ref-type="bibr" rid="B33">1992</xref>; Montague et al., <xref ref-type="bibr" rid="B38">1996</xref>; Hollerman and Schultz, <xref ref-type="bibr" rid="B24">1998</xref>). Both functional imaging and electrophysiological recording suggest that cached values are represented in striatum (Samejima et al., <xref ref-type="bibr" rid="B47">2005</xref>; Tobler et al., <xref ref-type="bibr" rid="B56">2007</xref>). The hypothesis that these brain structures implement reinforcement learning has helped to link a theoretical understanding of behavior with neurophysiological experiments.</p>
<p>If certain brain structures implement reinforcement learning, then reinforcement learning models may also be able to make predictions about the neurophysiology of precommitment. For example, our model implies that if precommitment is the preferred strategy, then selecting the choice option should produce a pause in dopamine firing (at the transition from P to C, the average &#x003B4; is <inline-formula><mml:math id="M30"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula>, which is negative because <inline-formula><mml:math id="M31"><mml:mrow><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0003C;</mml:mo><mml:mover accent='true'><mml:mi>V</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>There is also some evidence that the brain may implement distributed discounting. A sum of exponentials may in some cases be a statistically better fit to the time courses of human forgetting (Rubin and Wenzel, <xref ref-type="bibr" rid="B46">1996</xref>; Rubin et al., <xref ref-type="bibr" rid="B45">1999</xref>), suggesting the possibility of a distributed learning and memory process. During delay discounting tasks, there is a distribution across the striatum of areas correlated with different discount rates (Tanaka et al., <xref ref-type="bibr" rid="B53">2004</xref>), consistent with the theory of a distributed set of agents exponentially discounting in parallel. For a more complete discussion of the evidence for distributed discounting, see Kurth-Nelson and Redish (<xref ref-type="bibr" rid="B32">2009</xref>).</p>
<p>Some have argued that the brain implements a decision-making system in which each reward is assigned a single hyperbolically discounted subjective value (Kable and Glimcher, <xref ref-type="bibr" rid="B29">2009</xref>), as suggested by fMRI correlates of hyperbolically discounted subjective value (Kable and Glimcher, <xref ref-type="bibr" rid="B28">2007</xref>). However, fMRI is spatially and temporally averaged, which could blur an underlying distribution of exponentials to look like a hyperbolic representation. Even if the fMRI data do reflect an underlying hyperbolic representation, this does not disprove the existence of multiple exponentials elsewhere in the brain. A distribution of exponentials would need to be averaged before taking an action, producing a hyperbolic representation downstream.</p>
<p>In order to precommit, hyperbolic discounting must function across multiple state transitions, which in general is not possible in a non-distributed system that estimates values using only local information. If the brain represents single hyperbolically discounted values at each state, then the decision-making system must use some type of non-local information, such as multiple variables to track the time until each possible reward, or a look-ahead system to anticipate future rewards.</p>
<p>To our knowledge, precommitment of the form described here has not been empirically studied in humans. If precommitment is a product of reinforcement learning, implemented in the basal ganglia, then interfering with these brain structures should prevent precommitment. For example, Parkinson&#x00027;s patients would be impaired in learning precommitment strategies. On the other hand, brain structures such as frontal cortex that are not necessary for basic operant conditioning would not be necessary for precommitment.</p>
</sec>
<sec>
<title>Multiple systems and cognitive precommitment</title>
<p>In this paper we have presented a model of precommitment arising from hyperbolic discounting in an automated (habitual, model-free) learning system, using cached values to decide on actions without planning or cognitive involvement. An alternative possibility is that precommitment may be produced by a cognitive (look-ahead, model-based) system. The role for an interaction between automated and cognitive systems has been discussed extensively, especially in the context of impulsive choice and addiction (Tiffany, <xref ref-type="bibr" rid="B55">1990</xref>; Bickel et al., <xref ref-type="bibr" rid="B8">2007</xref>; Redish et al., <xref ref-type="bibr" rid="B43">2008</xref>; Gl&#x000E4;scher et al., <xref ref-type="bibr" rid="B22">2010</xref>). The cognitive system might recognize that the automated system will make a suboptimal choice and precommit to effectively override the automated system (&#x0201C;If I go to the bar, I will drink. Therefore I will not go to the bar.&#x0201D;) (Ainslie, <xref ref-type="bibr" rid="B4">2001</xref>; Bernheim and Rangel, <xref ref-type="bibr" rid="B7">2004</xref>; Isoda and Hikosaka, <xref ref-type="bibr" rid="B25">2007</xref>; Johnson et al., <xref ref-type="bibr" rid="B28">2007</xref>; Redish et al., <xref ref-type="bibr" rid="B43">2008</xref>). Rats appear to project themselves mentally into the future when making a difficult decision (Johnson and Redish, <xref ref-type="bibr" rid="B26">2007</xref>). In humans, vividly imagining a delayed outcome slows the discounting to that outcome (Peters and B&#x000FC;chel, <xref ref-type="bibr" rid="B41">2010</xref>). Constructing a cognitive representation of the future may allow an individual to carefully weigh the possible outcomes, even when those outcomes have never been experienced. The cognitive resources needed for this deliberation may be depleted by placing demands on working memory (Hinson et al., <xref ref-type="bibr" rid="B23">2003</xref>) or self-control (Vohs et al., <xref ref-type="bibr" rid="B58">2008</xref>).</p>
<p>Both automated and cognitive systems are likely to play a role in precommitment. To begin to identify the role of each system, we should look at the empirically distinguishable properties of precommitment produced by each system. First, if precommitment comes from an automated system, it should not be sensitive to attention or to cognitive load. A cognitive mechanism of precommitment would likely be sensitive to attention and cognitive load. Second, the automated system hypothesis predicts that precommitment would <italic>require learning</italic>; an individual must be repeatedly exposed to the outcomes and would not choose to precommit in novel situations. A cognitive model would predict that precommitment could occur in novel situations. Third, the automated model predicts that fast discounters are more likely than slow discounters to commit to long-range precommitment strategies. It seems likely that a cognitive model would predict the opposite. Fourth, the two hypotheses predict the involvement of different brain structures in precommitment. The automated hypothesis predicts that learning requires basal ganglia structures such as striatum and ventral tegmental area, while the cognitive hypothesis predicts the involvement of cortex and hippocampus. Testing these predictions should provide clues about the possible role of cognitive systems in precommitment.</p>
</sec>
</sec>
<sec>
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgment</title>
<p>This work was funded by National Institutes of Health Grant R01 DA024080.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ainslie</surname> <given-names>G.</given-names></name></person-group> (<year>1974</year>). <article-title>Impulse control in pigeons</article-title>. <source>J. Exp. Anal. Behav.</source> <volume>21</volume>, <fpage>485</fpage>&#x02013;<lpage>489</lpage>.<pub-id pub-id-type="doi">10.1901/jeab.1974.21-485</pub-id><pub-id pub-id-type="pmid">16811760</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ainslie</surname> <given-names>G.</given-names></name></person-group> (<year>1975</year>). <article-title>Specious reward: a behavioral theory of impulsiveness and impulse control</article-title>. <source>Psychol. Bull.</source> <volume>82</volume>, <fpage>463</fpage>&#x02013;<lpage>496</lpage>.<pub-id pub-id-type="doi">10.1037/h0076860</pub-id><pub-id pub-id-type="pmid">1099599</pub-id></citation></ref>
<ref id="B3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ainslie</surname> <given-names>G.</given-names></name></person-group> (<year>1992</year>). <source>Picoeconomics</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B4"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ainslie</surname> <given-names>G.</given-names></name></person-group> (<year>2001</year>). <source>Breakdown of Will</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alexander</surname> <given-names>W. H.</given-names></name> <name><surname>Brown</surname> <given-names>J. W.</given-names></name></person-group> (<year>2010</year>). <article-title>Hyperbolically discounted temporal difference learning</article-title>. <source>Neural. Comput.</source> <volume>22</volume>, <fpage>1511</fpage>&#x02013;<lpage>1527</lpage>.<pub-id pub-id-type="doi">10.1162/neco.2010.08-09-1080</pub-id><pub-id pub-id-type="pmid">20100071</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bellman</surname> <given-names>R.</given-names></name></person-group> (<year>1958</year>). <article-title>On a routing problem</article-title>. <source>Q. J. Appl. Math.</source> <volume>16</volume>, <fpage>87</fpage>&#x02013;<lpage>90</lpage>.</citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernheim</surname> <given-names>B. D.</given-names></name> <name><surname>Rangel</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Addiction and cue-triggered decision processes</article-title>. <source>Am. Econ. Rev.</source> <volume>94</volume>, <fpage>1558</fpage>&#x02013;<lpage>1590</lpage>.<pub-id pub-id-type="doi">10.1257/0002828043052222</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bickel</surname> <given-names>W. K.</given-names></name> <name><surname>Miller</surname> <given-names>M. L.</given-names></name> <name><surname>Yi</surname> <given-names>R.</given-names></name> <name><surname>Kowal</surname> <given-names>B. P.</given-names></name> <name><surname>Lindquist</surname> <given-names>D. M.</given-names></name> <name><surname>Pitcock</surname> <given-names>J. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Behavioral and neuroeconomics of drug addiction: competing neural systems and temporal discounting processes</article-title>. <source>Drug Alcohol Depend.</source> <volume>90</volume>, S85&#x02013;S91.<pub-id pub-id-type="doi">10.1016/j.drugalcdep.2006.09.016</pub-id><pub-id pub-id-type="pmid">17101239</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bickel</surname> <given-names>W. K.</given-names></name> <name><surname>Odum</surname> <given-names>A. L.</given-names></name> <name><surname>Madden</surname> <given-names>G. J.</given-names></name></person-group> (<year>1999</year>). <article-title>Impulsivity and cigarette smoking: delay discounting in current, never, and ex-smokers</article-title>. <source>Psychopharmacology (Berl.)</source> <volume>146</volume>, <fpage>447</fpage>&#x02013;<lpage>454</lpage>.<pub-id pub-id-type="doi">10.1007/PL00005490</pub-id><pub-id pub-id-type="pmid">10550495</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coffey</surname> <given-names>S. F.</given-names></name> <name><surname>Gudleski</surname> <given-names>G. D.</given-names></name> <name><surname>Saladin</surname> <given-names>M. E.</given-names></name> <name><surname>Brady</surname> <given-names>K. T.</given-names></name></person-group> (<year>2003</year>). <article-title>Impulsivity and rapid discounting of delayed hypothetical rewards in cocaine-dependent individuals</article-title>. <source>Exp. Clin. Psychopharmacol.</source> <volume>11</volume>, <fpage>18</fpage>&#x02013;<lpage>25</lpage>.<pub-id pub-id-type="doi">10.1037/1064-1297.11.1.18</pub-id><pub-id pub-id-type="pmid">12622340</pub-id></citation></ref>
<ref id="B11"><citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Daw</surname> <given-names>N. D.</given-names></name></person-group> (<year>2003</year>). <article-title>Reinforcement Learning Models of the Dopamine System and Their Behavioral Implications</article-title>. Ph.D. thesis, <publisher-name>Carnegie Mellon University</publisher-name>, <publisher-name>Pittsburgh</publisher-name>.</citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>Touretzky</surname> <given-names>D. S.</given-names></name></person-group> (<year>2000</year>). <article-title>Behavioral considerations suggest an average reward TD model of the dopamine system</article-title>. <source>Neurocomputing</source> <fpage>32</fpage>&#x02013;<lpage>33</lpage>, <fpage>679</fpage>&#x02013;<lpage>684</lpage>.<pub-id pub-id-type="doi">10.1016/S0925-2312(00)00232-0</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Wit</surname> <given-names>H.</given-names></name> <name><surname>Enggasser</surname> <given-names>J. L.</given-names></name> <name><surname>Richards</surname> <given-names>J. B.</given-names></name></person-group> (<year>2002</year>). <article-title>Acute administration of d-amphetamine decreases impulsivity in healthy volunteers</article-title>. <source>Neuropsychopharmacology</source> <volume>27</volume>, <fpage>813</fpage>&#x02013;<lpage>825</lpage>.<pub-id pub-id-type="doi">10.1016/S0893-133X(02)00343-3</pub-id><pub-id pub-id-type="pmid">12431855</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dezfouli</surname> <given-names>A.</given-names></name> <name><surname>Piray</surname> <given-names>P.</given-names></name> <name><surname>Keramati</surname> <given-names>M. M.</given-names></name> <name><surname>Ekhtiari</surname> <given-names>H.</given-names></name> <name><surname>Lucas</surname> <given-names>C.</given-names></name> <name><surname>Mokri</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Aneurocomputational model for cocaine addiction</article-title>. <source>Neural Comput.</source> <volume>21</volume>, <fpage>2869</fpage>&#x02013;<lpage>2893</lpage>.<pub-id pub-id-type="doi">10.1162/neco.2009.10-08-882</pub-id><pub-id pub-id-type="pmid">19635010</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dixon</surname> <given-names>M. R.</given-names></name> <name><surname>Jacobs</surname> <given-names>E. A.</given-names></name> <name><surname>Sanders</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>Contextual control of delay discounting by pathological gamblers</article-title>. <source>J. Appl. Behav. Anal.</source> <volume>39</volume>, <fpage>413</fpage>&#x02013;<lpage>422</lpage>.<pub-id pub-id-type="doi">10.1901/jaba.2006.173-05</pub-id><pub-id pub-id-type="pmid">17236338</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dom</surname> <given-names>G.</given-names></name> <name><surname>D&#x00027;haene</surname> <given-names>P.</given-names></name> <name><surname>Hulstijn</surname> <given-names>W.</given-names></name> <name><surname>Sabbe</surname> <given-names>B.</given-names></name></person-group> (<year>2006</year>). <article-title>Impulsivity in abstinent early- and late-onset alcoholics: differences in self-report measures and a discounting task</article-title>. <source>Addiction</source> <volume>101</volume>, <fpage>50</fpage>&#x02013;<lpage>59</lpage>.<pub-id pub-id-type="doi">10.1111/j.1360-0443.2005.01270.x</pub-id><pub-id pub-id-type="pmid">16393191</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doya</surname> <given-names>K.</given-names></name></person-group> (<year>1999</year>). <article-title>What are the computations of the cerebellum, the basal ganglia, and the cerebral cortex?</article-title> <source>Neural Netw.</source> <volume>12</volume>, <fpage>961</fpage>&#x02013;<lpage>974</lpage>.<pub-id pub-id-type="doi">10.1016/S0893-6080(99)00046-5</pub-id><pub-id pub-id-type="pmid">12662639</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dripps</surname> <given-names>D. A.</given-names></name></person-group> (<year>1993</year>). <article-title>Precommitment, prohibition, and the problem of dissent</article-title>. <source>J. Legal Stud.</source> <volume>22</volume>, <fpage>255</fpage>&#x02013;<lpage>263</lpage>.<pub-id pub-id-type="doi">10.1086/468165</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Evenden</surname> <given-names>J. L.</given-names></name></person-group> (<year>1999</year>). <article-title>Varieties of impulsivity</article-title>. <source>Psychopharmacology</source> <volume>146</volume>, <fpage>348</fpage>&#x02013;<lpage>361</lpage>.<pub-id pub-id-type="doi">10.1007/PL00005481</pub-id><pub-id pub-id-type="pmid">10550486</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fishburn</surname> <given-names>P. C.</given-names></name> <name><surname>Rubinstein</surname> <given-names>A.</given-names></name></person-group> (<year>1982</year>). <article-title>Time preference</article-title>. <source>Int. Econ. Rev.</source> <volume>23</volume>, <fpage>677</fpage>&#x02013;<lpage>694</lpage>.<pub-id pub-id-type="doi">10.2307/2526382</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frederick</surname> <given-names>S.</given-names></name> <name><surname>Loewenstein</surname> <given-names>G.</given-names></name> <name><surname>O&#x00027;Donoghue</surname> <given-names>T.</given-names></name></person-group> (<year>2002</year>). <article-title>Time discounting and time preference: a critical review</article-title>. <source>J. Econ. Lit.</source> <volume>40</volume>, <fpage>351</fpage>&#x02013;<lpage>401</lpage>.<pub-id pub-id-type="doi">10.1257/002205102320161311</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gl&#x000E4;scher</surname> <given-names>J.</given-names></name> <name><surname>Daw</surname> <given-names>N.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>O&#x00027;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2010</year>). <article-title>States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning</article-title>. <source>Neuron</source> <volume>66</volume>, <fpage>585</fpage>&#x02013;<lpage>595</lpage>.<pub-id pub-id-type="doi">10.1016/j.neuron.2010.04.016</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinson</surname> <given-names>J. M.</given-names></name> <name><surname>Jameson</surname> <given-names>T. L.</given-names></name> <name><surname>Whitney</surname> <given-names>P.</given-names></name></person-group> (<year>2003</year>). <article-title>Impulsive decision making and working memory</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>29</volume>, <fpage>298</fpage>&#x02013;<lpage>306</lpage>.<pub-id pub-id-type="doi">10.1037/0278-7393.29.2.298</pub-id><pub-id pub-id-type="pmid">12696817</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollerman</surname> <given-names>J. R.</given-names></name> <name><surname>Schultz</surname> <given-names>W.</given-names></name></person-group> (<year>1998</year>). <article-title>Dopamine neurons report an error in the temporal prediction of reward during learning</article-title>. <source>Nat.</source> <source>Neurosci.</source> <volume>1</volume>, <fpage>304</fpage>&#x02013;<lpage>309</lpage>.<pub-id pub-id-type="doi">10.1038/1124</pub-id><pub-id pub-id-type="pmid">10195164</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Isoda</surname> <given-names>M.</given-names></name> <name><surname>Hikosaka</surname> <given-names>O.</given-names></name></person-group> (<year>2007</year>). <article-title>Switching from automatic to controlled action by monkey medial frontal cortex</article-title>. <source>Nat. Neurosci.</source> <volume>10</volume>, <fpage>240</fpage>&#x02013;<lpage>248</lpage>.<pub-id pub-id-type="doi">10.1038/nn1830</pub-id><pub-id pub-id-type="pmid">17237780</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>A.</given-names></name> <name><surname>Redish</surname> <given-names>A. D.</given-names></name></person-group> (<year>2007</year>). <article-title>Neural ensembles in CA3 transiently encode paths forward of the animal at a decision point</article-title>. <source>J. Neurosci.</source> <volume>27</volume>, <fpage>12176</fpage>&#x02013;<lpage>12189</lpage>.<pub-id pub-id-type="doi">10.1523/JNEUROSCI.3761-07.2007</pub-id><pub-id pub-id-type="pmid">17989284</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>A.</given-names></name> <name><surname>van der Meer</surname> <given-names>M. A.</given-names></name> <name><surname>Redish</surname> <given-names>A. D.</given-names></name></person-group> (<year>2007</year>). <article-title>Integrating hippocampus and striatum in decision-making</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>17</volume>, <fpage>692</fpage>&#x02013;<lpage>697</lpage>.<pub-id pub-id-type="doi">10.1016/j.conb.2008.01.003</pub-id><pub-id pub-id-type="pmid">18313289</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kable</surname> <given-names>J. W.</given-names></name> <name><surname>Glimcher</surname> <given-names>P. W.</given-names></name></person-group> (<year>2007</year>). <article-title>The neural correlates of subjective value during intertemporal choice</article-title>. <source>Nat. Neurosci</source>. <volume>10</volume>, <fpage>1625</fpage>&#x02013;<lpage>1633</lpage>.<pub-id pub-id-type="doi">10.1038/nn2007</pub-id><pub-id pub-id-type="pmid">17982449</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kable</surname> <given-names>J. W.</given-names></name> <name><surname>Glimcher</surname> <given-names>P. W.</given-names></name></person-group> (<year>2009</year>). <article-title>The neurobiology of decision: consensus and controversy</article-title>. <source>Neuron</source> <volume>63</volume>, <fpage>733</fpage>&#x02013;<lpage>745</lpage>.<pub-id pub-id-type="doi">10.1016/j.neuron.2009.09.003</pub-id><pub-id pub-id-type="pmid">19778504</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kendall</surname> <given-names>P. C.</given-names></name> <name><surname>Wilcox</surname> <given-names>L. E.</given-names></name></person-group> (<year>1980</year>). <article-title>Cognitive-behavioral treatment for impulsivity: concrete versus conceptual training in non-self-controlled problem children</article-title>. <source>J. Consult. Clin. Psychol.</source> <volume>48</volume>, <fpage>80</fpage>&#x02013;<lpage>91</lpage>.<pub-id pub-id-type="doi">10.1037/0022-006X.48.1.80</pub-id><pub-id pub-id-type="pmid">7365047</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koopmans</surname> <given-names>T. C.</given-names></name></person-group> (<year>1960</year>). <article-title>Stationary ordinal utility and impatience</article-title>. <source>Econometrica</source> <volume>28</volume>, <fpage>287</fpage>&#x02013;<lpage>309</lpage>.<pub-id pub-id-type="doi">10.2307/1907722</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kurth-Nelson</surname> <given-names>Z.</given-names></name> <name><surname>Redish</surname> <given-names>A. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Temporal-difference reinforcement learning with distributed representations</article-title>. <source>PLoS One</source> <volume>4</volume>:<fpage>e7362</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0007362</pub-id><pub-id pub-id-type="pmid">19841749</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ljungberg</surname> <given-names>T.</given-names></name> <name><surname>Apicella</surname> <given-names>P.</given-names></name> <name><surname>Schultz</surname> <given-names>W.</given-names></name></person-group> (<year>1992</year>). <article-title>Responses of monkey dopamine neurons during learning of behavioral reactions</article-title>. <source>J. Neurophysiol.</source> <volume>67</volume>, <fpage>145</fpage>&#x02013;<lpage>163</lpage>.<pub-id pub-id-type="pmid">1552316</pub-id></citation></ref>
<ref id="B34"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Madden</surname> <given-names>G. J.</given-names></name> <name><surname>Bickel</surname> <given-names>W. K.</given-names></name></person-group> (<year>2010</year>). <source>Impulsivity: The Behavioral and Neurological Science of Discounting</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>American Psychological Association</publisher-name>.</citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Madden</surname> <given-names>G. J.</given-names></name> <name><surname>Petry</surname> <given-names>N. M.</given-names></name> <name><surname>Badger</surname> <given-names>G. J.</given-names></name> <name><surname>Bickford</surname> <given-names>W. K.</given-names></name></person-group> (<year>1997</year>). <article-title>Impulsive and self-control choices in opioid-dependent patients and non-drug-using control patients: drug and monetary rewards</article-title>. <source>Exp. Clin. Psychopharmacol.</source> <volume>5</volume>, <fpage>256</fpage>&#x02013;<lpage>262</lpage>.<pub-id pub-id-type="doi">10.1037/1064-1297.5.3.256</pub-id><pub-id pub-id-type="pmid">9260073</pub-id></citation></ref>
<ref id="B36"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mazur</surname> <given-names>J.</given-names></name></person-group> (<year>1987</year>). <article-title>&#x0201C;An adjusting procedure for studying delayed reinforcement,&#x0201D;</article-title> in <source>Quantitative Analysis of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value</source>, eds <person-group person-group-type="editor"><name><surname>Commons</surname> <given-names>M. L.</given-names></name> <name><surname>Mazur</surname> <given-names>J. F.</given-names></name> <name><surname>Nevin</surname> <given-names>J. A.</given-names></name> <name><surname>Rachlin</surname> <given-names>H.</given-names></name></person-group> (<publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Erlbaum</publisher-name>), <fpage>55</fpage>&#x02013;<lpage>73</lpage>.</citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mobini</surname> <given-names>S.</given-names></name> <name><surname>Chiang</surname> <given-names>T. J.</given-names></name> <name><surname>Al-Ruwaitea</surname> <given-names>A. S.</given-names></name> <name><surname>Ho</surname> <given-names>M. Y.</given-names></name> <name><surname>Bradshaw</surname> <given-names>C. M.</given-names></name> <name><surname>Szabadi</surname> <given-names>E.</given-names></name></person-group> (<year>2000</year>). <article-title>Effect of central 5-hydroxytryptamine depletion on inter-temporal choice: a quantitative analysis</article-title>. <source>Psychopharmacology</source> <volume>149</volume>, <fpage>313</fpage>&#x02013;<lpage>318</lpage>.<pub-id pub-id-type="doi">10.1007/s002130000385</pub-id><pub-id pub-id-type="pmid">10823413</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Montague</surname> <given-names>P. R.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>1996</year>). <article-title>A framework for mesencephalic dopamine systems based on predictive Hebbian learning</article-title>. <source>J. Neurosci.</source> <volume>16</volume>, <fpage>1936</fpage>&#x02013;<lpage>1947</lpage>.<pub-id pub-id-type="pmid">8774460</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nelson</surname> <given-names>W. M.</given-names></name> <name><surname>Behler</surname> <given-names>J. J.</given-names></name></person-group> (<year>1989</year>). <article-title>Cognitive impulsivity training: the effects of peer teaching</article-title>. <source>J. Behav. Ther. Exp. Psychiatry</source> <volume>20</volume>, <fpage>303</fpage>&#x02013;<lpage>309</lpage>.<pub-id pub-id-type="doi">10.1016/0005-7916(89)90061-X</pub-id><pub-id pub-id-type="pmid">2636234</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Donoghue</surname> <given-names>T.</given-names></name> <name><surname>Rabin</surname> <given-names>M.</given-names></name></person-group> (<year>1999</year>). <article-title>Doing it now or later</article-title>. <source>Am. Econ. Rev.</source> <volume>89</volume>, <fpage>103</fpage>&#x02013;<lpage>124</lpage>.<pub-id pub-id-type="doi">10.1257/aer.89.1.103</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peters</surname> <given-names>J.</given-names></name> <name><surname>B&#x000FC;chel</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Episodic future thinking reduces reward delay discounting through an enhancement of prefrontal-mediotemporal interactions</article-title>. <source>Neuron</source> <volume>66</volume>, <fpage>138</fpage>&#x02013;<lpage>148</lpage>.<pub-id pub-id-type="doi">10.1016/j.neuron.2010.03.026</pub-id><pub-id pub-id-type="pmid">20399735</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rachlin</surname> <given-names>H.</given-names></name> <name><surname>Green</surname> <given-names>L.</given-names></name></person-group> (<year>1972</year>). <article-title>Commitment, choice, and self-control</article-title>. <source>J. Exp. Anal. Behav.</source> <volume>17</volume>, <fpage>15</fpage>&#x02013;<lpage>22</lpage>.<pub-id pub-id-type="doi">10.1901/jeab.1972.17-15</pub-id><pub-id pub-id-type="pmid">16811561</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Redish</surname> <given-names>A. D.</given-names></name> <name><surname>Jensen</surname> <given-names>S.</given-names></name> <name><surname>Johnson</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>A unified framework for addiction: vulnerabilities in the decision process</article-title>. <source>Behav. Brain Sci.</source> <volume>31</volume>, <fpage>415</fpage>&#x02013;<lpage>437</lpage>.<pub-id pub-id-type="doi">10.1017/S0140525X0800472X</pub-id><pub-id pub-id-type="pmid">18662461</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reynolds</surname> <given-names>B.</given-names></name> <name><surname>Ortengren</surname> <given-names>A.</given-names></name> <name><surname>Richards</surname> <given-names>J. B.</given-names></name> <name><surname>Wit</surname> <given-names>H. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Dimensions of impulsive behavior: Personality and behavioral measures</article-title>. <source>Pers. Individ. Dif.</source> <volume>40</volume>, <fpage>305</fpage>&#x02013;<lpage>315</lpage>.<pub-id pub-id-type="doi">10.1016/j.paid.2005.03.024</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rubin</surname> <given-names>D. C.</given-names></name> <name><surname>Hinton</surname> <given-names>S.</given-names></name> <name><surname>Wenzel</surname> <given-names>A.</given-names></name></person-group> (<year>1999</year>). <article-title>The precise time course of retention</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>25</volume>, <fpage>1161</fpage>&#x02013;<lpage>1176</lpage>.<pub-id pub-id-type="doi">10.1037/0278-7393.25.5.1161</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rubin</surname> <given-names>D. C.</given-names></name> <name><surname>Wenzel</surname> <given-names>A. E.</given-names></name></person-group> (<year>1996</year>). <article-title>One hundred years of forgetting: a quantitative description of retention</article-title>. <source>Psychol. Rev.</source> <volume>103</volume>, <fpage>734</fpage>&#x02013;<lpage>760</lpage>.<pub-id pub-id-type="doi">10.1037/0033-295X.103.4.734</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samejima</surname> <given-names>K.</given-names></name> <name><surname>Ueda</surname> <given-names>Y.</given-names></name> <name><surname>Doya</surname> <given-names>K.</given-names></name> <name><surname>Kimura</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>Representation of action-specific reward values in the striatum</article-title>. <source>Science</source> <volume>310</volume>, <fpage>1337</fpage>&#x02013;<lpage>1340</lpage>.<pub-id pub-id-type="doi">10.1126/science.1115270</pub-id><pub-id pub-id-type="pmid">16311337</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samuelson</surname> <given-names>P. A.</given-names></name></person-group> (<year>1937</year>). <article-title>A note on measurement of utility</article-title>. <source>Rev. Econ. Stud.</source> <volume>4</volume>, <fpage>155</fpage>&#x02013;<lpage>161</lpage>.<pub-id pub-id-type="doi">10.2307/2967612</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schweighofer</surname> <given-names>N.</given-names></name> <name><surname>Bertin</surname> <given-names>M.</given-names></name> <name><surname>Shishida</surname> <given-names>K.</given-names></name> <name><surname>Okamoto</surname> <given-names>Y.</given-names></name> <name><surname>Tanaka</surname> <given-names>S. C.</given-names></name> <name><surname>Yamawaki</surname> <given-names>S.</given-names></name> <name><surname>Doya</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <article-title>Low-serotonin levels increase delayed reward discounting in humans</article-title>. <source>J. Neurosci.</source> <volume>28</volume>, <fpage>4528</fpage>&#x02013;<lpage>4532</lpage>.<pub-id pub-id-type="doi">10.1523/JNEUROSCI.4982-07.2008</pub-id><pub-id pub-id-type="pmid">18434531</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sozou</surname> <given-names>P. D.</given-names></name></person-group> (<year>1998</year>). <article-title>On hyperbolic discounting and uncertain hazard rates</article-title>. <source>Proc. Biol. Sci.</source> <volume>265</volume>, <fpage>2015</fpage>&#x02013;<lpage>2020</lpage>.<pub-id pub-id-type="doi">10.1098/rspb.1998.0534</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Strotz</surname> <given-names>R. H.</given-names></name></person-group> (<year>1955</year>). <article-title>Myopia and inconsistency in dynamic utility maximization</article-title>. <source>Rev. Econ. Stud.</source> <volume>23</volume>, <fpage>165</fpage>&#x02013;<lpage>180</lpage>.<pub-id pub-id-type="doi">10.2307/2295722</pub-id></citation></ref>
<ref id="B52"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source>Reinforcement Learning: An Introduction</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>S. C.</given-names></name> <name><surname>Doya</surname> <given-names>K.</given-names></name> <name><surname>Okada</surname> <given-names>G.</given-names></name> <name><surname>Ueda</surname> <given-names>K.</given-names></name> <name><surname>Okamoto</surname> <given-names>Y.</given-names></name> <name><surname>Yamawaki</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>Prediction of immediate and future rewards differentially recruits cortico-basal ganglia loops</article-title>. <source>Nat. Neurosci.</source> <volume>7</volume>, <fpage>887</fpage>&#x02013;<lpage>893</lpage>.<pub-id pub-id-type="doi">10.1038/nn1279</pub-id><pub-id pub-id-type="pmid">15235607</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>S. C.</given-names></name> <name><surname>Schweighofer</surname> <given-names>N.</given-names></name> <name><surname>Asahi</surname> <given-names>S.</given-names></name> <name><surname>Shishida</surname> <given-names>K.</given-names></name> <name><surname>Okamoto</surname> <given-names>Y.</given-names></name> <name><surname>Yamawaki</surname> <given-names>S.</given-names></name> <name><surname>Doya</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <article-title>Serotonin differentially regulates short- and long-term prediction of rewards in the ventral and dorsal striatum</article-title>. <source>PLoS One</source> <volume>2</volume>:<fpage>e1333</fpage>.<pub-id pub-id-type="doi">10.1371/journal.pone.0001333</pub-id><pub-id pub-id-type="pmid">18091999</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tiffany</surname> <given-names>S. T.</given-names></name></person-group> (<year>1990</year>). <article-title>A cognitive model of drug urges and drug-use behavior: role of automatic and nonautomatic processes</article-title>. <source>Psychol. Rev.</source> <volume>97</volume>, <fpage>147</fpage>&#x02013;<lpage>168</lpage>.<pub-id pub-id-type="doi">10.1037/0033-295X.97.2.147</pub-id><pub-id pub-id-type="pmid">2186423</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tobler</surname> <given-names>P. N.</given-names></name> <name><surname>O&#x00027;Doherty</surname> <given-names>J. P.</given-names></name> <name><surname>Dolan</surname> <given-names>R. J.</given-names></name> <name><surname>Schultz</surname> <given-names>W.</given-names></name></person-group> (<year>2007</year>). <article-title>Reward value coding distinct from risk attitude-related uncertainty coding in human reward systems</article-title>. <source>J. Neurophysiol</source>. <volume>97</volume>, <fpage>1621</fpage>&#x02013;<lpage>1632</lpage>.<pub-id pub-id-type="doi">10.1152/jn.00745.2006</pub-id><pub-id pub-id-type="pmid">17122317</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tsitsiklis</surname> <given-names>J. N.</given-names></name> <name><surname>Van Roy</surname> <given-names>B.</given-names></name></person-group> (<year>1999</year>). <article-title>Average cost temporal-difference learning</article-title>. <source>Automatica</source> <volume>35</volume>, <fpage>1799</fpage>&#x02013;<lpage>1808</lpage>.<pub-id pub-id-type="doi">10.1016/S0005-1098(99)00099-0</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vohs</surname> <given-names>K. D.</given-names></name> <name><surname>Baumeister</surname> <given-names>R. F.</given-names></name> <name><surname>Schmeichel</surname> <given-names>B. J.</given-names></name> <name><surname>Twenge</surname> <given-names>J. M.</given-names></name> <name><surname>Nelson</surname> <given-names>N. M.</given-names></name> <name><surname>Tice</surname> <given-names>D. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Making choices impairs subsequent self-control: a limited-resource account of decision making, self-regulation, and active initiative</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>94</volume>, <fpage>883</fpage>&#x02013;<lpage>898</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.94.5.883</pub-id><pub-id pub-id-type="pmid">18444745</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup>To be precise, this is the economic notion of impulsivity; impulsivity can also refer to making a decision without waiting for sufficient information, inability to stop a prepotent action, and other related phenomena. But these other phenomena entail different mechanisms and are dissociable from delay discounting (Evenden, <xref ref-type="bibr" rid="B19">1999</xref>; Reynolds et al., <xref ref-type="bibr" rid="B44">2006</xref>).</p></fn>
<fn id="fn2"><p><sup>2</sup>Note that in this paper, a state contains a delay that is preceded by value of that state and followed by the reward (if any) of the state; this sequence is slightly different from Kurth-Nelson and Redish (<xref ref-type="bibr" rid="B32">2009</xref>), where reward comes at the &#x0201C;beginning&#x0201D; of the state. This difference has no effect on the behavior of the model.</p></fn>
<fn id="fn3"><p><sup>3</sup>This is because the last C state transitioned to both the first SS state and the first LL state. In temporal difference learning, the value of a state with multiple transitions going out will converge to a weighted average of the discounted next states&#x02019; value plus reward. The average is weighted by the relative frequency of making each transition, which in this case is determined by the agent&#x00027;s choice. Thus, the model is &#x0201C;sophisticated&#x0201D; (O&#x00027;Donoghue and Rabin, <xref ref-type="bibr" rid="B40">1999</xref>): the knowledge that it will choose SS is encoded in the value of the C state.</p></fn>
</fn-group>
</back>
</article>
