<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2022.784604</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Reinforcement Learning Model With Dynamic State Space Tested on Target Search Tasks for Monkeys: Extension to Learning Task Events</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Sakamoto</surname> <given-names>Kazuhiro</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1219106/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yamada</surname> <given-names>Hinata</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Kawaguchi</surname> <given-names>Norihiko</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Furusawa</surname> <given-names>Yoshito</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Saito</surname> <given-names>Naohiro</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Mushiake</surname> <given-names>Hajime</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Neuroscience, Faculty of Medicine, Tohoku Medical and Pharmaceutical University</institution>, <addr-line>Sendai</addr-line>, <country>Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Physiology, Tohoku University School of Medicine</institution>, <addr-line>Sendai</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jun Ota, The University of Tokyo, Japan</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Tetsuro Funato, The University of Electro-Communications, Japan; Yuichi Kobayashi, Shizuoka University, Japan</p></fn>
<corresp id="c001">&#x002A;Correspondence: Kazuhiro Sakamoto, <email>sakamoto@tohoku-mpu.ac.jp</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>784604</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Sakamoto, Yamada, Kawaguchi, Furusawa, Saito and Mushiake.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Sakamoto, Yamada, Kawaguchi, Furusawa, Saito and Mushiake</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Learning is a crucial basis for biological systems to adapt to environments. Environments include various states or episodes, and episode-dependent learning is essential in adaptation to such complex situations. Here, we developed a model for learning a two-target search task used in primate physiological experiments. In the task, the agent is required to gaze one of the four presented light spots. Two neighboring spots are served as the correct target alternately, and the correct target pair is switched after a certain number of consecutive successes. In order for the agent to obtain rewards with a high probability, it is necessary to make decisions based on the actions and results of the previous two trials. Our previous work achieved this by using a dynamic state space. However, to learn a task that includes events such as fixation to the initial central spot, the model framework should be extended. For this purpose, here we propose a &#x201C;history-in-episode architecture.&#x201D; Specifically, we divide states into episodes and histories, and actions are selected based on the histories within each episode. When we compared the proposed model including the dynamic state space with the conventional SARSA method in the two-target search task, the former performed close to the theoretical optimum, while the latter never achieved target-pair switch because it had to re-learn each correct target each time. The reinforcement learning model including the proposed history-in-episode architecture and dynamic state scape enables episode-dependent learning and provides a basis for highly adaptable learning systems to complex environments.</p>
</abstract>
<kwd-group>
<kwd>reinforcement learning</kwd>
<kwd>target search task</kwd>
<kwd>dynamic state space</kwd>
<kwd>episode-dependent learning</kwd>
<kwd>history-in-episode architecture</kwd>
</kwd-group>
<contract-num rid="cn001">17K07060</contract-num>
<contract-num rid="cn001">20K07726</contract-num>
<contract-num rid="cn002">26120703</contract-num>
<contract-num rid="cn002">20H05478</contract-num>
<contract-num rid="cn002">22H04780</contract-num>
<contract-num rid="cn002">15H05879</contract-num>
<contract-sponsor id="cn001">Japan Society for the Promotion of Science<named-content content-type="fundref-id">10.13039/501100001691</named-content></contract-sponsor>
<contract-sponsor id="cn002">Ministry of Education, Culture, Sports, Science and Technology<named-content content-type="fundref-id">10.13039/501100001700</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="0"/>
<equation-count count="12"/>
<ref-count count="41"/>
<page-count count="12"/>
<word-count count="9088"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Learning is a fundamental process that is crucial for biological systems to adapt to the real world. Real environments have diverse states, and situation-dependent learning is indispensable to adapt successfully to such complexity. A good example of situation-dependent learning in humans is a baseball game: to win, the batter needs to bat according to the situation of the game and batting order, i.e., according to whether the previous batter got a hit and got on base. However, the batter also needs to consider of his own episode, that is, how he played against the pitcher last few times to predict what kind of ball the pitcher will throw next. An episode, which is also referred to as context in the field of neuroscience, is defined as a state (or framework) of the environment in which an agent gains experience and makes decisions or predictions (<xref ref-type="bibr" rid="B16">Maren et al., 2013</xref>; <xref ref-type="bibr" rid="B40">Yonelinas et al., 2019</xref>). Studies on episode-dependent learning provide a basis for understanding the high adaptability of living systems to real environments, and applying this to engineering.</p>
<p>The two-target search task used in our non-human primate neurophysiological experiments has advantages for building models that learn behaviors based on the sequence of episodes and history of each individual episode (<xref ref-type="bibr" rid="B13">Kawaguchi et al., 2013</xref>, <xref ref-type="bibr" rid="B14">2015</xref>). The episodes of one trial of the task are shown in <xref ref-type="fig" rid="F1">Figure 1A</xref> (i.e., the sequence of task events): the central fixation spot is presented, and the animal fixates on it (2nd episode); during fixation, four light spots appear around the fixation spot (3rd episode); the disappearance of the fixation spot is used as a go signal for gaze shift to one of the four spots. If the correct light spot is fixed on, a reward is given (4th episode). To be successful in the 4th episode, an action based on the history must be selected. In the task, two adjacent light points (the target pair) among the four should be alternately selected (<xref ref-type="fig" rid="F1">Figure 1B</xref>). However, after a certain number of consecutive correct responses (exploitation phase), the target pair is switched without an instruction signal, and the animal must identify a new target pair through trial and error (exploration phase). To achieve a high correct response rate in this task, action selection must be based on the history of actions and outcomes of the previous two trials.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The target search task. <bold>(A)</bold> The event sequence of the target search task. <bold>(B)</bold> A similar illustration of a valid-pair switch in the two-target search task. Green circle: correct target; arrow: choice; dashed line: valid pair. Note that the subjects were not instructed to move their eyes by the green spot before gaze shift.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g001.tif"/>
</fig>
<p>The first model of choice for learning action while inferring what cannot be directly observed, such as a target pair, would be a reinforcement learning model using a partially observable Markov decision process (POMDP; <xref ref-type="bibr" rid="B11">Jaakkola et al., 1995</xref>; <xref ref-type="bibr" rid="B39">Thrun et al., 2005</xref>). However, applied to a two-target search task, learning models using a POMDP have <italic>a priori</italic> knowledge of the target pairs. Models that require such knowledge will not be able to learn unassumed tasks, as our previous studies have shown (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>). Some models do not require prior knowledge and make decisions based on history, including models involving infinite hidden Markov processes, such as the hierarchical Dirichlet process (<xref ref-type="bibr" rid="B4">Beal et al., 2002</xref>; <xref ref-type="bibr" rid="B38">Teh et al., 2006</xref>; <xref ref-type="bibr" rid="B17">Mochihashi and Sumita, 2007</xref>; <xref ref-type="bibr" rid="B18">Mochihashi et al., 2009</xref>; <xref ref-type="bibr" rid="B22">Pfau et al., 2010</xref>; <xref ref-type="bibr" rid="B8">Doshi-Velez et al., 2015</xref>). However, models using such processes do not exhibit stable performance, because they generate many useless action-value functions due to a lack of criteria regarding the appropriateness of history length required for decision-making (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>).</p>
<p>The reinforcement learning model with a dynamic state space that we demonstrated in our previous study does not require prior knowledge of target pairs, and adheres to criteria regarding appropriate history length, and when that length should be increased for decision-making. The model showed high performance in a two-target search task, suggesting excellent generality (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>). However, in the model described above, one trial is equal to one-time step. Thus, it cannot learn appropriate behavior in a case involving a sequence of episodes (i.e., the task event sequence shown in <xref ref-type="fig" rid="F1">Figure 1A</xref>).</p>
<p>In this study, we developed a reinforcement learning model with a dynamic state space to enable episode-dependent learning. Specifically, we added a &#x201C;dynamic-state-within-episode,&#x201D; or &#x201C;history-in-episode,&#x201D; architecture to the model. The model architecture dynamically generates a memory set when encountering a novel episode, namely, a task event (<xref ref-type="fig" rid="F2">Figures 2A,B</xref>). Furthermore, the dynamic state space was used to generate a <italic>Q</italic>-table (action value function) for each episode (<xref ref-type="fig" rid="F2">Figure 2C</xref>), according to the aforementioned criteria for appropriateness of determining state expansion: the experience saturation and decision uniqueness of action selection. These two mechanisms enable episode-dependent learning in the two-target search task. That is, the model autonomously determines that the previous state in the relevant episode is the last two trials (we refer to this as the &#x201C;history&#x201D; in this paper), and can find the correct new target pair in a short time without significant re-learning, resulting in high performance comparable to that exhibited by monkeys. Such learning greatly contributes to our understanding of the high adaptability of living systems to complex real environments and could lead to engineering applications.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Schematic diagram of model operations as training progresses. <bold>(A)</bold> Schematic diagram of the model after the most elementary task, the fixation task, has been completed. Some states and corresponding <italic>Q</italic>-tables are generated by reflecting on the previous action and its outcome in each episode. <bold>(B)</bold> Schematic diagram of the model in the first trial in which the fixed one-target task was performed after the fixation task. Since the fixed one-target task includes task events 3 and 4 that the fixation task did not include, episode-dependent memory sets corresponding to task events 3 and 4 are newly generated. <bold>(C)</bold> Schematic diagram of the model while it is learning the two-target search task. The figure illustrates that the number of states in task event 4 is still increasing, while those in task events 1 to 3 have already stopped increasing. w.m.: working memory; arrow: choice; ob: observation; a: action.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g002.tif"/>
</fig>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2.SS1">
<title>Model Architecture</title>
<p>Our proposed model has two types of time steps/sequences (<xref ref-type="fig" rid="F2">Figure 2</xref>). The first type is the sequence of episodes, <italic>t</italic>, on which the changes of episode, <italic>E</italic><sub><italic>t</italic></sub>, depend. In this study, an episode is defined as a task event, specifically a display presented to an animal, rather than a sequence of events. Hence, episodes are explicit and directly observable, in the sense that the agent does not need to make any particular inferences. Temporally neighboring episodes interact when calculating reward prediction error (see below for details). The other type of step pertains to the history within an episode. Within this framework, the history at the <italic>N<sup>th</sup></italic> trial denotes the experience with the same task event, <italic>e</italic><sub><italic>i</italic></sub>, accrued over previous trials, and is represented as <italic>H</italic><sub><italic>N</italic></sub>(<italic>e</italic><sub><italic>i</italic></sub>). A given history, <italic>h</italic><sub><italic>j</italic></sub>, is a state composed of a sequence of action&#x2013;outcome pairs. Each history can include an arbitrary number of trials; however, they have a length of one trial when learning begins. Herein, we refer to this temporal structure of the model as history-in-episode architecture. The model generates a new set of memories, consisting of working memory and a dynamic <italic>Q</italic>-table (<xref ref-type="fig" rid="F2">Figure 2A</xref>; episode-dependent memory set), when a novel task event is encountered (<xref ref-type="fig" rid="F2">Figure 2B</xref>). Since our goal was to develop a learning model for a two-target search task consisting of a discrete sequence of events, the model has a simple mechanism to generate a memory set with probability 1 when a new display is exhibited. This history-in-episode architecture enables behaviors to be learned in each task period; this was not possible using the one trial-one time-step model in our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>), in which one trial had one time unit and only the fourth task period in <xref ref-type="fig" rid="F1">Figure 1A</xref> was considered.</p>
<p>Each episode-dependent memory set in the proposed model contained the same dynamic states as the one proposed in our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>). The basic structure of the episode-dependent memory set was grounded in the conventional temporal difference (TD) learning (<xref ref-type="bibr" rid="B37">Sutton and Barto, 1998</xref>). The action value function, <italic>Q</italic><sub><italic>N</italic></sub>(<italic>E</italic><sub><italic>t</italic></sub> = <italic>e</italic><sub><italic>i</italic></sub>, <italic>H</italic><sub><italic>N</italic></sub> = <italic>h</italic><sub><italic>j</italic></sub>, <italic>A<sub><italic>N</italic></sub> = a<sub><italic>k</italic></sub></italic>) for the set of a particular episode, <italic>e</italic><sub><italic>i</italic></sub>, history, <italic>h</italic><sub><italic>j</italic></sub>, and an action, <italic>a</italic><sub><italic>k</italic></sub>, at the <italic>N</italic>th trial were updated by the following equation:</p>
<disp-formula id="S2.E1">
<label>(1)</label>
<mml:math id="M1">
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>N</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x2190;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03B1;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x03B4;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03B1; is the learning rate, set to 0.1 in the range that showed desirable results revealed by the parameter search. <italic>&#x03B4;<sub><italic>t,N</italic></sub></italic> is the reward prediction error, given by</p>
<disp-formula id="S2.E2">
<label>(2)</label>
<mml:math id="M2">
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x03B4;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x2261;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mpadded>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03B3;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>t</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:msup>
<mml:mi>i</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:msup>
<mml:mi>j</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:msup>
<mml:mi>k</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>-</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>r</italic><sub><italic>t</italic></sub> is the reward delivered for <italic>A</italic><sub><italic>N</italic></sub> taken at <italic>E</italic><sub><italic>t</italic></sub> and <italic>H</italic><sub><italic>N</italic></sub> at time <italic>t</italic> in the <italic>N</italic>th trial, and the discount factor &#x03B3; was set to 0.7 decided empirically. If the correct spot was selected, a reward <italic>r</italic> = 1 was delivered, otherwise <italic>r</italic> = 0 was given. <italic>A</italic><sub><italic>t,N</italic></sub> was selected according to the stochastic function, <italic>P</italic><sup>&#x03C0;</sup>(<italic>A<sub><italic>t,N</italic></sub> = a<sub><italic>k</italic></sub></italic> | <italic>E<sub><italic>t</italic></sub> = e<sub><italic>i</italic></sub></italic>, <italic>H<sub><italic>N</italic></sub> = h<sub><italic>j</italic></sub></italic>), under the policy &#x03C0;. We used a softmax function for <italic>P</italic><sup>&#x03C0;</sup>, defined by</p>
<disp-formula id="S2.E3">
<label>(3)</label>
<mml:math id="M3">
<mml:mrow>
<mml:msup>
<mml:mi>P</mml:mi>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>|</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2261;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext>exp</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03B2;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mn>5</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mtext>exp</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03B2;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where the parameter &#x03B2;, termed the inverse-temperature, was set to 100 in the range that provided desirable results. 5 is the number of actions that the model can take. For action selection, the <italic>Q</italic>-table that refers to the longest history among generated <italic>Q</italic>-tables was used.</p>
<p>Our model was designed to avoid the need for stochastic decisions as much as possible. Specifically, when the model did not have a value function for a particular action that required a much larger value compared with others following extensive experience with the episode and history, it expanded the <italic>Q</italic>-table of the episode backward in sequence of trial (<xref ref-type="fig" rid="F2">Figure 2C</xref>). We illustrate the algorithm of this expansion in <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1A</xref>.</p>
<p>The initial <italic>Q-table</italic> was set as the one of a particular combination of the five possible actions, namely gazing at the right-up (RU), left-up (LU), left-down (LD), right-down (RD) spot, or center (C), which are represented by arrows and a black dot, and the outcome (correct or error), denoted by o and x in <xref ref-type="fig" rid="F3">Figure 3A</xref> and <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1B</xref>. The initial <italic>Q</italic>-value for each action was set to 0.5. The model monitored the stochastic mean policy for each episode <italic>e</italic><sub><italic>i</italic></sub> and history <italic>h</italic><sub><italic>j</italic></sub>, given by</p>
<disp-formula id="S2.E4">
<label>(4)</label>
<mml:math id="M4">
<mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x2261;</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mfrac>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>m</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:munderover>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Calculation of temporary difference learning by the proposed and control models. <bold>(A)</bold> A history-in-episode architecture using the dynamic state model (proposed method) as an example. The action value function is selected according to the history of each episode, and the reward prediction error is calculated. Fixed 5-, 10- and 10 by 10-state models also have a history-in-episode architecture, i.e., a <italic>Q</italic>-table generated for each episode. However, its size does not change dynamically. <bold>(B)</bold> The conventional SARSA model, which is the simplest control model. Since the previous actions are irrelevant in this model, the state value function <italic>V</italic> is used here. This model does not have a history-in-episode architecture, as only one state-value function can be assigned to each episode.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g003.tif"/>
</fig>
<p>where <italic>N<sub><italic>update</italic>, e<italic>i</italic>,hj</sub></italic> is the number of times that the <italic>Q</italic>-values for the episode <italic>e</italic><sub><italic>i</italic></sub> and history <italic>h</italic><sub><italic>j</italic></sub> were updated. Then, the information gain or the Kullback-Leibler divergence (KLD) obtained by updating the stochastic policy (step 1 in <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1A</xref>) is calculated:</p>
<disp-formula id="S2.Ex1">
<mml:math id="M5">
<mml:mrow>
<mml:mpadded>
<mml:mi>U</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex2">
<mml:math id="M6">
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.E5">
<label>(5)</label>
<mml:math id="M7">
<mml:mrow>
<mml:mi/>
<mml:mo>&#x2261;</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mn>5</mml:mn>
</mml:munderover>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>o</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>We referred to this as the Update_KLD. <italic>N<sub><italic>update</italic>,<italic>ei</italic>,hj</sub></italic> &#x2013; 1 indicates the number of trials since the model last encountered episode <italic>e</italic><sub><italic>i</italic></sub> and history <italic>h</italic><sub><italic>j</italic></sub> and calculated the mean <italic>P</italic><sup>&#x03C0;</sup>(<bold><italic>a</italic></bold>| <italic>e<sub><italic>i</italic></sub>, h<sub><italic>j</italic></sub></italic>).</p>
<p>Next, the model judged whether the Update_KLD of the episode <italic>e</italic><sub><italic>i</italic></sub> and history <italic>h</italic><sub><italic>j</italic></sub>, fell below the criterion for experience saturation, &#x03B6; (step 2),</p>
<disp-formula id="S2.E6">
<label>(6)</label>
<mml:math id="M8">
<mml:mrow>
<mml:mrow>
<mml:mi>U</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mmultiscripts>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:none/>
<mml:mi>j</mml:mi>
<mml:none/>
</mml:mmultiscripts>
</mml:msub>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">&#x2264;</mml:mo>
<mml:mi mathvariant="normal">&#x03B6;</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>indicating that information can no longer be gained by updating. The value of &#x03B6; was determined to be 10<sup>&#x2013;2</sup> in the range that showed desirable results. When the Update_KLD<italic><sub><italic>ei,hj</italic></sub></italic> was &#x003C; &#x03B6;, the distribution of <inline-formula><mml:math id="INEQ1"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> was compared with <inline-formula><mml:math id="INEQ2"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>. <inline-formula><mml:math id="INEQ3"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> is the action selection probability that only one action will be selected and was obtained as follows. First, the ideal policy, <italic>Q</italic><sub><italic>ideal</italic></sub>(<bold><italic>a</italic></bold>| <italic>e</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>), was obtained by setting the largest value within <italic>Q</italic>(<bold><italic>a</italic></bold>| <italic>e</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>) to 1 and the other values to zero. For example, if the <italic>Q</italic>(<bold><italic>a</italic></bold>| <italic>e</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>) were, {0.1, 0.4, 0.1, 0.2, 0.1}, the <italic>Q</italic><sub><italic>ideal</italic></sub>(<bold><italic>a</italic></bold>| <italic>e</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>), would be set to {0, 1, 0, 0, 0}.</p>
<p>Thereafter, the <inline-formula><mml:math id="INEQ4"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>e</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> was calculated from <italic>Q</italic><sub><italic>ideal</italic></sub>(<bold><italic>a</italic></bold>| <italic>e</italic><sub><italic>i</italic></sub>, <italic>h</italic><sub><italic>j</italic></sub>) using the softmax function in Eq. 3. For comparison, another KLD was calculated, as described below (step 3):</p>
<disp-formula id="S2.Ex3">
<mml:math id="M9">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mi>K</mml:mi>
<mml:mi>L</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">|</mml:mo>
<mml:mo stretchy="false">|</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo stretchy="false">|</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.E7">
<label>(7)</label>
<mml:math id="M10">
<mml:mrow>
<mml:mi/>
<mml:mo>&#x2261;</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>l</mml:mi>
</mml:munder>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>o</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">&#x03C0;</mml:mi>
</mml:msubsup>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>We called this the Decision-uniqueness KLD (D_KLD). When the D_KLD was below the criterion for a preference for deterministic action selection, &#x03B7; (step 4),</p>
<disp-formula id="S2.E8">
<label>(8)</label>
<mml:math id="M11">
<mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi/>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">&lt;</mml:mo>
<mml:mi mathvariant="normal">&#x03B7;</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>the agent had uniquely selected an action for the episode <italic>e</italic><sub><italic>i</italic></sub> and history <italic>h</italic><sub><italic>j</italic></sub>, and the <italic>Q</italic>-table was not expanded any further. &#x03B7; was set to 2 within the range which produced fair performance revealed by the parameter search. These two criteria, &#x03B6; and &#x03B7;, guaranteed the appropriateness of state (history) expansion: the former is for the appropriate timing of expansion; the latter is for whether the <italic>Q</italic>-table should be expanded or not (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>). When the D_KLD did not meet the criterion, it was also compared to the parent D_KLD (step 5), defined as the D_KLD of the parent history from which the current history <italic>h</italic><sub><italic>j</italic></sub> had been expanded (e.g., <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1B</xref>). In step 6, when the D_KLD is judged to be less than its corresponding parent D_KLD, as in Eq. 9,</p>
<disp-formula id="S2.E9">
<label>(9)</label>
<mml:math id="M12">
<mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mmultiscripts>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:none/>
<mml:mi>j</mml:mi>
<mml:none/>
</mml:mmultiscripts>
</mml:msub>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">&lt;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+5pt">
<mml:mi>t</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi mathvariant="normal">_</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mmultiscripts>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:none/>
<mml:mi>j</mml:mi>
<mml:none/>
</mml:mmultiscripts>
</mml:msub>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>the D_KLD value is saved as the parent D_KLD, and the history is expanded as depicted in the <italic>Q</italic>-table of <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1B</xref> (step 7). That is, the new history (child history) is the combination of the parent history and the history of one more previous trial to which the parent history refers. In the schematic example in <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1B</xref>, a new history is generated from one in which the agent looked at LD and was rewarded one trial ago; this is changed to one in which it looked at LD and was rewarded one trial ago after it looked at RD and was rewarded two trials ago. The initial <italic>Q</italic>-value for each action is set to 0.5. On the other hand, if Eq. 9 does not hold, the current history being processed (see flowchart in <xref ref-type="supplementary-material" rid="FS1">Supplementary Figure 1A</xref>) is pruned (step 6&#x2019;). When the current history consists of only the previous one trial, it is not erased because there is no parent history with which it could be compared. The bias is set to be &#x2212;1 in all calculation.</p>
<p>In the current study, we compared the proposed model, including the dynamic state space, to several models with fixed state-space using the two-target search task and related simpler tasks. However, these control models also generated a new episode-dependent memory set when they encountered a novel episode or task event. The models were classified depending on the type of fixed <italic>Q</italic>-table in the generated episode-dependent memory set. The fixed 10-state model had a <italic>Q</italic>-table of size 5 by 10 in each episode, meaning that it had five action choices in each of the 10 states (histories), which were the combinations of five actions and their outcomes in the previous trial. The fixed 10 by 10-state model had states consisting of the combinations of the actions and outcomes of the two previous trials, i.e., fixed 10 by 10 states (histories). The results for this model are not shown in the current study, but this model is the optimal model when created with prior knowledge of the task structure of the two-target search task. Our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>) showed its performance as a fixed 8 by 8-model. The fixed 5-state model obviously had five states for each episode, corresponding to the actions in the previous trial. In other words, this model did not explicitly include the result of the previous trial in the state. This model is an instrumental learning model, the results for which are omitted from the current study. The conventional SARSA model had only one value function (<italic>V</italic>-table, since the state was independent of the agent&#x2019;s action) for each episode, and selected one action among the five choices based on the <italic>V</italic>-table. Therefore, this model did not include &#x201C;history.&#x201D; That is, while the other models contained a history-in-episode architecture (<xref ref-type="fig" rid="F3">Figure 3A</xref>), the conventional SARSA model did not have that architecture (<xref ref-type="fig" rid="F3">Figure 3B</xref>). It should also be noted that the conventional SARSA model is a Pavlovian learning model in which each task event serves as a CS.</p>
</sec>
<sec id="S2.SS2">
<title>Behavioral Tasks and Simulation Framework</title>
<p>The target search task included the four task events &#x201C;trial start,&#x201D; &#x201C;fixation spot on,&#x201D; &#x201C;peripheral spots on,&#x201D; and &#x201C;go &#x0026; gaze shift&#x201D; (<xref ref-type="fig" rid="F1">Figure 1A</xref>). During the &#x201C;fixation spot on&#x201D; period, the agent was required to fixate on the central spot (C). In the subsequent &#x201C;peripheral spots on&#x201D; period, the agent was required to keep fixating on C without being distracted by the four spots presented around it: left-up (LU), right-up (RU), left-down (LD), and right-down (RD). When C disappeared at the beginning of the &#x201C;go &#x0026; gaze shift&#x201D; period, the agent shifted its focus to one of the four surrounding spots, and if it focused on the correct target spot, it was rewarded. Note that, in <xref ref-type="fig" rid="F1">Figure 1B</xref> and <xref ref-type="supplementary-material" rid="FS2">Supplementary Figure 2</xref>, the correct target is shown in green to help readers identify the currently correct target. In actual calculations, the agent only observe correct or error after gaze shift and cannot directly observe the true target. If the agent chose the wrong target spot, the trial was repeated under the same condition, i.e., the correct target stayed the same. The duration of each task period in the experiments with primates was 500 ms (<xref ref-type="bibr" rid="B13">Kawaguchi et al., 2013</xref>, <xref ref-type="bibr" rid="B14">2015</xref>). In our simulations, the time step for calculation was set to one task period for simplicity.</p>
<p>The one-target search task (<xref ref-type="supplementary-material" rid="FS2">Supplementary Figure 2</xref>) was easier than the two-target task, and was used as a pretraining task for monkeys. In this task, one out of four spots served as the correct target until the target was switched to another spot after seven successive successes without the provision of additional instructions. After the target switch, the subject was required to search for the new correct target.</p>
<p>In the two-target search task (<xref ref-type="fig" rid="F1">Figure 1B</xref>), two neighboring spots, referred to as a valid pair, were used as correct targets alternately. A valid pair was switched after seven consecutive successes without additional instructions, followed by an exploration phase for the new valid pair. Details are described elsewhere (<xref ref-type="bibr" rid="B13">Kawaguchi et al., 2013</xref>, <xref ref-type="bibr" rid="B14">2015</xref>).</p>
<p>We also tested fixed one- and two-target tasks, in which the correct target or valid pair was fixed throughout the simulations, respectively, to evaluate each learning model.</p>
</sec>
<sec id="S2.SS3">
<title>Animal Behavior</title>
<p>Our animal research was performed in accordance with National Institutes of Health guidelines and the guidelines of Tohoku University. All experimental protocols were approved by the Animal Care and Use Committee, Tohoku University (Permit No. ido-74). Two Japanese monkeys (<italic>Macaca fuscata</italic>; monkey K: 6.5 kg, monkey G: 6.1 kg) were trained to perform the two-target search task. The monkeys were kept in individual primate cages in an air-conditioned room with food available <italic>ad libitum</italic>. During the experiments, the monkeys sat in a primate chair with their heads restrained and faced a screen on which visual stimuli were presented. Eye position was monitored with an infrared corneal reflection system sampling at 250 Hz. Details are described elsewhere (<xref ref-type="bibr" rid="B13">Kawaguchi et al., 2013</xref>, <xref ref-type="bibr" rid="B14">2015</xref>).</p>
</sec>
</sec>
<sec id="S3" sec-type="results">
<title>Results</title>
<p>We tested the proposed dynamic state model using several behavioral tasks related to the two-target search task and compared it to other models with fixed sets of states or value functions. This comparison revealed fundamental differences between the compared models.</p>
<p>First, we tested all models using a fixed one-target search task with only one correct target spot during the entire simulation. All models exhibited almost perfect performance (<xref ref-type="fig" rid="F4">Figure 4A</xref>; data not shown for the fixed 10 by 10- and 5-state models. The same applies to the following results). However, it is noteworthy that the simplest model, i.e., the conventional SARSA model, learned the fastest.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Comparison of the performance on each task between the proposed and control models. <bold>(A)</bold> Evolution of the correct response rate in the fixed one-target task. <bold>(B)</bold> The fixed two-target task. <bold>(C)</bold> One-target search task. Dashed line: ideal performance. <bold>(D)</bold> Evolution of the number of target switches. <bold>(E)</bold> Two-target search task. Dashed line: ideal performance. <bold>(F)</bold> Number of valid-pair switches. All calculations started at the initial state.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g004.tif"/>
</fig>
<p><xref ref-type="fig" rid="F4">Figure 4B</xref> shows the results of the fixed two-target task. In this task, the correct valid target pair was not changed during the entire simulation, but two targets in the pair were the correct target alternately. This setup created additional difficulty since the correct strategy in the previous trial is not valid, and the models had to switch their behavior alternatively depending on the state, i.e., the history. Under these conditions, we expected the conventional SARSA model to exhibit poor performance because it was not able to make decisions based on the previous actions. As expected, all models except the conventional SARSA model showed almost perfect performance.</p>
<p>The one-target search task revealed additional differences between the tested models (<xref ref-type="fig" rid="F4">Figures 4C,D</xref>). This task required the agent to adapt to a switched correct target after every seven consecutive successes. This requirement forced the conventional SARSA model, as well as the fixed 5-state model (data not shown), to re-learn the correct target after each switch. As a result, they exhibited much lower correct response rates (<xref ref-type="fig" rid="F4">Figure 4C</xref>) and numbers of target switch (<xref ref-type="fig" rid="F4">Figure 4D</xref>) than the dynamic state, fixed 10- and 10 by 10-state models. These superior models, in contrast, learned how to explore in the exploration phase after a target switch, because the state, i.e., history, explicitly included the previous outcome as well as the action, which led to almost ideal performance (dashed line in <xref ref-type="fig" rid="F4">Figure 4C</xref>), although some delay in the increase in correct response rate was observed for the fixed 10-state model.</p>
<p>Finally, we tested all models on the two-target search task (<xref ref-type="fig" rid="F4">Figures 4E,F</xref>). As expected, our dynamic state model reproduced the results of our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>), and showing nearly ideal performance (dashed line in <xref ref-type="fig" rid="F4">Figure 4E</xref>) and a high number of pair switches (<xref ref-type="fig" rid="F4">Figure 4F</xref>); the same performance was obtained for the fixed 10 by 10-state model (data not shown), which was created as an ideal model with prior knowledge of the task structure. As for the fixed 10-state model, although it performed well for the one-target search task, its performance for the two-target search task was much worse than the ideal performance. This poor performance was expected because the model included only one previous trial in its history, while the ideal performance required inclusion of the two previous trials in its history. The fixed 5-state model showed similar performance to the fixed 10-state model. The conventional SARSA model exhibited a lower correct response rate than in the one-target search task and achieved no pair switch. Re-learning to focus on each of the spots of the valid pair never allowed the conventional SARSA model to achieve a pair switch.</p>
<p>The proposed model performed as well as a monkey in the two-target search task (<xref ref-type="fig" rid="F5">Figure 5</xref>). The monkey quickly located new valid pairs after valid-pair switches (<xref ref-type="fig" rid="F5">Figure 5A</xref>). Since valid pairs were switched without any explicit instruction, he inevitably gazed at the target of the previously valid pair in the first trial of the exploration phase (dark blue line in <xref ref-type="fig" rid="F5">Figure 5A</xref>), whereas he was highly likely to gaze at the new pair target after the first trial (red line in <xref ref-type="fig" rid="F5">Figure 5A</xref>). The rapid switching to the new pair displayed by the monkey was also seen in the proposed model (<xref ref-type="fig" rid="F5">Figure 5B</xref>). Furthermore, to examine the exploratory behaviors of the monkey and model in detail, gaze directions in the second trials of the exploration phase were analyzed, and we found that both the monkey (<xref ref-type="fig" rid="F5">Figure 5C</xref>) and model (<xref ref-type="fig" rid="F5">Figure 5D</xref>) were highly likely to gaze at the target diagonal to the one in the first trial (orange circles in <xref ref-type="fig" rid="F5">Figures 5C,D</xref>). These results indicate that the early detection of new pairs is achieved by sophisticated, non-random exploratory behavior.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Comparisons of model performance and the exploratory behavior of a monkey after valid-pair switches in the two-target search task. <bold>(A)</bold> Monkey G&#x2019;s exploratory behavior while recording the activity of a neuron (20509U1a2) after completing the training; there were 23 pair switches. <bold>(B)</bold> Exploration behavior of the model, with 774 pair switches. The gaze direction distributions of the monkey <bold>(C)</bold> and model <bold>(D)</bold> in the second trial, after switching the valid pairs.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g005.tif"/>
</fig>
<p>Previously, we showed that good performance can be achieved over a wide range of meta-parameter, i.e., the learning rate, inverse temperature, threshold of experience saturation, and threshold of decision uniqueness, through parameter search (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>). Here, we examine model performance while varying the discount factor Eq. 2, which was not included in our previous one trial-one time-step model (<xref ref-type="fig" rid="F6">Figure 6</xref>). When the high default value of 0.7 was reduced to 0.4, the model achieved a high correct response rate, although learning was relatively slow. However, when the default value was reduced further, the performance deteriorated rapidly (<xref ref-type="fig" rid="F6">Figure 6A</xref>). This deterioration was not due only to the selection of the correct target in task event 4, but also to the inability to maintain fixation in the preceding task events. When the discount factor was reduced, the fixation error rate in each task event, i.e., the percentage of trials in the task event of interest that had fixation errors relative to the total number of trials on which task performance was maintained up to that task event, increased. In addition, the error rate in task event 3 was lower than that in task event 2, which is remote from task event 4 (in which the reward is actually delivered; <xref ref-type="fig" rid="F6">Figure 6B</xref>). This implies that a high discount factor is required to learn a task involving a long sequence of events with a reward given only at the end of a trial.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>Effects of varying the discount factor on performance in the two-target search task. <bold>(A)</bold> Changes in the percentage of correct trials with learning. The incorrect response rates during task events 2 <bold>(B)</bold> The incorrect response rates during task events 2 and 3 when varying the discount factor.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g006.tif"/>
</fig>
<p>Executing the two-target search task with a high correct response rate requires making decisions based on the actions of the previous two trials and their outcomes. However, this is only true for the action selection during task event 4. Other task events require the agent to only fixate to the central spot. The dynamic state model learns to execute the task while increasing the states consisting of actions and their outcomes. However, when learning to focus on only one spot regardless of the previous actions, learning using a single state, i.e., Pavlovian learning, might not only be sufficient, but could even speed up learning.</p>
<p>To test this idea, we implemented a hybrid model in which we used a single <italic>Q</italic>-table for task events 1 to 3 and a dynamic <italic>Q</italic>-table for only task event 4 (<xref ref-type="fig" rid="F7">Figure 7A</xref>). <xref ref-type="fig" rid="F7">Figure 7B</xref> compares performances between the dynamic state and hybrid models on the two-target search task. Almost ideal correct response rates were obtained (dashed line in <xref ref-type="fig" rid="F7">Figure 7B</xref>); however, the performance of the hybrid model increased earlier than that of the dynamic state model. These results support our idea that, by minimizing the number of states when learning how to fixate on the center spot, the hybrid model speeds up its learning during the first three task events.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Configuration and performance of the hybrid model. <bold>(A)</bold> Calculation of reward prediction error in the hybrid model. Dynamic state space is given only to task event 4. <bold>(B)</bold> Performance comparison between the hybrid and full dynamic state models (<xref ref-type="fig" rid="F3">Figure 3A</xref>) in the two-target search task. All calculations started at the initial state.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-784604-g007.tif"/>
</fig>
<p>To further confirm this, we developed a parallel model in which Pavlovian, fixed 5-state, and dynamic state models were calculated in parallel for each episode and an action was selected based on the <italic>Q</italic>-table exhibiting the highest decision uniqueness among the three models. After executing 10,000 trials, we examined the model used in each task event and found that the dynamic model was used in task event 4, while the Pavlovian model was used in the other three task events. This result indicated that the most appropriate learning model changes depending on the task requirements.</p>
</sec>
<sec id="S4" sec-type="discussion">
<title>Discussion</title>
<p>In this study, we proposed a history-in-episode architecture to extend a reinforcement learning model, enabling episode-dependent learning. In addition, we built a model that also included the dynamic state space proposed in our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>), and tested its performance in a two-target search task. By having episode and history, the model was able to learn the appropriate action for each event in one trial based on the history of recent trials. The proposed model, which includes the dynamic state space and the history-in-episode architecture, is expected to be further developed and applied as a pioneering learning model with high adaptability to complex real environments, since it learns appropriate behaviors under various circumstances.</p>
<p>As shown in our previous paper (<xref ref-type="bibr" rid="B12">Katakura et al., 2022</xref>), the dynamic state model had a sufficient range of well-behaved meta-parameters for its intrinsic parameters, such as experience saturation and decision uniqueness, as well as conventional parameters such as learning rate and inverse temperature of the softmax function for action selection. This robustness was also true for the model with the history-in-episode architecture presented in the current study. Unlike the one trial-one time-step model in our previous paper, the model including the history-in-episode architecture uses TD learning to learn the task events. In TD learning, the discount factor is used as a coefficient that is multiplied by the reward prediction at the next time step in calculating the reward prediction error in Eq. 2. The model exhibited desirable performance in a sufficiently wide range of discount factors as shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. When the discount rate was too low, TD learning was unsuccessful and the model did not learn to take any action, specifically not during the earlier task events. The desirability of a high discount rate is also consistent with Go and Shogi models (<xref ref-type="bibr" rid="B35">Silver et al., 2016</xref>, <xref ref-type="bibr" rid="B36">2017</xref>), which learn behavior for long and complex orders of steps.</p>
<p>In recent years, machine learning and artificial intelligence (AI), as exemplified by learning models for Go and Shogi, have outperformed humans in some tasks (<xref ref-type="bibr" rid="B35">Silver et al., 2016</xref>, <xref ref-type="bibr" rid="B36">2017</xref>). However, it is questionable whether these models can be implemented in field robots working in real environments. Although the models can outperform humans in a single task, they lack some basic structures that are crucial for flexible learning in a real environment with complex situations and multiple goals. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, when multiple targets must be achieved (fixed two-target task), or when targets are frequently switched (one-target search task), the conventional SARSA or Pavlovian model or exhibited poor performance. In contrast, the fixed 10-state model with history-in-episode architecture achieved high performance for these two tasks, although state space was fixed. Even the fixed 5-state, i.e., the conventional instrumental learning model, which did not explicitly include the outcome of the previous trial, showed high performance in the fixed two-target task by choosing an action depending on the action in the fourth task period of the previous trial. Therefore, the history-in-episode architecture proposed in the current study provided a framework for achieving multiple goals, which has recently been a research issue (<xref ref-type="bibr" rid="B2">Bai et al., 2019</xref>; <xref ref-type="bibr" rid="B6">Colas et al., 2019</xref>; <xref ref-type="bibr" rid="B41">Zhao et al., 2019</xref>; <xref ref-type="bibr" rid="B24">Pitis et al., 2020</xref>; <xref ref-type="bibr" rid="B34">Shantia et al., 2021</xref>).</p>
<p>However, that does not mean that Pavlovian learning is always inferior. Our proposed model dynamically generated states and corresponding <italic>Q</italic>-tables based on combinations of actions and their outcomes. However, in task events 1 to 3 of the two-target search task, it is sufficient to simply learn to fixate, and having multiple states in each task event seems redundant. The hybrid model, which learns in a Pavlovian fashion in all task periods except the fourth using only a single <italic>V</italic>-table, learned the task faster than the full dynamic model (<xref ref-type="fig" rid="F7">Figure 7B</xref>). A similar observation is shown in <xref ref-type="fig" rid="F4">Figure 4A</xref>: for the fixed one-target task, the conventional SARSA model with Pavlovian learning during all task periods learned the task faster than the other models. These computational examples show that when states are redundant, the frequency with which each state is encountered decreases, resulting in slower learning. These arguments are related to the debate about whether Pavlovian or instrumental learning is better (<xref ref-type="bibr" rid="B25">Rescorla and Solomon, 1967</xref>), and how they can be used differently (<xref ref-type="bibr" rid="B5">Cartoni et al., 2016</xref>). <xref ref-type="bibr" rid="B7">Dorfman and Gershman (2019)</xref> developed a model in which either Pavlovian or instrumental conditioning predominated, depending on the degree to which an action can control the reward. We also generated a parallel model that included Pavlovian, instrumental, and dynamic state models, computed them in parallel, and let it select an action via the model exhibiting the highest decision uniqueness. We found that in the two-target search task, the Pavlovian model was used in task events 1 to 3, which are independent of the previous action. These observations suggest that learning models that are as simple as possible, i.e., having only the necessary states, are preferable. Choosing a resource-saving learning method according to the task requirements can avoid the curse of dimensionality problem in reinforcement learning (<xref ref-type="bibr" rid="B37">Sutton and Barto, 1998</xref>) and increase the learning speed.</p>
<p>If an action and its outcome are not uniquely predicted, it is desirable to increase the number of states so that the action and outcome can be uniquely expected by incorporating new clues. When presented with an ambiguous CS, i.e., when a US follows a CS in an episode or experimental condition but not in another condition, rats can uniquely predict the US by considering the information available under each condition, i.e., some clues in the environment or the configuration between them (<xref ref-type="bibr" rid="B9">Fanselow, 1990</xref>). The first brain region that contributes to such episode-dependent learning is the hippocampus. For example, hippocampal lesions in rodents produce deficits in freezing behavior during exposure to a shock-paired condition (<xref ref-type="bibr" rid="B33">Selden et al., 1991</xref>; <xref ref-type="bibr" rid="B15">Kim and Fanselow, 1992</xref>; <xref ref-type="bibr" rid="B23">Phillips and LeDoux, 1992</xref>). The structure and function of the hippocampus should be taken into account when developing our proposed model into one more in line with the structure of the real brain.</p>
<p>Some readers may find similarities between assigning a different <italic>Q-</italic>table to each episode in our model and learning sub-tasks in hierarchical reinforcement learning (HRL) models (<xref ref-type="bibr" rid="B3">Barto and Mahadevan, 2003</xref>; <xref ref-type="bibr" rid="B10">Hengst, 2010</xref>; <xref ref-type="bibr" rid="B1">Al-Emran, 2015</xref>; <xref ref-type="bibr" rid="B21">Pateria et al., 2021</xref>). However, since the two-target search task has temporally discrete task events, we need only generate a new episode-dependent memory set when a new task event is presented, and avoid the difficult problem of generating sub-tasks by deciding how to divide a continuous scene, which is one of the main issues for HRL. Moreover, our model is not hierarchical in the same sense of HRL. That is, our model does not include a supervisor that overlooks the units learning the sub-tasks, and gives them sub-goals. For these reasons, our model is not meant to be considered alongside or compared with HRL models. Rather, the proposed model includes two types of time steps i.e., the time step across different task events and the dynamic history in the episode of interest, and has a structure that generates memory sets or <italic>Q</italic>-tables as required in each time direction, especially in the case of history, where the state is generated dynamically to refer to multiple steps in the past. We consider these to be two novel points of the proposed model, and to be indispensable for learning the two-target search task. In our previous physiological studies, we observed neuronal activities in the lateral prefrontal cortex of monkeys that reflected sub-task generation (<xref ref-type="bibr" rid="B26">Saito et al., 2005</xref>; <xref ref-type="bibr" rid="B19">Mushiake et al., 2006</xref>; <xref ref-type="bibr" rid="B31">Sakamoto et al., 2008</xref>, <xref ref-type="bibr" rid="B28">2013</xref>, <xref ref-type="bibr" rid="B32">2020a</xref>). In the future, HRL will have to be considered when modeling those neural activities.</p>
<p>Other route to improvement of our proposed model is the involvement of finer time and space increments. To train a monkey to perform the two-target search task, it is necessary to start with the fixation task, in which the monkey is required to fixate on a single point in a continuous wide field of view and then complete simple tasks, such as the one-target search task used in the current study. In addition, it is necessary to gradually increase the length of each task period and gradually decrease the number of trials required to switch between targets or valid pairs in the pretraining trials. In contrast, in our computer simulations, the dynamic state model was able to learn the two-target search task without any pretraining. This is because we made the conditions of the simulations as simple as possible: one task period corresponded to one time step in the calculation and there were only five discrete choices of actions. In future, when the model becomes applicable to finer increments of time and space, the training of the model will require steps compatible to the training of monkeys.</p>
<p>However, it remains difficult to determine appropriate training steps. We have trained monkeys to perform many advanced behavioral tasks (<xref ref-type="bibr" rid="B20">Mushiake et al., 2001</xref>; <xref ref-type="bibr" rid="B31">Sakamoto et al., 2008</xref>, <xref ref-type="bibr" rid="B28">2013</xref>, <xref ref-type="bibr" rid="B30">2015</xref>, <xref ref-type="bibr" rid="B32">2020a</xref>,<xref ref-type="bibr" rid="B29">b</xref>) and obtained empirical knowledge regarding appropriate training steps. This knowledge is crucial for effective training. A major focus for future work will be to determine appropriate training steps for the model. We will aim to develop a &#x201C;coach&#x201D; that outputs task parameters, such as the length of the task period or complexity of the task, depending on the task conditions and learner&#x2019;s behavior, etc. The coach, which also needs to be equipped with a model involving a dynamic state space and a history-in-episode architecture, and the learner then start co-learning. In developing such a coach, the main challenge will be formulating the task complexity and generating a new training step depending on the progress of the learner. However, if we can generate such a coach model, it will likely be the prototype of a new type of AI that co-develops with humans and draws out our potential abilities rather than &#x201C;confronting&#x201D; us. Such a system could be referred to as hyper-adaptable, and we hope to use such systems to create a new discipline called neuro-coaching (<xref ref-type="bibr" rid="B27">Sakamoto, 2019</xref>).</p>
</sec>
<sec id="S5" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="S6">
<title>Ethics Statement</title>
<p>The animal study was reviewed and approved by the Animal Care and Use Committee, Tohoku University (Permit No. ido-74).</p>
</sec>
<sec id="S7">
<title>Author Contributions</title>
<p>KS designed the research, analyzed the data, wrote the first draft of the manuscript, edited the manuscript, and wrote the manuscript. KS and HY performed the research. NS and HM designed the two-target search task. YF obtained the behavioral data. KS and NK analyzed the behavioral data. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S8" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by the JSPS KAKENHI Grant Numbers 17K07060 and 20K07726 (Kiban C), MEXT KAKENHI Grant Number 15H05879 (Non-linear Neuro-oscillology), 26120703 (Prediction and Decision Making), and 20H05478 and 22H04780 (Hyper&#x2013;Adaptability).</p>
</sec>
<ack><p>We thank Y. Matsuzaka and Y. Nishimura of Tohoku Medical and Pharmaceutical University for advice and suggestions.</p>
</ack>
<sec id="S10" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fncom.2022.784604/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fncom.2022.784604/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Image_1.TIFF" id="FS1" mimetype="image/tiff" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 1</label>
<caption><p>Expansion and contraction of the history. <bold>(A)</bold> Flowchart of the expansion and contraction process. <bold>(B)</bold> An example of expansion of a history derived from the parent history in <italic>Q</italic>-table. The direction of the arrow represents the target that the agent looked at, and o and x represent the correct answer and error, respectively. The example in the figure shows that a new history is generated from the history that the agent looked at LD and was rewarded one trial ago, to the history that it looked at LD and was rewarded one trial ago after it looked at RD and was rewarded two trials ago. The numbers in the <italic>Q</italic>-table represent <italic>Q</italic>-values. The initial <italic>Q</italic>-value for each action is set to 0.5.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Image_2.TIFF" id="FS2" mimetype="image/tiff" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 2</label>
<caption><p>A schematic example of a target switch in the one-target search task. The format is the same as the task shown in <xref ref-type="fig" rid="F1">Figure 1B</xref>.</p></caption>
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Al-Emran</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Hierarchical reinforcement learning: a survey.</article-title> <source><italic>Int. J. Comput. Dig. Syst.</italic></source> <volume>4</volume>:<issue>2</issue>. <pub-id pub-id-type="doi">10.12785/IJCDS/040207</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bai</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>P.</given-names></name> <name><surname>Zhao</surname> <given-names>W.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Guided goal generation for hindsight multi-goal reinforcement learning.</article-title> <source><italic>Neurocomput</italic></source> <volume>359</volume> <fpage>353</fpage>&#x2013;<lpage>367</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.06.022</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barto</surname> <given-names>A. G.</given-names></name> <name><surname>Mahadevan</surname> <given-names>S.</given-names></name></person-group> (<year>2003</year>). <article-title>Recent advances in hierarchical reinforcement learning.</article-title> <source><italic>Discr. Event Dyn. Syst.</italic></source> <volume>13</volume> <fpage>41</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1023/A:1025696116075</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beal</surname> <given-names>M. J.</given-names></name> <name><surname>Ghahramani</surname> <given-names>Z.</given-names></name> <name><surname>Rasmussen</surname> <given-names>C.</given-names></name></person-group> (<year>2002</year>). <article-title>The infinite hidden Markov model.</article-title> <source><italic>Adv. Neural Inform. Proc. Sys.</italic></source> <volume>14</volume> <fpage>577</fpage>&#x2013;<lpage>584</lpage>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cartoni</surname> <given-names>E.</given-names></name> <name><surname>Balleine</surname> <given-names>B.</given-names></name> <name><surname>Baldassarre</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Appetitive Pavlovian-instrumental transfer: A review.</article-title> <source><italic>Neurosci. Biobehav. Rev.</italic></source> <volume>71</volume> <fpage>829</fpage>&#x2013;<lpage>848</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2016.09.020</pub-id> <pub-id pub-id-type="pmid">27693227</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colas</surname> <given-names>C.</given-names></name> <name><surname>Fournier</surname> <given-names>P.</given-names></name> <name><surname>Chetouani</surname> <given-names>M.</given-names></name> <name><surname>Sigaud</surname> <given-names>O.</given-names></name> <name><surname>Oudeyer</surname> <given-names>P.-Y.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>CURIOUS: Intrinsically motivated modular multi-goal reinforcement learning</article-title>&#x201D;. <source><italic>Proceedings of the 36th International Conference on Machine Learning</italic></source>, <publisher-loc>Long Beach</publisher-loc>: <publisher-name>ICML</publisher-name>. <volume>97</volume> <fpage>1331</fpage>&#x2013;<lpage>1340</lpage>.</citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dorfman</surname> <given-names>H. M.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Controllability governs the balance between Pavlovian and instrumental action selection.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>10</volume>:<issue>5826</issue>. <pub-id pub-id-type="doi">10.1038/s41467-019-13737-7</pub-id> <pub-id pub-id-type="pmid">31862876</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doshi-Velez</surname> <given-names>F.</given-names></name> <name><surname>Pfau</surname> <given-names>D.</given-names></name> <name><surname>Wood</surname> <given-names>F.</given-names></name> <name><surname>Roy</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Bayesian nonparametric methods for partially-observable reinforcement learning.</article-title> <source><italic>IEEE Trans. Patt. Anal. Mach. Intell.</italic></source> <volume>37</volume> <fpage>394</fpage>&#x2013;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2013.191</pub-id> <pub-id pub-id-type="pmid">26353250</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fanselow</surname> <given-names>M. S.</given-names></name></person-group> (<year>1990</year>). <article-title>Factors governing one trial contextual conditioning.</article-title> <source><italic>Anim. Learn. Behav.</italic></source> <volume>18</volume> <fpage>264</fpage>&#x2013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.3758/BF03205285</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hengst</surname> <given-names>B.</given-names></name></person-group> (<year>2010</year>). <source><italic>Hierarchical reinforcement learning. In Encyclopedia of Machine Learning.</italic></source> <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>495</fpage>&#x2013;<lpage>502</lpage>. <pub-id pub-id-type="doi">10.1007/978-0-387-30164-8_363</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaakkola</surname> <given-names>T.</given-names></name> <name><surname>Singh</surname> <given-names>S. P.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>1995</year>). <article-title>Reinforcement learning algorithm for partially observable Markov decision problems.</article-title> <source><italic>Adv. Neural Inf. Proc. Syst.</italic></source> <volume>7</volume> <fpage>345</fpage>&#x2013;<lpage>352</lpage>.</citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katakura</surname> <given-names>T.</given-names></name> <name><surname>Yoshida</surname> <given-names>M.</given-names></name> <name><surname>Hisano</surname> <given-names>H.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name></person-group> (<year>2022</year>). <article-title>Reinforcement learning model with dynamic state space tested on target search tasks for monkeys: Self-determination of previous states based on experience saturation and decision uniqueness.</article-title> <source><italic>Front. Comput. Neurosci.</italic></source> <volume>15</volume>:<issue>784592</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2021.784592</pub-id> <pub-id pub-id-type="pmid">35185502</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kawaguchi</surname> <given-names>N.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Furusawa</surname> <given-names>Y.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <article-title>Dynamic information processing in the frontal association areas of monkeys during hypothesis testing behavior.</article-title> <source><italic>Adv. Cogn. Neurodynam.</italic></source> <volume>4</volume> <fpage>691</fpage>&#x2013;<lpage>698</lpage>. <pub-id pub-id-type="doi">10.1007/978-94-007-4792-0_92</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kawaguchi</surname> <given-names>N.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Furusawa</surname> <given-names>Y.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name> <name><surname>Aoki</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Surprise signals in the supplementary eye field: rectified prediction errors drive exploration&#x2013;exploitation transitions.</article-title> <source><italic>J. Neurophysiol.</italic></source> <volume>113</volume> <fpage>1001</fpage>&#x2013;<lpage>1014</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00128.2014</pub-id> <pub-id pub-id-type="pmid">25411455</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Fanselow</surname> <given-names>M.</given-names></name></person-group> (<year>1992</year>). <article-title>Modality-specific retrograde amnesia of fear.</article-title> <source><italic>Science</italic></source> <volume>256</volume> <fpage>675</fpage>&#x2013;<lpage>677</lpage>. <pub-id pub-id-type="doi">10.1126/science.1585183</pub-id> <pub-id pub-id-type="pmid">1585183</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maren</surname> <given-names>S.</given-names></name> <name><surname>Phan</surname> <given-names>K. L.</given-names></name> <name><surname>Liberzon</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>The contextual brain: implication for fear conditioning, extinction and psychopathology.</article-title> <source><italic>Nat. Rev. Neurosci.</italic></source> <volume>14</volume> <fpage>417</fpage>&#x2013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3492</pub-id> <pub-id pub-id-type="pmid">23635870</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mochihashi</surname> <given-names>D.</given-names></name> <name><surname>Sumita</surname> <given-names>E.</given-names></name></person-group> (<year>2007</year>). <article-title>The infinite Markov model.</article-title> <source><italic>Adv. Neural Inform. Proc. Syst.</italic></source> <volume>20</volume> <fpage>1017</fpage>&#x2013;<lpage>1024</lpage>.</citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mochihashi</surname> <given-names>D.</given-names></name> <name><surname>Tamada</surname> <given-names>T.</given-names></name> <name><surname>Ueda</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>Bayesian unsupervised word segmentation with nested Pitman-Yor language modeling</article-title>,&#x201D; in <source><italic>Proc. 47th Annual Meeting ACL 4th IJCNLP AFNLP</italic></source>, <publisher-loc>Singapore</publisher-loc>. <fpage>100</fpage>&#x2013;<lpage>108</lpage>.</citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Itoyama</surname> <given-names>Y.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Activity in the lateral prefrontal cortex reflects multiple steps of future events in action plans.</article-title> <source><italic>Neuron</italic></source> <volume>50</volume> <fpage>631</fpage>&#x2013;<lpage>641</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2006.03.045</pub-id> <pub-id pub-id-type="pmid">16701212</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Sato</surname> <given-names>Y.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>Visually based path planning by Japanese monkeys.</article-title> <source><italic>Cogn. Brain Res.</italic></source> <volume>11</volume> <fpage>165</fpage>&#x2013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1016/S0926-6410(00)00067-7</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pateria</surname> <given-names>S.</given-names></name> <name><surname>Subagdja</surname> <given-names>B.</given-names></name> <name><surname>Tan</surname> <given-names>A.-H.</given-names></name> <name><surname>Quek</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>Hierarchical reinforcement learning: a comprehensive survey.</article-title> <source><italic>ACM Comput. Surv.</italic></source> <volume>54</volume>:<issue>109</issue>. <pub-id pub-id-type="doi">10.1145/3453160</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pfau</surname> <given-names>D.</given-names></name> <name><surname>Bartlett</surname> <given-names>N.</given-names></name> <name><surname>Wood</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>Probabilistic deterministic infinite automata.</article-title> <source><italic>Adv. Neural Inform. Proc. Syst.</italic></source> <volume>23</volume> <fpage>1930</fpage>&#x2013;<lpage>1938</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.1982.4767292</pub-id> <pub-id pub-id-type="pmid">21869067</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Phillips</surname> <given-names>R.</given-names></name> <name><surname>LeDoux</surname> <given-names>J.</given-names></name></person-group> (<year>1992</year>). <article-title>Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning.</article-title> <source><italic>Behav. Neurosci.</italic></source> <volume>106</volume> <fpage>274</fpage>&#x2013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1037/0735-7044.106.2.274</pub-id> <pub-id pub-id-type="pmid">1590953</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pitis</surname> <given-names>S.</given-names></name> <name><surname>Chan</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>S.</given-names></name> <name><surname>Stadie</surname> <given-names>B.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Maximum entropy gain exploration for long horizon multi-goal reinforcement learning</article-title>&#x201D;. <source><italic>Proceedings of the 37th International Conference on Machine Learning</italic></source>, <publisher-loc>Paris</publisher-loc>: <publisher-name>PMLR</publisher-name>. <volume>119</volume> <fpage>7750</fpage>&#x2013;<lpage>7761</lpage>.</citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rescorla</surname> <given-names>R. A.</given-names></name> <name><surname>Solomon</surname> <given-names>R. L.</given-names></name></person-group> (<year>1967</year>). <article-title>Two-process learning theory: relationships between Pavlovian conditioning and instrumental learning.</article-title> <source><italic>Psychol. Rev.</italic></source> <volume>74</volume> <fpage>151</fpage>&#x2013;<lpage>182</lpage>. <pub-id pub-id-type="doi">10.1037/h0024475</pub-id> <pub-id pub-id-type="pmid">5342881</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Itoyama</surname> <given-names>Y.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>Representation of immediate and final behavioral goals in the monkey prefrontal cortex during an instructed delay period.</article-title> <source><italic>Cereb. Cor.</italic></source> <volume>15</volume> <fpage>1535</fpage>&#x2013;<lpage>1546</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhi032</pub-id> <pub-id pub-id-type="pmid">15703260</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <source><italic>Brain science of creativity: beyond the complex systems theory of biological systems.</italic></source> <publisher-loc>Tokyo</publisher-loc>: <publisher-name>Univ Tokyo Press</publisher-name>.</citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Katori</surname> <given-names>Y.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Yoshida</surname> <given-names>S.</given-names></name> <name><surname>Aihara</surname> <given-names>K.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <article-title>Increased firing irregularity as an emergent property of neural-state transition in monkey prefrontal cortex.</article-title> <source><italic>PLoS One</italic></source> <volume>8</volume>:<issue>e80906</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0080906</pub-id> <pub-id pub-id-type="pmid">24349020</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Kawaguchi</surname> <given-names>N.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name></person-group> (<year>2020b</year>). <article-title>Differences in task-phase-dependent time-frequency patterns of local field potentials in the dorsal and ventral regions of the monkey lateral prefrontal cortex.</article-title> <source><italic>Neurosci. Res.</italic></source> <volume>156</volume> <fpage>41</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1016/j.neures.2019.12.016</pub-id> <pub-id pub-id-type="pmid">31923449</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Kawaguchi</surname> <given-names>N.</given-names></name> <name><surname>Yagi</surname> <given-names>K.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>Spatiotemporal patterns of current source density in the prefrontal cortex of a behaving monkey.</article-title> <source><italic>Neural Netw.</italic></source> <volume>62</volume> <fpage>67</fpage>&#x2013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2014.06.009</pub-id> <pub-id pub-id-type="pmid">25027732</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Aihara</surname> <given-names>K.</given-names></name> <name><surname>Yano</surname> <given-names>M.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Discharge synchrony during the transition of behavioral goal representations encoded by discharge rates of prefrontal neurons.</article-title> <source><italic>Cereb. Cor.</italic></source> <volume>18</volume> <fpage>2036</fpage>&#x2013;<lpage>2045</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhm234</pub-id> <pub-id pub-id-type="pmid">18252744</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Yoshida</surname> <given-names>S.</given-names></name> <name><surname>Mushiake</surname> <given-names>H.</given-names></name></person-group> (<year>2020a</year>). <article-title>Dynamic axis-tuned cells in the monkey lateral prefrontal cortex during a path-planning task.</article-title> <source><italic>J. Neurosci.</italic></source> <volume>40</volume> <fpage>203</fpage>&#x2013;<lpage>219</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2526-18.2019</pub-id> <pub-id pub-id-type="pmid">31719167</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Selden</surname> <given-names>N.</given-names></name> <name><surname>Everitt</surname> <given-names>B.</given-names></name> <name><surname>Jarrard</surname> <given-names>L.</given-names></name> <name><surname>Robbins</surname> <given-names>T.</given-names></name></person-group> (<year>1991</year>). <article-title>Complementary roles for the amygdala and hippocampus in aversive conditioning to explicit and contextual cues.</article-title> <source><italic>Neurosci</italic></source> <volume>42</volume> <fpage>335</fpage>&#x2013;<lpage>350</lpage>. <pub-id pub-id-type="doi">10.1016/0306-4522(91)90379-3</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shantia</surname> <given-names>A.</given-names></name> <name><surname>Timmers</surname> <given-names>R.</given-names></name> <name><surname>Chong</surname> <given-names>Y.</given-names></name> <name><surname>Kuiper</surname> <given-names>C.</given-names></name> <name><surname>Bidoia</surname> <given-names>F.</given-names></name> <name><surname>Schomaker</surname> <given-names>L.</given-names></name><etal/></person-group> (<year>2021</year>). <article-title>Two-stage visual navigation by deep neural networks and multi-goal reinforcement learning.</article-title> <source><italic>Robot. Autonom. Syst.</italic></source> <volume>138</volume>:<issue>103731</issue>. <pub-id pub-id-type="doi">10.1016/j.robot.2021.103731</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>A.</given-names></name> <name><surname>Maddison</surname> <given-names>C. J.</given-names></name> <name><surname>Guez</surname> <given-names>A.</given-names></name> <name><surname>Sifre</surname> <given-names>L.</given-names></name> <name><surname>van den Driessche</surname> <given-names>G.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Mastering the game of Go with deep neural networks and tree search.</article-title> <source><italic>Nature</italic></source> <volume>529</volume> <fpage>484</fpage>&#x2013;<lpage>489</lpage>. <pub-id pub-id-type="doi">10.1038/nature16961</pub-id> <pub-id pub-id-type="pmid">26819042</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Schrittwieser</surname> <given-names>J.</given-names></name> <name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Antonoglou</surname> <given-names>I.</given-names></name> <name><surname>Huang</surname> <given-names>A.</given-names></name> <name><surname>Guez</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2017</year>). <article-title>Mastering the game of Go without human knowledge.</article-title> <source><italic>Nature</italic></source> <volume>550</volume> <fpage>354</fpage>&#x2013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1038/nature24270</pub-id> <pub-id pub-id-type="pmid">29052630</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source><italic>Reinforcement learning: An introduction.</italic></source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teh</surname> <given-names>Y. W.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name> <name><surname>Beal</surname> <given-names>M. J.</given-names></name> <name><surname>Blei</surname> <given-names>D. M.</given-names></name></person-group> (<year>2006</year>). <article-title>Hierarchical Dirichlet processes.</article-title> <source><italic>J. Amer. Statist. Assoc.</italic></source> <volume>101</volume> <fpage>1566</fpage>&#x2013;<lpage>1581</lpage>. <pub-id pub-id-type="doi">10.1198/016214506000000302</pub-id> <pub-id pub-id-type="pmid">12611515</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thrun</surname> <given-names>S.</given-names></name> <name><surname>Burgard</surname> <given-names>W.</given-names></name> <name><surname>Fox</surname> <given-names>D.</given-names></name></person-group> (<year>2005</year>). <source><italic>Probabilistic Robotics.</italic></source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yonelinas</surname> <given-names>A. P.</given-names></name> <name><surname>Ranganath</surname> <given-names>C.</given-names></name> <name><surname>Ekstrom</surname> <given-names>A. D.</given-names></name> <name><surname>Wiltgen</surname> <given-names>B. J.</given-names></name></person-group> (<year>2019</year>). <article-title>A contextual binding theory of episodic memory: systems consolidation reconsidered.</article-title> <source><italic>Nat. Rev. Neurosci</italic>.</source> <volume>20</volume> <fpage>364</fpage>&#x2013;<lpage>375</lpage>. <pub-id pub-id-type="doi">10.1038/s41583-019-0150-4</pub-id> <pub-id pub-id-type="pmid">30872808</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Tresp</surname> <given-names>V.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Maximum entropy-regularized multi-goal reinforcement learning</article-title>&#x201D;. <source><italic>Proceedings of the 36th International Conference on Machine Learning.</italic></source> <publisher-loc>Long Beach</publisher-loc>. <volume>97</volume> <fpage>7553</fpage>&#x2013;<lpage>7562</lpage>.</citation></ref>
</ref-list>
</back>
</article>