<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2024.1397860</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Intrinsic motivation in cognitive architecture: intellectual curiosity originated from pattern discovery</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Nagashima</surname> <given-names>Kazuma</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2337111/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Morita</surname> <given-names>Junya</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/880917/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Takeuchi</surname> <given-names>Yugo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2514302/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Information Science and Technology, Graduate School of Science and Technology, Shizuoka University</institution>, <addr-line>Hamamatsu</addr-line>, <country>Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Behavior Informatics, Faculty of Informatics, Shizuoka University</institution>, <addr-line>Hamamatsu</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Simone Belli, Complutense University of Madrid, Spain</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Raymond Lee, Beijing Normal University-Hong Kong Baptist University United International College, China</p>
<p>Christian Lebiere, Carnegie Mellon University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Kazuma Nagashima <email>nagashima.kazuma.16&#x00040;shizuoka.ac.jp</email></corresp>
<corresp id="c002">Junya Morita <email>j-morita&#x00040;inf.shizuoka.ac.jp</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>10</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>7</volume>
<elocation-id>1397860</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>03</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>09</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Nagashima, Morita and Takeuchi.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Nagashima, Morita and Takeuchi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Studies on reinforcement learning have developed the representation of curiosity, which is a type of intrinsic motivation that leads to high performance in a certain type of tasks. However, these studies have not thoroughly examined the internal cognitive mechanisms leading to this performance. In contrast to this previous framework, we propose a mechanism of intrinsic motivation focused on pattern discovery from the perspective of human cognition. This study deals with intellectual curiosity as a type of intrinsic motivation, which finds novel compressible patterns in the data. We represented the process of continuation and boredom of tasks driven by intellectual curiosity using &#x0201C;pattern matching,&#x0201D; &#x0201C;utility,&#x0201D; and &#x0201C;production compilation,&#x0201D; which are general functions of the adaptive control of thought-rational (ACT-R) architecture. We implemented three ACT-R models with different levels of thinking to navigate multiple mazes of different sizes in simulations, manipulating the intensity of intellectual curiosity. The results indicate that intellectual curiosity negatively affects task completion rates in models with lower levels of thinking, while positively impacting models with higher levels of thinking. In addition, comparisons with a model developed by a conventional framework of reinforcement learning (intrinsic curiosity module: ICM) indicate the advantage of representing the agent&#x00027;s intention toward a goal in the proposed mechanism. In summary, the reported models, developed using functions linked to a general cognitive architecture, can contribute to our understanding of intrinsic motivation within the broader context of human innovation driven by pattern discovery.</p></abstract>
<kwd-group>
<kwd>cognitive modeling</kwd>
<kwd>ACT-R</kwd>
<kwd>intrinsic motivation</kwd>
<kwd>intellectual curiosity</kwd>
<kwd>pattern discovery</kwd>
</kwd-group>
<contract-sponsor id="cn001">Japan Society for the Promotion of Science<named-content content-type="fundref-id">10.13039/501100001691</named-content></contract-sponsor>
<counts>
<fig-count count="10"/>
<table-count count="0"/>
<equation-count count="7"/>
<ref-count count="60"/>
<page-count count="17"/>
<word-count count="11814"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>AI for Human Learning and Behavior Change</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>According to Baron-Cohen (<xref ref-type="bibr" rid="B8">2020</xref>), human evolution and the development of civilization are associated with &#x0201C;systematizing mechanisms,&#x0201D; which are achieved by discovering, combining, and using patterns of cause-and-effect relationships in an environment. He also stated that the ability of humans to think systematically has evolved by using the &#x0201C;if-and-then&#x0201D; logic to combine patterns, resulting in inventions and innovations that lead to our modern society.</p>
<p>Several studies have reported that such an ability of pattern discovery is associated with fun, a personal feeling leading to intrinsic motivation (Caillois, <xref ref-type="bibr" rid="B15">1958</xref>; Csikszentmihalyi, <xref ref-type="bibr" rid="B19">1990</xref>; Huizinga, <xref ref-type="bibr" rid="B27">1939</xref>; Koster, <xref ref-type="bibr" rid="B31">2013</xref>). The other researchers (Aubret et al., <xref ref-type="bibr" rid="B6">2019</xref>; Schmidhuber, <xref ref-type="bibr" rid="B49">2010</xref>) have also explored the computational realization of intrinsic motivation employing the framework of reinforcement learning (Sutton and Barto, <xref ref-type="bibr" rid="B53">1998</xref>). However, these studies have not explored the link between intrinsic motivation and primitive cognitive functions related to pattern discovery. Therefore, further analysis of the computational mechanisms of intrinsic motivation in terms of agents&#x00027; internal processing is needed.</p>
<p>The aforementioned problem can be addressed by using a cognitive architecture that integrates the basic cognitive functions involved in various tasks. Despite the existence of several cognitive architectures, the architectural differences have been reduced over the years and integrated into a common structure (Laird et al., <xref ref-type="bibr" rid="B33">2017</xref>). The representative architecture adopting such a structure is adaptive control of thought-rational (ACT-R), developed by Anderson (<xref ref-type="bibr" rid="B2">2007</xref>). According to Kotseruba and Tsotsos (<xref ref-type="bibr" rid="B32">2020</xref>)&#x00027;s comprehensive review of the topic, ACT-R is one of the most widely used cognitive architectures, including a greater number of features compared to the other architectures.</p>
<p>In this study, we propose a mechanism of intrinsic motivation based on pattern discovery by integrating primitive cognitive functions of ACT-R. The main advantage of the proposed approach is its interpretability. Based on commonly used building blocks in the architecture, our proposed mechanism can provide a foundation for understanding intrinsic motivation from the perspective of human cognition. Furthermore, this study presents a simulation experiment to explore the conditions of stimulating intrinsic motivation and the learning process driven by stimulated intellectual curiosity. Our analysis confirmed that the proposed mechanism can represent the role of intellectual motivation in human learning at diverse levels of thinking and task difficulty. Additionally, we examined the relationship between the proposed mechanism and an existing mechanism of intrinsic motivation based on reinforcement learning.</p>
<p>The remainder of this paper is organized as follows. Section 2 summarizes the existing studies related to this concept. Section 3 introduces the proposed mechanism, which is developed based on pattern discovery. The effectiveness of the mechanism is discussed based on simulations in Section 4. Finally, Section 5 summarizes the findings and indicates directions for future investigations.</p>
</sec>
<sec id="s2">
<title>2 Related works</title>
<p>The objective of this study is to represent a mechanism of intrinsic motivation based on the discovery of patterns. This section focuses on three directions of previous studies, namely, human curiosity, machine curiosity, and cognitive models with cognitive architectures.</p>
<sec>
<title>2.1 Human curiosity</title>
<p>Numerous studies have attempted to systematize intrinsic motivation as a driving factor to continue activities in a wide range of fields, including education, entertainment, healthcare, sports, and work. For instance, Malone (<xref ref-type="bibr" rid="B36">1981</xref>), who tried to systematize this concept in entertainment fields, categorized intrinsic motivation into three types, namely, &#x0201C;challenge,&#x0201D; which originates from goals of appropriate difficulty; &#x0201C;fantasy,&#x0201D; which leads to the imagination of unrealistic experiences; and &#x0201C;curiosity,&#x0201D; which is stimulated by a surprising, interesting, or fun activity. Here, curiosity is related to the discussion presented in Section 1 that pattern discovery accompanying the feeling of fun has led to human innovations. However, we believe that the first type of intrinsic motivation, challenge, is inseparable from curiosity. Rather than treating those as independent factors, we assume that curiosity is a mechanism of intrinsic motivation, stimulated by the appropriate difficulty (challenge) of a task.</p>
<p>The above assumption is supported by several authors who reported the relationship between the levels of task difficulty, the preferred level of thinking, and intrinsic motivation. The theory behind this is referred to as the optimal level of intrinsic motivation (Csikszentmihalyi, <xref ref-type="bibr" rid="B19">1990</xref>; Yerkes and Dodson, <xref ref-type="bibr" rid="B60">1908</xref>). According to this theory, intrinsic motivation is effectively stimulated when the task difficulty level matches the preferred level of thinking of a person. Furthermore, the level of thinking can be located on an axis with at least two levels. These include a shallow automatic level without careful thinking (fast process) and a deep deliberative level that requires time to carefully think (slow process) (Brooks, <xref ref-type="bibr" rid="B12">1986</xref>; Evans, <xref ref-type="bibr" rid="B22">2003</xref>; Kahneman, <xref ref-type="bibr" rid="B30">2011</xref>).</p>
</sec>
<sec>
<title>2.2 Machine curiosity</title>
<p>Based on the aforementioned discussion, we assumed a close relationship between curiosity and the feeling of fun involved in the discovery of patterns. This relationship was computationally theorized by Schmidhuber (<xref ref-type="bibr" rid="B49">2010</xref>), wherein the discovery of patterns is defined as identifying and compressing recurring canonical patterns in data. Schmidhuber also related compressing data or obtaining compressible data to fun by assuming that the agent aims to maximize fun as a reward. This idea was based on the prediction error theory (Friston, <xref ref-type="bibr" rid="B23">2010</xref>), which considers curiosity to be caused by the difference between prior predictions and the current situation (Bayesian surprise). In Schmidhuber&#x00027;s theory, prediction implies applying already compressed data; here, surprise occurs when identifying a pattern that can be newly compressed.</p>
<p>Schmidhuber&#x00027;s proposal can be discussed as an extension of conventional reinforcement learning. Typically, agents in reinforcement learning receive rewards from the environment and intend to maximize them over time. Sutton and Barto (<xref ref-type="bibr" rid="B53">1998</xref>) distinguished the boundaries between the agent and the environment from the physical boundaries between the body and the environment. Based on this idea, Singh et al. (<xref ref-type="bibr" rid="B50">2005</xref>) proposed a framework of intrinsically motivated reinforcement learning (IMRL), which divides the environment into external and internal segments. In contrast to conventional reinforcement learning, which directly receives a reward from the external environment, rewards in IMRL are determined depending on the state of the internal environment, such as stimulating curiosity for an unexpected response.</p>
<p>Since the proposal of IMRL, the framework of reinforcement learning has significantly progressed by integrating deep learning techniques. The preliminary framework was referred to as Deep Q Network (DQN) (Mnih et al., <xref ref-type="bibr" rid="B38">2015</xref>), which combined Q learning with a convolutional neural network (CNN). Subsequently, several researchers introduced the concept of intrinsic motivation in deep reinforcement learning. For instance, Bellemare et al. (<xref ref-type="bibr" rid="B10">2016</xref>) developed count-based exploration methods, wherein visit counts were used to guide an agent&#x00027;s behavior toward reducing uncertainty. In their research, calculating prediction errors as internal rewards led the agents to search for novel states and ultimately outperformed DQN. Following this idea, Pathak et al. (<xref ref-type="bibr" rid="B43">2017</xref>) proposed the intrinsic curiosity module (ICM), which regarded the difference between the predicted state of an agent and the situation obtained from the pixel information on the screen as curiosity. Herein, ICM was integrated with the asynchronous actor-critic model (A3C) (Mnih et al., <xref ref-type="bibr" rid="B37">2016</xref>). Based on this method, Burda et al. (<xref ref-type="bibr" rid="B13">2018</xref>) implemented an approach to explore the environment using only internal rewards regarding curiosity. Moreover, Burda et al. (<xref ref-type="bibr" rid="B14">2019</xref>) proposed a method named random network distillation, which made it possible to learn tasks that were difficult to accomplish with the previous methods.</p>
</sec>
<sec>
<title>2.3 Cognitive models with cognitive architectures</title>
<p>Although the aforementioned studies successfully represented curiosity in reinforcement learning, their integration with cognitive functions has not been sufficiently explored. As explained in Section 2.1, curiosity is associated with the discovery of patterns. Therefore, the computational representation should include basic human cognitive functions behind pattern discovery; this can be achieved using ACT-R. The subsequent section explains the representation of individual cognitive functions in ACT-R and the type of learning realized by combining cognitive functions. Herein, we predominantly focus on the cognitive functions of ACT-R involved in this study. Further information on ACT-R can be obtained from Anderson (<xref ref-type="bibr" rid="B2">2007</xref>), the ACT-R manual (Bothell, <xref ref-type="bibr" rid="B11">2020</xref>), and other reviews (Ritter et al., <xref ref-type="bibr" rid="B47">2019</xref>).</p>
<sec>
<title>2.3.1 Structure of ACT-R modules</title>
<p>ACT-R comprises modules corresponding to brain regions, as indicated in <xref ref-type="fig" rid="F1">Figure 1</xref>. The mapping between the modules and regions has been discussed based on neuroscientific findings (Stocco et al., <xref ref-type="bibr" rid="B52">2021</xref>). The principal assumption of this structure is that one module takes responsibility for a set of functions. For instance, the declarative module comprises functions for storing experience and knowledge, the imaginal module contains functions to create new knowledge by combining multiple internal representations, and the function of the goal module is to maintain the current status of tasks to manage the process of the model. The state of each module at each time point (e.g., the declarative knowledge being recalled and the state of the current goal) is expressed using a symbol referred to as a <italic>chunk</italic>, which is stored in a buffer for each module. The chunks stored in the buffer are evaluated using a type of procedural knowledge, called &#x0201C;productions,&#x0201D; comprising IF (conditions) and THEN (actions) clauses in the production module. The productions transmit chunks describing commands to modules as actions, such as searching for knowledge that satisfies the conditions and updating the current state of the task.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Overview of the adaptive control of thought-rational (ACT-R) modules. Modules excluded in this study are grayed out. This figure is created with reference to Anderson et al. (<xref ref-type="bibr" rid="B3">2004</xref>) and Ritter et al. (<xref ref-type="bibr" rid="B47">2019</xref>). VLPFC, ventrolateral prefrontal cortex; ACC, anterior cingulate cortex; DLPFC, Dorsolateral prefrontal cortex; PPC, posterior parietal cortex.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0001.tif"/>
</fig>
<p>Therefore, the declarative and production modules in ACT-R contain different types of knowledge. The retrieval cost of declarative knowledge (chunks) in the declarative module is greater than that of procedural knowledge (productions) in the production module. The cost in ACT-R corresponds to the processing time, which simulates human reaction times (van der Velde et al., <xref ref-type="bibr" rid="B56">2022</xref>). A single production can be executed in 50 ms, whereas the retrieval of declarative knowledge requires longer as various factors are involved. Moreover, declarative knowledge is not automatically retrieved from the goal module or the external environment as it is always used by applying two productions; one for retrieving declarative knowledge and the other for applying the retrieved knowledge to change the states of buffers (e.g., goal or perceived external environment).</p>
<p>Biologically, the ACT-R theory assumes that the two types of knowledge are connected through the cortico-basal ganglia loop. As depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>, the productions are assumed to be executed in the basal ganglia; however, the ones used for retrieving declarative knowledge require the prefrontal cortex as well. <xref ref-type="fig" rid="F2">Figure 2</xref> illustrates a simple example of retrieving and using declarative knowledge through two productions. In the figure, variables &#x0201C;var1&#x0201D; and &#x0201C;var2&#x0201D; in the productions are bound to numerical values, such as 1 and 2, stored in the declarative module. This mechanism is referred to as &#x0201C;pattern matching&#x0201D; and is assumed to be executed intentionally in the prefrontal cortex [particularly in the ventrolateral prefrontal cortex (VLPFC) indicated in <xref ref-type="fig" rid="F1">Figure 1</xref>]. Therefore, we considered the pattern matching between the current situation (buffers) and knowledge in the declarative module as the criterion for distinguishing the aforementioned levels of thinking (Section 2.1). In this framework, the shallow level of thinking involved fewer pattern-matching scenarios than the deliberative level of thinking.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Simple example of pattern matching in the adaptive control of thought-rational (ACT-R) architecture. This example illustrates the flow of a declarative knowledge query to the declarative module (DM) in the THEN clause of the previous production. The variables are bound in the IF clause of the subsequent production.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0002.tif"/>
</fig>
</sec>
<sec>
<title>2.3.2 Learning in ACT-R</title>
<p>The existence of pattern matching also makes a distinction between two types of learning in ACT-R: learning with pattern matching and learning without pattern matching. The latter type uses &#x0201C;utility learning,&#x0201D; which corresponds to reinforcement learning. Specifically, it changes the selection probability of productions that conflict with each other by receiving rewards from the environment. Many studies have used this type of learning in ACT-R modeling (Anderson et al., <xref ref-type="bibr" rid="B4">1993</xref>; Balaji et al., <xref ref-type="bibr" rid="B7">2023</xref>; Ceballos et al., <xref ref-type="bibr" rid="B16">2020</xref>; Xu and Stocco, <xref ref-type="bibr" rid="B58">2021</xref>). For example, Fu and Anderson (<xref ref-type="bibr" rid="B24">2006</xref>) developed a model to solve the repeated maze task by applying procedural knowledge representing up-down and left-right movements. The model received positive rewards for actions that led to the achievement of the current goal and negative rewards for actions that failed to achieve the goal. As a result of their simulation, the model was able to learn optimal behavior in the maze search by repeating the rewarding trials.</p>
<p>The other type of learning in ACT-R involves pattern matching to retrieve chunks in the declarative module, which is called instance-based learning (IBL) (Gonzalez et al., <xref ref-type="bibr" rid="B25">2003</xref>). This framework accumulates past problem-solving instances in the declarative module and uses it for future task trials. Several studies show that IBL outperforms conventional utility learning. Relating to this method, Reitter and Lebiere (<xref ref-type="bibr" rid="B46">2010</xref>) constructed a model to solve maze like Fu and Anderson (<xref ref-type="bibr" rid="B24">2006</xref>), but unlike them, by combining path-finding with backtracking and instance-based inference. In their model, location information of the maze was represented as declarative knowledge to construct a topological map (graph-like structure representing geological locations). In addition to conventional knowledge-search algorithms (e.g., depth-first search), an instance-based inference was applied by using stored maze-solving experience in the declarative module. By conducting simulations using the strategies of maze search, they demonstrated the advantage of this memory-based search.</p>
<p>Furthermore, ACT-R contains another learning function that uses the two aforementioned functions. This function is the &#x0201C;compilation,&#x0201D; which combines two productions into a single production (Taatgen and Lee, <xref ref-type="bibr" rid="B54">2003</xref>). During the task execution, this function integrates a repeatedly selected series of productions and reduces the number of productions used in the task as learning progresses. Typically, the target series of compilation involves pattern matching to retrieve declarative knowledge (<xref ref-type="fig" rid="F2">Figure 2</xref>). The function replaces the variables present in the production with instantiated values in the declarative knowledge. Additionally, the conflicting conditions for pre-compiled and post-compiled productions are resolved using the utility learning. The post-compiled production inherits higher utility from those associated with the two pre-compiled productions. Furthermore, the utility of post-compiled production increases with the compilation of the same series of productions. This process increases the probability of selecting a post-compiled production to represent a routine and automatic operation (procedural knowledge) in a task.</p>
</sec>
<sec>
<title>2.3.3 Emotion in ACT-R</title>
<p>The subject of the present study, motivation, is considered part of the emotional or affective phenomena in the recently emerging field of affective science.<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> In this field, researchers have repeatedly pointed to the relations between emotions, cognition, and body (Barrett, <xref ref-type="bibr" rid="B9">2017</xref>; Damasio, <xref ref-type="bibr" rid="B20">2003</xref>; LeDoux and Pine, <xref ref-type="bibr" rid="B35">2016</xref>), underscoring the importance of incorporating emotional and physiological responses into cognitive models.</p>
<p>Following these trends in affective science, several researchers have constructed ACT-R models that represent the interactions between cognition, emotion, and the body. For example, van Vugt and van der Velde (<xref ref-type="bibr" rid="B57">2018</xref>) constructed a model explaining depression based on the proportion of memories accompanied by emotional moods. Similarly, Juvina et al. (<xref ref-type="bibr" rid="B29">2018</xref>) considered the relationship between emotional memories and reward functions. In addition to these links between emotion and cognition, researchers have included psychophysiological factors such as fatigue (Atashfeshan and Razavi, <xref ref-type="bibr" rid="B5">2017</xref>; Gunzelmann et al., <xref ref-type="bibr" rid="B26">2009</xref>) and stress (Dancy et al., <xref ref-type="bibr" rid="B21">2015</xref>) in ACT-R. Based on these models of emotions, several ACT-R models of motivations have been developed (Nishikawa et al., <xref ref-type="bibr" rid="B42">2022</xref>; Nagashima et al., <xref ref-type="bibr" rid="B41">2022</xref>; Yang and Stocco, <xref ref-type="bibr" rid="B59">2024</xref>). Furthermore, in recent discussions on the common cognitive model, Rosenbloom et al. (<xref ref-type="bibr" rid="B48">2024</xref>) proposed an architecture including metacognitive modules to represent interactions between cognition and emotion.</p>
<p>However, to implement such emotional processes, all the aforementioned studies developed novel modules or functions of ACT-R. By contrast, the current study aims to model intrinsic motivation using the existing built-in functions of ACT-R. While we recognize the importance of developing new modules to create a neurally faithful structure, we believe that, in line with the philosophy of cognitive architecture (Anderson et al., <xref ref-type="bibr" rid="B3">2004</xref>), it is preferable to represent various cognitive processes by integrating a small set of core functions.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3 Mechanism of intellectual curiosity based on pattern discovery</title>
<p>This section proposes a mechanism of intrinsic motivation. Before presenting details of the mechanism, the basic idea behind our proposal is introduced.</p>
<sec>
<title>3.1 Basic idea</title>
<p>The mechanism proposed here focuses on intellectual curiosity among the types of intrinsic motivation. We used the modifier &#x0201C;intellectual&#x0201D; based on the discussion reported by Malone (<xref ref-type="bibr" rid="B36">1981</xref>). According to him, the curiosity derived from higher-cognitive functions is distinguished from that derived from sensory perceptions. He further argued that the former initiates &#x0201C;a desire to bring better form to one&#x00027;s knowledge structures.&#x0201D; This discussion is consistent with the principle of fun discussed by Schmidhuber (<xref ref-type="bibr" rid="B49">2010</xref>), who claimed that discovering the compressible structure of data would be beneficial to organize knowledge structures in the agent.</p>
<p>We developed a mechanism of intellectual curiosity by associating ACT-R pattern-matching computation. As explained earlier, pattern matching of ACT-R is a core mechanism for understanding higher-order cognitive processes with the discovery of structures (patterns) that map data (declarative knowledge) to a current situation (buffer states of the module) according to a pattern of variables in the production. This mechanism has been considered essential for achieving cognitive flexibility that adapts changing environments by leveraging existing knowledge in novel forms (Spiro et al., <xref ref-type="bibr" rid="B51">2012</xref>). In fact, Anderson (<xref ref-type="bibr" rid="B2">2007</xref>) demonstrated that ACT-R could model human-specific cognitive functions, such as linguistic processing, metacognition, and analogical reasoning by using a certain type of pattern matching.<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> More importantly, ACT-R pattern matching is involved in the learning mechanisms, as described in Section 2. In the following part of this section, the learning mechanisms of ACT-R are combined into a general framework of intrinsic motivation.</p>
</sec>
<sec>
<title>3.2 Components of intellectual curiosity</title>
<p>To understand the role of intellectual curiosity in general cognitive processes, we first discuss its decay (boredom) process. Typically, boredom is caused by stimulus saturation and is related to learning processes as suggested by Csikszentmihalyi (<xref ref-type="bibr" rid="B19">1990</xref>). According to his theory, boredom occurs when the person extensively learns a particular task and it becomes less challenging. Raffaelli et al. (<xref ref-type="bibr" rid="B45">2018</xref>) reviewed research confirming such a process based on subjective and physiological indices, which sometimes showed complex interactions between cognitive and physiological processes. Based on these discussions especially about the relation between learning and boredom, we used the &#x0201C;utility learning&#x0201D; and &#x0201C;production compilation&#x0201D; to represent the decay of intellectual curiosity. Although the general concept of these mechanisms has already been discussed, the subsequent sections focus on the technical details of the modules as ingredients of our integrated mechanism of intellectual curiosity.</p>
<sec>
<title>3.2.1 Motivation as utility for task continuation</title>
<p>We used utility learning in this study as a mechanism for determining whether a task should be continued or terminated. As mentioned in Section 2.3, the utility learning corresponds to reinforcement learning (Fu and Anderson, <xref ref-type="bibr" rid="B24">2006</xref>). When multiple productions (i.e., the production for task continuation and the production for task termination) match the current situation, the probability of selecting a production can be calculated as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>s</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>s</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>e</italic> denotes the base of the natural logarithm, <italic>s</italic> indicates the parameter that determines the variance of noise according to the logistic distribution, and <italic>j</italic> distinguishes the conflicting productions. Additionally, <italic>U</italic> representing the utility of controlled production can be estimated as</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here, &#x003B1; represents the learning rate, and <italic>Ri</italic>(<italic>n</italic>) denotes the reward obtained by production <italic>i</italic> at time <italic>n</italic>. In general, rewards occur when a production associated with the goal of the task is executed. Typically, rewards are backpropagated to the productions that are executed before the reward is triggered. Each time a production is rewarded, the utility values of all productions that have been executed since the last update (<italic>n</italic>&#x02212;1) are updated using <xref ref-type="disp-formula" rid="E2">Equation 2</xref>. In this study, events related to intellectual curiosity and boredom are represented by assigning positive and negative rewards, respectively.</p>
</sec>
<sec>
<title>3.2.2 Reduction of pattern matching through production compilation</title>
<p>As mentioned earlier, production compilation compresses productions to reduce frequencies of pattern matching. Therefore, the fun generated by pattern matching (identifying structures in the data) was considered decayed by the compression accompanied with production compilation.</p>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> depicts the traces of an ACT-R model in a maze task used in simulations performed in this study. The vertical axis indicates time, and each column indicates an event in a module. The left-hand side trace represents the process of identifying the path from the declarative knowledge using pattern matching. The trace on the right-hand side expresses the search for a path without pattern matching or retrieving paths from the declarative module; in other words, it represents the processing after production compilation.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Example illustrating the before and after of learning using the production compilation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0003.tif"/>
</fig>
</sec>
</sec>
<sec>
<title>3.3 Integrated mechanism of task continuation based on intellectual curiosity</title>
<p>We propose a mechanism for determining the continuation or termination of a task based on intellectual curiosity. <xref ref-type="fig" rid="F4">Figure 4</xref> illustrates the procedure of task continuation when executing general tasks. At the beginning of each round (unit related to the continuation of a task), the model determines whether to continue or terminate the task based on the conflict resolution between the two productions (<italic>stop</italic> and <italic>continue productions</italic>). The model proceeds with the round by firing various productions, such as searching the map, after deciding to continue the task. When the model encounters a condition that terminates the round, a new round is initiated, and the model again determines whether the task should be continued or terminated.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Flowchart of the task continuation model. The model generates positive rewards when pattern matching occurs.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0004.tif"/>
</fig>
<p>In the aforementioned process, the assigned initial values of utilities are higher in the continue production than in the stop production. At the beginning of the task, it can be assumed that agents intend to continue the task. The process of experiencing boredom from this initial state can be modeled by assigning a trigger of negative reward to the production recognizing the end of each round.<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> The utility of the production decreases when a negative reward is generated by the continue production at the end of the round, which in turn increases the firing probability of the stop production.</p>
<p>To deter boredom and continue the task, positive rewards corresponding to &#x0201C;fun&#x0201D; are necessary. This study associates the occurrence of pattern matching with the feeling of fun. We consider that this association is consistent with the definition of fun reported by Schmidhuber (<xref ref-type="bibr" rid="B49">2010</xref>) because it involves the discovery of patterns in the environment. However, repeated application of the same production causes habituation (production compression) and increases the opportunity to generate negative rewards to the continue production at the end of a round. In other words, the factor that ensures task continuation in the mechanism is the continued stimulation of intellectual curiosity through the discovery of declarative knowledge (data), which is the target of pattern matching.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Simulation</title>
<p>We performed simulations to verify the proposed mechanism of intellectual curiosity. This section explains the purpose of the simulations, the employed task, model details, and other settings involved in the simulations. Finally, the obtained results are summarized.<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref></p>
<sec>
<title>4.1 Aims and indicators</title>
<p>To examine the mechanism of intellectual curiosity based on pattern discovery, we address the following questions.</p>
<list list-type="order">
<list-item><p>What type of environment stimulates intellectual curiosity?</p></list-item>
<list-item><p>How does stimulated intellectual curiosity affect task learning?</p></list-item>
<list-item><p>What is the relationship between the proposed mechanism and the curiosity represented in existing reinforcement learning models?</p></list-item>
</list>
<p>The first question was answered by distinguishing between <italic>external</italic> and <italic>internal</italic> environments surrounding the model. Here, based on the previous discussion on IMRL (Singh et al., <xref ref-type="bibr" rid="B50">2005</xref>), we adopted the term internal environment to explore the individual differences surrounding and affecting intellectual curiosity. In this context, the external environment was manipulated by varying the complexity (difficulty level) of the learning environment, while the internal environment was defined as the strategy employed by the model to explore the external environment.</p>
<p>Furthermore, we examined how the factors of the internal and external environment affect intellectual curiosity using the indicators</p>
<list list-type="simple">
<list-item><p>(a) up-time ratio (percentage of time the model was running relative to the time limit of one run); and</p></list-item>
<list-item><p>(b) number of rounds (frequency of firing the task continuation production, depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>).</p></list-item>
</list>
<p>These indicators represent the extent to which the model engaged with the task. By definition, if the model obtains strong motivation, these indicators are assumed to be increased. We explored the internal and external environments that fostered this effect. As an internal environment, we manipulated the depth of thinking when searching external environments. According to the discussion presented in Section 2.1, this factor is expected to affect the effect of intrinsic motivation of the model via the interactions with the difficulty level of the external environment.</p>
<p>To answer the second question, the effect of stimulated intellectual curiosity on learning in the task was examined using the indicators</p>
<list list-type="simple">
<list-item><p>(c) entropy (variety of behavior patterns in the environment search);</p></list-item>
<list-item><p>(d) goal rate (the goal achievement rate); and</p></list-item>
<list-item><p>(e) the number of newly generated productions (frequency of occurrence of production compilation).</p></list-item>
</list>
<p>These indicators quantify the effect of intellectual curiosity on three aspects, namely, the behavior pattern (c), learning outcome (d), and internal states (e). We assumed that these indicators would increase with higher intellectual curiosity. In other words, the higher the motivation, the more opportunities the model has to explore the map. Moreover, as the model is extensively exploring the map, entropy (c) and the goal rates (d) increase while the model discovers more patterns in the external environment (e).</p>
<p>The complexity of model behavior (c) was computed as the entropy normalized for the frequency of occurrence of states of the task as follows:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mstyle mathvariant="italic"><mml:mtext>n</mml:mtext></mml:mstyle></mml:mrow></mml:munder></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here, <italic>x</italic><sub><italic>i</italic></sub> denotes a particular state in an environment, and <italic>n</italic> represents the total number of states in an environment. This index increased when the model extensively explored the environment and the value decreased during local behaviors.</p>
<p>Finally, to address the last question, we used these indicators to examine whether the proposed behavior of curiosity was consistent with the previous models of curiosity. Among several models, we focused on the ICM model (Pathak et al., <xref ref-type="bibr" rid="B43">2017</xref>) as the representative mechanism of deep reinforcement learning with curiosity and compared it with the ACT-R models with various internal and external environments.</p>
</sec>
<sec>
<title>4.2 Task: manipulation of the external environment</title>
<p>Based on the previous reports (Fu and Anderson, <xref ref-type="bibr" rid="B24">2006</xref>; Reitter and Lebiere, <xref ref-type="bibr" rid="B46">2010</xref>) on ACT-R explained in Section 2.3.2, we adopted the task of searching mazes. To systematically manipulate the difficulty level of the task, we applied a maze generation algorithm<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> to grids of sizes 5 &#x000D7; 5, 7 &#x000D7; 7, and 9 &#x000D7; 9, with 10 different maps prepared for each size; <xref ref-type="fig" rid="F5">Figure 5</xref> depicts an example of the maps. As indicated in the figure, the created mazes are loop-less structures with the starting location at the top-leftmost corner, and the goal location at the corner where the maximum number of corner points is traversed from the start point. In other words, two corner points with the highest number of hops were selected as the start and goal locations. The difficulty level of this task corresponded to the size of the maps. As described in Section 2.2, we assumed that an appropriate level of difficulty stimulates intrinsic motivation. Therefore, the factors that stimulate the proposed intellectual curiosity were examined by comparing different sizes of the external environments.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Manipulation of external environments.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0005.tif"/>
</fig>
<p>This task was implemented in ACT-R using a simplified method to obtain stable results over numerous runs. Rather than presenting a visual representation of the map to the model, we included chunks representing the structure of the map in the declarative module of the model.<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref> In other words, the task corresponded to a situation where the model performed path planning without actually moving the body with respect to the topologically represented declarative knowledge of the environment.</p>
<p>The topological map provided to the model comprised chunks representing nodes (corner points) and paths (connections between the corner points) of the maze. When the task was executed, the model stored a node chunk in the goal module that indicated the currently focused corner point. From this state of the goal module, the model attempted to discover the chunk of paths stored in the declarative modules by matching them with patterns of variables embedded in the productions. When the chunk containing the current node was retrieved from the declarative module, the other node associated with the corresponding path chunk was newly stored in the goal module. This process was repeated until the model reached the goal point or the designated time was elapsed.</p>
</sec>
<sec>
<title>4.3 Search strategy: manipulation of the internal environment</title>
<p>To examine the internal environment that stimulates intellectual curiosity, we manipulated the strategy of exploring the external environment in terms of different levels of thinking (Brooks, <xref ref-type="bibr" rid="B12">1986</xref>; Evans, <xref ref-type="bibr" rid="B22">2003</xref>; Kahneman, <xref ref-type="bibr" rid="B30">2011</xref>). As explained in Section 2.1, human mental activities are traditionally divided into at least two levels despite a continuous debate on the simple separation. This study follows the discussion reported by Conway-Smith and West (<xref ref-type="bibr" rid="B18">2022</xref>), suggesting that individual mental process is characterized by a spectrum between the fast automatic and slow deliberate processes. According to them, the levels in this spectrum can determine the amount of mental effort (computational cost) required for the task. Among several types of computational costs, we focused on the effort of retrieving declarative knowledge. As described in Section 2.3.1, retrieval of declarative knowledge in ACT-R can be hypothesized to increase prefrontal cortex activity. Therefore, it can be reasonably assumed that deliberative levels of thinking, which affect the optimal level of intrinsic motivation (Csikszentmihalyi, <xref ref-type="bibr" rid="B19">1990</xref>; Yerkes and Dodson, <xref ref-type="bibr" rid="B60">1908</xref>), are estimated from the amount of declarative knowledge retrieved during the task execution.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> depicts the manipulation of the levels of thinking in this study focusing on the maze search task. The process of the model became complex from left to right, and the amount of declarative knowledge used in the task was assumed to increase. These models were developed based on the authors&#x00027; previous work (Nagashima et al., <xref ref-type="bibr" rid="B40">2021</xref>) with two modifications; more complex pattern matching in the path retrieval and leveraging all pattern matching as triggers of intrinsic reward. In the previous research, the smallest number of variables in the productions was only one, so there was no pattern in the rule. Also, the previous research limited the triggers of intrinsic rewards only when the maze searching rules were fired, omitting rewards generated from pattern matching that occurred by other productions during the task.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Manipulation of internal environments. DFS, depth-first search; IBL, instance-based learning.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0006.tif"/>
</fig>
<p>These changes were made to ensure the model&#x00027;s consistency with our theoretical assumptions. There may be debate about assuming that every pattern match triggers intrinsic rewards. For example, it might be possible to prioritize pattern matching based on complexity or to select productions for positive rewards by setting certain criteria. However, in this study, we prioritized a simpler setting to verify the basic idea, avoiding any arbitrariness. The next sections explain the specific process of the model in each internal environment.</p>
<sec>
<title>4.3.1 Random model</title>
<p>The model with the lowest level of thinking randomly transitioned to the current location stored in the goal module. The model repeated the following process during each round until the goal was achieved or the time limit was reached.</p>
<list list-type="order">
<list-item><p>Path search: the model retrieved the declarative knowledge related to the paths adjacent to the current location. To retrieve the declarative knowledge, the model used productions in which the current location was bound to a variable.</p></list-item>
<list-item><p>Move:
<list list-type="simple">
<list-item><p>(a) If the path retrieval was successful (pattern matching occurred), the model updated the state of the goal module according to the retrieved path, and the model returned to Step (1).</p></list-item>
<list-item><p>(b) If the path retrieval failed, the model returned to Step (1) without modifying the state of the goal module.</p></list-item>
</list></p></list-item>
</list>
<p>While the model explored the maze using this search strategy, the productions that were used for the successful retrieval of the path were compiled. The model was assumed to have a few opportunities for pattern matching because production compilation occurred only when the stored path was retrieved. Therefore, stimulating intellectual curiosity in this model was considered as difficult.</p>
</sec>
<sec>
<title>4.3.2 Stochastic Depth-first Search (DFS) model</title>
<p>To include higher cognitive functions (declarative module), we constructed a probabilistic DFS model, which backtracked to search the environment based on the study by Reitter and Lebiere (<xref ref-type="bibr" rid="B46">2010</xref>). As indicated in <xref ref-type="fig" rid="F7">Figure 7</xref>, the model exhibited a stacked structure with chunks generated by the imaginal module of ACT-R. The push function in the stack was realized by storing a chunk that contained the name of the previous chunk in the <italic>Link</italic> slot. Additionally, the pop function in the stack was realized by returning this slot value to the previous slot value. We implemented all these processes using only ACT-R productions without defining any external functions written in other programming languages, such as LISP.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Construction of the stack structure using chunks of adaptive control of thought-rational (ACT-R) architecture. This stack was implemented using the imaginal module of ACT-R.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0007.tif"/>
</fig>
<p>Similar to the random model, the stochastic DFS model compiled productions that could retrieve declarative knowledge of paths and backtrack to learn new productions that did not contain variables. The specific model behavior can be summarized as follows.</p>
<list list-type="order">
<list-item><p>Path search: the model determined the destination by retrieving the declarative knowledge associated with the path, similar to the random model. The IF clause included the current location stored in the goal buffer and five variables, corresponding to the current location and the directions (west, north, east, and south), which were flags indicating whether the direction was already searched or not.</p></list-item>
<list-item><p>Move:
<list list-type="simple">
<list-item><p>(a) If the knowledge retrieval was successful (pattern matching occurred), the model flagged the retrieved direction as &#x0201C;searched,&#x0201D; created a new chunk using the imaginal module, and stored the chunk as declarative memory, as depicted in <xref ref-type="fig" rid="F7">Figure 7</xref>. Simultaneously, the model updated the current location of the goal buffer according to the retrieved path. At this point, the searched flag in the goal buffer was reset, whereas the searched flag in the direction opposite to the direction of movement was set to prevent its return to the previous location. After this procedure, the model returned to Step (1).</p></list-item>
<list-item><p>(b) The backtracking process was executed if the path retrieval failed, returning the model to the previous state by popping chunks in the stack; eventually, the model returned to Step (1).</p></list-item>
</list></p></list-item>
</list>
<p>The model repeated this behavior until the goal was achieved or the time limit was reached. Contrary to the random model, the DFS model used the stack when the path search failed. Therefore, the model required more rounds to compress (compile the production) the declarative knowledge of the paths.</p>
</sec>
<sec>
<title>4.3.3 Stochastic DFS plus Instance-based Learning (IBL) model</title>
<p>This model combined the stochastic DFS with the IBL, which leverages past memories to solve current tasks (Gonzalez et al., <xref ref-type="bibr" rid="B25">2003</xref>; Lebiere et al., <xref ref-type="bibr" rid="B34">2007</xref>). In this task, the model held all the retrieved paths in the stack from the beginning of each round until the goal was reached. After the model attained the goal, the path chunks in the stack were retrieved one by one, and the chunks labeled &#x0201C;correct path&#x0201D; were generated. During each round, the model repeated the following two steps until the goal was achieved or the time limit was reached.</p>
<list list-type="order">
<list-item><p>Determining strategies: at the beginning of each round, the model decided between the stochastic DFS and the IBL strategies by retrieving chunks associated with the current location and labeled &#x0201C;correct path.&#x0201D;</p></list-item>
<list-item><p>Move:
<list list-type="simple">
<list-item><p>(a) When the DFS strategy was employed (failed to retrieve the &#x0201C;correct path&#x0201D;), the model behaved as a stochastic DFS model.</p></list-item>
<list-item><p>(b) When the model successfully retrieved the &#x0201C;correct path,&#x0201D; the model updated the current location according to the retrieved path chunk. Subsequently, the model returned to Step (1).</p></list-item>
</list></p></list-item>
</list>
<p>The model behaved as the stochastic DFS model in the early stages of the task. With the repetition of rounds and the increase in the number of instances with the &#x0201C;correct path,&#x0201D; the model effectively reached the goal. Here, IBL was a time-consuming process in comparison with the DFS strategy. This was because the model had to retrieve the path in the stack at the end of the round to assign a label to a path. Moreover, retrieval trials for past successful rounds at the beginning of each round resulted in additional time, which was not included in the other models. We hypothesized that similar to the stochastic DFS model, this model is likely to stimulate intellectual curiosity, and the IBL function would positively affect the learning of the task.</p>
</sec>
<sec>
<title>4.3.4 Deep reinforcement learning model based on curiosity</title>
<p>To explore the relationship between the aforementioned ACT-R models and previous models of intrinsic motivation using deep reinforcement learning, we constructed an ICM model based on the report by Pathak et al. (<xref ref-type="bibr" rid="B43">2017</xref>). The ICM model in this study explored the maze using the policy &#x003C0; in actor-critic model.<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref> This search resulted in a policy that maximized the rewards represented as</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Thus, the reward of the model was calculated as the sum of the internal reward (<italic>r</italic><sub><italic>i</italic></sub>) and external reward (<italic>r</italic><sub><italic>e</italic></sub>). Based on this equation, the model explored the environment by balancing the two types of rewards.</p>
<p>In this study, following Pathak et al. (<xref ref-type="bibr" rid="B43">2017</xref>), the internal reward was determined by</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here, state <italic>s</italic> is defined as pixel data in deep reinforcement learning. In this study, the maze situation (players, walls, and paths) was converted into a grayscale image (42 &#x000D7; 42) and served as input to a CNN, whose parameters were represented as &#x003D5;. By subtracting the predicted and actual outputs of CNN, prediction errors were computed and weighted using the coefficient &#x003B7;. This coefficient was regarded as the intensity of curiosity.</p>
<p>By contrast, the external reward was defined as</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>if&#x000A0;failed&#x000A0;to&#x000A0;move;</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mn>0</mml:mn></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>if&#x000A0;succeeded&#x000A0;to&#x000A0;move;</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>if&#x000A0;the&#x000A0;goal&#x000A0;was&#x000A0;achieved.</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>In each action, the model attempted to select one of the directions, namely, west, north, east, or south, and to transition the state from one corner point to another. If the model selected a direction that did not lead to a path, the action was considered a failure. If the model reached the goal owing to its movement, it was rewarded for its success; subsequently, the task moved on to the next round.</p>
<p>The search for the maze was terminated when the condition</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:mn>500</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>s</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>was satisfied. Here, <italic>th</italic> denotes the threshold value, and <italic>egs</italic> indicates noise. The model search was terminated when the internal reward was less than the threshold.<xref ref-type="fn" rid="fn0008"><sup>8</sup></xref></p>
</sec>
</sec>
<sec>
<title>4.4 Simulation settings</title>
<sec>
<title>4.4.1 Setting for ACT-R models</title>
<p>As parameters relevant to the general model of task continuation (<xref ref-type="fig" rid="F4">Figure 4</xref>), the simulation assigned the initial utility values of the <italic>continue</italic> and <italic>stop productions</italic> to 10 and 5, respectively. Additionally, we assigned the triggers of the negative reward (<italic>r</italic> &#x0003D; 0) to productions that recognized the end of the round, which was either reaching the goal or recognizing that the time limit of each round was elapsed. Conversely, the triggers of the positive reward were assigned to productions that included pattern matching, which corresponded to intellectual curiosity. We manipulated the intensity of the model&#x00027;s intellectual curiosity by sampling the positive reward at five equal intervals, ranging from 2 to 18. For parameters not directly related to our proposed mechanism, we adopted values from previous studies. Following Anderson et al. (<xref ref-type="bibr" rid="B3">2004</xref>), the activation noise level (ANS), which represents the noise in memory recall, was set to 0.4, and the production noise level (EGS), which reflects the noise in comparing utilities for continuing or terminating productions, was set to 0.5.</p>
<p>To enable the above setting of rewarding by pattern matching, we made small modifications to the original source code of ACT-R (Ver. 7.21). We first modified the source code of ACT-R to assign the reward trigger at any time point after the production compilation occurred. Subsequently, we modified the code to not inherit those triggers after the compilation. In the original ACT-R source code, the compiled production inherits the reward triggers from the original production. We redefined this function to represent boredom caused by the lack of new production compilation.</p>
<p>Simulations based on these settings were run 10 times for each map and each positive reward setting. The limits in the ACT-R simulation time for each round and run were set to 180 and 3,600 s, respectively.</p>
</sec>
<sec>
<title>4.4.2 Setting for ICM model</title>
<p>The ICM model was implemented using PyTorch (ver. 1.9.0), with parameters set to match those of the ACT-R, wherein the simulations were run on 30 maps and the proportion of the internal reward (&#x003B7;) for each run was divided into five samples with equal intervals, ranging from 0.1 to 0.9. We compared the sum of the internal reward (<italic>r</italic><sub><italic>i</italic></sub>) and the noise (<italic>egs</italic>) with the threshold (<italic>th</italic> &#x0003D; 5) in <xref ref-type="disp-formula" rid="E7">Equation 7</xref> to determine whether the task was continued or terminated. Furthermore, we set 156 and 3,130 steps as the limit of the action in each round and run, respectively. These steps were set to be equivalent to the time limit set at the ACT-R models. One step of the ICM model was equivalent to the rule transition time of 1.15 s in the default random model. The ICM model was run 100 times for each reward setting as it could run faster than the ACT-R models.</p>
</sec>
</sec>
<sec>
<title>4.5 Simulation results</title>
<p><xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref> illustrate the simulation results as a function of the internal reward for each of the indicators discussed in Section 4.1. Each graph depicts the average value, which was <italic>n</italic> &#x0003D; 100 (10 times &#x000D7; 10 maps) for the ACT-R models and <italic>n</italic> &#x0003D; 1, 000 (100 times &#x000D7; 10 maps) for the ICM model, aggregated for each internal and external environment condition with respect to the map size. The influence of the maps of the external environment was examined by comparing the three series in each graph, whereas the influence of the internal environment (random, DFS, DFS &#x0002B; IBL) of the model was analyzed based on the difference between the graphs aligned in the horizontal direction. The subsequent sections discuss the obtained results based on the three questions posed as objectives of the simulation.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Simulation results. The numbers in the horizontal line distinguish the models (1&#x02013;3: adaptive control of thought-rational (ACT-R) models; 4: intrinsic curiosity module (ICM) model), and the vertical alphabet differentiates the indicators (<bold>A</bold>: Up-time ratio; <bold>B</bold>: Number of rounds). The error bars in each graph indicate the mean value (<italic>n</italic> &#x0003D; 10) of the standard deviations (ACT-R: <italic>n</italic> &#x0003D; 10, ICM: <italic>n</italic> &#x0003D; 100) obtained for each map when multiplied by 1/10.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0008.tif"/>
</fig>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Simulation results. The numbers in the horizontal line distinguish the models (1&#x02013;3: adaptive control of thought-rational (ACT-R) models; 4: intrinsic curiosity module (ICM) model), and the vertical alphabet differentiates the indicators (<bold>C</bold>: Entropy; <bold>D</bold>: Goal rate; <bold>E</bold>: Number of productions). The error bars in each graph indicate the mean value (<italic>n</italic> &#x0003D; 10) of the standard deviations (ACT-R: <italic>n</italic> &#x0003D; 10, ICM: <italic>n</italic> &#x0003D; 100) obtained for each map when multiplied by 1/10.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0009.tif"/>
</fig>
<sec>
<title>4.5.1 Environment that stimulates intellectual curiosity</title>
<p>In Section 4.1, the first question posed was &#x0201C;What type of environment stimulates intellectual curiosity?&#x0201D; To address this question, we focused on the up-time ratio (<xref ref-type="fig" rid="F8">Figure 8A</xref>) and number of rounds (<xref ref-type="fig" rid="F8">Figure 8B</xref>) as the behavior indicators of intellectual curiosity. These indicators increased continuously with the increase in the strength of intellectual curiosity for every series (map size) in every graph (levels of thinking) of <xref ref-type="fig" rid="F8">Figure 8</xref>. This general trend suggested that the implemented intellectual curiosity actually enhanced the motivation for the task and ensured task continuation.</p>
<p>In terms of the difference in the external environment, the up-time ratio (<xref ref-type="fig" rid="F8">Figure 8A</xref>) increased as the map became more complex (9 &#x000D7; 9&#x0003E;7 &#x000D7; 7&#x0003E;5 &#x000D7; 5). However, the number of rounds (<xref ref-type="fig" rid="F8">Figure 8B</xref>) presented a reverse trend, wherein the simpler external environment increased the number of rounds (9 &#x000D7; 9 &#x0003C; 7 &#x000D7; 7 &#x0003C; 5 &#x000D7; 5), except for the DFS model (<xref ref-type="fig" rid="F8">Figure 8B2</xref>). The discrepancy between the two indices of motivation was caused by the time limit of the simulation (3,600 s). The model could complete the simple map faster, resulting in a greater number of rounds within the time limit (<xref ref-type="fig" rid="F8">Figure 8B</xref>). However, as indicated by <xref ref-type="fig" rid="F8">Figure 8A</xref>, the simple map enabled the model to terminate the task early owing to the lack of new patterns for production compilation. This result implies that the complex external environment stimulates intellectual curiosity.</p>
<p>Furthermore, we determined the difference between the internal environment, which was not expected in advance. When comparing the three horizontally aligned ACT-R models, we observed that the models with high levels of thinking (DFS &#x0002B; IBL and DFS) had smaller indicators of motivation than the random model. The reason for this difference could be the advantage of the random model with less thinking time and more trials and errors. In this condition without physical constraints, the random model had a better chance of receiving positive rewards by identifying novel paths than the other models.</p>
</sec>
<sec>
<title>4.5.2 Effect of task continuation on model learning</title>
<p>The second question posed was &#x0201C;How does stimulated intellectual curiosity affect task learning?&#x0201D; <xref ref-type="fig" rid="F9">Figure 9</xref> presents the three learning indices, namely, the changes in behavior (<xref ref-type="fig" rid="F9">Figure 9C</xref>: entropy), the learning outcome (<xref ref-type="fig" rid="F9">Figure 9D</xref>: goal rate), and the changes in internal state (<xref ref-type="fig" rid="F9">Figure 9E</xref>: the number of productions).</p>
<p>Based on the analysis of <xref ref-type="fig" rid="F8">Figure 8</xref>, we confirmed that all conditions of the internal and external environments were stimulated by intellectual curiosity. However, <xref ref-type="fig" rid="F9">Figure 9</xref> indicates that the effect of intellectual curiosity on task learning differs depending on the internal environment. The intensity of intellectual curiosity affected positively for higher levels of thinking. In the highest level of thinking (the DFS &#x0002B; IBL model), all learning indices (<xref ref-type="fig" rid="F9">Figures 9C1</xref>, <xref ref-type="fig" rid="F9">D1</xref>, <xref ref-type="fig" rid="F9">E1</xref>) increased with the intensity of intellectual curiosity. In the middle level (the DFS model) increased the number of productions (<xref ref-type="fig" rid="F9">Figure 9C2</xref>) while maintaining the entropy (<xref ref-type="fig" rid="F9">Figure 9C2</xref>) and goal rate (<xref ref-type="fig" rid="F9">Figure 9D2</xref>). In the case of the lowest level (the random model), the intrinsic motivation decreased all indices (<xref ref-type="fig" rid="F9">Figures 9C3</xref>, <xref ref-type="fig" rid="F9">D3</xref>, <xref ref-type="fig" rid="F9">E3</xref>). These trends indicated that the DFS &#x0002B; IBL model exhibited a goal-oriented behavior because of the learning effect of the IBL strategy, whereas the behavior of the DFS model had to search the entire map. Furthermore, the random model did not lead to the goal; this was because the model reinforced unfavorable behavior by repeatedly visiting the same location without expanding the search.</p>
<p>In terms of the effect of the challenge of the task (task difficulty), the entropy (<xref ref-type="fig" rid="F9">Figure 9C</xref>) and the goal rate (<xref ref-type="fig" rid="F9">Figure 9D</xref>) were greater on the small map, whereas the number of productions was higher on the large map. These differences may be attributed to the fact that the small map was easier to explore, which in turn increased the entropy and the goal rate. Conversely, the large map exhibited more pattern-matching opportunities, leading to more accumulated knowledge by frequent compilation.</p>
<p>In summary, intellectual curiosity promoted learning in the DFS &#x0002B; IBL model, which exhibited the highest level of thinking. By contrast, learning in the DFS and random models was not promoted by intellectual curiosity. Moreover, the effect of intellectual curiosity negatively impacted the learning environment in the random model.</p>
</sec>
<sec>
<title>4.5.3 ACT-R curiosity vs. ICM curiosity</title>
<p>Finally, we compared the ICM and ACT-R models in <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref>. Similar to all ACT-R models, the ICM model was stimulated by stronger intellectual curiosity (<xref ref-type="fig" rid="F8">Figure 8</xref>). However, the effect of intellectual curiosity for task learning was specifically similar to the random ACT-R model that exhibited decreasing trends of the entropy (<xref ref-type="fig" rid="F9">Figure 9C4</xref>) and the goal rate (<xref ref-type="fig" rid="F9">Figure 9D4</xref>) with the increase in the strength of intellectual curiosity. With respect to the effect of the external environment, the ICM model was also similar to the random ACT-R model; the up-time ratio (<xref ref-type="fig" rid="F8">Figure 8A4</xref>) and the number of rounds (<xref ref-type="fig" rid="F8">Figure 8B4</xref>) were greater for larger maps, whereas the entropy (<xref ref-type="fig" rid="F9">Figure 9C4</xref>) and the goal rate (<xref ref-type="fig" rid="F9">Figure 9D4</xref>) were greater for smaller maps.</p>
<p>This comparison confirmed commonalities and differences between the developed ACT-R curiosity model and the existing curiosity model in deep reinforcement learning. The proposed ACT-R curiosity mechanism can represent similar learning to the existing model by including a simple internal environment (random search strategy). At the same time, it can incorporate goal-directed behavior by including &#x0201C;explicit use of success memory.&#x0201D; Such an explicit nature of the proposed mechanism also leads to a direct examination of the model&#x00027;s internal learning. The analysis of <xref ref-type="fig" rid="F9">Figure 9E</xref> clearly shows this advantage of interpretability made by the proposed approach.</p>
</sec>
<sec>
<title>4.5.4 Cases of paths discovered by the models</title>
<p>To compare detailed behaviors between models, <xref ref-type="fig" rid="F10">Figure 10</xref> illustrates example paths in a 5 &#x000D7; 5 map. The map depicts start and goal positions at the top left and bottom right corners respectively. The circles&#x00027; colors and line thickness represent visit frequencies during runs. The random model exhibited diagonal movement and movement through walls, a result of compiling multiple movement rules. To gather these examples, we conducted 10 runs for each model across three levels of intrinsic rewards, selecting the runs with the lowest and highest performance for analysis.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Trajectories of the model runs [<bold>(left)</bold>: low-performance runs, <bold>(right)</bold>: high-performance runs]. The columns indicate the strength of intellectual curiosity, while the rows correspond to each model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1397860-g0010.tif"/>
</fig>
<p>These figures reveal distinct behavioral characteristics of each model. The random model predominantly exhibits localized movements within specific areas, often distant from the goal. On the other hand, the DFS model explores the map evenly but does not necessarily move directly toward the goal. In contrast, the DFS &#x0002B; IBL model demonstrates deliberate behaviors aimed at reaching the goal, particularly under high-reward conditions. In terms of localized movement patterns, the ICM model was more similar to the random and DFS models than the DFS &#x0002B; IBL model. Thus, the results suggested that the DFS &#x0002B; IBL model had a greater effect on curiosity strength than the other models regarding directionality toward the goal. These observations support the findings observed in the quantitative results shown in <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref>.</p>
</sec>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5 Conclusions</title>
<p>The objective of this study was to develop a mechanism for intrinsic motivation based on pattern discovery by combining basic modules of ACT-R. This section summarizes the significance of the proposed mechanism and presents the potential future lines of investigation.</p>
<sec>
<title>5.1 Summary and implications</title>
<p>The proposed mechanism was based on the assumption that pattern discovery is associated with the feeling of fun and is the source of intellectual curiosity. Additionally, its attenuation was expressed by the learning mechanism incorporated in ACT-R. To support this proposal, we implemented multiple external environments (challenges in the task) and strategies for exploring the external environment (levels of thinking) and examined the role of intellectual curiosity in each situation. The simulation results indicated that the rewards associated with pattern discovery exhibited different effects on models at different levels of thinking. The model with the lowest level of thinking (random) and that with the middle level of thinking (DFS) had negative and neutral effects of intellectual curiosity on performance, respectively. The only model that benefited from intellectual curiosity was the one with the highest level of thinking (DFS &#x0002B; IBL), which comprised a function that enabled it to remember previous experiences that led to the goal.</p>
<p>These results are partially consistent with the past arguments made for human intrinsic motivation. Particularly, the effectiveness of intrinsic motivation in the DFS &#x0002B; IBL model is consistent with a discussion, in which intrinsic motivation operates well with deliberative thinking, which requires &#x0201C;autonomy,&#x0201D; &#x0201C;mastery,&#x0201D; and &#x0201C;purpose&#x0201D; (Pink, <xref ref-type="bibr" rid="B44">2011</xref>). Furthermore, consistent with our negative results in the random model, several reports exist on behavioral addictions caused by the negative effects of intrinsic motivation (Alter, <xref ref-type="bibr" rid="B1">2017</xref>). For instance, people often forget their goals and become engrossed in exploratory tasks, such as browsing the internet, resulting in poor performance. This irrational behavior might also relate to <italic>computational psychiatry</italic> (Huys et al., <xref ref-type="bibr" rid="B28">2016</xref>).</p>
<p>In addition to the above discussion on the internal environment, we identified a connection with a previous discussion on the external environment. Consistent with the discussion on &#x0201C;challenge&#x0201D; (Malone, <xref ref-type="bibr" rid="B36">1981</xref>), we determined that the up-time ratio in larger maps was greater than that in the smaller maps. This result indicates that a complex external environment stimulates intellectual curiosity. However, we also determined that difficult challenges generate ineffective learning on a wide map (<xref ref-type="fig" rid="F9">Figure 9</xref>). The aforementioned positive and negative effects of the task difficulty indicate the optimal level of challenge (Csikszentmihalyi, <xref ref-type="bibr" rid="B19">1990</xref>; Yerkes and Dodson, <xref ref-type="bibr" rid="B60">1908</xref>).</p>
<p>Furthermore, this study successfully corresponded with past studies on intrinsic motivation in reinforcement learning. The comparisons with the ICM model (Pathak et al., <xref ref-type="bibr" rid="B43">2017</xref>) confirmed that the developed ACT-R model, specifically the random model, is a succession of existing studies. Although we cannot claim its superiority as a learning algorithm based on the current simulation alone, the model with the higher level of thinking (DFS &#x0002B; IBL) exhibited characteristic behavior toward the ICM model. Future analysis of more extensively manipulating parameters, such as the balancing of <italic>r</italic><sub><italic>i</italic></sub> and <italic>r</italic><sub><italic>e</italic></sub> in <xref ref-type="disp-formula" rid="E4">Equation 4</xref> and designing the external environment stimulating curiosity (Burda et al., <xref ref-type="bibr" rid="B13">2018</xref>), could reveal further correspondence between the ACT-R model and reinforcement learning framework.</p>
<p>We believe that the comparisons of the previous model of reinforcement learning reveal the methodological advantage of using cognitive architecture. An integrated cognitive architecture, such as ACT-R, provides criteria to set numerical parameters (e.g., time limits and utilities) based on previous studies. Furthermore, ACT-R comprises neuroscientific modules that correspond to basic cognitive functions, such as declarative and procedural knowledge. Based on this relation, arguments associated with human intrinsic motivation can be developed. Therefore, this study contributes to the understanding of intrinsic motivation in a wide context of the relationship between human evolution and the development of civilization by mapping the discovery of patterns to intrinsic motivation (Baron-Cohen, <xref ref-type="bibr" rid="B8">2020</xref>).</p>
</sec>
<sec>
<title>5.2 Future work</title>
<p>The proposed mechanism of intellectual curiosity has the potential for several future studies. The primary focus among them is human experiments that manipulate the internal and external environments as in the simulation. A simulation study without data is nothing more than a demonstration derived deductively from theory. Therefore, the model&#x00027;s value must be proven by applying it to human scenarios.</p>
<p>One of the obstacles to conducting human experiments for the proposed mechanism is setting tasks to stimulate human curiosity. In this study, we adopted the maze task because several previous researchers based on ACT-R have constructed models for this task. However, setting experimental situations with human participants to exhibit intrinsic motivation for solving such simple tasks may be difficult. Therefore, in the future, we intend to explore tasks that both humans and developed models can execute with proper intrinsic motivation.</p>
<p>Other future work will focus on modeling the curiosity and motivation that was not explained in the current study. As we discussed in Section 3.1, this study targeted on intellectual curiosity relating &#x0201C;a desire to bring better form to one&#x00027;s knowledge structures&#x0201D; (Malone, <xref ref-type="bibr" rid="B36">1981</xref>) or &#x0201C;intrinsic desire to build a better model for the world&#x0201D; (Schmidhuber, <xref ref-type="bibr" rid="B49">2010</xref>). Therefore, we have not yet explained the sensory curiosity that drives us to acquire new knowledge from the world. These two types of curiosity are considered complementary, similar to the explore-exploit trade-off in reinforcement learning. Without including sensory curiosity in the model, we cannot explain how declarative knowledge is acquired for intellectual curiosity, nor how the initial utility settings of continuing the task exceed those of stopping.</p>
<p>The above future study possibly leads to a deeper exploration of levels of thinking. Conway-Smith et al. (<xref ref-type="bibr" rid="B17">2023</xref>) recently summarized the relationship between metacognition and levels of thinking, arguing that compilation of existing knowledge reduces the effort involved in metacognition, making it more automatic. Building on this, we can suggest that achieving such a metacognitive state as a result of higher-level thinking enabled with enough intellectual curiosity, exemplified by the DFS &#x0002B; IBL model.</p>
<p>On the contrary, we can assume the exploratory role of the lower level of thinking. As shown in <xref ref-type="fig" rid="F9">Figure 9</xref>, the random model showed a higher goal ratio with larger learning products in some conditions. These results suggest links between low-level thinking and sensory curiosity, leading to exploration of the environment. Our recent work (Nagashima and Morita, <xref ref-type="bibr" rid="B39">2024</xref>) provides support for the above interpretation. In the experiment, human participants observed the behaviors generated by the models in the current study and rated the random model as having the most curious features.</p>
<p>The final direction for future research is the generalization of the ideas presented in this paper to other tasks in real-world settings. We believe such tasks are linked to the earlier discussion on the civilization of society (Baron-Cohen, <xref ref-type="bibr" rid="B8">2020</xref>). As suggested by Toya and Hashimoto (<xref ref-type="bibr" rid="B55">2018</xref>), tool-making requires recursive compilation of intermediate products. Integrating this multi-agent simulation with the mechanisms proposed in the current study could offer a detailed explanation of the driving forces behind the evolution of civilization.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>KN: Software, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. JM: Conceptualization, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. YT: Conceptualization, Supervision, Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by JSPS KAKENHI Grant Number JP24KJ1215.</p>
</sec>
<ack><p>We acknowledge the use of ChatGPT-4.0 (Open AI, <ext-link ext-link-type="uri" xlink:href="https://chat.openai.com/">https://chat.openai.com/</ext-link>), DeepL (<ext-link ext-link-type="uri" xlink:href="https://www.deepl.com/">https://www.deepl.com/</ext-link>), and Grammarly (<ext-link ext-link-type="uri" xlink:href="https://app.grammarly.com/">https://app.grammarly.com/</ext-link>) for assistance in editing and proofreading for this article.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2024.1397860/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frai.2024.1397860/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.zip" id="SM1" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_2.zip" id="SM2" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_3.zip" id="SM3" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_4.zip" id="SM4" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>The mission of the society for affective science includes &#x0201C;motivated states.&#x0201D; See <ext-link ext-link-type="uri" xlink:href="https://society-for-affective-science.org/about-sas/">https://society-for-affective-science.org/about-sas/</ext-link>.</p></fn>
<fn id="fn0002"><p><sup>2</sup>Claimed by Anderson (<xref ref-type="bibr" rid="B2">2007</xref>) in the chapter featuring &#x0201C;dynamic pattern matching&#x0201D; in ACT-R.</p></fn>
<fn id="fn0003"><p><sup>3</sup>A similar mechanism of stopping a navigation task was presented by Anderson et al. (<xref ref-type="bibr" rid="B4">1993</xref>), although the exact equations of utility calculation were different because of the architectural difference.</p></fn>
<fn id="fn0004"><p><sup>4</sup>The models and data are included in <ext-link ext-link-type="uri" xlink:href="https://github.com/AcmlNagashima/CuriosityAgents">https://github.com/AcmlNagashima/CuriosityAgents</ext-link>.</p></fn>
<fn id="fn0005"><p><sup>5</sup><ext-link ext-link-type="uri" xlink:href="https://algoful.com/Archive/Algorithm/MazeExtend">https://algoful.com/Archive/Algorithm/MazeExtend</ext-link>.</p></fn>
<fn id="fn0006"><p><sup>6</sup>The exclusion of perceptual and motor processes in basic simulations is also recommended in the official ACT-R tutorial.</p></fn>
<fn id="fn0007"><p><sup>7</sup>Discount rate <italic>gamma</italic> &#x0003D; 0.99.</p></fn>
<fn id="fn0008"><p><sup>8</sup>A fixed value of 500 was tentatively multiplied because the scales of the internal reward in the ACT-R and ICM models were different.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alter</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <source>Irresistible: The Rise of Addictive Technology and The Business of Keeping Us Hooked</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Penguin</publisher-name>.</citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name></person-group> (<year>2007</year>). <source>How Can the Human Mind Occur in The Physical Universe</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>. <pub-id pub-id-type="doi">10.1093/acprof:oso/9780195324259.001.0001</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name> <name><surname>Bothell</surname> <given-names>D.</given-names></name> <name><surname>Byrne</surname> <given-names>M. D.</given-names></name> <name><surname>Douglass</surname> <given-names>S.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Qin</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>An integrated theory of the mind</article-title>. <source>Psychol. Rev</source>. <volume>111</volume>, <fpage>1036</fpage>&#x02013;<lpage>1060</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.111.4.1036</pub-id><pub-id pub-id-type="pmid">15482072</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name> <name><surname>Kushmerick</surname> <given-names>N.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name></person-group> (<year>1993</year>). <article-title>&#x0201C;Navigation and conflict resolution,&#x0201D;</article-title> in <source>Rules of The Mind</source> (<publisher-loc>Psychology Press</publisher-loc>), <fpage>93</fpage>&#x02013;<lpage>120</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atashfeshan</surname> <given-names>N.</given-names></name> <name><surname>Razavi</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Determination of the proper rest time for a cyclic mental task using ACT-R architecture</article-title>. <source>Hum. Factors</source> <volume>59</volume>, <fpage>299</fpage>&#x02013;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.1177/0018720816670767</pub-id><pub-id pub-id-type="pmid">27738278</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aubret</surname> <given-names>A.</given-names></name> <name><surname>Matignon</surname> <given-names>L.</given-names></name> <name><surname>Hassas</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>A survey on intrinsic motivation in reinforcement learning</article-title>. <source>arXiv</source> [Preprint]. arXiv:1908.06976. <pub-id pub-id-type="doi">10.48550/arXiv.1908.06976</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balaji</surname> <given-names>B.</given-names></name> <name><surname>Shahab</surname> <given-names>M. A.</given-names></name> <name><surname>Srinivasan</surname> <given-names>B.</given-names></name> <name><surname>Srinivasan</surname> <given-names>R.</given-names></name></person-group> (<year>2023</year>). <article-title>ACT-R based human digital twin to enhance operators&#x00027; performance in process industries</article-title>. <source>Front. Hum. Neurosci</source>. <volume>17</volume>:<fpage>18</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2023.1038060</pub-id><pub-id pub-id-type="pmid">36845875</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baron-Cohen</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <source>The Pattern Seekers: How Autism Drives Human Invention</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Basic Books</publisher-name>.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barrett</surname> <given-names>L. F.</given-names></name></person-group> (<year>2017</year>). <source>How Emotions Are Made: The Secret Life of the Brain</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Pan Macmillan</publisher-name>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <name><surname>Srinivasan</surname> <given-names>S.</given-names></name> <name><surname>Ostrovski</surname> <given-names>G.</given-names></name> <name><surname>Schaul</surname> <given-names>T.</given-names></name> <name><surname>Saxton</surname> <given-names>D.</given-names></name> <name><surname>Munos</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Unifying count-based exploration and intrinsic motivation,&#x0201D;</article-title> in <source>Proceedings of the 30th International Conference on Neural Information Processing Systems</source> (<publisher-loc>Barcelona</publisher-loc>), <fpage>1479</fpage>&#x02013;<lpage>1487</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bothell</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <source>ACT-R 7.21&#x0002B; Reference Manual</source>. Available at: <ext-link ext-link-type="uri" xlink:href="http://act-r.psy.cmu.edu/actr7.21/reference-manual.pdf">http://act-r.psy.cmu.edu/actr7.21/reference-manual.pdf</ext-link></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brooks</surname> <given-names>R.</given-names></name></person-group> (<year>1986</year>). <article-title>A robust layered control system for a mobile robot</article-title>. <source>IEEE J. Robot. Autom</source>. <volume>2</volume>, <fpage>14</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/JRA.1986.1087032</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burda</surname> <given-names>Y.</given-names></name> <name><surname>Edwards</surname> <given-names>H.</given-names></name> <name><surname>Pathak</surname> <given-names>D.</given-names></name> <name><surname>Storkey</surname> <given-names>A.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Large-scale study of curiosity-driven learning</article-title>. <source>arXiv</source> [Preprint]. arXiv:1808.04355. <pub-id pub-id-type="doi">10.48550/arXiv.1808.04355</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Burda</surname> <given-names>Y.</given-names></name> <name><surname>Edwards</surname> <given-names>H.</given-names></name> <name><surname>Storkey</surname> <given-names>A.</given-names></name> <name><surname>Klimov</surname> <given-names>O.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Exploration by random network distillation,&#x0201D;</article-title> in <source>7th International Conference on Learning Representations (ICLR 2019)</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>17</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caillois</surname> <given-names>R.</given-names></name></person-group> (<year>1958</year>). <source>Les Jeux et les Hommes: Le Masque et la Vertige</source>. <publisher-loc>Paris</publisher-loc>: <publisher-name>Gallimard</publisher-name>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ceballos</surname> <given-names>J. M.</given-names></name> <name><surname>Stocco</surname> <given-names>A.</given-names></name> <name><surname>Prat</surname> <given-names>C. S.</given-names></name></person-group> (<year>2020</year>). <article-title>The role of basal ganglia reinforcement learning in lexical ambiguity resolution</article-title>. <source>Top. Cogn. Sci</source>. <volume>12</volume>, <fpage>402</fpage>&#x02013;<lpage>416</lpage>. <pub-id pub-id-type="doi">10.1111/tops.12488</pub-id><pub-id pub-id-type="pmid">32023006</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Conway-Smith</surname> <given-names>B.</given-names></name> <name><surname>West</surname> <given-names>R. L.</given-names></name> <name><surname>Mylopoulos</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Metacognitive skill: how it is acquired,&#x0201D;</article-title> in <source>Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 45</source> (<publisher-loc>Sydney, NSW</publisher-loc>).</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Conway-Smith</surname> <given-names>K.</given-names></name> <name><surname>West</surname> <given-names>R. L.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Clarifying system 1 &#x00026; 2 through the common model of cognition,&#x0201D;</article-title> in <source>Proceedings of the 20th International Conference on Cognitive Modelling</source> (<publisher-loc>Toronto, ON</publisher-loc>), <fpage>40</fpage>&#x02013;<lpage>45</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Csikszentmihalyi</surname> <given-names>M.</given-names></name></person-group> (<year>1990</year>). <source>Flow: The Psychology of Optimal Experience</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Harper &#x00026; Row</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Damasio</surname> <given-names>A. R.</given-names></name></person-group> (<year>2003</year>). <source>Looking for Spinoza: Joy, Sorrow, and the Feeling Brain</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Houghton Mifflin Harcourt</publisher-name>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dancy</surname> <given-names>C. L.</given-names></name> <name><surname>Ritter</surname> <given-names>F. E.</given-names></name> <name><surname>Berry</surname> <given-names>K. A.</given-names></name> <name><surname>Klein</surname> <given-names>L. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Using a cognitive architecture with a physiological substrate to represent effects of a psychological stressor on cognition</article-title>. <source>Comput. Math. Organ. Theory</source> <volume>21</volume>, <fpage>90</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1007/s10588-014-9178-1</pub-id><pub-id pub-id-type="pmid">26162004</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Evans</surname> <given-names>J. S.</given-names></name></person-group> (<year>2003</year>). <article-title>In two minds: dual-process accounts of reasoning</article-title>. <source>Trends Cogn. Sci</source>. <volume>7</volume>, <fpage>454</fpage>&#x02013;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2003.08.012</pub-id><pub-id pub-id-type="pmid">14550493</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci</source>. <volume>11</volume>, <fpage>127</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id><pub-id pub-id-type="pmid">20068583</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>W.</given-names></name> <name><surname>Anderson</surname> <given-names>J. R.</given-names></name></person-group> (<year>2006</year>). <article-title>From recurrent choice to skill learning: a reinforcement-learning model</article-title>. <source>J. Exp. Psychol. Gen</source>. <volume>135</volume>, <fpage>184</fpage>&#x02013;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1037/0096-3445.135.2.184</pub-id><pub-id pub-id-type="pmid">16719650</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gonzalez</surname> <given-names>C.</given-names></name> <name><surname>Lerch</surname> <given-names>J. F.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name></person-group> (<year>2003</year>). <article-title>Instance-based learning in dynamic decision making</article-title>. <source>Cogn. Sci</source>. <volume>27</volume>, <fpage>591</fpage>&#x02013;<lpage>635</lpage>. <pub-id pub-id-type="doi">10.1207/s15516709cog2704_2</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gunzelmann</surname> <given-names>G.</given-names></name> <name><surname>Byrne</surname> <given-names>M. D.</given-names></name> <name><surname>Gluck</surname> <given-names>K. A.</given-names></name> <name><surname>Moore Jr</surname> <given-names>L. R.</given-names></name></person-group> (<year>2009</year>). <article-title>Using computational cognitive modeling to predict dual-task performance with sleep deprivation</article-title>. <source>Hum. Factors</source> <volume>51</volume>, <fpage>251</fpage>&#x02013;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1177/0018720809334592</pub-id><pub-id pub-id-type="pmid">19653487</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huizinga</surname> <given-names>J.</given-names></name></person-group> (<year>1939</year>). <source>Homo Ludens Versuch einer Bestimmung des Spielelementest der Kultur</source>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Pantheon</publisher-name>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huys</surname> <given-names>Q. J.</given-names></name> <name><surname>Maia</surname> <given-names>T. V.</given-names></name> <name><surname>Frank</surname> <given-names>M. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Computational psychiatry as a bridge from neuroscience to clinical applications</article-title>. <source>Nat. Neurosci</source>. <volume>19</volume>, <fpage>404</fpage>&#x02013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4238</pub-id><pub-id pub-id-type="pmid">26906507</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Juvina</surname> <given-names>I.</given-names></name> <name><surname>Larue</surname> <given-names>O.</given-names></name> <name><surname>Hough</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Modeling valuation and core affect in a cognitive architecture: the impact of valence and arousal on memory and decision-making</article-title>. <source>Cogn. Syst. Res</source>. <volume>48</volume>, <fpage>4</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogsys.2017.06.002</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kahneman</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). <source>Thinking, Fast and Slow</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Macmillan</publisher-name>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Koster</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <source>Theory of Fun for Game Design</source>. <publisher-loc>Sebastopol, CA</publisher-loc>: <publisher-name>O&#x00027;Reilly Media</publisher-name>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kotseruba</surname> <given-names>I.</given-names></name> <name><surname>Tsotsos</surname> <given-names>J. K.</given-names></name></person-group> (<year>2020</year>). <article-title>40 years of cognitive architectures: core cognitive abilities and practical applications</article-title>. <source>Artif. Intell. Rev</source>. <volume>53</volume>, <fpage>17</fpage>&#x02013;<lpage>94</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-018-9646-y</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laird</surname> <given-names>J. E.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Rosenbloom</surname> <given-names>P. S.</given-names></name></person-group> (<year>2017</year>). <article-title>A standard model of the mind: Toward a common computational framework across artificial intelligence, cognitive science, neuroscience, and robotics</article-title>. <source>AI Mag</source>. <volume>38</volume>, <fpage>13</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1609/aimag.v38i4.2744</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Gonzalez</surname> <given-names>C.</given-names></name> <name><surname>Martin</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Instance-based decision making model of repeated binary choice,&#x0201D;</article-title> in <source>Proceedings of the 8th International Conference on Cognitive Modelling</source> (<publisher-loc>Ann Arbor, MI</publisher-loc>), <fpage>67</fpage>&#x02013;<lpage>72</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeDoux</surname> <given-names>J. E.</given-names></name> <name><surname>Pine</surname> <given-names>D. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Using neuroscience to help understand fear and anxiety: a two-system framework</article-title>. <source>Am. J. Psychiatry</source> <volume>173</volume>, <fpage>1083</fpage>&#x02013;<lpage>1093</lpage>. <pub-id pub-id-type="doi">10.1176/appi.ajp.2016.16030353</pub-id><pub-id pub-id-type="pmid">27609244</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malone</surname> <given-names>T. W.</given-names></name></person-group> (<year>1981</year>). <article-title>Toward a theory of intrinsically motivating instruction</article-title>. <source>Cogn. Sci</source>. <volume>5</volume>, <fpage>333</fpage>&#x02013;<lpage>369</lpage>. <pub-id pub-id-type="doi">10.1016/S0364-0213(81)80017-1</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Badia</surname> <given-names>A. P.</given-names></name> <name><surname>Mirza</surname> <given-names>M.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Lillicrap</surname> <given-names>T.</given-names></name> <name><surname>Harley</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>&#x0201C;Asynchronous methods for deep reinforcement learning,&#x0201D;</article-title> in <source>Proceedings of The 33rd International Conference on Machine Learning</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>1928</fpage>&#x02013;<lpage>1937</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source> <volume>518</volume>, <fpage>529</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nagashima</surname> <given-names>K.</given-names></name> <name><surname>Morita</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Trait inference on cognitive model of curiosity: relationship between perceived intelligence and levels of processing,&#x0201D;</article-title> in <source>Proceedings of the 22th International Conference on Cognitive Modelling</source> (<publisher-loc>Tilburg</publisher-loc>).</citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nagashima</surname> <given-names>K.</given-names></name> <name><surname>Morita</surname> <given-names>J.</given-names></name> <name><surname>Takeuchi</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Curiosity as pattern matching: Simulating the effects of intrinsic rewards on the levels of processing,&#x0201D;</article-title> in <source>Proceedings of the 19th International Conference on Cognitive Modelling</source>, <fpage>197</fpage>&#x02013;<lpage>203</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nagashima</surname> <given-names>K.</given-names></name> <name><surname>Nishikawa</surname> <given-names>J.</given-names></name> <name><surname>Yoneda</surname> <given-names>R.</given-names></name> <name><surname>Morita</surname> <given-names>J.</given-names></name> <name><surname>Terada</surname> <given-names>T.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Modeling optimal arousal by integrating basic cognitive components,&#x0201D;</article-title> in <source>Proceedings of the 20th International Conference on Cognitive Modeling</source> (<publisher-loc>Toronto, ON</publisher-loc>), <fpage>196</fpage>&#x02013;<lpage>202</lpage>.<pub-id pub-id-type="pmid">35666011</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nishikawa</surname> <given-names>J.</given-names></name> <name><surname>Nagashima</surname> <given-names>K.</given-names></name> <name><surname>Yoneda</surname> <given-names>R.</given-names></name> <name><surname>Morita</surname> <given-names>J.</given-names></name> <name><surname>Terada</surname> <given-names>T.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Representing motivation in a simple perceptual and motor coordination task based on a goal activation mechanism,&#x0201D;</article-title> in <source>Advances in Cognitive Systems 2022 (ACS 2022)</source> (<publisher-loc>Chicago, IL</publisher-loc>), <fpage>102</fpage>&#x02013;<lpage>120</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pathak</surname> <given-names>D.</given-names></name> <name><surname>Agrawal</surname> <given-names>P.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Curiosity-driven exploration by self-supervised prediction,&#x0201D;</article-title> in <source>In Proceedings of the 34th International Conference on Machine Learning</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>2778</fpage>&#x02013;<lpage>2787</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2017.70</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pink</surname> <given-names>D. H.</given-names></name></person-group> (<year>2011</year>). <source>Drive: The Surprising Truth about What Motivates Us</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Penguin</publisher-name>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Raffaelli</surname> <given-names>Q.</given-names></name> <name><surname>Mills</surname> <given-names>C.</given-names></name> <name><surname>Christoff</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>The knowns and unknowns of boredom: a review of the literature</article-title>. <source>Exp. Brain Res</source>. <volume>236</volume>, <fpage>2451</fpage>&#x02013;<lpage>2462</lpage>. <pub-id pub-id-type="doi">10.1007/s00221-017-4922-7</pub-id><pub-id pub-id-type="pmid">28352947</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reitter</surname> <given-names>D.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>A cognitive model of spatial path-planning</article-title>. <source>Comput. Math. Organ. Theory</source> <volume>16</volume>, <fpage>220</fpage>&#x02013;<lpage>245</lpage>. <pub-id pub-id-type="doi">10.1007/s10588-010-9073-3</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ritter</surname> <given-names>F. E.</given-names></name> <name><surname>Tehranchi</surname> <given-names>F.</given-names></name> <name><surname>Oury</surname> <given-names>J. D.</given-names></name></person-group> (<year>2019</year>). <article-title>ACT-R: a cognitive architecture for modeling cognition</article-title>. <source>Wiley Interdiscip. Rev. Cogn. Sci</source>. <volume>10</volume>:<fpage>e1488</fpage>. <pub-id pub-id-type="doi">10.1002/wcs.1488</pub-id><pub-id pub-id-type="pmid">30536740</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rosenbloom</surname> <given-names>P.</given-names></name> <name><surname>Laird</surname> <given-names>J.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Stocco</surname> <given-names>A.</given-names></name> <name><surname>Granger</surname> <given-names>R.</given-names></name> <name><surname>Huyck</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>&#x0201C;A proposal for extending the common model of cognition to emotion,&#x0201D;</article-title> in <source>Proceedings of the 22th International Conference on Cognitive Modeling</source> (<publisher-loc>Tilburg</publisher-loc>).</citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Formal theory of creativity, fun, and intrinsic motivation (1990-2010)</article-title>. <source>IEEE Trans. Auton. Ment. Dev</source>. <volume>2</volume>, <fpage>230</fpage>&#x02013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1109/TAMD.2010.2056368</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name> <name><surname>Chentanez</surname> <given-names>N.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;Intrinsically motivated reinforcement learning,&#x0201D;</article-title> in <source>Proceedings of the 34th International Conference on Machine Learning</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>2778</fpage>&#x02013;<lpage>2787</lpage>. <pub-id pub-id-type="doi">10.21236/ADA440280</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spiro</surname> <given-names>R. J.</given-names></name> <name><surname>Feltovich</surname> <given-names>P. J.</given-names></name> <name><surname>Jacobson</surname> <given-names>M. J.</given-names></name> <name><surname>Coulson</surname> <given-names>R. L.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Cognitive flexibility, constructivism, and hypertext: random access instruction for advanced knowledge acquisition in ill-structured domains,&#x0201D;</article-title> in <source>Constructivism in Education</source>, eds. L. P. Steffe, and J. Gale (New York, NY: Routledge), <fpage>85</fpage>&#x02013;<lpage>107</lpage>.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stocco</surname> <given-names>A.</given-names></name> <name><surname>Sibert</surname> <given-names>C.</given-names></name> <name><surname>Steine-Hanson</surname> <given-names>Z.</given-names></name> <name><surname>Koh</surname> <given-names>N.</given-names></name> <name><surname>Laird</surname> <given-names>J. E.</given-names></name> <name><surname>Lebiere</surname> <given-names>C. J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Analysis of the human connectome data supports the notion of a &#x0201C;Common Model of Cognition&#x0201D; for human and human-like intelligence across domains</article-title>. <source>Neuroimage</source> <volume>235</volume>:<fpage>118035</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2021.118035</pub-id><pub-id pub-id-type="pmid">33838264</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source>Reinforcement Learning: An Introduction</source>. Cambridge: MIT Press. <pub-id pub-id-type="doi">10.1109/TNN.1998.712192</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taatgen</surname> <given-names>N. A.</given-names></name> <name><surname>Lee</surname> <given-names>F. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Production compilation: a simple mechanism to model complex skill acquisition</article-title>. <source>Hum. Factors</source> <volume>45</volume>, <fpage>61</fpage>&#x02013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.1518/hfes.45.1.61.27224</pub-id><pub-id pub-id-type="pmid">12916582</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toya</surname> <given-names>G.</given-names></name> <name><surname>Hashimoto</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Recursive combination has adaptability in diversifiability of production and material culture</article-title>. <source>Front. Psychol</source>. <volume>9</volume>:<fpage>1512</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2018.01512</pub-id><pub-id pub-id-type="pmid">30283369</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van der Velde</surname> <given-names>M.</given-names></name> <name><surname>Sense</surname> <given-names>F.</given-names></name> <name><surname>Borst</surname> <given-names>J. P.</given-names></name> <name><surname>van Maanen</surname> <given-names>L.</given-names></name> <name><surname>Van Rijn</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>Capturing dynamic performance in a cognitive model: estimating ACT-R memory parameters with the linear ballistic accumulator</article-title>. <source>Top. Cogn. Sci</source>. <volume>14</volume>, <fpage>889</fpage>&#x02013;<lpage>903</lpage>. <pub-id pub-id-type="doi">10.1111/tops.12614</pub-id><pub-id pub-id-type="pmid">35531959</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Vugt</surname> <given-names>M. K.</given-names></name> <name><surname>van der Velde</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>How does rumination impact cognition? a first mechanistic model</article-title>. <source>Top. Cogn. Sci</source>. <volume>10</volume>, <fpage>175</fpage>&#x02013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1111/tops.12318</pub-id><pub-id pub-id-type="pmid">29383884</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Stocco</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Recovering reliable idiographic biological parameters from noisy behavioral data: the case of basal ganglia indices in the probabilistic selection task</article-title>. <source>Comput. Brain Behav</source>. <volume>4</volume>, <fpage>318</fpage>&#x02013;<lpage>334</lpage>. <pub-id pub-id-type="doi">10.1007/s42113-021-00102-5</pub-id><pub-id pub-id-type="pmid">33782661</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y. C.</given-names></name> <name><surname>Stocco</surname> <given-names>A.</given-names></name></person-group> (<year>2024</year>). <article-title>Allocating mental effort in cognitive tasks: a model of motivation in the ACT-R cognitive architecture</article-title>. <source>Top. Cogn. Sci</source>. <volume>16</volume>, <fpage>74</fpage>&#x02013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1111/tops.12711</pub-id><pub-id pub-id-type="pmid">37986131</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yerkes</surname> <given-names>R. M.</given-names></name> <name><surname>Dodson</surname> <given-names>J. D.</given-names></name></person-group> (<year>1908</year>). <article-title>The relation of strength of stimulus to rapidity of habit-formation</article-title>. <source>J. Comp. Neurol. Psychol</source>. <volume>18</volume>, <fpage>459</fpage>&#x02013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1002/cne.920180503</pub-id></citation>
</ref>
</ref-list>
</back>
</article>