<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2023.1099593</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A neural active inference model of perceptual-motor learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Yang</surname> <given-names>Zhizhuo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1871536/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Diaz</surname> <given-names>Gabriel J.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1195907/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Fajen</surname> <given-names>Brett R.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/21163/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bailey</surname> <given-names>Reynold</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/927836/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ororbia</surname> <given-names>Alexander G.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2181318/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Golisano College of Computing and Information Sciences, Rochester Institute of Technology</institution>, <addr-line>Rochester, NY</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Chester F. Carlson Center for Imaging Science, Rochester Institute of Technology</institution>, <addr-line>Rochester, NY</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Cognitive Science, Rensselaer Polytechnic Institute</institution>, <addr-line>Troy, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mario Senden, Maastricht University, Netherlands</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Thomas Parr, University College London, United Kingdom; Francisco Barcel&#x000F3;, University of the Balearic Islands, Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Zhizhuo Yang &#x02709; <email>zy8981&#x00040;rit.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1099593</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Yang, Diaz, Fajen, Bailey and Ororbia.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Yang, Diaz, Fajen, Bailey and Ororbia</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The active inference framework (AIF) is a promising new computational framework grounded in contemporary neuroscience that can produce human-like behavior through reward-based learning. In this study, we test the ability for the AIF to capture the role of anticipation in the visual guidance of action in humans through the systematic investigation of a visual-motor task that has been well-explored&#x02014;that of intercepting a target moving over a ground plane. Previous research demonstrated that humans performing this task resorted to anticipatory changes in speed intended to compensate for semi-predictable changes in target speed later in the approach. To capture this behavior, our proposed &#x0201C;neural&#x0201D; AIF agent uses artificial neural networks to select actions on the basis of a very short term prediction of the information about the task environment that these actions would reveal along with a long-term estimate of the resulting cumulative expected free energy. Systematic variation revealed that anticipatory behavior emerged only when required by limitations on the agent&#x00027;s movement capabilities, and only when the agent was able to estimate accumulated free energy over sufficiently long durations into the future. In addition, we present a novel formulation of the <italic>prior mapping function</italic> that maps a multi-dimensional world-state to a uni-dimensional distribution of free-energy/reward. Together, these results demonstrate the use of AIF as a plausible model of anticipatory visually guided behavior in humans.</p>
</abstract>
<kwd-group>
<kwd>interception</kwd>
<kwd>locomotion</kwd>
<kwd>active inference</kwd>
<kwd>learning</kwd>
<kwd>anticipation</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="0"/>
<equation-count count="7"/>
<ref-count count="55"/>
<page-count count="13"/>
<word-count count="10580"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The active inference framework (AIF) (Friston et al., <xref ref-type="bibr" rid="B24">2009</xref>) is an emerging theory of neural encoding and processing that captures a wide range of cognitive, perceptual, and motor phenomena, while also offering a neurobiologically plausible means of conducting reward-based learning through the capacity to predict sensory information. The behavior of an AIF agent involves the selection of action-plans that span into the near future and centers around the learning of a probabilistic generative model of the world through interaction with the environment. Ultimately, the agent must take action such that it is making progress toward its goals (goal-seeking behavior) while also balancing the drive to explore and understand its environment (information maximizing behavior), adjusting the internal states of its world to better account for the evidence that it acquires over time. As a result, AIF unifies perception, action, and learning by framing them as processes that result from approximate Bayesian inference.</p>
<p>The AIF framework has been used to study a variety of reinforcement learning (RL) tasks, including the inverted pendulum problem (<italic>CartPole</italic>) (Millidge, <xref ref-type="bibr" rid="B36">2020</xref>; Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>), the mountain car problem (<italic>MountainCar</italic>) (Friston et al., <xref ref-type="bibr" rid="B24">2009</xref>; Ueltzh&#x000F6;ffer, <xref ref-type="bibr" rid="B49">2018</xref>; &#x000C7;atal et al., <xref ref-type="bibr" rid="B5">2020</xref>; Tschantz et al., <xref ref-type="bibr" rid="B47">2020a</xref>; Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>) and the frozen lake problem (<italic>Frozen Lake</italic>) (Sajid et al., <xref ref-type="bibr" rid="B43">2021</xref>). Each task places different demands on motor and cognitive abilities. For instance, <italic>CartPole</italic> requires online control of a paddle to balance a pole upright, whereas <italic>MountainCar</italic> requires intelligent exploration of the task environment; a simple &#x0201C;greedy&#x0201D; policy (typical of many modern-day RL approaches) would fail to solve the problem. The popular <italic>Frozen Lake</italic> requires skills related to spatial navigation and planning if the agent is to find the goal while avoiding unsafe states.</p>
<p>One fundamental aspect of human and animal behavior that has so far not been sufficiently studied from an active inference perspective is the on-line visual guidance of locomotion. On-line visual guidance comprises a class of ecologically important behaviors for which movements of the body are continuously regulated based on currently available visual information seen from the first-person perspective. Some of the most extensively studied tasks include steering toward a goal (Warren et al., <xref ref-type="bibr" rid="B52">2010</xref>), negotiating complex terrain on foot (Matthis and Fajen, <xref ref-type="bibr" rid="B35">2013</xref>; Diaz et al., <xref ref-type="bibr" rid="B10">2018</xref>), intercepting moving targets (Fajen and Warren, <xref ref-type="bibr" rid="B14">2007</xref>), braking to avoid a collision (Yilmaz and Warren, <xref ref-type="bibr" rid="B54">1995</xref>; Fajen and Devaney, <xref ref-type="bibr" rid="B13">2006</xref>), and intercepting a fly ball (Chapman, <xref ref-type="bibr" rid="B6">1968</xref>; Fajen et al., <xref ref-type="bibr" rid="B12">2008</xref>). For each of these tasks, researchers have formulated control strategies that capture the coupling of visual information and action.</p>
<p>One aspect of on-line visual guidance that AIF might be particularly well-suited to capture is anticipation. To successfully perform any of these kinds of tasks, actors must be able to regulate their actions in anticipation of future events. One approach to capturing anticipation in visual guidance is to identify sources of visual information that specify how the actor should move at the current instant in order to reach the goal in the future. For example, when running to intercept a moving target, the sufficiency of the interceptor&#x00027;s current speed is specified by the rate of change in the exocentric visual direction of the target, or <italic>bearing angle</italic> (<xref ref-type="fig" rid="F1">Figure 1</xref>). If the interceptor is able to move so as to maintain a constant bearing angle (CBA), then an interception is guaranteed. Such accounts of anticipation are appealing because they avoid the need for planning on the basis of predictions or extrapolations of the agent&#x00027;s or target&#x00027;s motion, thereby presumably requiring fewer cognitive resources for task execution. Similar accounts of anticipation in the context of locomotor control have been developed for fly ball catching (Chapman, <xref ref-type="bibr" rid="B6">1968</xref>) and braking (Lee, <xref ref-type="bibr" rid="B34">1976</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>A top-down view of the interception problem. The agent (triangle) and target (circle) approach the invisible interception point (square) by going straight ahead. &#x003A8; denotes the exocentric direction of the target (bearing angle) and &#x003B1; denotes the target&#x00027;s approach angle. Image adapted from Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0001.tif"/>
</fig>
<p>However, there are other aspects of anticipatory control that are more difficult to capture based on currently available information alone. For example, moving targets sometimes change speeds and directions in ways that are somewhat predictable, allowing actors to alter their movement in advance in anticipation of the most likely change in target motion. This was demonstrated in a previously published study in which subjects were instructed to adjust their self-motion speed while moving along a linear path in order to intercept a moving target that changed speed partway through each episode (Diaz et al., <xref ref-type="bibr" rid="B11">2009</xref>). Note that episode refers to a single, complete course of interception for the agent and the target to be compatible with the conventions used by the reinforcement learning community. In contrast, Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>) uses the word <italic>trial</italic>. The final target speed randomly varied between episodes such that the target usually accelerated but occasionally decelerated. In response, subjects quickly learned to adjust their speed during the first part of the episode in anticipation of the change in target speed that was most likely given past experience and the initial conditions of that episode.</p>
<p>Active inference offers a potentially useful framework for understanding and modeling this kind of anticipatory behavior. The behavior of an AIF agent involves the selection of action plans (or policies) that span into the near future. These plans are selected based on <italic>expected free energy</italic> (EFE), i.e., a reward signal that takes into account both the action&#x00027;s contribution to reaching a desired goal state (i.e., an <italic>instrumental</italic> component), and the new information gained by the action (i.e., an <italic>epistemic</italic> component). This method of action selection is ideal for the study of predictive and anticipatory behavior in that it allows for the selection of action plans that do not immediately contribute to task completion, but that reveal to the agent something previously unknown about how the agent&#x00027;s action affects the environment. Similarly, in the task presented in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>), the human participants learned that success required increasing speed early in the episode in order to increase the likelihood of an interception after the target&#x00027;s semi-predictable change in speed. Critically, this early change in speed was not motivated by currently available visual information, but rather by the positive reinforcement of actions selected in the process of task exploration.</p>
<p>In contrast to reinforcement learning methods, active inference (AIF) formulates action-driven learning and inference from a Bayesian, belief-based perspective (Parr and Friston, <xref ref-type="bibr" rid="B42">2019</xref>; Sajid et al., <xref ref-type="bibr" rid="B43">2021</xref>). Generally, AIF offers: (1) flexibility to define a prior preference (or preferred outcome) over the observation space (which pushes the agent to uncover goal-orienting policies), which provides an alternative to designing a reward function, (2) a principled treatment for epistemic exploration as a means of uncertainty reduction, information gain, and intrinsic motivation (Parr and Friston, <xref ref-type="bibr" rid="B40">2017</xref>, <xref ref-type="bibr" rid="B42">2019</xref>; Schwartenbeck et al., <xref ref-type="bibr" rid="B44">2019</xref>), and (3) an encompassing uncertainty or precision over the beliefs that the generative model of the AIF agent computes as a natural part of then agent&#x00027;s belief updating (Parr and Friston, <xref ref-type="bibr" rid="B40">2017</xref>). Despite being a popular and powerful framework of perception, action (Friston, <xref ref-type="bibr" rid="B17">2009</xref>, <xref ref-type="bibr" rid="B18">2010</xref>; Buckley et al., <xref ref-type="bibr" rid="B4">2017</xref>; Friston K. et al., <xref ref-type="bibr" rid="B21">2017</xref>), decision-making and planning (Kaplan and Friston, <xref ref-type="bibr" rid="B30">2018</xref>; Parr and Friston, <xref ref-type="bibr" rid="B41">2018</xref>) with biological plausibility, AIF has been mostly applied to problems with a low-dimensionality and often discrete state space and actions (Friston et al., <xref ref-type="bibr" rid="B23">2012</xref>, <xref ref-type="bibr" rid="B22">2015</xref>, <xref ref-type="bibr" rid="B26">2018</xref>; Friston K. et al., <xref ref-type="bibr" rid="B21">2017</xref>; Friston K. J. et al., <xref ref-type="bibr" rid="B25">2017</xref>). One of the key limitations is that calculation of the EFE values for all policies starting from the current time step is needed in order to select the optimal action at the immediate time step. The exact EFE calculation becomes intractable quickly as the size of the action space <inline-formula><mml:math id="M1"><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">A</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula> and the planning time horizon <italic>H</italic> grows (Millidge, <xref ref-type="bibr" rid="B36">2020</xref>; Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>). We refer to Da Costa et al. (<xref ref-type="bibr" rid="B7">2020</xref>) for a comprehensive review on AIF.</p>
<p>The present study makes several specific contributions to the understanding of visually guided action and active inference:</p>
<list list-type="bullet">
<list-item><p>We present a novel model for locomotor interception of a target that changes speeds semi-predictably, as in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). This model is a scaled-up version of AIF where EFE is treated as a negative value function in reinforcement learning (RL) (Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>) and deep RL methodology is utilized to scale AIF to solve tasks such as locomotor interception with continuous state spaces. Specifically, our method predicts action-conditioned EFE values with a <italic>joint</italic> network (see Section 2.4.2) and by bootstrapping on the continuous observation space over a long time horizon. This allows the agent to account for the long-term effects of its current chosen action(s).</p></list-item>
<list-item><p>To calculate the <italic>instrumental</italic> value, we designed a problem-specific <italic>prior mapping function</italic> to convert the original observations into a one-dimensional prior space where a prior preference can be (more easily) specified. This allows us to inject domain knowledge into the <italic>instrumental</italic> reward. The <italic>instrumental</italic> measurements in prior space simultaneously promote interpretability as well as computationally efficient task performance.</p></list-item>
<list-item><p>We present a comparison of task performance of a baseline deep-Q network (DQN) agent, or an AIF agent in which EFE is computed using only the <italic>instrumental</italic> signal/component, with a full AIF agent in which EFE is computed using both <italic>instrumental</italic> and <italic>epistemic</italic> signals/components.</p></list-item>
<list-item><p>We demonstrate behavioral differences among our full AIF agent under the influence of two varying parameters: the discount factor &#x003B3;, which describes the weight on future accumulated quantities when calculating EFE value at each time step, and pedal lag coefficient <italic>K</italic>, which specifies how responsive changes in pedal position is reflected on agent&#x00027;s speed (or the amount of inertia that is associated with the agent&#x00027;s vehicle).</p></list-item>
<list-item><p>We interpret our findings as a model for anticipation in the context of visually guided action as well as in terms of specific contributions to the active inference and machine learning communities.</p></list-item>
</list>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and methods</title>
<p>Our aim in this study was to develop an agent that selects from a set of discrete actions in order to perform the task of interception. In this section, we describe the task that we aim to solve as well as formally describe the AIF model designed to tackle it. We start with the problem formulation and brief notation and definitions, then move on to describe our proposed agent&#x00027;s inference and learning dynamics.</p>
<sec>
<title>2.1. The perceptual-motor problem: Intercepting a moving target</title>
<p>We designed and simulated a perception-motor problem based on the human interception task used by Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). In the original study, subjects sat in front of a large rear-projection screen depicting an open field with a heavily textured ground plane. The subject&#x00027;s task was to intercept a moving spherical target by controlling the speed of self-movement along a linear trajectory with a foot pedal, the position of which was mapped onto speed according to a first-order lag. Subjects began each episode from a stationary position at an initial distance sampled uniformly from between 25 and 30 m from the interception point. The spherical target approached the subject&#x00027;s path at one of three initial speeds (11.25, 9.47, and 8.18 m/s). Between 2.5 and 3.25 s after the episode began, the target changed speeds linearly by an amount that was sampled from a normal distribution of possible final speeds. The mean of the distribution was 15 m/s such that target speed usually increased, but occasionally decreased (standard deviation was 5 m/s, final speed is truncated by one standard deviation from the mean). The change of target speed takes exactly 500 ms.</p>
<p>This interception problem is difficult because a human or agent that is purely reactive to the likely change in target speed will often arrive at the interception point after the target (e.g., they will be too slow). The problem is exacerbated when the agent&#x00027;s vehicle is less responsive. In Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>), subjects were found to increase their speed during the early part of the episode in order to anticipate the most likely change in target speed, which helped them perform at near optimal levels. Differences between the behavior of subjects and the ideal pursuer were also found under some conditions. Findings in the original study further yielded insight into the strategies that humans adopt when dealing with uncertainty in realistic interception tasks.</p>
</sec>
<sec>
<title>2.2. Notation</title>
<p>We next define the notation and mathematical operators that we will use throughout the rest of this paper. &#x02299; indicates a Hadamard product, &#x000B7; indicates a matrix/vector multiplication (or dot product if the two objects it is applied to are vectors of the same shape), and (<bold>v</bold>)<sup><italic>T</italic></sup> denotes the transpose. ||<bold>v</bold>||<sub><italic>p</italic></sub> is used to represent the <italic>p</italic>-norm where <italic>p</italic> &#x0003D; 2 results in the 2-norm or Euclidean (L2) distance.</p>
</sec>
<sec>
<title>2.3. Action and input space specification</title>
<p>To simplify the problem for this work, we assume that the mapping between environmental (latent) states and observations is the identity matrix. Furthermore, we formulate the problem as a Markov Decision Process (MDP) with a discrete action space. The action space <bold>a</bold><sub><italic>t</italic></sub> (action vector at time <italic>t</italic>) is defined as a one-hot vector <bold>a</bold> &#x02208; {0, 1}<sup>6&#x000D7;1</sup>, where each dimension corresponds to a unique action and the actions are mutually exclusive. Each dimension corresponds to one of the pedal speeds (m/s) in {2, 4, 8, 10, 12, 14} respectively. Once a pedal speed is selected, the agent will change its own speed by the amount of &#x00394;<italic>V</italic> &#x0003D; <italic>K</italic>&#x0002A;(<italic>V</italic><sub><italic>p</italic></sub>&#x02212;<italic>V</italic><sub><italic>s</italic></sub>) in one time step where <italic>V</italic><sub><italic>p</italic></sub> is pedal speed, <italic>V</italic><sub><italic>s</italic></sub> is current subject speed and <italic>K</italic> is a constant lag coefficient. In this study, we experiment with 2 variants of peal lag coefficient, i.e., <italic>K</italic> &#x0003D; 1.0<italic>K&#x00027;</italic> and <italic>K</italic> &#x0003D; 0.5<italic>K&#x00027;</italic>. <italic>K&#x00027;</italic> is set to 0.017 to be consistent with the original study (Diaz et al., <xref ref-type="bibr" rid="B11">2009</xref>) and provides a smooth relationship between the pedal movement and vehicle speed change. Similar to Tschantz et al. (<xref ref-type="bibr" rid="B48">2020b</xref>), we assume that the control state vector (which, in AIF, control states are originally treated separately from action states) lines up one-to-one with the action vector, meaning that it too is a vector of the form <bold>u</bold> &#x02208; {0, 1}<sup>6&#x000D7;1</sup>. We define the observation/state space (<inline-formula><mml:math id="M2"><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>4</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) to be a 4-dimensional vector <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which corresponds to target distance, target speed, subject distance and subject speed. All distances aforementioned are with respect to the invisible interception point.</p>
</sec>
<sec>
<title>2.4. Neural active inference</title>
<p>Active inference (AIF) is a Bayesian computational framework that brings together perception and action under one single imperative: minimizing <italic>free energy</italic>. It accounts for how self-organizing agents operate in dynamic, non-stationary environments (Friston, <xref ref-type="bibr" rid="B19">2019</xref>), offering an alternative to standard, reward function-centric reinforcement learning (RL). In this study, we craft a simple AIF agent that resembles Q-learning (Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>) where the <italic>expected free energy</italic> (EFE) serves the role of a negative action-value function in RL. We frame the definition of EFE in the context of a stochastic policy and cast action-conditioned EFE as a negative action-value using a policy &#x003D5; &#x0003D; &#x003D5;(<italic>a</italic><sub><italic>t</italic></sub>|<bold>s</bold><sub><italic>t</italic></sub>) (where <bold>s</bold><sub><italic>t</italic></sub> &#x0003D; <bold>o</bold><sub><italic>t</italic></sub> as per our assumption earlier). The same policy &#x003D5; is used for each future time step &#x003C4;, and the probability distribution over the first-step action is separated from &#x003D5; resulting in a substitution distribution <italic>q</italic>(<italic>a</italic><sub><italic>t</italic></sub>) for &#x003D5;(<italic>a</italic><sub><italic>t</italic></sub>). Therefore, the one-step substituted EFE can be interpreted as the EFE of a policy of (<italic>q</italic>(<italic>a</italic><sub><italic>t</italic></sub>), &#x003D5;(<italic>a</italic><sub><italic>t</italic>&#x0002B;1</sub>), &#x02026;, &#x003D5;(<italic>a</italic><sub><italic>T</italic></sub>))).</p>
<p>Following Shin et al. (<xref ref-type="bibr" rid="B45">2022</xref>), we consider the deterministic optimal policy &#x003D5;<sup>&#x0002A;</sup> which always seeks an action with a minimum EFE and obtain the following EFE definition:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>According to Shin et al. (<xref ref-type="bibr" rid="B45">2022</xref>), the equation above is quite similar to the Bellman optimality equation, where <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> corresponds to the state-value function <inline-formula><mml:math id="M6"><mml:msup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:munder><mml:msup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> corresponds to the action-value function <inline-formula><mml:math id="M8"><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Then the first term <inline-formula><mml:math id="M9"><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> can be treated as a one-step negative reward and thus EFE can be treated as a negative value function. This term can then be further decomposed into two components:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mstyle displaystyle="true"><mml:munder accentunder="false"><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x0FE38;</mml:mo></mml:munder></mml:mstyle></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">Instrumental</mml:mtext></mml:mrow></mml:munder></mml:mstyle><mml:mo>&#x0002B;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mstyle displaystyle="true"><mml:munder accentunder="false"><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x0FE38;</mml:mo></mml:munder></mml:mstyle></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">Epistemic</mml:mtext></mml:mrow></mml:munder></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>To connect the formulation above back to AIF, with the term rephrased as <italic>R</italic><sub><italic>t, i</italic></sub> is the instrumental (also known as <italic>extrinsic, pragmatic</italic> or <italic>goal-directed</italic>) component (Tschantz et al., <xref ref-type="bibr" rid="B48">2020b</xref>), which measures the similarity between the future outcome following the policy &#x003D5; and preferred outcome (or prior preference). The term rephrased as <italic>R</italic><sub><italic>t, e</italic></sub> is known as the epistemic (also known as <italic>intrinsic, uncertainty-reducing</italic> or <italic>information-seeking</italic>) component (Tschantz et al., <xref ref-type="bibr" rid="B48">2020b</xref>), which measures the prediction error between the estimation of the future state by the transition model and the state predicted by the encoder given the actual observation from the environment.</p>
<p>Ultimately, we simplify and approximate the search for optimal EFE values by adapting an estimation approach based on the Bellman equation, arriving at a Q-learning bootstrap scheme. We assume that the outcome/observation can be set equal to state variables and, as a result, our generative model is designed with respect to fully observed environment (Tschantz et al., <xref ref-type="bibr" rid="B47">2020a</xref>). Following the active inference literature, we adopt the Laplace assumption and mean-field approximation. Therefore, a fixed identity covariance matrix is used for the likelihood distribution <italic>p</italic>(<italic>o</italic>|<italic>s</italic>). Our model (the function approximator) outputs the mean of states, which encodes the belief that there is a direct mapping between outcomes and states. Similarly, our model outputs the mean of estimated EFE values. Following (Mnih et al., <xref ref-type="bibr" rid="B37">2015</xref>), we integrated an experience replay as well as a target network in order to facilitate learning and improve sample efficiency. Note that the Q-learning style framing of negative EFE estimation is referred to as G-learning. Our model estimates the EFE for each possible action that it could take in the immediate next time step (i.e., time <italic>t</italic>&#x0002B;1) then selects the action that corresponds to the minimal EFE value. This, in effect, corresponds to only explicitly calculating the EFE over a horizon of 1 (whereas as planning over horizons &#x0003E;1 quickly become prohibitive, requiring expensive search methods such as Monte Carlo tree search) but incorporates a bootstrap estimate of future EFE values via the G-learning setup. Our definition in Equation (1) is similar to the EFE definition in Friston et al. (<xref ref-type="bibr" rid="B20">2021</xref>) in the sense that EFE is formulated recursively in both works. However, differences between our method and sophisticated inference (Friston et al., <xref ref-type="bibr" rid="B20">2021</xref>) still exist. For instance, our method works with continuous state space whereas sophisticated inference works with a discrete state space. Our method displays a connection to Q-learning, thus it is able to plan over a trajectory of arbitrary length in principle using bootstrap estimation, whereas sophisticated inference terminates the evaluation of recursive EFE whenever an action is found as unlikely or an outcome is implausible. We utilize the AIF framework within the G-learning framing for the interception task and modify the framework to fit the interception task, see <xref ref-type="fig" rid="F2">Figure 2</xref>. Spatial variables, i.e., distance and speed, will serve as the inputs to our framework and, as mentioned before, an identity mapping is assumed to connect the observation directly to the state variables (allowing us to avoid having to learn additional parameterized encoder/decoder functions). As a result, the AIF agent we designed for this paper&#x00027;s experiments consists of two major components: a <italic>prior mapping function</italic> and a multi-headed joint neural model.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Our neural AIF architecture for the interception task. The joint model is a two-headed artificial neural network which consists of shared hidden layers, an EFE (estimation) head, and a transition dynamics (prediction) head. The EFE head estimates EFE values for all possible actions given the current (latent/hidden) state. An action which is associated with maximum EFE value is selected and executed in the environment and the resulting observation is fed into the <italic>prior mapping function</italic> where the <italic>instrumental</italic> value <italic>R</italic><sub><italic>t,i</italic></sub> is calculated in prior space. Meanwhile, the transition dynamics head predicts the resulting observation given the current (latent/hidden) state. The error between the predicted and actual observation at <italic>t</italic> &#x0002B; 1 forms the <italic>epistemic</italic> value <italic>R</italic><sub><italic>t,e</italic></sub>. The summation of <italic>R</italic><sub><italic>t,i</italic></sub> and <italic>R</italic><sub><italic>t,e</italic></sub> results in the final EFE (target) value.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0002.tif"/>
</fig>
<p>Notably, our particular proposed joint model works jointly as a function approximator of EFE values as well as a forward dynamics predictor. It takes in the current observation <bold>o</bold><sub><italic>t</italic></sub> as input and then conducts, jointly, action selection and next-state prediction (as well as <italic>epistemic</italic> value estimation). The selected action is executed and the resulting observation is returned by the environment. The <italic>prior mapping function</italic> itself takes in as input the next observation <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub>, the consequence/result of the agent&#x00027;s currently selected action, and calculates the log likelihood of the preferred/prior distribution (set according to expert knowledge related to the problem), or the <italic>instrumental</italic> term <italic>R</italic><sub><italic>t, i</italic></sub>. The squared difference between the outcome of the selected action <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub> and its estimation <inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> (as per the generative transition component of our model) forms the <italic>epistemic</italic> term <italic>R</italic><sub><italic>t, e</italic></sub> as shown in Equation 7. The summation of the <italic>instrumental</italic> and <italic>epistemic</italic> terms forms the G-value (or negative EFE value) which is ultimately used to train / adapt the joint model. Formally, <inline-formula><mml:math id="M12"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. We explain each component in detail below.</p>
<sec>
<title>2.4.1. The prior mapping function and prior space</title>
<p>With the ability and freedom of designing a prior preference (or distribution over problem goal states or preferred outcomes) afforded by AIF, we integrate domain knowledge of the interception task into the design of a <italic>prior mapping function</italic>. In essence, our designed <italic>prior mapping function</italic> transforms the original observation vector <bold>o</bold><sub><italic>t</italic></sub> to a lower-dimensional space (the prior space) where a semantically meaningful variable is calculated and prior preference distribution is specified over this new variable&#x02014;in our case, this is set to be the <italic>speed difference</italic>, as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The <italic>speed difference</italic> represents the difference between the agent&#x00027;s speed after taking the selected action and the speed required for successful interception, i.e., <italic>speed difference</italic> &#x0003D; <italic>speed</italic><sub><italic>agent</italic></sub>&#x02212;<italic>speed</italic><sub><italic>required</italic></sub>. Given the current observation, the required speed is calculated as the agent&#x00027;s distance to the interception point divided by the first-order target time-to-contact (TTC). We define TTC as the duration for the target or agent to reach the theoretical interception point from the current time step regardless of the success of the actual interception task. Then, target&#x00027;s first-order TTC is the amount of time that it would take for the target to reach the interception point assuming that target speed does not change throughout the episode, i.e., <italic>TTC</italic><sub><italic>first</italic>&#x02212;<italic>order</italic></sub> &#x0003D; <italic>x</italic><sub><italic>t</italic></sub>/<italic>v</italic><sub><italic>constant</italic></sub> where <italic>x</italic><sub><italic>t</italic></sub> is the target distance and <italic>v</italic><sub><italic>constant</italic></sub> is the target constant speed. In our neural AIF framework, the <italic>instrumental</italic> values are calculated given all possible actions (blue circles in <xref ref-type="fig" rid="F3">Figure 3</xref>) and a prior distribution over <italic>speed difference</italic>. The smaller the absolute <italic>speed difference</italic> associated with a particular action, the higher the <italic>instrumental</italic> value <italic>prior mapping function</italic> assigns.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The prior preference specified in the prior space where each action corresponds to a different <italic>instrumental</italic> value. Circles correspond to pedal positions to choose from.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0003.tif"/>
</fig>
<p>Note that the agent might not have enough time to adjust its speed later in the interception task if it only follows the guidance of this <italic>prior mapping function</italic> without anticipating the likely future speed change of target, since this <italic>prior mapping function</italic> only accounts/embodies first-order information. To overcome this limitation, we investigated the effects of discounted long-term EFE value on the behavior of the agent in Section 3.4.</p>
</sec>
<sec>
<title>2.4.2. Joint model</title>
<p>Our proposed joint model embodies two key functionalities: EFE estimation and transition dynamics prediction, which are typically implemented as separate artificial neural networks (ANNs) in earlier AIF studies (Shin et al., <xref ref-type="bibr" rid="B45">2022</xref>) (in contrast, we found that, during preliminary experimentation, that a joint, fused architecture improved both the agent&#x00027;s overall generalization ability as well as its training stability). Concretely, we implement the joint model as a multi-headed ANN with an EFE head and a transition head (see <xref ref-type="fig" rid="F2">Figure 2</xref>). The system takes in the current observation <bold>o</bold><sub><italic>t</italic></sub> and predicts: (1) the EFE values for all possible actions, and (2) a future observation at a distance <bold>o</bold><sub><italic>t</italic>&#x0002B;<italic>D</italic></sub> (in this work, we fix the temporal distance to be one step, i.e., <italic>D</italic> &#x0003D; 1). Within the joint model, the current observation <bold>o</bold><sub><italic>t</italic></sub> is taken as input and a latent hidden activity vector <bold>z</bold><sub><italic>t</italic></sub> is produced, which is then provided to both output heads as input. The transition head <italic>p</italic>(<bold>o</bold><sub><italic>t</italic>&#x0002B;<italic>D</italic></sub>|<bold>z</bold><sub><italic>t</italic></sub>) serves as a generative model (or a forward dynamics model) and the EFE head <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> represents an approximation over the EFE values. As a result, EFE module and transition modules are wired together such that the prediction of the future observation <bold>o</bold><sub><italic>t</italic>&#x0002B;<italic>D</italic></sub> and the estimation of EFE values <bold>G</bold><sub><italic>t</italic>&#x0002B;<italic>D</italic></sub> are driven by the shared encoding from the topmost (hidden) layer of the joint model. This enables the sharing of underlying knowledge between the module selecting actions and the module predicting the outcome(s) of an action. Our intuition is that we humans tend to evaluate the &#x0201C;value&#x0201D; of an action by the consequences that it produces.</p>
<p>We next formally describe the dynamics of our joint model, including both its inference and learning processes.</p>
<p><bold>Inference</bold> In general, our agent is meant to produce an action conditioned on observations (or states) sampled from the environment at particular time-steps. Specifically, within any given <italic>T</italic>-step episode, our agent receives as input the observation <inline-formula><mml:math id="M14"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, where <italic>D</italic> is the dimensionality of the observation space <bold>o</bold><sub><italic>t</italic></sub> (<italic>D</italic> &#x0003D; 4 for the problem investigated in this study). The agent then produces a set of approximate free energy values, one for each action (similar in spirit to Q-values) as well as a prediction of the next observation that it is to receive from its environment (i.e., the perceptual consequence of its selected action).</p>
<p>Formally, in this work, the outputs described above are ultimately produced by a multi-output function <inline-formula><mml:math id="M15"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, implemented as a multi-layer perceptron (MLP), where <inline-formula><mml:math id="M16"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> contains estimated expected free energy values (one per discrete action) while <inline-formula><mml:math id="M17"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is the generative component&#x00027;s estimation of the next incoming observation <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub>. Note that we denote only outputting an action value set from this model as <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> (using only the action output head) and only outputting an observation prediction as <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> (using only the state prediction head). This MLP is parameterized by a set of synaptic weight matrices <inline-formula><mml:math id="M20"><mml:mi>&#x00398;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, that operates according to the following:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M21"><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mn>1</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mn>1</mml:mn></mml:msup><mml:mo>&#x000B7;</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>0</mml:mn></mml:mstyle></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x000B7;</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>1</mml:mn></mml:mstyle></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M22"><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>a</mml:mi><mml:mn>3</mml:mn></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>a</mml:mi><mml:mn>3</mml:mn></mml:msubsup><mml:mo>&#x000B7;</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>o</mml:mi><mml:mn>3</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>W</mml:mi></mml:mstyle><mml:mi>o</mml:mi><mml:mn>3</mml:mn></mml:msubsup><mml:mo>&#x000B7;</mml:mo><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:math></disp-formula>
<p>Where <inline-formula><mml:math id="M23"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (the input layer to our model is the observation at <italic>t</italic>). Note that a single discrete action is read out/chosen from our agent function&#x00027;s action output head as: <inline-formula><mml:math id="M24"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo class="qopname">argmax</mml:mo></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. The linear rectifier &#x003D5;<sub><italic>z</italic></sub>(<bold>v</bold>) &#x0003D; max(0, <bold>v</bold>) was chosen to be the activation function applied to the internal layers of our model while &#x003D5;<sub><italic>a</italic></sub>(<bold>v</bold>) &#x0003D; <bold>v</bold> (the identity) is the function specifically applied to the action neural activity layer <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and &#x003D5;<sub><italic>o</italic></sub>(<bold>v</bold>) &#x0003D; <bold>v</bold> is the function applied to predicted observation layer neurons. Note that the first hidden layer <inline-formula><mml:math id="M26"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> contains <italic>J</italic><sub>1</sub> neurons and <inline-formula><mml:math id="M27"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> contains <italic>J</italic><sub>2</sub> neurons, respectively. The action output layer <inline-formula><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> contains <italic>A</italic> neurons (<italic>A</italic> &#x0003D; 6 for the problem investigated in this study), one neuron per discrete action (out of <italic>A</italic> total possible actions as defined by the environment/problem), while the observation prediction layer <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> contains <italic>D</italic> &#x0003D; 4 neurons, making it the same dimensionality/shape as the observation space.</p>
<p><bold>Learning</bold> While there are many possible ways to adjust the values inside of &#x00398;, we opted to design a cost function and calculate the gradients of this objective with respect to the synaptic weight matrices of our model for the sake of simulation speed. The cost function that we designed to train our full agent was multi-objective in nature and is defined in the following manner:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(6)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M32"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where the target value for the action output head is calculated as <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> while the target action vector is computed as <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>a</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>a</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02299;</mml:mo><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. In the above set of equations, we see that the MLP model&#x00027;s weights are adjusted so as to minimize the linear combination of two terms, the cost associated with the difference between a target vector <bold>t</bold>, which contains the bootstrap-estimated of the EFE values, and the agent&#x00027;s original estimate <inline-formula><mml:math id="M35"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> as well as the cost associated with how far off the agent&#x00027;s prediction/expectation <inline-formula><mml:math id="M36"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> of its environment is from the actual observation <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub>. In this study, the standard deviation coefficients associated with both output layers are set to one, i.e., &#x003C3;<sub><italic>a</italic></sub> &#x0003D; &#x003C3;<sub><italic>o</italic></sub> &#x0003D; 1 (highlighting that we assume unit variance for our model&#x00027;s free energy estimates and its environmental state predictions&#x02014;note that a dynamic variance could be modeled by adding an additional output head responsible for computing the aleatoric uncertainty associated with <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub>).</p>
<p>Updating the parameters &#x00398; of the neural system then consists of computing the gradient <inline-formula><mml:math id="M37"><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>&#x00398;</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula> using reverse-mode differentiation and adjusting their values using a method such as stochastic gradient descent or variants, e.g., Adam (Kingma and Ba, <xref ref-type="bibr" rid="B31">2014</xref>), RMSprop (Tieleman et al., <xref ref-type="bibr" rid="B46">2012</xref>). Specifically, at each time step of any simulated episode, our agent first stores the current transition of the form (<bold>o</bold><sub><italic>t</italic></sub>, <bold>a</bold><sub><italic>t</italic></sub>, <italic>r</italic><sub><italic>t</italic></sub>, <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub>) into an episodic memory replay buffer (Mnih et al., <xref ref-type="bibr" rid="B37">2015</xref>) and then immediately calculates <inline-formula><mml:math id="M38"><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mo>;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>&#x00398;</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula> from a batch of observation/transition data (uniformly) sampled from the replay buffer, which stores up to 10<sup>5</sup> transitions. We will demonstrate the benefit of this design empirically in the results section.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Hypotheses for interception strategies</title>
<p>Given the fact that the target changes its speed during an episode in our interception task, the agent / human subject could gain advantage by anticipating the target speed change prior to the change of target speed. To select an optimal action early within the trail, the agent needs to take into consideration the initial target speed in the current episode and make adjustments based on the experience acquired from previous episodes. So, the question becomes: how does the agent adapt its behavior on the basis of current episode&#x00027;s observation of target speed/distance from the interception point and the learned statistics across episodes?</p>
</sec>
<sec>
<title>3.2. Experimental setup</title>
<p>We implemented the interception task as an environment in Python based on the OpenAI gym (Brockman et al., <xref ref-type="bibr" rid="B3">2016</xref>) library. This integration provides the full functionality and usability of the gym environment, which means that the environment can work / be used with any RL algorithm and is made accessible to the machine learning community as well. Our AIF agents and baseline algorithm DQN are implemented with the Tensorflow2 (Abadi et al., <xref ref-type="bibr" rid="B1">2015</xref>) library. Experimental data and code will be made publicly available upon acceptance.</p>
</sec>
<sec>
<title>3.3. Task performance</title>
<p>We compare AIF agents with and without the <italic>epistemic</italic> component and a baseline algorithm, i.e., a deep-Q network (DQN) (Mnih et al., <xref ref-type="bibr" rid="B37">2015</xref>). We define a trial as a computational experiment where the agent performs the interception task sequentially for a number of episodes. We run a number of trials and then calculate the mean and standard deviation across trials in order to obtain a statistically valid results. The simulations in our study set the update frequency of the task environment to be 60<italic>Hz</italic> in order to match the exact frequency of the original human study by Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). During each episode, the joint model receives an observation each time step at 60<italic>Hz</italic> and estimates the EFE value for each possible action. Finally, an action is selected based on the estimated EFE values and executed in the environment. This process corresponds to Section 2.4.2. Experiments are conducted for 20 trials where each trial contains 3000 episodes. The task performance of agents is shown as curves plotting window-averaged rewards (with a window size of 100 episodes) in <xref ref-type="fig" rid="F4">Figure 4</xref>, where the solid line depicts the mean value across trials and the shaded area represents standard deviation. We conducted a set of experiments where the discount factor &#x003B3; of the models and the pedal lag coefficient <italic>K</italic> were varied (note that, in AIF and RL research, &#x003B3; is typically fixed to a value between 0.9 and 1 to enable the model to account for long term returns). In order to compare the performance of our agents to that of human subjects, we apply the original pedal lag coefficient in one set of our experiments (specifically shown in <xref ref-type="fig" rid="F4">Figure 4C</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>(A&#x02013;D)</bold> Window-averaged reward measurements of agent performance on the interception task. <italic>DQN</italic>_<italic>Reward</italic> represents a DQN agent that utilized the sparse reward signal and &#x003F5;&#x02212;<italic>greedy</italic> exploration; <italic>AIF</italic>_<italic>InstOnly</italic> represents our AIF agent with only <italic>instrumental</italic> component which is defined by the <italic>prior mapping function</italic>; <italic>AIF</italic>_<italic>InstEpst</italic> represents an AIF agent that consists of both <italic>instrumental</italic> and <italic>epistemic</italic> components. Discount factor is denoted by &#x003B3;, pedal lag coefficient is denoted by <italic>K</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0004.tif"/>
</fig>
<p>Observe that our AIF agents are able to reach around a 90% success rate stably with very low variance. This beats human performance with 47% (std = 11.31) on average and 54.9% in the final block of experiments reported in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). The baseline DQN agent, which learns from the problem&#x00027;s sparse reward signal at the end of each episode, yields an average success rate of 22% at test time. Similarly, the AIF agent with both <italic>instrumental</italic> and <italic>epistemic</italic> components achieves a 90% mean success rate.</p>
<p>Note that the DQN agent is outperformed by the AIF agents trained with our customized prior preference function by a large margin. This reveals that the flexibility of injecting prior knowledge is crucial for solving complex tasks more efficiently and validates our motivation of applying AIF to cognitive tasks. In our preliminary experiments, we tested an AIF agent which consists of an EFE network and a transition network separately. This AIF agent is out-performed by the AIF agent with joint model in terms of windowed mean rewards and stability. Furthermore, the AIF agent with joint model has lower model complexity. Specifically, AIF agent with joint model has only 66.8% of the parameter counts of that of AIF agent with separate models. This supports our intuition that combining the EFE model with the transition model yields an overall better model agent.</p>
<p>Interestingly, the AIF agent with only <italic>instrumental</italic> component was able to nearly reach the same level of performance as the full AIF agent. However, success rate of this agent exhibited a larger variance than the full AIF agent. Based on comparison between agents with and without <italic>epistemic</italic> component, we argue that <italic>epistemic</italic> component serves, at least in the context of the interception task we investigate, as a regularizer for the AIF models, providing improved robustness. Since we apply experience replay and bootstrapping to train the AIF models, it is possible that a local minimum is reached in the optimization process because the replay buffer is filled up with samples which come from the same subspace as the state space. Therefore, with the help of <italic>epistemic</italic> component, the agent is encouraged to explore the environment more often and adjusts its prediction of future observations such that it has a higher chance of escaping poorer local optima. Our proposed AIF agent reaches a plateau in performance after about 1, 000 episodes and stabilizes more after 1, 500 episodes. Note that, in contrast, human subjects were able to perform the task at an average success rate after 9 episodes of initial practice (Diaz et al., <xref ref-type="bibr" rid="B11">2009</xref>).</p>
</sec>
<sec>
<title>3.4. Anticipatory behavior of AIF agents</title>
<p>Do the AIF agents exhibit a similar capacity for anticipatory behavior as humans do? To answer this question and to compare the strategy used by our AIF agents to that of human subjects, we record the Time-To-Contact (TTC) from trained AIF agents at the onset of the target&#x00027;s speed change in each episode. We then calculate, at the same time: 1) the target&#x00027;s TTC using first-order information, and 2) target&#x00027;s TTC with the assumption that the target would change its speed at the most likely time and reach an averaged final speed. Finally, we compose these three types of TTC data grouped by target initial speed into a single boxplot in <xref ref-type="fig" rid="F5">Figure 5</xref>. Following the assumptions made in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>), we expect that the agent would adjust its speed in a way such that its first-order TTC will equal the target&#x00027;s first-order TTC before it learns enough from experience to realize that the target almost always accelerates. The target&#x00027;s actual TTC with the interception point would be less than the first-order TTC if the target accelerates midway through. If the agent is able to anticipate the target&#x00027;s acceleration later in the episode, it should accelerate even before the target does in order to match the target&#x00027;s actual TTC with the interception point.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>(A&#x02013;D)</bold> TTC values taken at the onset of target&#x00027;s speed change. In each subplot, the target&#x00027;s first-order TTC, the target&#x00027;s actual mean TTC, and the agent&#x00027;s TTC are shown in different colors, with data grouped by target initial speed. The discount factor is denoted by &#x003B3; while the pedal lag coefficient is denoted by <italic>K</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0005.tif"/>
</fig>
<p>In our experimental analysis, we found that the discount factor &#x003B3; plays a big role in forming different behavior patterns within AIF agents. All variants of AIF agents were trained with the <italic>instrumental</italic> value computed using our first-order <italic>prior mapping function</italic>. Intuitively, the agent&#x00027;s behavior should conform to a reactive agent who uses only the first-order information and acts to match its own TTC to the target&#x00027;s first-order TTC, just like what has been observed in <xref ref-type="fig" rid="F5">Figure 5A</xref> (please see that the green box is nearly identical to the blue box under all target initial conditions). The AIF agent depicted in <xref ref-type="fig" rid="F5">Figure 5A</xref> is set to use a discount factor of 0, which means that the agent only seeks to maximize its immediate reward without considering the long-term impact of the action(s) that it selects. Such an agent converges to a reactive behavior. However, when we increase the discount factor to 0.99 (which is a common practice in RL literature), the AIF agent starts to behave more interestingly. In <xref ref-type="fig" rid="F5">Figure 5C</xref>, the agent&#x00027;s TTC (green box) lies in between target&#x00027;s first-order TTC (blue box) and target&#x00027;s actual mean TTC (orange box), which suggests that the AIF agent tends to move faster than a pure-reactive, first-order agent would in the early phase of interception. In other words, the agent tends to anticipate the likely target speed change in the future and adjusts its action selection policy. This behavioral pattern can be explained as exploiting the benefits provided by estimating long-term accumulated <italic>instrumental</italic> reward signal (when the discount factor value is increased). Given a higher discount factor, in this case &#x003B3; &#x0003D; 0.99, the AIF agent estimates the summation of <italic>instrumental</italic> values from its current (time) step in the task until the end of the interception using discounting. This leads to an agent who seeks to maximize long-term benefits in terms of reaching the goal when selecting actions.</p>
</sec>
<sec>
<title>3.5. Effect of vehicle dynamics on agent behavior</title>
<p>To test how anticipatory behavior is affected when simple reactive behavior is no longer sufficient, we increased the inertia on the agent&#x00027;s vehicle by changing the pedal lag coefficient <italic>K</italic>. Given the same discount factor &#x003B3; &#x0003D; 0.99, we compare two different pedal lag coefficients <italic>K</italic> &#x0003D; 1.0 in <xref ref-type="fig" rid="F5">Figure 5C</xref> and <italic>K</italic> &#x0003D; 0.5 in <xref ref-type="fig" rid="F5">Figure 5D</xref>, where lower <italic>K</italic> indicates less responsive vehicle dynamics. With the same discount factor, the AIF agent performing the task under a lower pedal lag coefficient in <xref ref-type="fig" rid="F5">Figure 5D</xref> has a lower success rate in intercepting the target. This is due to the fact that the agent&#x00027;s ability to manipulate its own speed is limited, therefore there is less room left for error. However, the AIF agent in this condition yields TTC values that are closer to the target&#x00027;s actual mean TTC. Note that, when the target initial speed is 11.25 <italic>m</italic>/<italic>s</italic> (<xref ref-type="fig" rid="F5">Figure 5D</xref>), the median of agent&#x00027;s TTC value is actually smaller than target&#x00027;s actual mean TTC. This supports our hypothesis that purely reactive behavior is not sufficient for successful interception and anticipatory behavior is emergent when the vehicle becomes less responsive.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>Variations of an AIF agent were trained to manipulate the speed of movement so as to intercept a target moving across the ground plane, and eventually across the agent&#x00027;s linear path of travel. On each episode, the target would change in speed on most episodes to a value that was selected from a Gaussian distribution of final speeds. The results demonstrate that the AIF framework is able to model both on-line visual and anticipatory control strategies in an interception task, as was previously demonstrated by humans performing the same task (Diaz et al., <xref ref-type="bibr" rid="B11">2009</xref>). The agent&#x00027;s anticipatory behavior aimed to maximize the cumulative expected free energy in the duration that follows action selection. Variation of the agent&#x00027;s discount factor modified the length of this duration. At lower discount factors, the agent behaved in a reactive manner throughout the approach, consistent with the constant bearing angle strategy of interception. At higher values, actions that were selected before the predictable change in speed took into account the most likely change in target speed that would occur later in the episode. Anticipatory behavior was also influenced by the agent&#x00027;s capabilities for action.This anticipatory behavior was most apparent when the pedal lag coefficient was set to lower values, which had the effect of changing the agent&#x00027;s movement dynamics so that purely reactive control was insufficient for interception behavior.</p>
<p>Despite the agent&#x00027;s demonstration of qualitatively human-like prediction, careful comparison of the agent&#x00027;s behavior to the human performance and learning rates demonstrated in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>) reveals notable differences. Analysis of participant behavior in the fourth and final block of Experiment 1 in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>) reveals that subject TTC at the onset of the target&#x00027;s change in speed was well matched to the most likely time and magnitude of the target&#x00027;s likely change in speed (i.e., the mean actual target TTC in <xref ref-type="fig" rid="F6">Figure 6</xref>). In contrast, the AIF agent with an equivalent pedal lag (<italic>K</italic> = 1.0; i.e., the <italic>matched</italic> agent) demonstrated only partial matching of its TTC to the likely change in target speed (the target&#x00027;s mean actual TTC in <xref ref-type="fig" rid="F5">Figure 5C</xref>). Although one might attribute this to under-training of the agent, it is notable that the agent achieved a hit rate exceeding 80% by the end of training, while human participants in the original study consistently improved in performance until reaching 55% hit rate at the end of the experiment.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Human subject data from Exp 1. of Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). TTCs were taken at the onset of target&#x00027;s speed change. Dotted line represents the mean of target&#x00027;s first-order TTC, solid line represents the mean of target&#x00027;s actual TTC, black disk represents the mean of subject&#x00027;s TTC with a bar indicating 95% confidence interval of the mean.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-17-1099593-g0006.tif"/>
</fig>
<p>To better understand the potential causes of these differences between agent and human performance, it is helpful to consider how the agent&#x00027;s mechanism for anticipation differs from that of humans. The agent chooses actions on the basis of a weighted combination of reward-based reinforcement (instrumental reward) and short model-based prediction (epistemic reward), both of which are computed within the two-headed joint model. EFE values are computed in the EFE head, which is responsible for selecting the action (i.e., pedal position) that it estimates would produce the lowest expected free energy later in the agent&#x00027;s approach. The estimate of EFE associated with each pedal position does not involve an explicit process of model-based prediction, but is learned retrospectively, through the use of an experience replay buffer. Following action selection, visual feedback provides an indication of the cumulative EFE over the duration of the replay buffer. The values of EFE within this buffer are weighted by their temporal distance from the selected action in accordance with the parameter of discount factor. This is similar to both reward-based learning and is often compared to the dopaminergic reward system in humans (Holroyd and Coles, <xref ref-type="bibr" rid="B29">2002</xref>; Haruno, <xref ref-type="bibr" rid="B27">2004</xref>; Lee et al., <xref ref-type="bibr" rid="B33">2012</xref>; Momennejad et al., <xref ref-type="bibr" rid="B38">2017</xref>). The epistemic component of the EFE reward signal is thought to drive exploration toward uncertain world states, and it relies on predictions made in the transition head. This component of the model relies on the hidden states provided by the shared neural layers in the joint model and predicts an observation at next time step <inline-formula><mml:math id="M39"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle class="text"><mml:mtext mathvariant="bold">o</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. The estimated observation at next time step is then compared to the ground truth observation <bold>o</bold><sub><italic>t</italic>&#x0002B;1</sub> and the difference between them generates the epistemic signal <italic>R</italic><sub><italic>t, e</italic></sub>. The role of the transition head is in many ways consistent with a &#x0201C;strong model-based&#x0201D; form of prediction (Zhao and Warren, <xref ref-type="bibr" rid="B55">2015</xref>), whereby predictive behaviors are planned on the basis of an internal model of world states and dynamics that facilitate continuous extrapolation. In summary, whereas the EFE head is consistent with reward based learning, the transition head is consistent with relatively short-term model based prediction.</p>
<p>How does this account of anticipation demonstrated by our agent compare with what we know about anticipation in humans? As discussed in the introduction, empirical data on the quality of model-based prediction suggests that it degrades sufficiently quickly that it cannot explain behaviors of the sort demonstrated here, by our agent, or by the humans in Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>). In contrast, a common theory in motor control and learning relies upon a comparison of a very short-term prediction (e.g., milliseconds) of self-generated action with immediate sensory feedback (Hoist et al., <xref ref-type="bibr" rid="B28">1950</xref>; Wade, <xref ref-type="bibr" rid="B50">1994</xref>; Wolpert et al., <xref ref-type="bibr" rid="B53">1995</xref>; Blakemore et al., <xref ref-type="bibr" rid="B2">1998</xref>). However, this similarity is weakened by the observation that, in the context of motor-learning, short-term prediction is thought to rely upon access to an efferent copy of the motor signal used to generate the action. For this reason, it is problematic that the AIF agent is predicting both its own future state (<italic>x</italic><sub><italic>s</italic></sub>, <italic>v</italic><sub><italic>s</italic></sub>) and the future state of the target (<italic>x</italic><sub><italic>t</italic></sub>, <italic>v</italic><sub><italic>t</italic></sub>), for which there is no efferent copy or analogous information concerning movement dynamics. Although research on eye movements has revealed evidence for the short-term prediction of future object position and trajectory (Ferrera and Barborica, <xref ref-type="bibr" rid="B15">2010</xref>; Diaz et al., <xref ref-type="bibr" rid="B8">2013a</xref>,<xref ref-type="bibr" rid="B9">b</xref>), it remains unclear whether these behaviors are the result of predictive models of object dynamics or representation-minimal heuristics.</p>
<p>Another possible contribution to the observed differences between agent and human performance is the perceptual input. When considering potential causes for the difference between agent and human anticipatory behavior, it is notable that the agent relies upon an observation vector defined by agent&#x00027;s and target&#x00027;s position and velocity measured in meters, and meters per second, respectively. However, in the natural context, these spatial variables must be recovered or estimated on the basis of perceptual sources of information, such as the rate of global optic flow due to translation over the ground plane, the exocentric direction of the target, the instantaneous angular size of the target, or the looming rate of the target during the agent&#x00027;s approach. It is possible that by depriving the agent of these optical variables, we are also depriving the agent of opportunities to exploit task-relevant relationships between the agent and environment, such as the bearing angle. It is also notable that some perceptual variables may provide redundant information about a particular spatial variable (e.g., both change in bearing angle and rate of change in angular size may be informative about an objects approach speed). However, redundant variables will differ in reliability by virtue of sensory thresholds and resolutions. For these reasons, a more complete and comprehensive model of human visually guided action and anticipation would take as input potential sources of information and learn to weight them according to context-dependent reliability and variability.</p>
<p>Another potential contributor to differences between human and agent performance is the notable lack of visuo-motor delays within the agent&#x00027;s architecture. In contrast, human visuo-motor delay has been estimated to be on the order of 100&#x02013;200 ms between the arrival of new visual information and the modification or execution of an action (Nijhawan, <xref ref-type="bibr" rid="B39">2008</xref>; Le Runigo et al., <xref ref-type="bibr" rid="B32">2010</xref>). Because uncompensated delays would have devastating consequences on human visual and motor control, they are often cited as evidence that humans must have some form of predictive mechanism that acts in compensation (Wolpert et al., <xref ref-type="bibr" rid="B53">1995</xref>). Future attempts to make this model&#x00027;s anticipatory behavior more human-like in nature may do so by imposing similar length delay between the agent&#x00027;s choice of motor plan on the basis of the observed world-state and the time that this motor plan is executed (Walsh et al., <xref ref-type="bibr" rid="B51">2009</xref>). Finally, note that our proposed architecture is &#x0201C;flat&#x0201D; in the temporal sense. It other words, EFE values are calculated and actions are planned in a single linear time scale. In contrast, a deep/hierarchical temporal model would imply that policies are inferred, learned, and ultimately operate at different time scales (Friston et al., <xref ref-type="bibr" rid="B26">2018</xref>). We believe that our approach is sufficient for the given task of this study. However, if one intended to extend the problem to more sophisticated settings where higher level cognitive functions are separated from lower-level motor control, a deep temporal model could be a more suitable/useful approach.</p>
<p>Due to limited computation resources that we have access to and the high computational cost of the full Bayesian inference framework (which, in the context of neural networks, requires formulating each neural network as a Bayesian neural network where training, typically to obtain good-quality performance, requires Markov chain Monte Carlo), we simplify the Bayesian inference by assuming a uniform prior (or uninformative prior) on the parameters of our model, similar to Tschantz et al. (<xref ref-type="bibr" rid="B47">2020a</xref>). Maximum likelihood estimation (MLE), in our setup, is generally equivalent to maximum a posteriori (MAP) estimation while assuming the priors to be uniform distributions. More general forms of Bayesian inference with different prior assumptions could be examined in future work. Also, note that the Laplace approximation applied in this work leads to the expected free energy reducing to a KL-divergence (i.e., KL control).</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>We present a novel scaled-up version of active inference framework (AIF) model for studying online visually guided locomotion using an interception task where a moving target changes its speeds in a semi-predictable manner. In order to drive the agent toward the goal more effectively, we devised a problem-specific <italic>prior mapping function</italic>, improving the agent&#x00027;s computational efficiency and interpretability. Notably, we found that our proposed AIF agent exhibits better task performance when compared to a commonly used RL agent, i.e., the deep-Q network (DQN). The full AIF agent, containing both <italic>instrumental</italic> and <italic>epistemic</italic> components, exhibited slightly better task performance and lower variance compared to the AIF agent with only an <italic>instrumental</italic> component. Furthermore, we demonstrated behavioral differences among our full AIF agents given different discount factor &#x003B3; values as well as levels of the agent&#x00027;s action-to-speed responsiveness. Finally, we analyzed the anticipatory behavior demonstrated by our agent and examined the differences between the agent&#x00027;s behavior and human behavior. While our results are promising, future work should address the following limitations&#x02014;first, inputs to our agent are defined in a simplified vector space whereas sensory inputs to the humans that actually perform the interception task are visual in nature (i.e., the model should work directly with unstructured sensory data such as pixel values). We remark that a vision-based approach could facilitate the extraction of additional information and features that are useful for solving the interception task more reliably. Second, our simulations do not account for visuo-motor delays inherent to the human visual and motor systems, and that might be modeled using techniques like delayed Markov decision process formulations (Walsh et al., <xref ref-type="bibr" rid="B51">2009</xref>; Firoiu et al., <xref ref-type="bibr" rid="B16">2018</xref>).</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="ethics-statement" id="s7">
<title>Ethics statement</title>
<p>The Diaz et al. (<xref ref-type="bibr" rid="B11">2009</xref>) study was reviewed and approved by Rensselaer Polytechnic Institute IRB. The patients/participants provided their written informed consent to participate in that study. Ethical review and approval was not required for the the present study in accordance with the local legislation and institutional requirements. Written informed consent to participate in the current study was not required in accordance with the local legislation and institutional requirements.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>AO aided ZY in preliminary simulation/testing and they both devised the neural AIF algorithm. ZY implemented the experimental simulations as well as collected and analyzed the results. All authors contributed to the experimental design and the project&#x00027;s development, data interpretation, drafting of the manuscript, and approval of the final version of the manuscript for submission.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This material was based upon work supported by the National Science Foundation (NSF) under Award No. DGE-2125362 and CPS-2225354.</p>
</sec>
<ack>
<p>The authors would like to thank Tim Johnson for contributing to the initial software development of the interception task simulation environment(s) used in this project.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Author disclaimer</title>
<p>Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Abadi</surname> <given-names>M.</given-names></name> <name><surname>Agarwal</surname> <given-names>A.</given-names></name> <name><surname>Barham</surname> <given-names>P.</given-names></name> <name><surname>Brevdo</surname> <given-names>E.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Citro</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2015</year>). <source>TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.tensorflow.org">https://www.tensorflow.org</ext-link></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blakemore</surname> <given-names>S.-J.</given-names></name> <name><surname>Wolpert</surname> <given-names>D. M.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name></person-group> (<year>1998</year>). <article-title>Central cancellation of self-produced tickle sensation</article-title>. <source>Nat. Neurosci</source>. <volume>1</volume>, <fpage>635</fpage>&#x02013;<lpage>640</lpage>. <pub-id pub-id-type="doi">10.1038/2870</pub-id><pub-id pub-id-type="pmid">10196573</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brockman</surname> <given-names>G.</given-names></name> <name><surname>Cheung</surname> <given-names>V.</given-names></name> <name><surname>Pettersson</surname> <given-names>L.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Schulman</surname> <given-names>J.</given-names></name> <name><surname>Tang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>OpenAi gym</article-title>. <source>arXiv preprint</source> arXiv:1606.01540. <pub-id pub-id-type="doi">10.48550/arXiv.1606.01540</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buckley</surname> <given-names>C. L.</given-names></name> <name><surname>Kim</surname> <given-names>C. S.</given-names></name> <name><surname>McGregor</surname> <given-names>S.</given-names></name> <name><surname>Seth</surname> <given-names>A. K.</given-names></name></person-group> (<year>2017</year>). <article-title>The free energy principle for action and perception: a mathematical review</article-title>. <source>J. Math. Psychol</source>. <volume>81</volume>, <fpage>55</fpage>&#x02013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2017.09.004</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x000C7;atal</surname> <given-names>O.</given-names></name> <name><surname>Wauthier</surname> <given-names>S.</given-names></name> <name><surname>De Boom</surname> <given-names>C.</given-names></name> <name><surname>Verbelen</surname> <given-names>T.</given-names></name> <name><surname>Dhoedt</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Learning generative state space models for active inference</article-title>. <source>Front. Comput. Neurosci</source>. <volume>14</volume>, <fpage>574372</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2020.574372</pub-id><pub-id pub-id-type="pmid">33304260</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chapman</surname> <given-names>S.</given-names></name></person-group> (<year>1968</year>). <article-title>Catching a baseball</article-title>. <source>Am. J. Phys</source>. <volume>36</volume>, <fpage>868</fpage>&#x02013;<lpage>870</lpage>. <pub-id pub-id-type="doi">10.1119/1.1974297</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Da Costa</surname> <given-names>L.</given-names></name> <name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Sajid</surname> <given-names>N.</given-names></name> <name><surname>Veselic</surname> <given-names>S.</given-names></name> <name><surname>Neacsu</surname> <given-names>V.</given-names></name> <name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Active inference on discrete state-spaces: a synthesis</article-title>. <source>J. Math. Psychol</source>. <volume>99</volume>, <fpage>102447</fpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2020.102447</pub-id><pub-id pub-id-type="pmid">33343039</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diaz</surname> <given-names>G.</given-names></name> <name><surname>Cooper</surname> <given-names>J.</given-names></name> <name><surname>Hayhoe</surname> <given-names>M.</given-names></name></person-group> (<year>2013a</year>). <article-title>Memory and prediction in natural gaze control</article-title>. <source>Philos. Trans. R. Soc. B Biol. Sci</source>. <volume>368</volume>, <fpage>20130064</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2013.0064</pub-id><pub-id pub-id-type="pmid">24018726</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diaz</surname> <given-names>G.</given-names></name> <name><surname>Cooper</surname> <given-names>J.</given-names></name> <name><surname>Rothkopf</surname> <given-names>C.</given-names></name> <name><surname>Hayhoe</surname> <given-names>M.</given-names></name></person-group> (<year>2013b</year>). <article-title>Saccades to future ball location reveal memory-based prediction in a virtual-reality interception task</article-title>. <source>J. Vis</source>. <volume>13</volume>, <fpage>20</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1167/13.1.20</pub-id><pub-id pub-id-type="pmid">23325347</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diaz</surname> <given-names>G. J.</given-names></name> <name><surname>Parade</surname> <given-names>M. S.</given-names></name> <name><surname>Barton</surname> <given-names>S. L.</given-names></name> <name><surname>Fajen</surname> <given-names>B. R.</given-names></name></person-group> (<year>2018</year>). <article-title>The pickup of visual information about size and location during approach to an obstacle</article-title>. <source>PLoS ONE</source> <volume>13</volume>, <fpage>e0192044</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0192044</pub-id><pub-id pub-id-type="pmid">29401511</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diaz</surname> <given-names>G. J.</given-names></name> <name><surname>Phillips</surname> <given-names>F.</given-names></name> <name><surname>Fajen</surname> <given-names>B. R.</given-names></name></person-group> (<year>2009</year>). <article-title>Intercepting moving targets: a little foresight helps a lot</article-title>. <source>Exp. Brain Res</source>. <volume>195</volume>, <fpage>345</fpage>&#x02013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.1007/s00221-009-1794-5</pub-id><pub-id pub-id-type="pmid">19396594</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fajen</surname> <given-names>B.</given-names></name> <name><surname>Diaz</surname> <given-names>G.</given-names></name> <name><surname>Cramer</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>Reconsidering the role of action in perceiving the catchability of fly balls</article-title>. <source>J. Vis</source>. <volume>8</volume>, <fpage>621</fpage>&#x02013;<lpage>621</lpage>. <pub-id pub-id-type="doi">10.1167/8.6.621</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fajen</surname> <given-names>B. R.</given-names></name> <name><surname>Devaney</surname> <given-names>M. C.</given-names></name></person-group> (<year>2006</year>). <article-title>Learning to control collisions: the role of perceptual attunement and action boundaries</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform</source>. <volume>32</volume>, <fpage>300</fpage>&#x02013;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.32.2.300</pub-id><pub-id pub-id-type="pmid">16634672</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fajen</surname> <given-names>B. R.</given-names></name> <name><surname>Warren</surname> <given-names>W. H.</given-names></name></person-group> (<year>2007</year>). <article-title>Behavioral dynamics of intercepting a moving target</article-title>. <source>Exp. Brain Res</source>. <volume>180</volume>, <fpage>303</fpage>&#x02013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1007/s00221-007-0859-6</pub-id><pub-id pub-id-type="pmid">17273872</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrera</surname> <given-names>V. P.</given-names></name> <name><surname>Barborica</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Internally generated error signals in monkey frontal eye field during an inferred motion task</article-title>. <source>J. Neurosci</source>. <volume>30</volume>, <fpage>11612</fpage>&#x02013;<lpage>11623</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2977-10.2010</pub-id><pub-id pub-id-type="pmid">20810882</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Firoiu</surname> <given-names>V.</given-names></name> <name><surname>Ju</surname> <given-names>T.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>At human speed: deep reinforcement learning with action delay</article-title>. <source>arXiv preprint</source> arXiv:1810.07286. <pub-id pub-id-type="doi">10.48550/arXiv.1810.07286</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>The free-energy principle: a rough guide to the brain?</article-title> <source>Trends Cogn. Sci</source>. <volume>13</volume>, <fpage>293</fpage>&#x02013;<lpage>301</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2009.04.005</pub-id><pub-id pub-id-type="pmid">19559644</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci</source>. <volume>11</volume>, <fpage>127</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id><pub-id pub-id-type="pmid">20068583</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>A free energy principle for a particular physics</article-title>. <source>arXiv preprint</source> arXiv:1906.10184. <pub-id pub-id-type="doi">10.48550/arXiv.1906.10184</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Da Costa</surname> <given-names>L.</given-names></name> <name><surname>Hafner</surname> <given-names>D.</given-names></name> <name><surname>Hesp</surname> <given-names>C.</given-names></name> <name><surname>Parr</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>Sophisticated inference</article-title>. <source>Neural Comput</source>. <volume>33</volume>, <fpage>713</fpage>&#x02013;<lpage>763</lpage>. <pub-id pub-id-type="doi">10.1162/neco_a_01351</pub-id><pub-id pub-id-type="pmid">33626312</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>FitzGerald</surname> <given-names>T.</given-names></name> <name><surname>Rigoli</surname> <given-names>F.</given-names></name> <name><surname>Schwartenbeck</surname> <given-names>P.</given-names></name> <name><surname>Pezzulo</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Active inference: a process theory</article-title>. <source>Neural Comput</source>. <volume>29</volume>, <fpage>1</fpage>&#x02013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00912</pub-id><pub-id pub-id-type="pmid">27870614</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Rigoli</surname> <given-names>F.</given-names></name> <name><surname>Ognibene</surname> <given-names>D.</given-names></name> <name><surname>Mathys</surname> <given-names>C.</given-names></name> <name><surname>Fitzgerald</surname> <given-names>T.</given-names></name> <name><surname>Pezzulo</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Active inference and epistemic value</article-title>. <source>Cogn. Neurosci</source>. <volume>6</volume>, <fpage>187</fpage>&#x02013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1080/17588928.2015.1020053</pub-id><pub-id pub-id-type="pmid">25689102</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name> <name><surname>Samothrakis</surname> <given-names>S.</given-names></name> <name><surname>Montague</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>Active inference and agency: optimal control without cost functions</article-title>. <source>Biol. Cybern</source>. <volume>106</volume>, <fpage>523</fpage>&#x02013;<lpage>541</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-012-0512-8</pub-id><pub-id pub-id-type="pmid">22864468</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Daunizeau</surname> <given-names>J.</given-names></name> <name><surname>Kiebel</surname> <given-names>S. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Reinforcement learning or active inference?</article-title> <source>PLoS ONE</source> <volume>4</volume>, <fpage>e6421</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0006421</pub-id><pub-id pub-id-type="pmid">19641614</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Lin</surname> <given-names>M.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name> <name><surname>Pezzulo</surname> <given-names>G.</given-names></name> <name><surname>Hobson</surname> <given-names>J. A.</given-names></name> <name><surname>Ondobaka</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Active inference, curiosity and insight</article-title>. <source>Neural Comput</source>. <volume>29</volume>, <fpage>2633</fpage>&#x02013;<lpage>2683</lpage>. <pub-id pub-id-type="doi">10.1162/neco_a_00999</pub-id><pub-id pub-id-type="pmid">28777724</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Rosch</surname> <given-names>R.</given-names></name> <name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Price</surname> <given-names>C.</given-names></name> <name><surname>Bowman</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep temporal models and active inference</article-title>. <source>Neurosci. Biobehav. Rev</source>. <volume>90</volume>, <fpage>486</fpage>&#x02013;<lpage>501</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2018.04.004</pub-id><pub-id pub-id-type="pmid">29747865</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haruno</surname> <given-names>M.</given-names></name></person-group> (<year>2004</year>). <article-title>A neural correlate of reward-based behavioral learning in caudate nucleus: a functional magnetic resonance imaging study of a stochastic decision task</article-title>. <source>J. Neurosci</source>. <volume>24</volume>, <fpage>1660</fpage>&#x02013;<lpage>1665</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3417-03.2004</pub-id><pub-id pub-id-type="pmid">14973239</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoist</surname> <given-names>E. v</given-names></name> <name><surname>Mittelstaedt</surname> <given-names>H.</given-names></name> <name><surname>Martin</surname> <given-names>R.</given-names></name></person-group> (<year>1950</year>). <article-title>Das reafferenzpr&#x000EC;nzip. wechselwirkung zwischen zentralnervensystem und peripherie</article-title>. <source>Die Naturwissenschaften</source> <volume>37</volume>, <fpage>464</fpage>. <pub-id pub-id-type="doi">10.1007/BF00622503</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holroyd</surname> <given-names>C. B.</given-names></name> <name><surname>Coles</surname> <given-names>M. G. H.</given-names></name></person-group> (<year>2002</year>). <article-title>The neural basis of human error processing: Reinforcement learning, dopamine, and the error-related negativity</article-title>. <source>Psychol. Rev</source>. <volume>109</volume>, <fpage>679</fpage>&#x02013;<lpage>709</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.109.4.679</pub-id><pub-id pub-id-type="pmid">12374324</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaplan</surname> <given-names>R.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2018</year>). <article-title>Planning and navigation as active inference</article-title>. <source>Biol. Cybern</source>. <volume>112</volume>, <fpage>323</fpage>&#x02013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-018-0753-2</pub-id><pub-id pub-id-type="pmid">29572721</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>arXiv preprint</source> arXiv:1412.6980. <pub-id pub-id-type="doi">10.48550/arXiv.1412.6980</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le Runigo</surname> <given-names>C.</given-names></name> <name><surname>Benguigui</surname> <given-names>N.</given-names></name> <name><surname>Bardy</surname> <given-names>B. G.</given-names></name></person-group> (<year>2010</year>). <article-title>Visuo-motor delay, information-movement coupling, and expertise in ball sports</article-title>. <source>J. Sports Sci</source>. <volume>28</volume>, <fpage>327</fpage>&#x02013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1080/02640410903502782</pub-id><pub-id pub-id-type="pmid">20131141</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>D.</given-names></name> <name><surname>Seo</surname> <given-names>H.</given-names></name> <name><surname>Jung</surname> <given-names>M. W.</given-names></name></person-group> (<year>2012</year>). <article-title>Neural basis of reinforcement learning and decision making</article-title>. <source>Annu. Rev. Neurosci</source>. <volume>35</volume>, <fpage>287</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-neuro-062111-150512</pub-id><pub-id pub-id-type="pmid">22462543</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>D. N.</given-names></name></person-group> (<year>1976</year>). <article-title>A theory of visual control of braking based on information about time-to-collision</article-title>. <source>Perception</source> <volume>5</volume>, <fpage>437</fpage>&#x02013;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.1068/p050437</pub-id><pub-id pub-id-type="pmid">1005020</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matthis</surname> <given-names>J. S.</given-names></name> <name><surname>Fajen</surname> <given-names>B. R.</given-names></name></person-group> (<year>2013</year>). <article-title>Humans exploit the biomechanics of bipedal gait during visually guided walking over complex terrain</article-title>. <source>Proc. R. Soc. B Biol. Sci</source>. <volume>280</volume>, <fpage>20130700</fpage>. <pub-id pub-id-type="doi">10.1098/rspb.2013.0700</pub-id><pub-id pub-id-type="pmid">23658204</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Millidge</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep active inference as variational policy gradients</article-title>. <source>J. Math. Psychol</source>. <volume>96</volume>, <fpage>102348</fpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2020.102348</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source> <volume>518</volume>, <fpage>529</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Momennejad</surname> <given-names>I.</given-names></name> <name><surname>Russek</surname> <given-names>E. M.</given-names></name> <name><surname>Cheong</surname> <given-names>J. H.</given-names></name> <name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2017</year>). <article-title>The successor representation in human reinforcement learning</article-title>. <source>Nat. Hum. Behav</source>. <volume>1</volume>, <fpage>680</fpage>&#x02013;<lpage>692</lpage>. <pub-id pub-id-type="doi">10.1038/s41562-017-0180-8</pub-id><pub-id pub-id-type="pmid">31024137</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nijhawan</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Visual prediction: Psychophysics and neurophysiology of compensation for time delays</article-title>. <source>Behav. Brain Sci</source>. <volume>31</volume>, <fpage>179</fpage>&#x02013;<lpage>198</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X08003804</pub-id><pub-id pub-id-type="pmid">18479557</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Uncertainty, epistemics and active inference</article-title>. <source>J. R. Soc. Interface</source> <volume>14</volume>, <fpage>20170376</fpage>. <pub-id pub-id-type="doi">10.1098/rsif.2017.0376</pub-id><pub-id pub-id-type="pmid">29167370</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2018</year>). <article-title>The anatomy of inference: generative models and brain structure</article-title>. <source>Front. Comput. Neurosci</source>. <volume>12</volume>, <fpage>90</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2018.00090</pub-id><pub-id pub-id-type="pmid">30483088</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Generalised free energy and active inference</article-title>. <source>Biol. Cybern</source>. <volume>113</volume>, <fpage>495</fpage>&#x02013;<lpage>513</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-019-00805-w</pub-id><pub-id pub-id-type="pmid">31562544</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sajid</surname> <given-names>N.</given-names></name> <name><surname>Ball</surname> <given-names>P. J.</given-names></name> <name><surname>Parr</surname> <given-names>T.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2021</year>). <article-title>Active inference: demystified and compared</article-title>. <source>Neural Comput</source>. <volume>33</volume>, <fpage>674</fpage>&#x02013;<lpage>712</lpage>. <pub-id pub-id-type="doi">10.1162/neco_a_01357</pub-id><pub-id pub-id-type="pmid">33400903</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwartenbeck</surname> <given-names>P.</given-names></name> <name><surname>Passecker</surname> <given-names>J.</given-names></name> <name><surname>Hauser</surname> <given-names>T. U.</given-names></name> <name><surname>FitzGerald</surname> <given-names>T. H.</given-names></name> <name><surname>Kronbichler</surname> <given-names>M.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Computational mechanisms of curiosity and goal-directed exploration</article-title>. <source>Elife</source> <volume>8</volume>, <fpage>e41703</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.41703</pub-id><pub-id pub-id-type="pmid">31074743</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shin</surname> <given-names>J. Y.</given-names></name> <name><surname>Kim</surname> <given-names>C.</given-names></name> <name><surname>Hwang</surname> <given-names>H. J.</given-names></name></person-group> (<year>2022</year>). <article-title>Prior preference learning from experts: designing a reward with active inference</article-title>. <source>Neurocomputing</source> <volume>492</volume>, <fpage>508</fpage>&#x02013;<lpage>515</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.12.042</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tieleman</surname> <given-names>T.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Lecture 6.5-rmsprop: divide the gradient by a running average of its recent magnitude</article-title>. <source>Coursera Neural Netw. Mach. Learn</source>. <volume>4</volume>, <fpage>26</fpage>&#x02013;<lpage>31</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tschantz</surname> <given-names>A.</given-names></name> <name><surname>Baltieri</surname> <given-names>M.</given-names></name> <name><surname>Seth</surname> <given-names>A. K.</given-names></name> <name><surname>Buckley</surname> <given-names>C. L.</given-names></name></person-group> (<year>2020a</year>). <article-title>&#x0201C;Scaling active inference,&#x0201D;</article-title> in <source>2020 International Joint Conference on Neural Networks (IJCNN)</source> (<publisher-loc>Glasgow, UK</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tschantz</surname> <given-names>A.</given-names></name> <name><surname>Seth</surname> <given-names>A. K.</given-names></name> <name><surname>Buckley</surname> <given-names>C. L.</given-names></name></person-group> (<year>2020b</year>). <article-title>Learning action-oriented models through active inference</article-title>. <source>PLoS Comput. Biol</source>. <volume>16</volume>, <fpage>e1007805</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1007805</pub-id><pub-id pub-id-type="pmid">32324758</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ueltzh&#x000F6;ffer</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep active inference</article-title>. <source>Biol. Cybern</source>. <volume>112</volume>, <fpage>547</fpage>&#x02013;<lpage>573</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-018-0785-7</pub-id><pub-id pub-id-type="pmid">30350226</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wade</surname> <given-names>N. J.</given-names></name></person-group> (<year>1994</year>). <article-title>Hermann von helmholtz (1821&#x02013;1894)</article-title>. <source>Perception</source> <volume>23</volume>, <fpage>981</fpage>&#x02013;<lpage>999</lpage>. <pub-id pub-id-type="doi">10.1068/p230981</pub-id><pub-id pub-id-type="pmid">7899051</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Walsh</surname> <given-names>T. J.</given-names></name> <name><surname>Nouri</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Littman</surname> <given-names>M. L.</given-names></name></person-group> (<year>2009</year>). <article-title>Learning and planning in environments with delayed feedback</article-title>. <source>Auton. Agent Multi Agent. Syst</source>. <volume>18</volume>, <fpage>83</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1007/s10458-008-9056-7</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Warren</surname> <given-names>W.</given-names></name> <name><surname>Fajen</surname> <given-names>B.</given-names></name> <name><surname>Belcher</surname> <given-names>D.</given-names></name></person-group> (<year>2010</year>). <article-title>Behavioral dynamics of steering, obstacle avoidance, and route selection</article-title>. <source>J. Vis</source>. <volume>1</volume>, <fpage>184</fpage>&#x02013;<lpage>184</lpage>. <pub-id pub-id-type="doi">10.1167/1.3.184</pub-id><pub-id pub-id-type="pmid">12760620</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolpert</surname> <given-names>D. M.</given-names></name> <name><surname>Ghahramani</surname> <given-names>Z.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>1995</year>). <article-title>An internal model for sensorimotor integration</article-title>. <source>Science</source> <volume>269</volume>, <fpage>1880</fpage>&#x02013;<lpage>1882</lpage>. <pub-id pub-id-type="doi">10.1126/science.7569931</pub-id><pub-id pub-id-type="pmid">7569931</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yilmaz</surname> <given-names>E. H.</given-names></name> <name><surname>Warren</surname> <given-names>W. H.</given-names></name></person-group> (<year>1995</year>). <article-title>Visual control of braking: a test of the hypothesis</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform</source>. <volume>21</volume>, <fpage>996</fpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.21.5.996</pub-id><pub-id pub-id-type="pmid">7595250</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Warren</surname> <given-names>W. H.</given-names></name></person-group> (<year>2015</year>). <article-title>On-line and model-based approaches to the visual control of action</article-title>. <source>Vis. Res</source>. <volume>110</volume>, <fpage>190</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2014.10.008</pub-id><pub-id pub-id-type="pmid">25454700</pub-id></citation></ref>
</ref-list> 
</back>
</article>
