<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2017.00112</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Computational Foundations of Natural Intelligence</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>van Gerven</surname> <given-names>Marcel</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/28677/overview"/>
</contrib>
</contrib-group>
<aff><institution>Computational Cognitive Neuroscience Lab, Department of Artificial Intelligence, Donders Institute for Brain, Cognition and Behaviour, Radboud University</institution>, <addr-line>Nijmegen</addr-line>, <country>Netherlands</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Florentin W&#x000F6;rg&#x000F6;tter, University of G&#x000F6;ttingen, Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Sebastian Herzog, Max Planck Institute for Dynamics and Self Organization (MPG), Germany; Carme Torras, Consejo Superior de Investigaciones Cient&#x000ED;ficas (CSIC), Spain</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Marcel van Gerven <email>m.vangerven&#x00040;donders.ru.nl</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>12</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>112</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>08</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>11</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 van Gerven.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>van Gerven</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>New developments in AI and neuroscience are revitalizing the quest to understanding natural intelligence, offering insight about how to equip machines with human-like capabilities. This paper reviews some of the computational principles relevant for understanding natural intelligence and, ultimately, achieving strong AI. After reviewing basic principles, a variety of computational modeling approaches is discussed. Subsequently, I concentrate on the use of artificial neural networks as a framework for modeling cognitive processes. This paper ends by outlining some of the challenges that remain to fulfill the promise of machines that show human-like intelligence.</p></abstract>
<kwd-group>
<kwd>natural intelligence</kwd>
<kwd>strong AI</kwd>
<kwd>cognition</kwd>
<kwd>artificial neural networks</kwd>
<kwd>machine learning</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="0"/>
<equation-count count="11"/>
<ref-count count="384"/>
<page-count count="24"/>
<word-count count="21055"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Understanding how mind emerges from matter is one of the great remaining questions in science. How is it possible that organized clumps of matter such as our own brains give rise to all of our beliefs, desires and intentions, ultimately allowing us to contemplate ourselves as well as the universe from which we originate? This question has occupied cognitive scientists who study the computational basis of the mind for decades. It also occupies other breeds of scientists. For example, ethologists and psychologists focus on the complex behavior exhibited by animals and humans whereas cognitive, computational and systems neuroscientists wish to understand the mechanistic basis of processes that give rise to such behavior.</p>
<p>The ambition to understand natural intelligence as encountered in biological organisms can be contrasted with the motivation to build intelligent machines, which is the subject matter of artificial intelligence (AI). Wouldn&#x00027;t it be amazing if we could build synthetic brains that are endowed with the same qualities as their biological cousins? This desire to mimic human-level intelligence by creating artificially intelligent machines has occupied mankind for many centuries. For instance, mechanical men and artificial beings appear in Greek mythology and realistic human automatons had already been developed in Hellenic Egypt (McCorduck, <xref ref-type="bibr" rid="B218">2004</xref>). The engineering of machines that display human-level intelligence is also referred to as strong AI (Searle, <xref ref-type="bibr" rid="B306">1980</xref>) or artificial general intelligence (AGI) (Adams et al., <xref ref-type="bibr" rid="B4">2012</xref>), and was the original motivation that gave rise to the field of AI (Newell, <xref ref-type="bibr" rid="B242">1991</xref>; Nilsson, <xref ref-type="bibr" rid="B245">2005</xref>).</p>
<p>Excitingly, major advances in various fields of research now make it possible to attack the problem of understanding natural intelligence from multiple angles. From a theoretical point of view we have a solid understanding of the computational problems that are solved by our own brains (Dayan and Abbott, <xref ref-type="bibr" rid="B69">2005</xref>). From an empirical point of view, technological breakthroughs allow us to probe and manipulate brain activity in unprecedented ways, generating new neuroscientific insights into brain structure and function (Chang, <xref ref-type="bibr" rid="B53">2015</xref>). From an engineering perspective, we are finally able to build machines that learn to solve complex tasks, approximating and sometimes surpassing human-level performance (Jordan and Mitchell, <xref ref-type="bibr" rid="B158">2015</xref>). Still, these efforts have not yet provided a full understanding of natural intelligence, nor did they give rise to machines whose reasoning capacity parallels the generality and flexibility of cognitive processing in biological organisms.</p>
<p>The core thesis of this paper is that natural intelligence can be better understood by the coming together of multiple complementary scientific disciplines (Gershman et al., <xref ref-type="bibr" rid="B113">2015</xref>). This thesis is referred to as <italic>the great convergence</italic>. The advocated approach is to endow artificial agents with synthetic brains (i.e., cognitive architectures, Sun, <xref ref-type="bibr" rid="B329">2004</xref>) that mimic the thought processes that give rise to ethologically relevant behavior in their biological counterparts. A motivation for this approach is given by Braitenberg&#x00027;s law of uphill analysis and downhill invention, which states that it is much easier to understand a complex system by assembling it from the ground up, rather than by reverse engineering it from observational data (Braitenberg, <xref ref-type="bibr" rid="B40">1986</xref>). These synthetic brains, which can be put to use in virtual or real-world environments, can then be validated against neuro-behavioral data and analyzed using a multitude of theoretical tools. This approach not only elucidates our understanding of human brain function but also paves the way for the development of artificial agents that show truly intelligent behavior (Hassabis et al., <xref ref-type="bibr" rid="B134">2017</xref>).</p>
<p>The aim of this paper is to sketch the outline of a research program which marries the ambitions of neuroscientists to understand natural intelligence and AI researchers to achieve strong AI (Figure <xref ref-type="fig" rid="F1">1</xref>). Before embarking on our quest to build synthetic brains as models of natural intelligence, we need to formalize what problems are solved by biological brains. That is, we first need to understand how adaptive behavior ensues in animals and humans.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Understanding natural intelligence and achieving strong AI are seen as relying on the same theoretical foundations and require the convergence of multiple scientific and engineering disciplines.</p></caption>
<graphic xlink:href="fncom-11-00112-g0001.tif"/>
</fig>
</sec>
<sec id="s2">
<title>2. Adaptive behavior in biological agents</title>
<p>Ultimately, organisms owe their existence to the fact that they promote survival of their constituent genes; the basic physical and functional units of heredity that code for an organism (Dawkins, <xref ref-type="bibr" rid="B67">2016</xref>). At evolutionary time scales, organisms developed a range of mechanisms which ensure that they live long enough such as to produce offspring. For example, single-celled protozoans already show rather complex ingestive, defensive and reproductive behavior, which is regulated by molecular signaling (Swanson, <xref ref-type="bibr" rid="B335">2012</xref>; Sterling and Laughlin, <xref ref-type="bibr" rid="B326">2016</xref>).</p>
<sec>
<title>2.1. Why do we need a brain?</title>
<p>About 3.5 billion years ago, multicellular organisms started to appear. Multicellularity offers several competitive advantages over unicellularity. It allows organisms to increase in size without the limitations set by unicellularity and permits increased complexity by allowing cellular differentiation. It also increases life span since an organism can live beyond the demise of a single cell. At the same time, due to their increased size and complexity, multicellular organisms require more intricate mechanisms for signaling and regulation.</p>
<p>In multicellular organisms, behavior is regulated at multiple scales, ranging from intracellular molecular signaling all the way up to global regulation via the interactions between different organ systems. Hence, the nervous system allows for fast responses via electrochemical signaling and for slow responses by acting on the endocrine system. Nervous systems are found in almost all multicellular animals, but vary greatly in complexity. For example, the nervous system of the nematode roundworm <italic>Caenorhabditis elegans</italic> (<italic>C. elegans</italic>) is made up of 302 neurons and 7,000 synaptic connections (White et al., <xref ref-type="bibr" rid="B364">1986</xref>; Varshney et al., <xref ref-type="bibr" rid="B355">2011</xref>). In contrast, the human brain contains about 20 billion neocortical neurons that are wired together via as many as 0.15 quadrillion synapses (Pakkenberg and Gundersen, <xref ref-type="bibr" rid="B255">1997</xref>; Pakkenberg et al., <xref ref-type="bibr" rid="B256">2003</xref>).</p>
<p>In vertebrates, the nervous system can be partitioned into the central nervous system (CNS), consisting of the brain and the spinal cord, and the peripheral nervous system (PNS), which connects the CNS to every other part of the body. The brain allows for centralized control and efficient information transmission. It can be partitioned into the forebrain, midbrain and hindbrain, each of which contain dedicated neural circuits that allow for integration of information and generation of coordinated activity. The spinal cord connects the brain to the body by allowing sensory and motor information to travel back and forth between the brain and the body. It also coordinates certain reflexes that bypass the brain altogether.</p>
<p>The interplay between the nervous system, the body and the environment is nicely captured by Swanson&#x00027;s four system model of nervous system organization (Swanson, <xref ref-type="bibr" rid="B334">2000</xref>), as shown in Figure <xref ref-type="fig" rid="F2">2</xref>. Briefly, the brain exerts centralized control on the body by sending commands to the motor system based on information received via the sensory system. It exerts this control by way of the cognitive system, which drives voluntary initiation of behavior, as well as the state system, which refers to the intrinsic activity that controls global behavioral state. The motor system can also be influenced directly by the sensory system via spinal cord reflexes. Output of the motor system induces visceral responses that affect bodily state as well as somatic responses that act on the environment. It is also able to drive the secretion of hormones that act more globally on the body. Both the body and the environment generate sensations that are processed by the sensory system. This closed-loop system, tightly coupling sensation, thought and action, is known as the <italic>perception-action cycle</italic> (Dewey, <xref ref-type="bibr" rid="B77">1896</xref>; Sperry, <xref ref-type="bibr" rid="B320">1952</xref>; Fuster, <xref ref-type="bibr" rid="B106">2004</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The four system model of nervous system organization. CO, Cognitive system; EN, Environment; ES, Environmental stimuli; MO, Motor system; SE, Sensory system; SR, Somatic responses; ST, Behavioral state system; VR, Visceral responses; VS, Visceral stimuli. Solid arrows show influences pertaining to the nervous system. Dashed arrows show interactions produced by the body or the environment<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>.</p></caption>
<graphic xlink:href="fncom-11-00112-g0002.tif"/>
</fig>
<p>Summarizing, the brain, together with the spinal cord and the peripheral nervous system, can be seen as an organ that exploits sensory input such as to generate adaptive behavior through motor outputs. This ensures an organism&#x00027;s long-term survival in a world that is dominated by uncertainty, as a result of partial observability, noise and stochasticity. The upshot of this interpretation is that action, which drives the generation of adaptive behavior, is the ultimate reason why we have a brain in the first place. Citing Sperry (<xref ref-type="bibr" rid="B320">1952</xref>): &#x0201C;the entire output of our thinking machine consists of nothing but patterns of motor coordination.&#x0201D; To understand how adaptive behavior ensues, we therefore need to identify the ultimate causes that determine an agent&#x00027;s actions (Tolman, <xref ref-type="bibr" rid="B348">1932</xref>).</p>
</sec>
<sec>
<title>2.2. What makes us tick?</title>
<p>In biology, ultimately, all evolved traits must be connected to an organism&#x00027;s survival. This implies that, from the standpoint of evolutionary psychology, natural selection favors those behaviors and thought processes that provide the organism with a selective advantage under ecological pressure (Barkow et al., <xref ref-type="bibr" rid="B20">1992</xref>). Since causal links between behavior and long-term survival cannot be sensed or controlled directly, an agent needs to rely on other, directly accessible, ways to promote its survival. This can take the form of (1) evolving optimal sensors and effectors that allow it to maximize its control given finite resources and (2) evolving a behavioral repertoire that maximizes the information gained from the environment and generates optimal actions based on available sensory information.</p>
<p>In practice, behavior is the result of multiple competing needs that together provide an evolutionary advantage. These needs arise because they provide particular rewards to the organism. We distinguish <italic>primary rewards, intrinsic rewards</italic> and <italic>extrinsic rewards</italic>.</p>
<sec>
<title>Primary rewards</title>
<p>Primary rewards are those necessary for the survival of one&#x00027;s self and offspring, which includes homeostatic and reproductive rewards. Here, homeostasis refers to the maintenance of optimal settings of various biological parameters (e.g., temperature regulation) (Cannon, <xref ref-type="bibr" rid="B49">1929</xref>). A slightly more sophisticated concept is <italic>allostasis</italic>, which refers to the predictive regulation of biological parameters in order to prevent deviations rather than correcting them <italic>post hoc</italic> (Sterling, <xref ref-type="bibr" rid="B325">2012</xref>). An organism can use its nervous system (muscle signaling) or endocrine system (endocrine signaling) to globally control or adjust the activities of many systems simultaneously. This allows for visceral responses that ensure proper functioning of an agent&#x00027;s internal organs as well as basic drives such as ingestion, defense and reproduction that help ensure an agent&#x00027;s survival (Tinbergen, <xref ref-type="bibr" rid="B344">1951</xref>).</p>
</sec>
<sec>
<title>Intrinsic rewards</title>
<p>Intrinsic rewards are unconditioned rewards that are attractive and motivate behavior because they are inherently pleasurable (e.g., the experience of joy). The phenomenon of intrinsic motivation was first identified in studies of animals engaging in exploratory, playful and curiosity-driven behavior in the absence of external rewards or punishments (White, <xref ref-type="bibr" rid="B363">1959</xref>).</p>
</sec>
<sec>
<title>Extrinsic rewards</title>
<p>Extrinsic rewards are conditioned rewards that motivate behavior but are not inherently pleasurable (e.g., praise or monetary reward). They acquire their value through learned association with intrinsic rewards. Hence, extrinsic motivation refers to our tendency to perform activities for known external rewards, whether they be tangible or psychological in nature (Brown, <xref ref-type="bibr" rid="B47">2007</xref>).</p>
<p>Summarizing, the continual competition between multiple drives and incentives that have adaptive value to the organism and are realized by dedicated neural circuits is what ultimately generates behavior (Davies et al., <xref ref-type="bibr" rid="B65">2012</xref>). In humans, the evolutionary and cultural pressures that shaped our own intrinsic and extrinsic motivations have allowed us to reach great achievements, ranging from our mastery of the laws of nature to expressions of great beauty as encountered in the liberal arts. The question remains how we can gain an understanding of how our brains generate the rich behavioral repertoire that can be observed in nature.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3. Understanding natural intelligence</title>
<p>In a way, the recipe for understanding natural intelligence and achieving strong AI is simple. If we can construct synthetic brains that mimic the adaptive behavior displayed by biological brains in all its splendor then our mission has succeeded. This entails equipping synthetic brains with the same special purpose computing machinery encountered in real brains, solving those problems an agent may be faced with. In practice, of course, this is easier said than done given the incomplete state of our knowledge and the daunting complexity of biological systems.</p>
<sec>
<title>3.1. Levels of analysis</title>
<p>The neural circuits that make up the human brain can be seen as special-purpose devices that together guarantee the selection of (near-)optimal actions. David Marr in particular advocated the view that the nervous system should be understood as a collection of information processing systems that solve particular problems an organism is faced with (Marr, <xref ref-type="bibr" rid="B210">1982</xref>). His work gave rise to the field of computational neuroscience and has been highly influential in shaping ideas about neural information processing (Willshaw et al., <xref ref-type="bibr" rid="B368">2015</xref>). Marr and Poggio (<xref ref-type="bibr" rid="B211">1976</xref>) proposed that an understanding of information processing systems should take place at distinct levels of analysis, namely the <italic>computational level</italic>, which specifies what problem the system solves, the <italic>algorithmic level</italic>, which specifies how the system solves the problem, and the <italic>implementational level</italic>, which specifies how the system is physically realized.</p>
<p>A canonical example of a three-level analysis is prey localization in the barn owl (Grothe, <xref ref-type="bibr" rid="B125">2003</xref>). At the computational level, the owl needs to use auditory information to localize its prey. At the algorithmic level, this can be implemented by circuits composed of delay lines and coincidence detectors that detect inter-aural time differences (Jeffress, <xref ref-type="bibr" rid="B154">1948</xref>). At the implementational level, neurons in the nucleus laminaris have been shown to act as coincidence detectors (Carr and Konishi, <xref ref-type="bibr" rid="B51">1990</xref>).</p>
<p>Marr&#x00027;s levels of analysis sidestep one important point, namely how a system gains the ability to solve a computational problem in the first place. That is, it is also crucial to understand how an organism (or species as a whole) is able to learn and evolve the computations and representations that allow it to survive in the natural world (Poggio, <xref ref-type="bibr" rid="B266">2012</xref>). Learning itself takes place at the level of the individual organism as well as of the species. In the individual, one can observe lasting changes in the brain throughout its lifetime, which is referred to as neural plasticity. At the species level, natural selection is responsible for evolving the mechanisms that are involved in neural plasticity (Poggio, <xref ref-type="bibr" rid="B266">2012</xref>). As argued by Poggio, an understanding at the level of learning in the individual and the species is sufficiently powerful to solve a problem and can thereby act as an explanation of natural intelligence. To illustrate the relevance of this revised model, in the prey localization example it would be imperative to understand how owls are able to adapt to changes in their environment (Huo and Murray, <xref ref-type="bibr" rid="B149">2009</xref>), as well as how owls were equipped with such machinery during evolution.</p>
<p>Sun et al. (<xref ref-type="bibr" rid="B330">2005</xref>) propose an alternative organization of levels of cognitive modeling. They distinguish sociological, psychological, componential and physiological levels. The sociological level refers to the collective behavior of agents, including interactions between agents as well as their environment. It stresses the importance of socio-cultural processes in shaping cognition. The psychological level covers individual behaviors, beliefs, concepts, and skills. The componential level describes inter-agent processes specified in terms of Marr&#x00027;s computational and algorithmic levels. Finally, the physiological level describes the biological substrate which underlies the generation of adaptive behavior, corresponding to Marr&#x00027;s implementational level. It can provide valuable input about important computations and plausible architectures at a higher level of abstraction.</p>
<p>Figure <xref ref-type="fig" rid="F3">3</xref> visualizes the different interpretations of levels of analysis. Without committing to a definitive stance on levels of analysis, all described levels provide important complementary perspectives concerning the modeling and understanding of natural intelligence.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Levels of analysis. Left column shows Poggio&#x00027;s extension of Marr&#x00027;s levels of analysis, emphasizing learning at various timescales. Right column shows Sun&#x00027;s levels of analysis, emphasizing individual beliefs and socio-cultural processes.</p></caption>
<graphic xlink:href="fncom-11-00112-g0003.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Modeling approaches</title>
<p>The previous section suggests that different approaches to understanding natural intelligence and developing cognitive architectures can be taken depending on the levels of analysis one considers. We briefly review a number of core approaches.</p>
<sec>
<title>Artificial life</title>
<p>Artificial life is a broad area of research encompassing various different modeling strategies which all have in common that they aim to explain the emergence of life and, ultimately, cognition in a bottom-up manner (Steels, <xref ref-type="bibr" rid="B323">1993</xref>; Bedau, <xref ref-type="bibr" rid="B26">2003</xref>).</p>
<p>A canonical example of an artificial life system is the cellular automaton, first introduced by von Neumann (<xref ref-type="bibr" rid="B359">1966</xref>) as an approach to understand the fundamental properties of living systems. Cellular automata operate within a universe consisting of cells, whose states change over multiple generations based on simple local rules. They have been shown to be capable of acting as universal Turing machines, thereby giving them the capacity to compute any fixed partial computable function (Wolfram, <xref ref-type="bibr" rid="B371">2002</xref>).</p>
<p>A famous example of a cellular automaton is Conway&#x00027;s Game of Life. Here, every cell can assume an &#x0201C;alive&#x0201D; or a &#x0201C;dead&#x0201D; state. State changes are determined by its interactions with its eight direct neighbors. At each time step, a live cell with fewer than two or more than three live neighbors dies and a dead cell with exactly three live neighbors will become alive. Figure <xref ref-type="fig" rid="F4">4</xref> shows an example of a breeder pattern which produces Gosper guns in the Game of Life. Gosper guns have been used to prove that the game of life is Turing complete (Gardner, <xref ref-type="bibr" rid="B109">2001</xref>). SmoothLife (Rafler, <xref ref-type="bibr" rid="B271">2011</xref>), as a continuous-space extension of the Game of Life, shows emerging structures that bear some superficial resemblance to biological structures.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Example of the Game of Life, where each cell state evolves according to a set of deterministic rules that depend on the states of neighboring cells. Depicted is a <italic>breeder</italic> pattern that moves across the universe (here from left to right), leaving behind debris. The breeder produces <italic>Gosper guns</italic> which periodically emit <italic>gliders</italic>; the small patterns that together form the triangular shape on the left-hand side.</p></caption>
<graphic xlink:href="fncom-11-00112-g0004.tif"/>
</fig>
<p>In principle, by virtue of their universality, cellular automata offer the capacity to explain how self-replicating adaptive (autopoeietic, Maturana and Varela, <xref ref-type="bibr" rid="B213">1980</xref>) systems emerge from basic rules. This bottom-up approach is also taken by physicists who aim to explain life and, ultimately, cognition purely from thermodynamic principles (Dewar, <xref ref-type="bibr" rid="B75">2003</xref>, <xref ref-type="bibr" rid="B76">2005</xref>; Grinstein and Linsker, <xref ref-type="bibr" rid="B124">2007</xref>; Wissner-Gross and Freer, <xref ref-type="bibr" rid="B370">2013</xref>; Perunov et al., <xref ref-type="bibr" rid="B263">2014</xref>; Fry, <xref ref-type="bibr" rid="B101">2017</xref>).</p>
</sec>
<sec>
<title>Biophysical modeling</title>
<p>A more direct way to model natural intelligence is to presuppose the existence of the building blocks of life which can be used to create realistic simulations of organisms <italic>in silico</italic>. The reasoning is that biophysically realistic models can eventually mimic the information processing capabilities of biological systems. An example thereof is the OpenWorm project which has as its ambition to understand how the behavior of <italic>C. elegans</italic> emerges from its underlying physiology purely via bottom-up biophysical modeling (Szigeti et al., <xref ref-type="bibr" rid="B337">2014</xref>) (Figure <xref ref-type="fig" rid="F5">5A</xref>). It also acknowledges the importance of including not only a model of the worm&#x00027;s nervous system but also of its body and environment in the simulation. That is, adaptive behavior depends on the organism being both embodied and embedded in the world (Anderson, <xref ref-type="bibr" rid="B12">2003</xref>). If successful, then this project would constitute the first example of a digital organism.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Biophysical modeling. <bold>(A)</bold> Body plan of <italic>C. elegans</italic><xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>. The OpenWorm project aims to provide an accurate bottom-up simulation of the worm acting in its environment. <bold>(B)</bold> Example of action potential generation via the Hodgkin-Huxley equations in the presence of a constant input current.</p></caption>
<graphic xlink:href="fncom-11-00112-g0005.tif"/>
</fig>
<p>It is a long stretch from the worm&#x00027;s 302 neurons to the 86 billion neurons that comprise the human brain (Herculano-Houzel and Lent, <xref ref-type="bibr" rid="B138">2005</xref>). Still, researchers have set out to develop large-scale models of the human brain. Biophysical modeling can be used to create detailed models of neurons and their processes using coupled systems of differential equations. For example, action potential generation can be described in terms of the Hodgkin-Huxley equations (Figure <xref ref-type="fig" rid="F5">5B</xref>) and the flow of electric current along neuronal fibers can be modeled using cable theory (Dayan and Abbott, <xref ref-type="bibr" rid="B69">2005</xref>). This approach is used in the Blue Brain project (Markram, <xref ref-type="bibr" rid="B207">2006</xref>) and its successor, the Human Brain Project (HBP) (Amunts et al., <xref ref-type="bibr" rid="B10">2016</xref>). See de Garis et al. (<xref ref-type="bibr" rid="B71">2010</xref>) for a review of various artificial brain projects.</p>
</sec>
<sec>
<title>Connectionism</title>
<p>Connectionism refers to the explanation of cognition as arising from the interplay between basic (sub-symbolic) processing elements (Smolensky, <xref ref-type="bibr" rid="B316">1987</xref>; Bechtel, <xref ref-type="bibr" rid="B25">1993</xref>). It has close links to cybernetics, which focuses on the development of control structures from which intelligent behavior emerges (Rid, <xref ref-type="bibr" rid="B279">2016</xref>).</p>
<p>Connectionism came to be equated with the use of artificial neural networks that abstract away from the details of biological neural networks. An artificial neural network (ANN) is a computational model which is loosely inspired by the human brain as it consists of an interconnected network of simple processing units (artificial neurons) that learns from experience by modifying its connections. Alan Turing was one of the first to propose the construction of computing machinery out of trainable networks consisting of neuron-like elements (Copeland and Proudfoot, <xref ref-type="bibr" rid="B58">1996</xref>). Marvin Minsky, one of the founding fathers of AI, is credited for building the first trainable ANN, called SNARC, out of tubes, motors, and clutches (Seising, <xref ref-type="bibr" rid="B308">2017</xref>).</p>
<p>Artificial neurons can be considered abstractions of (populations of) neurons while the connections are taken to be abstractions of modifiable synaptic connections (Figure <xref ref-type="fig" rid="F6">6</xref>). The behavior of an artificial neuron is fully determined by the connection strengths as well as how input is transformed into output. Contrary to detailed biophysical models, ANNs make use of basic matrix operations and nonlinear transformations as their fundamental operations. In its most basic incarnation, an artificial neuron simply transforms its input <bold>x</bold> into a response <italic>y</italic> through an activation function <italic>f</italic>, as shown in Figure <xref ref-type="fig" rid="F6">6</xref>. The activation function operates on an input activation which is typically taken to be the inner product between the input <bold>x</bold> and the parameters (weight vector) <bold>w</bold> of the artificial neuron. The weights are interpreted as synaptic strengths that determine how presynaptic input is translated into postsynaptic firing rate. This yields a simple linear-nonlinear mapping of the form
<disp-formula id="E1"><label>(1)</label><mml:math id="M26"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>w</mml:mtext></mml:mstyle><mml:mtext style="font-family:sans-serif">T</mml:mtext></mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
By connecting together multiple neurons, one obtains a neural network that implements some non-linear function <bold>y</bold> &#x0003D; <bold>f</bold>(<bold>x</bold>; <italic><bold>&#x003B8;</bold></italic>), where the <italic>f</italic><sub><italic>i</italic></sub> are nonlinear transformations and <italic><bold>&#x003B8;</bold></italic> stands for the network parameters (i.e., weight vectors). After training a neural network, representations become encoded in a distributed manner as a pattern which manifests itself across all its neurons (Hinton et al., <xref ref-type="bibr" rid="B141">1986</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Artificial neural networks (Yuste, <xref ref-type="bibr" rid="B379">2015</xref>). <bold>(A)</bold> Feedforward neural networks map inputs to outputs using nonlinear transformations. <bold>(B)</bold> Recurrent neural networks implement dynamical systems by feeding back output activity to the input layer, where it is combined with external input.</p></caption>
<graphic xlink:href="fncom-11-00112-g0006.tif"/>
</fig>
<p>Throughout the course of their history ANNs have fallen in and out of favor multiple times. At the same time, each next generation of neural networks has yielded new insights about how complex behavior may emerge through the collective action of simple processing elements. Modern neural networks perform so well on several benchmark problems that they obliterate all competition in, e.g., object recognition (Krizhevsky et al., <xref ref-type="bibr" rid="B177">2012</xref>), natural language processing (Sutskever et al., <xref ref-type="bibr" rid="B332">2014</xref>), game playing (Mnih et al., <xref ref-type="bibr" rid="B229">2015</xref>; Silver et al., <xref ref-type="bibr" rid="B312">2017</xref>) and robotics (Levine et al., <xref ref-type="bibr" rid="B192">2015</xref>), often matching and sometimes surpassing human-level performance (LeCun et al., <xref ref-type="bibr" rid="B186">2015</xref>). Their success relies on combining classical ideas (Widrow and Lehr, <xref ref-type="bibr" rid="B366">1990</xref>; Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B144">1997</xref>; LeCun et al., <xref ref-type="bibr" rid="B187">1998</xref>) with new algorithmic developments (Hinton et al., <xref ref-type="bibr" rid="B142">2006</xref>; Srivastava et al., <xref ref-type="bibr" rid="B321">2014</xref>; He et al., <xref ref-type="bibr" rid="B136">2015</xref>; Ioffe and Szegedy, <xref ref-type="bibr" rid="B151">2015</xref>; Zagoruyko and Komodakis, <xref ref-type="bibr" rid="B380">2017</xref>), while using high-performance graphical processing units (GPUs) to massively speed up training of ANNs on big datasets (Raina et al., <xref ref-type="bibr" rid="B273">2009</xref>).</p>
</sec>
<sec>
<title>Cognitivism</title>
<p>A conceptually different approach to the explanation of cognition as emerging from bottom-up principles is the view that cognition should be understood in terms of formal symbol manipulation. This computationalist view is associated with the cognitivist program which arose in response to earlier behaviorist theories. It embraces the notion that, in order to understand natural intelligence, one should study internal mental processes rather than just externally observable events. That is, cognitivism asserts that cognition should be defined in terms of formal symbol manipulation, where reasoning involves the manipulation of symbolic representations that refer to information about the world as acquired by perception.</p>
<p>This view is formalized by the physical symbol system hypothesis (Newell and Simon, <xref ref-type="bibr" rid="B243">1976</xref>), which states that &#x0201C;a physical symbol system has the necessary and sufficient means for intelligent action.&#x0201D; This hypothesis implies that artificial agents, when equipped with the appropriate symbol manipulation algorithms, will be capable of displaying intelligent behavior. As Newell and Simon (<xref ref-type="bibr" rid="B243">1976</xref>) wrote, the physical symbol system hypothesis also implies that &#x0201C;the symbolic behavior of man arises because he has the characteristics of a physical symbol system.&#x0201D; This also suggests that the specifics of our nervous system are not relevant for explaining adaptive behavior (Simon, <xref ref-type="bibr" rid="B314">1996</xref>).</p>
<p>Cognitivism gave rise to cognitive science as well as artificial intelligence, and spawned various cognitive architectures such as ACT-R (Anderson et al., <xref ref-type="bibr" rid="B11">2004</xref>) (see Figure <xref ref-type="fig" rid="F7">7</xref>) and SOAR (Laird, <xref ref-type="bibr" rid="B180">2012</xref>) that employ rule-based approaches in the search for a unified theory of cognition (Newell, <xref ref-type="bibr" rid="B242">1991</xref>).<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref></p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>ACT-R as an example cognitive architecture which employs symbolic reasoning. ACT-R interfaces with different modules through buffers. Cognition unfolds as a succession of activations of production rules as mediated by pattern matching and execution<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref>.</p></caption>
<graphic xlink:href="fncom-11-00112-g0007.tif"/>
</fig>
</sec>
<sec>
<title>Probabilistic modeling</title>
<p>Modern cognitive science still embraces the cognitivist program but has since taken a probabilistic approach to the modeling of cognition. As stated by Griffiths et al. (<xref ref-type="bibr" rid="B123">2010</xref>), this probabilistic approach starts from the notion that the challenges faced by the mind are often of an inductive nature, where the observed data are not sufficient to unambiguously identify the process that generated them. This precludes the use of approaches that are founded on mathematical logic and requires a quantification of the state of the world in terms of degrees of belief as afforded by probability theory (Jaynes, <xref ref-type="bibr" rid="B153">1988</xref>). The probabilistic approach operates by identifying a hypothesis space representing solutions to the inductive problem. It then prescribes how an agent should revise her belief in the hypotheses given the information provided by observed data. Hypotheses are typically formulated in terms of probabilistic graphical models that capture the independence structure between random variables of interest (Koller and Friedman, <xref ref-type="bibr" rid="B175">2009</xref>). An example of such a graphical model is shown in Figure <xref ref-type="fig" rid="F8">8</xref>.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Example of a probabilistic graphical model capturing the statistical relations between random variables of interest. This particular plate model describes a smoothed version of latent Dirichlet allocation as used in topic modeling (Blei et al., <xref ref-type="bibr" rid="B34">2003</xref>). Here, &#x003B1; and &#x003B2; are hyper-parameters, &#x003B8;<sub><italic>m</italic></sub> is the topic distribution for document <italic>m</italic>, &#x003D5;<sub><italic>k</italic></sub> is the word distribution for topic <italic>k</italic>, <italic>z</italic><sub><italic>nm</italic></sub> is the topic for the <italic>n</italic>-th word in document <italic>m</italic> and <italic>w</italic><sub><italic>mn</italic></sub> is a specific word. Capital letters <italic>K</italic>, <italic>M</italic> and <italic>N</italic> denote the number of topics, documents and words, respectively. The goal is to discover abstract topics from observed words. This general approach of inferring posteriors over latent variables from observed data is common to the probabilistic approach.</p></caption>
<graphic xlink:href="fncom-11-00112-g0008.tif"/>
</fig>
<p>Belief updating in the probabilistic sense is realized by solving a statistical inference problem. Consider a set of of hypotheses <inline-formula><mml:math id="M2"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> that might explain the observed data. Let <italic>p</italic>(<italic>h</italic>) denote our belief in a hypothesis <inline-formula><mml:math id="M3"><mml:mi>h</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula>, reflecting the state of the world, before observing any data (known as the <italic>prior</italic>). Let <italic>p</italic>(<bold>x</bold> &#x02223; <italic>h</italic>) indicate the probability of observing data <bold>x</bold> if <italic>h</italic> were true (known as the <italic>likelihood</italic>). Bayes&#x00027; rule tells us how to update our belief in a hypothesis after observing data. It states that the <italic>posterior probability p</italic>(<italic>h</italic> &#x02223; <bold>x</bold>) assigned to <italic>h</italic> after observing <bold>x</bold> should be
<disp-formula id="E2"><label>(2)</label><mml:math id="M27"><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x02223;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mo>&#x02223;</mml:mo><mml:mi>h</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>h</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mo>&#x02223;</mml:mo><mml:mi>h</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
where the denominator is a normalizing constant known as the <italic>evidence</italic> or <italic>marginal likelihood</italic><xref ref-type="fn" rid="fn0005"><sup>5</sup></xref>. Importantly, it can be shown that degrees of belief are coherent only if they satisfy the axioms of probability theory (Ramsey, <xref ref-type="bibr" rid="B275">1926</xref>).</p>
<p>The beauty of the probabilistic approach lies in its generality. It not only explains how our moment-to-moment percepts change as a function of our prior beliefs and incoming sensory data (Yuille and Kersten, <xref ref-type="bibr" rid="B378">2006</xref>) but also places learning, as the construction of internal models, under the same umbrella by viewing it as an inference problem (MacKay, <xref ref-type="bibr" rid="B201">2003</xref>). In the probabilistic framework, mental processes are modeled using algorithms for approximating the posterior (Koller and Friedman, <xref ref-type="bibr" rid="B175">2009</xref>) and neural processes are seen as mechanisms for implementing these algorithms (Gershman and Beck, <xref ref-type="bibr" rid="B112">2016</xref>).</p>
<p>The probabilistic approach also provides a basis for making optimal decisions under uncertainty. This is realized by extending probability theory with decision theory. According to decision theory, a rational agent ought to select that action which maximizes the expected utility (von Neumann and Morgenstern, <xref ref-type="bibr" rid="B360">1953</xref>). This is known as the maximum expected utility (MEU) principle. In real-life situations, biological (and artificial) agents need to operate under bounded resources, trading off precision for speed and effort when trying to attain their objectives (Gigerenzer and Goldstein, <xref ref-type="bibr" rid="B117">1996</xref>). This implies that MEU calculations may be intractable. Intractability issues have led to the development of algorithms that maximize a more general form of expected utility which incorporates the costs of computation. These algorithms can in turn be adapted so as to select the best approximation strategy in a given situation (Gershman et al., <xref ref-type="bibr" rid="B113">2015</xref>). Hence, at the algorithmic level, it has been postulated that brains use approximate inference algorithms (Andrieu et al., <xref ref-type="bibr" rid="B13">2003</xref>; Blei et al., <xref ref-type="bibr" rid="B33">2016</xref>) such as to produce good enough solutions for fast and frugal decision making.</p>
<p>Summarizing, by appealing to Bayesian statistics and decision theory, while acknowledging the constraints biological agents are faced with, cognitive science arrives at a theory of bounded rationality that agents should adhere to. Importantly, this normative view dictates that organisms must operate as Bayesian inference machines that aim to maximize expected utility. If they do not, then, under weak assumptions, they will perform suboptimally. This would be detrimental from an evolutionary point of view.</p>
</sec>
</sec>
<sec>
<title>3.3. Bottom-up emergence vs. top-down abstraction</title>
<p>The aforementioned modeling strategies each provide an alternative approach toward understanding natural intelligence and achieving strong AI. The question arises which of these strategies will be most effective in the long run.</p>
<p>While the strictly bottom-up approach used in artificial life research may lead to fundamental insights about the nature of self-replication and adaptability, in practice it remains an open question how emergent properties that derive from a basic set of rules can reach the same level of organization and complexity as can be found in biological organisms. Furthermore, running such simulations would be extremely costly from a computational point of view.</p>
<p>The same problem presents itself when using detailed biophysical models. That is, bottom-up approaches must either restrict model complexity or run simulations for limited periods of time in order to remain tractable (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B250">2012</xref>). Biophysical models additionally suffer from a lack of data. For example, the original aim of the Human Brain Project was to model the human brain within a decade (Markram et al., <xref ref-type="bibr" rid="B208">2011</xref>). This ambition may be hard to realize given the plethora of data required for model estimation. Furthermore, the resulting models may be difficult to link to cognitive function. Izhikevich, reflecting on his simulation of another large biophysically realistic brain model (Izhikevich and Edelman, <xref ref-type="bibr" rid="B152">2008</xref>), states: &#x0201C;Indeed, no significant contribution to neuroscience could be made by simulating one second of a model, even if it has the size of the human brain. However, I learned what it takes to simulate such a large-scale system<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>.&#x0201D;</p>
<p>Connectionist models, in contrast, abstract away from biophysical details, thereby making it possible to train large-scale models on large amounts of sensory data, allowing cognitively challenging tasks to be solved. Due to their computational simplicity, they are also more amenable to theoretical analysis (Hertz et al., <xref ref-type="bibr" rid="B139">1991</xref>; Bishop, <xref ref-type="bibr" rid="B32">1995</xref>). At the same time, connectionist models have been criticized for their inability to capture symbolic reasoning, their limitations when modeling particular cognitive phenomena, and their abstract nature, restricting their biological plausibility (Dawson and Shamanski, <xref ref-type="bibr" rid="B68">1994</xref>).</p>
<p>Cognitivism has been pivotal in the development of intelligent systems. However, it has also been criticized using the argument that systems which operate via formal symbol manipulation lack intentionality (Searle, <xref ref-type="bibr" rid="B306">1980</xref>)<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref>. Moreover, the representational framework that is used is typically constructed by a human designer. While this facilitates model interpretation, at the same time, this programmer-dependence may bias the system, leading to suboptimal solutions. That is, idealized descriptions may induce a semantic gap between perception and possible interpretation (Vernon et al., <xref ref-type="bibr" rid="B356">2007</xref>).</p>
<p>The probabilistic approach to cognition is important given its ability to define normative theories at the computational level. At the same time, it has also been criticized for its treatment of cognition as if it is in the business of selecting some statistical model. Proponents of connectionism argue that computation-level explanations of behavior that ignore mechanisms associated with bottom-up emergence are likely to fall short (McClelland et al., <xref ref-type="bibr" rid="B216">2010</xref>).</p>
<p>The different approaches provide complementary insights into the nature of natural intelligence. Artificial life informs about fundamental bottom-up principles, biophysical models make explicit how cognition is realized via specific mechanisms at the molecular and systems level, connectionist models show how problem solving capacities emerge from the interactions between basic processing elements, cognitivism emphasizes the importance of symbolic reasoning and probabilistic models inform how particular problems could be solved in an optimal manner.</p>
<p>Notwithstanding potential limitations, given their ability to solve complex cognitively challenging problems, connectionist models are taken to provide a promising starting point for understanding natural intelligence and achieving strong AI. They also naturally connect to the different modeling strategies. That is, they connect to artificial life principles by having network architectures emerge through evolutionary strategies (Real et al., <xref ref-type="bibr" rid="B277">2016</xref>; Salimans et al., <xref ref-type="bibr" rid="B287">2017</xref>) and connect to the biophysical level by viewing them as (rate-based) abstractions of biological neural networks (Dayan and Abbott, <xref ref-type="bibr" rid="B69">2005</xref>). They also connect to the computational level by grounding symbolic representations in real-world sensory states (Harnad, <xref ref-type="bibr" rid="B133">1990</xref>) and connect to the probabilistic approach through the observation that emergent computations effectively approximate Bayesian inference (Gal, <xref ref-type="bibr" rid="B107">2016</xref>; Orhan and Ma, <xref ref-type="bibr" rid="B252">2016</xref>; Ambrogioni et al., <xref ref-type="bibr" rid="B9">2017</xref>; Mandt et al., <xref ref-type="bibr" rid="B202">2017</xref>). It is for these reasons that, in the following, we will explore how ANNs, as canonical connectionist models, can be used to promote our understanding of natural intelligence.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Ann-based modeling of cognitive processes</title>
<p>We will now explore in more detail the ways in which ANNs can be used to understand and model aspects of natural intelligence. We start by addressing how neural networks can learn from data.</p>
<sec>
<title>4.1. Learning</title>
<p>The capacity of brains to behave adaptively relies on their ability to modify their own behavior based on changing circumstances. The appeal of neural networks stems from their ability to mimic this learning behavior in an efficient manner by updating network parameters <italic><bold>&#x003B8;</bold></italic> based on available data <inline-formula><mml:math id="M5"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, allowing the construction of large models that are able to solve complex cognitive tasks.</p>
<p>Learning proceeds by making changes to the network parameters <italic><bold>&#x003B8;</bold></italic> such that its output starts to agree more and more with the objectives of the agent at hand. This is formalized by assuming the existence of a cost function <inline-formula><mml:math id="M6"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> which measures the degree to which an agent deviates from its objectives. <inline-formula><mml:math id="M7"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow></mml:math></inline-formula> is computed by running a neural network in forward mode (from input to output) and comparing the predicted output with the desired output. During its lifetime, the agent obtains data from its environment (sensations) by sampling from a data-generating distribution <italic>p</italic><sub>data</sub>. The goal of an agent is to reduce the expected risk
<disp-formula id="E3"><label>(3)</label><mml:math id="M28"><mml:mrow><mml:msup><mml:mi mathvariant="-tex-caligraphic">J</mml:mi><mml:mo>&#x0002A;</mml:mo></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant='double-struck'>E</mml:mi><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>~</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mtext>data</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x02113;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
where &#x02113; is the incurred loss per datapoint <bold>z</bold>. In practice, an agent only has access to a finite number of datapoints which the agent experiences during its lifetime, yielding a training set <inline-formula><mml:math id="M9"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:math></inline-formula>. This training set can be represented in the form of an empirical distribution <inline-formula><mml:math id="M10"><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> which equals 1/<italic>N</italic> if <bold>z</bold> is equal to one of the <italic>N</italic> examples and zero otherwise. In practice, the aim therefore is to minimize the empirical risk
<disp-formula id="E4"><label>(4)</label><mml:math id="M29"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant='double-struck'>E</mml:mi><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>~</mml:mo><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x02113;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
as an approximation of <inline-formula><mml:math id="M12"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. In reality, the brain is thought to optimize a multitude of cost functions pertaining to the many objectives it aims to achieve in concert (Marblestone et al., <xref ref-type="bibr" rid="B204">2016</xref>).</p>
<p>Risk minimization can be accomplished by making use of a gradient descent procedure. Let <italic><bold>&#x003B8;</bold></italic> be the parameters of a neural network (i.e., the synaptic weights). We can define learning as a search for the optimal parameters <bold>&#x003B8;<sup>&#x0002A;</sup></bold> based on available training data <inline-formula><mml:math id="M13"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:math></inline-formula> such that
<disp-formula id="E5"><label>(5)</label><mml:math id="M30"><mml:mrow><mml:msup><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>&#x0002A;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:munder><mml:mrow><mml:mi>min</mml:mi></mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:munder><mml:mi mathvariant="-tex-caligraphic">J</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>A convenient way to approximate <bold>&#x003B8;<sup>&#x0002A;</sup></bold> is by measuring locally the change in slope of <inline-formula><mml:math id="M15"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> as a function of <italic><bold>&#x003B8;</bold></italic> and taking a step in the direction of steepest descent. This procedure, known as <italic>gradient descent</italic>, is based on the observation that if <inline-formula><mml:math id="M16"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow></mml:math></inline-formula> is defined and differentiable in the neighborhood of a point <italic><bold>&#x003B8;</bold></italic>, then <inline-formula><mml:math id="M17"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow></mml:math></inline-formula> decreases fastest if one goes from <italic><bold>&#x003B8;</bold></italic> in the direction of the negative gradient <inline-formula><mml:math id="M18"><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. In other words, if we use the update rule
<disp-formula id="E6"><label>(6)</label><mml:math id="M31"><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>&#x02190;</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003F5;</mml:mi><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:msub><mml:mi mathvariant="-tex-caligraphic">J</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula>
with small enough learning rate &#x003F5; then <italic><bold>&#x003B8;</bold></italic> is guaranteed to converge to a (local) minimum of <inline-formula><mml:math id="M20"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">J</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula><xref ref-type="fn" rid="fn0008"><sup>8</sup></xref>. Importantly, the gradient can be computed for arbitrary ANN architectures by running the network in backward mode (from output to input) and computing the gradient using automatic differentiation procedures. This forms the basis of the widely used backpropagation algorithm (Widrow and Lehr, <xref ref-type="bibr" rid="B366">1990</xref>).</p>
<p>One might argue that the backpropagation algorithm fails to connect to learning in biology due to implausible assumptions such as the fact that forward and backward passes use the same set of synaptic weights. There are a number of responses here. First, one might hold the view that backpropagation is just an efficient way to obtain effective network architectures, without committing to the biological plausibility of the learning algorithm <italic>per se</italic>. Second, if biologically plausible learning is the research objective then one is free to exploit other (Hebbian) learning schemes that may reflect biological learning more closely (Miconi, <xref ref-type="bibr" rid="B222">2017</xref>). Finally, researchers have started to put forward arguments that backpropagation may not be that biologically implausible after all (Roelfsema and van Ooyen, <xref ref-type="bibr" rid="B283">2005</xref>; Lillicrap et al., <xref ref-type="bibr" rid="B194">2016</xref>; Scellier and Bengio, <xref ref-type="bibr" rid="B292">2017</xref>).</p>
</sec>
<sec>
<title>4.2. Perceiving</title>
<p>One of the core skills any intelligent agent should possess is the ability to recognize patterns in its environment. The world around us consists of various objects that may carry significance. Being able to recognize edible food, places that provide shelter, and other agents will all aid survival.</p>
<p>Biological agents are faced with the problem that they need to be able to recognize objects from raw sensory input (vectors in &#x0211D;<sup><italic>n</italic></sup>). How can a brain use the incident sensory input to learn to recognize those things that are of relevance to the organism? Recall the artificial neuron formulation <italic>y</italic> &#x0003D; <italic>f</italic>(<bold>w</bold><inline-formula><mml:math id="M250"><mml:mrow><mml:mtext style="font-family:sans-serif">T</mml:mtext></mml:mrow></mml:math></inline-formula><bold>x</bold>). By learning proper weights <bold>w</bold>, this neuron can learn to distinguish different object categories. This is essentially equivalent to a classical model known as the perceptron (Rosenblatt, <xref ref-type="bibr" rid="B284">1958</xref>), which was used to solve simple pattern recognition problems via a simple error-correction mechanism. It also corresponds to a basic linear-nonlinear (LN) model which has been used extensively to model and estimate the receptive field of a neuron or a population of neurons (van Gerven, <xref ref-type="bibr" rid="B352">2017</xref>).</p>
<p>Single-layer ANNs such as the perceptron are capable of solving interesting learning problems. At the same time, they are limited in scope since they can only solve linearly separable classification problems (Minsky and Papert, <xref ref-type="bibr" rid="B226">1969</xref>). To overcome the limitations of the perceptron we can extend its capabilities by relaxing the constraint that the inputs are directly coupled to the outputs. A multilayer perceptron (MLP) is a feedforward network which generalizes the standard perceptron by having a hidden layer that resides between the input and the output layers. We can write an MLP with multiple output units as
<disp-formula id="E7"><label>(7)</label><mml:math id="M32"><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>y</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>g</mml:mtext></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>Wf</mml:mtext></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>Vx</mml:mtext></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
where <bold>V</bold> denotes the hidden layer weights and <bold>w</bold> denotes the output layer weights. By introducing a hidden layer, MLPs gain the ability to learn internal representations (Rumelhart et al., <xref ref-type="bibr" rid="B285">1986</xref>). Importantly, an MLP can approximate any continuous function to an arbitrary degree of accuracy, given a sufficiently large but finite number of hidden neurons (Cybenko, <xref ref-type="bibr" rid="B63">1989</xref>; Hornik, <xref ref-type="bibr" rid="B146">1991</xref>).</p>
<p>Complex systems tend to be hierarchical and modular in nature (Simon, <xref ref-type="bibr" rid="B313">1962</xref>). The nervous system itself can be thought of as a hierarchically organized system. This is exemplified by Felleman &#x00026; van Essen&#x00027;s hierarchical diagram of visual cortex (Felleman and Van Essen, <xref ref-type="bibr" rid="B89">1991</xref>), the proposed hierarchical organization of prefrontal cortex (Badre, <xref ref-type="bibr" rid="B18">2008</xref>), the view of the motor system as a behavioral control column (Swanson, <xref ref-type="bibr" rid="B334">2000</xref>) and the proposition that anterior and posterior cortex reflect hierarchically organized executive and perceptual systems (Fuster, <xref ref-type="bibr" rid="B105">2001</xref>). Representations at the top of these hierarchies correspond to highly abstract statistical invariances that occupy our ecological niche (Quian Quiroga et al., <xref ref-type="bibr" rid="B270">2005</xref>; Barlow, <xref ref-type="bibr" rid="B21">2009</xref>). A hierarchy can be modeled by a deep neural network (DNN) composed of multiple hidden layers (LeCun et al., <xref ref-type="bibr" rid="B186">2015</xref>), written as
<disp-formula id="E8"><label>(8)</label><mml:math id="M33"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>y</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>f</mml:mtext></mml:mstyle><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>w</mml:mtext></mml:mstyle><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>f</mml:mtext></mml:mstyle><mml:mi>L</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>w</mml:mtext></mml:mstyle><mml:mi>L</mml:mi></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>f</mml:mtext></mml:mstyle><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>w</mml:mtext></mml:mstyle><mml:mn>1</mml:mn></mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x022EF;</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>f</mml:mtext></mml:mstyle><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <bold>W</bold><sub><italic>l</italic></sub> is the weight matrix associated with layer <italic>l</italic>. Even though an MLP can already approximate any function to an arbitrary degree of precision, it has been shown that many classes of functions can be represented much more compactly using thin and deep neural networks compared to shallow and wide neural networks (Bengio and LeCun, <xref ref-type="bibr" rid="B29">2007</xref>; Bengio, <xref ref-type="bibr" rid="B27">2009</xref>; Le Roux and Bengio, <xref ref-type="bibr" rid="B185">2010</xref>; Delalleau and Bengio, <xref ref-type="bibr" rid="B72">2011</xref>; Mhaskar et al., <xref ref-type="bibr" rid="B221">2016</xref>).</p>
<p>A DNN corresponds to a stack of LN models, generalizing the concept of basic receptive field models. They have been shown to yield human-level performance on object categorization tasks (Krizhevsky et al., <xref ref-type="bibr" rid="B177">2012</xref>). The latest DNN incarnations are even capable of predicting the cognitive states of other agents. One example is the prediction of apparent personality traits from multimodal sensory input (G&#x000FC;&#x000E7;l&#x000FC;t&#x000FC;rk et al., <xref ref-type="bibr" rid="B131">2016</xref>). Deep architectures have been used extensively in neuroscience to model hierarchical processing (Selfridge, <xref ref-type="bibr" rid="B309">1959</xref>; Fukushima, <xref ref-type="bibr" rid="B102">1980</xref>, <xref ref-type="bibr" rid="B103">2013</xref>; Riesenhuber and Poggio, <xref ref-type="bibr" rid="B280">1999</xref>; Lehky and Tanaka, <xref ref-type="bibr" rid="B190">2016</xref>). Interestingly, it has been shown that the representations encoded in DNN layers correspond to the representations that are learned by areas that make up the sensory hierarchies of biological agents (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B128">2015</xref>, <xref ref-type="bibr" rid="B127">2017a</xref>; G&#x000FC;&#x000E7;l&#x000FC; et al., <xref ref-type="bibr" rid="B126">2016</xref>). Multiple reviews discuss this use of DNNs in sensory neuroscience (Cox and Dean, <xref ref-type="bibr" rid="B60">2014</xref>; Kriegeskorte, <xref ref-type="bibr" rid="B176">2015</xref>; Robinson and Rolls, <xref ref-type="bibr" rid="B282">2015</xref>; Marblestone et al., <xref ref-type="bibr" rid="B204">2016</xref>; Yamins and DiCarlo, <xref ref-type="bibr" rid="B374">2016</xref>; Kietzmann et al., <xref ref-type="bibr" rid="B169">2017</xref>; Peelen and Downing, <xref ref-type="bibr" rid="B262">2017</xref>; van Gerven, <xref ref-type="bibr" rid="B352">2017</xref>; Vanrullen, <xref ref-type="bibr" rid="B354">2017</xref>).</p>
</sec>
<sec>
<title>4.3. Remembering</title>
<p>Being able to perceive the environment also implies that agents can store and retrieve past knowledge about objects and events in their surroundings. In the feedforward networks considered in the previous section, this knowledge is encoded in the synaptic weights as a result of learning. Memories of the past can also be stored, however, in moment-to-moment neural activity patterns. This does require the availability of lateral or feedback connections in order to enable recurrent processing (Singer, <xref ref-type="bibr" rid="B315">2013</xref>; Maass, <xref ref-type="bibr" rid="B200">2016</xref>). Recurrent processing can be implemented by a recurrent neural network (RNN) (Jordan, <xref ref-type="bibr" rid="B157">1987</xref>; Elman, <xref ref-type="bibr" rid="B84">1990</xref>), defined by
<disp-formula id="E9"><label>(9)</label><mml:math id="M34"><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>y</mml:mtext></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>f</mml:mtext></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>w</mml:mtext></mml:mstyle><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>y</mml:mtext></mml:mstyle><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>U</mml:mtext></mml:mstyle><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
such that the neuronal activity at time <italic>n</italic> depends on the activity at time <italic>n</italic>&#x02212;1 as well as instantaneous bottom-up input. RNNs can be interpreted as numerical approximations of differential equations that describe rate-based neural models (Dayan and Abbott, <xref ref-type="bibr" rid="B69">2005</xref>) and have been shown to be universal approximators of dynamical systems (Funahashi and Nakamura, <xref ref-type="bibr" rid="B104">1993</xref>)<xref ref-type="fn" rid="fn0009"><sup>9</sup></xref>. Their parameters can be estimated using a variant of backpropagation, referred to as backpropagation through time (Mozer, <xref ref-type="bibr" rid="B234">1989</xref>).</p>
<p>When considering perception, feedforward architectures may seem sufficient. For example, the onset latencies of neurons in monkey inferior-temporal cortex during visual processing are about 100 ms (Thorpe and Fabre-Thorpe, <xref ref-type="bibr" rid="B341">2001</xref>), which means that there is ample time for the transmission of just a few spikes. This suggests that object recognition is largely an automatic feedforward process (Vanrullen, <xref ref-type="bibr" rid="B353">2007</xref>). However, recurrent processing is important in perception as well since it provides the ability to maintain state. This is important in detecting salient features in space and time (Joukes et al., <xref ref-type="bibr" rid="B159">2014</xref>), as well as for integrating evidence in noisy or ambiguous settings (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B251">2013</xref>). Moreover, perception is strongly influenced by top-down processes, as mediated by feedback connections (Gilbert and Li, <xref ref-type="bibr" rid="B118">2013</xref>). RNNs have also been used to model working memory (Miconi, <xref ref-type="bibr" rid="B222">2017</xref>) as well as hippocampal function, which is involved in a variety of memory-related processes (Willshaw et al., <xref ref-type="bibr" rid="B368">2015</xref>; Kumaran et al., <xref ref-type="bibr" rid="B179">2016</xref>).</p>
<p>A special kind of RNN is the Hopfield network (Hopfield, <xref ref-type="bibr" rid="B145">1982</xref>), where <bold>w</bold> is symmetric and <bold>U</bold> &#x0003D; <bold>0</bold>. Learning in a Hopfield net is based on a Hebbian learning scheme. Hopfield nets are attractor networks that converge to a state that is a local minimum of an energy function. They have been used extensively as models of associative memory (Wills et al., <xref ref-type="bibr" rid="B367">2005</xref>). It has even been postulated that dreaming can be seen as an unlearning process which gets rid of spurious minima in attractor networks, thereby improving their storage capacity (Crick and Mitchison, <xref ref-type="bibr" rid="B61">1983</xref>).</p>
</sec>
<sec>
<title>4.4. Acting</title>
<p>As already described, the ability to generate appropriate actions is what ultimately drives behavior. In real-world settings, such actions typically need to be inferred from reward signals <italic>r</italic><sub><italic>t</italic></sub> provided by the environment. This is the subject matter of reinforcement learning (RL) (Sutton and Barto, <xref ref-type="bibr" rid="B333">1998</xref>). Define a policy &#x003C0;(<italic>s, a</italic>) as the probability of selecting an action <italic>a</italic> given a state <italic>s</italic>. Let the return <inline-formula><mml:math id="M25"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> be the total reward accumulated in an episode, with &#x003B3; a discount factor that downweighs future rewards. The goal in RL is to identify an optimal policy &#x003C0;<sup>&#x0002A;</sup> that maximizes the expected return
<disp-formula id="E10"><label>(10)</label><mml:math id="M35"><mml:mrow><mml:msup><mml:mi>&#x003C0;</mml:mi><mml:mo>&#x0002A;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:munder><mml:mrow><mml:mi>max</mml:mi></mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:munder><mml:mi mathvariant='double-struck'>E</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x02223;</mml:mo><mml:mi>&#x003C0;</mml:mi><mml:mo stretchy='false'>]</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Reinforcement learning algorithms have been crucial in training neural networks that have the capacity to act. Such networks learn to generate suitable actions purely by observing the rewards entailed by previously generated actions. RL algorithms come in model-free and model-based variants. In the model-free setting, optimal actions are learned purely based on the reward that is gained by performing actions in the past. In the model-based setting, in contrast, an explicit model of the environment is used to predict the consequences of actions that are being executed. Importantly, model-free and model-based reinforcement learning approaches have clear correspondences with habitual and goal-directed learning in neuroscience (Daw, <xref ref-type="bibr" rid="B66">2012</xref>; Buschman et al., <xref ref-type="bibr" rid="B48">2014</xref>).</p>
<p>Various model-free reinforcement learning approaches have been used to develop a variety of neural networks for action generation. For example, Q-learning was used to train networks that play Atari games (Mnih et al., <xref ref-type="bibr" rid="B229">2015</xref>) and policy gradient methods have been used to play board games (Silver et al., <xref ref-type="bibr" rid="B312">2017</xref>) and solve problems in (simulated) robotics (Silver et al., <xref ref-type="bibr" rid="B311">2014</xref>; Schulman et al., <xref ref-type="bibr" rid="B303">2015</xref>), effectively closing the perception-action cycle. Evolutionary strategies are also proving to become an useful approach for solving challenging control problems (Salimans et al., <xref ref-type="bibr" rid="B287">2017</xref>). Similar successes have been achieved using model-based reinforcement learning approaches (Schmidhuber, <xref ref-type="bibr" rid="B298">2015</xref>; Mujika, <xref ref-type="bibr" rid="B236">2016</xref>; Santana and Hotz, <xref ref-type="bibr" rid="B288">2016</xref>).</p>
<p>Another important ingredient required for generating optimal actions is recurrent processing, as described in the previous section. Action generation must depend on the ability to integrate evidence over time since, otherwise, we are guaranteed to act suboptimally. That is, states that are qualitatively different can appear the same to the decision maker, leading to suboptimal policies. Consider for example the sensation of a looming object. The optimal decision depends crucially on whether this object is approaching or receding, which can only be determined by taking past sensations into account. This phenomenon is known as perceptual aliasing (Whitehead and Ballard, <xref ref-type="bibr" rid="B365">1991</xref>).</p>
<p>A key ability of biological organisms which requires recurrent processing is their ability to navigate in their environment, as mediated by the hippocampal formation (Moser et al., <xref ref-type="bibr" rid="B232">2015</xref>). Recent work shows that particular characteristics of hippocampal place cells, such as stable tuning curves that remap between environments, are recovered by training neural networks on navigation tasks (Kanitscheider and Fiete, <xref ref-type="bibr" rid="B162">2016</xref>). The ability to integrate evidence also allows agents to selectively sample the environment, such as to maximize the amount of information gained. This process, known as active sensing, is crucial for understanding perceptual processing in biology (Yarbus, <xref ref-type="bibr" rid="B377">1967</xref>; Regan and No&#x000EB;, <xref ref-type="bibr" rid="B278">2001</xref>; Friston et al., <xref ref-type="bibr" rid="B100">2010</xref>; Schroeder et al., <xref ref-type="bibr" rid="B302">2010</xref>; Gordon and Ahissar, <xref ref-type="bibr" rid="B120">2012</xref>). Active sensing, in the form of saccade planning, has been implemented using a variety of recurrent neural network architectures (Larochelle and Hinton, <xref ref-type="bibr" rid="B183">2010</xref>; Gregor et al., <xref ref-type="bibr" rid="B122">2014</xref>; Mnih et al., <xref ref-type="bibr" rid="B228">2014</xref>). RNNs that implement recurrent processing have also been used to model various other action-related processes such as timing (Laje and Buonomano, <xref ref-type="bibr" rid="B181">2013</xref>), sequence generation (Rajan et al., <xref ref-type="bibr" rid="B274">2015</xref>) and motor control (Sussillo et al., <xref ref-type="bibr" rid="B331">2015</xref>).</p>
<p>Recurrent processing and reinforcement learning are also essential in modeling higher-level processes, such as cognitive control as mediated by frontal brain regions (Fuster, <xref ref-type="bibr" rid="B105">2001</xref>; Miller and Cohen, <xref ref-type="bibr" rid="B224">2001</xref>). Examples are models of context-dependent processing (Mante et al., <xref ref-type="bibr" rid="B203">2013</xref>) and perceptual decision-making (Carnevale et al., <xref ref-type="bibr" rid="B50">2015</xref>). In general, RNNs that have been trained using RL on a variety of cognitive tasks have been shown to yield properties that are consistent with phenomena observed in biological neural networks (Song et al., <xref ref-type="bibr" rid="B319">2016</xref>; Miconi, <xref ref-type="bibr" rid="B222">2017</xref>).</p>
</sec>
<sec>
<title>4.5. Predicting</title>
<p>Modern theories of human brain function appeal to the idea that the brain can be viewed as a prediction machine, which is in the business of continuously generating top-down predictions that are integrated with bottom-up sensory input (Lee and Mumford, <xref ref-type="bibr" rid="B189">2003</xref>; Yuille and Kersten, <xref ref-type="bibr" rid="B378">2006</xref>; Clark, <xref ref-type="bibr" rid="B56">2013</xref>; Summerfield and de Lange, <xref ref-type="bibr" rid="B328">2014</xref>). This view of the brain as a prediction machine that performs unconscious inference has a long history, going back to the seminal work of Alhazen and Helmholtz (Hatfield, <xref ref-type="bibr" rid="B135">2002</xref>). Modern views cast this process in terms of Bayesian inference, where the brain is updating its internal model of the environment in order to explain away the data that impinge upon its senses, also referred to as the Bayesian brain hypothesis (Jaynes, <xref ref-type="bibr" rid="B153">1988</xref>; Doya et al., <xref ref-type="bibr" rid="B78">2006</xref>). The same reasoning underlies the free-energy principle, which assumes that biological systems minimize a free energy functional of their internal states that entail beliefs about hidden states in their environment (Friston, <xref ref-type="bibr" rid="B99">2010</xref>). Predictions can be seen as central to the generation of adaptive behavior, since anticipating the future will allow an agent to select appropriate actions in the present (Schacter et al., <xref ref-type="bibr" rid="B294">2007</xref>; Moulton and Kosslyn, <xref ref-type="bibr" rid="B233">2009</xref>).</p>
<p>Prediction is central in model-based RL approaches since it requires agents to plan their actions by predicting the outcomes of future actions (Daw, <xref ref-type="bibr" rid="B66">2012</xref>). This is strongly related to the notion of preplay of future events subserving path planning (Corneil and Gerstner, <xref ref-type="bibr" rid="B59">2015</xref>). Such preplay has been observed in hippocampal place cell sequences (Dragoi and Tonegawa, <xref ref-type="bibr" rid="B79">2011</xref>), giving further support to the idea that the hippocampal formation is involved in goal-directed navigation (Corneil and Gerstner, <xref ref-type="bibr" rid="B59">2015</xref>). Prediction also allows an agent to prospectively act on expected deviations from optimal conditions. This focus on error-correction and stability is also prevalent in the work of the cybernetic movement (Ashby, <xref ref-type="bibr" rid="B15">1952</xref>). Note further that predictive processing connects to the concept of allostasis, where the agent is actively trying to predict future states such as to minimize deviations from optimal homeostatic conditions. It is also central to optimal feedback control theory, which assumes that the motor system corrects only those deviations that interfere with task goals (Todorov and Jordan, <xref ref-type="bibr" rid="B347">2002</xref>).</p>
<p>The notion of predictive processing has been very influential in neural network research. For example, it provides the basis for predictive coding models that introduce specific neural network architectures in which feedforward connections are used to transmit the prediction errors that result from discrepancies between top-down predictions and bottom-up sensations (Rao and Ballard, <xref ref-type="bibr" rid="B276">1999</xref>; Huang and Rao, <xref ref-type="bibr" rid="B147">2011</xref>). It also led to the development of a wide variety of generative models that are able to predict their sensory states, also referred to as fantasies (Hinton, <xref ref-type="bibr" rid="B140">2013</xref>). Such fantasies may play a role in understanding cognitive processing involved in imagery, working memory and dreaming. In effect, these models aim to estimate a distribution over latent causes <bold>z</bold> in the environment that explain observed sensory data <bold>x</bold>. In this setting, the most probable explanation is given by
<disp-formula id="E11"><label>(11)</label><mml:math id="M36"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>&#x0002A;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:munder><mml:mrow><mml:mi>max</mml:mi></mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle></mml:munder><mml:mtext>&#x000A0;</mml:mtext><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo>&#x02223;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:munder><mml:mrow><mml:mi>max</mml:mi></mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle></mml:munder><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>x</mml:mtext></mml:mstyle><mml:mo>&#x02223;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo stretchy='false'>)</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mtext>z</mml:mtext></mml:mstyle><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Generative models also offer a way to perform unsupervised learning, since if a neural network is able to generate predictions then the discrepancy between predicted and observed stimuli can serve as a teaching signal. A canonical example is the Boltzmann machine, which is a stochastic variant of a Hopfield network that is able to discover regularities in the training data using a simple unsupervised learning algorithm (Hinton and Sejnowski, <xref ref-type="bibr" rid="B143">1983</xref>; Ackley et al., <xref ref-type="bibr" rid="B3">1985</xref>). Another classical example is the Helmholtz machine, which incorporates both bottom-up and top-down processing (Dayan et al., <xref ref-type="bibr" rid="B70">1995</xref>). Other, more recent examples of ANN-based generative models are deep belief networks (Hinton et al., <xref ref-type="bibr" rid="B142">2006</xref>), variational autoencoders (Kingma and Welling, <xref ref-type="bibr" rid="B171">2014</xref>) and generative adversarial networks (Goodfellow et al., <xref ref-type="bibr" rid="B119">2014</xref>). Recent work has started to use these models to predict future sensory states from current observations (Lotter et al., <xref ref-type="bibr" rid="B197">2016</xref>; Mathieu et al., <xref ref-type="bibr" rid="B212">2016</xref>; Xue et al., <xref ref-type="bibr" rid="B373">2016</xref>).</p>
</sec>
<sec>
<title>4.6. Reasoning</title>
<p>While ANNs are now able to solve complex tasks such as acting in natural environments or playing difficult board games, one could still argue that they are &#x0201C;just&#x0201D; performing sophisticated pattern recognition rather than showing the symbolic reasoning abilities that characterize our own brains. The question of whether connectionist systems are capable of symbolic reasoning has a long history, and has been debated by various researchers in the cognitivist (symbolic) program (Pinker and Mehler, <xref ref-type="bibr" rid="B265">1988</xref>). We will not settle this debate here but point out that efforts are underway to endow neural networks with sophisticated reasoning capabilities.</p>
<p>One example is the development of &#x0201C;differentiable computers&#x0201D; that learn to implement algorithms based on a finite amount of training data (Graves et al., <xref ref-type="bibr" rid="B121">2014</xref>; Weston et al., <xref ref-type="bibr" rid="B362">2015</xref>; Vinyals et al., <xref ref-type="bibr" rid="B358">2017</xref>). The resulting neural networks perform variable binding and are able to deal with variable length structures (Graves et al., <xref ref-type="bibr" rid="B121">2014</xref>), which are two objections that were originally raised against using ANNs to explain cognitive processing (Fodor and Pylyshyn, <xref ref-type="bibr" rid="B95">1988</xref>).</p>
<p>Another example is the development of neural networks that can answer arbitrary questions about text (Bordes et al., <xref ref-type="bibr" rid="B37">2015</xref>), images (Agrawal et al., <xref ref-type="bibr" rid="B7">2016</xref>) and movies (Tapaswi et al., <xref ref-type="bibr" rid="B338">2015</xref>), thereby requiring deep semantic knowledge about the experienced stimuli. Recent models have also been shown to be capable of compositional reasoning (Johnson et al., <xref ref-type="bibr" rid="B155">2017</xref>; Lake et al., <xref ref-type="bibr" rid="B182">2017</xref>; Yang et al., <xref ref-type="bibr" rid="B375">2017</xref>), which is an important ingredient for explaining the systematic nature of human thought (Fodor and Pylyshyn, <xref ref-type="bibr" rid="B95">1988</xref>). These architectures often make use of distributional semantics, where words are encoded as real vectors that capture word meaning (Mikolov et al., <xref ref-type="bibr" rid="B223">2013</xref>; Ferrone and Zanzotto, <xref ref-type="bibr" rid="B91">2017</xref>).</p>
<p>Several other properties characterize human thought processes, such as intuitive physics, intuitive psychology, relational reasoning and causal reasoning (Kemp and Tenenbaum, <xref ref-type="bibr" rid="B166">2008</xref>; Lake et al., <xref ref-type="bibr" rid="B182">2017</xref>). Another crucial hallmark of intelligent systems is that they are able to explain what they are doing (Brachman, <xref ref-type="bibr" rid="B39">2002</xref>). This requires agents to have a deep understanding of their world. These properties should be replicated in neural networks if they are to serve as accurate models of natural intelligence. New neural network architectures are slowly starting to take steps in this direction (e.g., Louizos et al., <xref ref-type="bibr" rid="B198">2017</xref>; Santoro et al., <xref ref-type="bibr" rid="B290">2017</xref>; Zhu et al., <xref ref-type="bibr" rid="B383">2017</xref>).</p>
</sec>
</sec>
<sec id="s5">
<title>5. Toward strong AI</title>
<p>We have reviewed the computational foundations of natural intelligence and outlined how ANNs can be used to model a variety of cognitive processes. However, our current understanding of natural intelligence remains limited and strong AI has not yet been attained. In the following, we will touch upon a number of important topics that will be of importance for eventually reaching these goals.</p>
<sec>
<title>5.1. Surviving in complex environments</title>
<p>Contemporary neural network architectures tend to excel at solving one particular problem well. However, in practice, we want to arrive at intelligent machines that are able to survive in complex environments. This requires the agent to deal with high-dimensional naturalistic input, be able to solve multiple tasks depending on context, and devise optimal strategies to ensure long-term survival.</p>
<p>The research community has embraced these desiderata by creating virtual worlds that allow development and testing of neural network architectures (e.g., Todorov et al., <xref ref-type="bibr" rid="B346">2012</xref>; Beattie et al., <xref ref-type="bibr" rid="B24">2016</xref>; Brockman et al., <xref ref-type="bibr" rid="B44">2016</xref>; Kempka et al., <xref ref-type="bibr" rid="B167">2016</xref>; Synnaeve et al., <xref ref-type="bibr" rid="B336">2016</xref>)<xref ref-type="fn" rid="fn0010"><sup>10</sup></xref>. While most work in this area has focused on environments with fully observable states, reward functions with low delay, and small action sets, research is shifting toward environments that are partially observable, require long-term planning, show complex dynamics and have noisy and high-dimensional control interfaces (Synnaeve et al., <xref ref-type="bibr" rid="B336">2016</xref>).</p>
<p>A particular challenge in these naturalistic environments is that networks need to be able to exhibit continual (life-long) learning (Thrun and Mitchell, <xref ref-type="bibr" rid="B342">1995</xref>), adapting continuously to the current state of affairs. This is difficult due to the phenomenon of catastrophic forgetting (McCloskey and Cohen, <xref ref-type="bibr" rid="B217">1989</xref>; French, <xref ref-type="bibr" rid="B97">1999</xref>), where previously acquired skills are overwritten by ongoing modification of synaptic weights. Recent algorithmic developments attenuate the detrimental effects of catastrophic forgetting (Kirkpatrick et al., <xref ref-type="bibr" rid="B172">2015</xref>; Zenke et al., <xref ref-type="bibr" rid="B382">2015</xref>), offering a (partial) solution to the stability vs. plasticity dilemma (Abraham and Robins, <xref ref-type="bibr" rid="B2">2005</xref>). Life-long learning is further complicated by the exploration-exploitation dilemma, where agents need to decide on whether to accrue either information or reward (Cohen et al., <xref ref-type="bibr" rid="B57">2007</xref>). Another challenge is the fact that reinforcement learning of complex actions is notoriously slow. Here, progress is being made using networks that make use of differentiable memories (Santoro et al., <xref ref-type="bibr" rid="B289">2016</xref>; Pritzel et al., <xref ref-type="bibr" rid="B269">2017</xref>). Survival in complex environments also requires that agents learn to perform multiple tasks well. This learning process can be facilitated through multitask learning (Caruana, <xref ref-type="bibr" rid="B52">1997</xref>) (also referred to as learning to learn Baxter, <xref ref-type="bibr" rid="B23">1998</xref> or transfer learning Pan and Fellow, <xref ref-type="bibr" rid="B258">2009</xref>), where learning of one task is facilitated by knowledge gained through learning to solve another task. Multitask learning has been shown to improve convergence speed and generalization to unseen data (Scholte et al., <xref ref-type="bibr" rid="B301">2017</xref>). Finally, effective learning also calls for agents that can generalize to cases that were not encountered before, which is known as zero-shot learning (Palatucci et al., <xref ref-type="bibr" rid="B257">2009</xref>), and can learn from rare events, which is known as one-shot learning (Fei-Fei et al., <xref ref-type="bibr" rid="B88">2006</xref>; Vinyals et al., <xref ref-type="bibr" rid="B357">2016</xref>; Kaiser and Roy, <xref ref-type="bibr" rid="B161">2017</xref>).</p>
<p>While the use of virtual worlds allows for testing the capabilities of artificial agents, it does not guarantee that the same agents are able to survive in the real world (Brooks, <xref ref-type="bibr" rid="B45">1992</xref>). That is, there may exist a reality gap, where skills acquired in virtual worlds do not carry over to the real world. In contrast to virtual worlds, acting in the real world requires the agent to deal with unforeseen circumstances resulting from the complex nature of reality, the agent&#x00027;s need for a physical body, as well as its engagement with a myriad of other agents (Anderson, <xref ref-type="bibr" rid="B12">2003</xref>). Moreover, the continuing interplay between an organism and its environment may itself shape and, ultimately, determine cognition (Gibson, <xref ref-type="bibr" rid="B116">1979</xref>; Maturana and Varela, <xref ref-type="bibr" rid="B214">1987</xref>; Brooks, <xref ref-type="bibr" rid="B46">1996</xref>; Edelman, <xref ref-type="bibr" rid="B83">2015</xref>). Effectively dealing with these complexities may not only require plasticity in individual agents but also the incorporation of developmental change, as well as learning at evolutionary time scales (Marcus, <xref ref-type="bibr" rid="B205">2009</xref>). From a developmental perspective, networks can be more effectively trained by presenting them with a sequence of increasingly complex tasks, instead of immediately requiring the network to solve the most complex task (Elman, <xref ref-type="bibr" rid="B86">1993</xref>). This process is known as curriculum learning (Bengio et al., <xref ref-type="bibr" rid="B30">2009</xref>) and is analogous to how a child learns by decomposing problems into simpler subproblems (Turing, <xref ref-type="bibr" rid="B350">1950</xref>). Evolutionary strategies have also been shown to be effective in learning to solve challenging control problems (Salimans et al., <xref ref-type="bibr" rid="B287">2017</xref>). Finally, to learn about the world, we may also turn toward cultural learning, where agents can offload task complexity by learning from each other (Bengio, <xref ref-type="bibr" rid="B28">2014</xref>).</p>
<p>As mentioned in section 2.2, adaptive behavior is the result of multiple competing drives and motivations that provide primary, intrinsic and extrinsic rewards. Hence, one strategy for endowing machines with the capacity to survive in the real world is to equip neural networks with drives and motivations that ensure their long-term survival<xref ref-type="fn" rid="fn0011"><sup>11</sup></xref>. In terms of primary rewards, one could conceivably provide artificial agents with the incentive to minimize computational resources or maximize offspring via evolutionary processes (Stanley and Miikkulainen, <xref ref-type="bibr" rid="B322">2002</xref>; Floreano et al., <xref ref-type="bibr" rid="B94">2008</xref>; Gauci and Stanley, <xref ref-type="bibr" rid="B111">2010</xref>). In terms of intrinsic rewards, one can think of various ways to equip agents with the drive to explore the environment (Oudeyer, <xref ref-type="bibr" rid="B253">2007</xref>). We briefly describe a number of principles that have been proposed in the literature. Artificial curiosity assumes that internal reward depends on how boring an environment is, with agents avoiding fully predictable and unpredictably random states (Schmidhuber, <xref ref-type="bibr" rid="B296">1991</xref>, <xref ref-type="bibr" rid="B297">2003</xref>; Pathak et al., <xref ref-type="bibr" rid="B261">2017</xref>). A related notion is that of information-seeking agents (Bachman et al., <xref ref-type="bibr" rid="B17">2016</xref>). The autotelic principle formalizes the concept of flow where an agent tries to maintain a state where learning is challenging, but not overwhelming (Csikszentmihalyi, <xref ref-type="bibr" rid="B62">1975</xref>; Steels, <xref ref-type="bibr" rid="B324">2004</xref>). The free-energy principle states that an agent seeks to minimize uncertainty by updating its internal model of the environment and selecting uncertainty-reducing actions (Friston, <xref ref-type="bibr" rid="B98">2009</xref>, <xref ref-type="bibr" rid="B99">2010</xref>). Empowerment is founded on information-theoretic principles and quantifies how much control an agent has over its environment, as well as its ability to sense this control (Klyubin et al., <xref ref-type="bibr" rid="B173">2005a</xref>,<xref ref-type="bibr" rid="B174">b</xref>; Salge et al., <xref ref-type="bibr" rid="B286">2013</xref>). In this setting, intrinsically motivated behavior is induced by the maximization of empowerment. Finally, various theories embrace the notion that optimal prediction of future states drives learning and behavior (Der et al., <xref ref-type="bibr" rid="B74">1999</xref>; Kaplan and Oudeyer, <xref ref-type="bibr" rid="B163">2004</xref>; Ay et al., <xref ref-type="bibr" rid="B16">2008</xref>). In terms of extrinsic rewards, one can think of imitation learning, where a teacher signal is used to inform the agent about its desired outputs (Schaal, <xref ref-type="bibr" rid="B293">1999</xref>; Duan et al., <xref ref-type="bibr" rid="B81">2017</xref>).</p>
</sec>
<sec>
<title>5.2. Bridging the gap between artificial and biological neural networks</title>
<p>To reduce the gap between artificial and biological neural networks, it makes sense to assess their operation on similar tasks. This can be done either by comparing the models at a neurobiological level or at a behavioral level. The former refers to comparing the internal structure or activation patterns of artificial and biological neural networks. The latter refers to comparing their behavioral outputs (e.g., eye movements, reaction times, high-level decisions). Moreover, comparisons can be made under changing conditions, i.e., during learning and development (Elman et al., <xref ref-type="bibr" rid="B87">1996</xref>). As such, ANNs can serve as explanatory mechanisms in cognitive neuroscience and behavioral psychology, embracing recent model-based approaches (Forstmann and Wagenmakers, <xref ref-type="bibr" rid="B96">2015</xref>).</p>
<p>From a psychological perspective, ANNs have been compared explicitly with their biological counterparts. Connectionist models were widely used in the 1980&#x00027;s to explain various psychological phenomena, particularly by the parallel distributed processing (PDP) movement, which stressed the parallel nature of neural processing and the distributed nature of neural representations (McClelland, <xref ref-type="bibr" rid="B215">2003</xref>). For example, neural networks have been used to explain grammar acquisition (Elman, <xref ref-type="bibr" rid="B85">1991</xref>), category learning (Kruschke, <xref ref-type="bibr" rid="B178">1992</xref>) and the organization of the semantic system (Ritter and Kohonen, <xref ref-type="bibr" rid="B281">1989</xref>). More recently, deep neural networks have been used to explain human similarity judgments (Peterson et al., <xref ref-type="bibr" rid="B264">2016</xref>). With new developments in cognitive and affective computing, where neural networks become more adept at solving high-level cognitive tasks, such as predicting people&#x00027;s (apparent) personality traits (G&#x000FC;&#x000E7;l&#x000FC;t&#x000FC;rk et al., <xref ref-type="bibr" rid="B131">2016</xref>), their use as a tool to explain psychological phenomena is likely to increase. This will also require embracing insights about how humans solve problems at a cognitive level (Tenenbaum et al., <xref ref-type="bibr" rid="B339">2011</xref>).</p>
<p>ANNs have also been related explicitly to brain function. For example, the perceptron has been used in the modeling of various neuronal systems, including sensorimotor learning in the cerebellum (Marr, <xref ref-type="bibr" rid="B209">1969</xref>) and associative memory in cortex (Gardner, <xref ref-type="bibr" rid="B108">1988</xref>), sparse coding has been used to explain receptive field properties (Olshausen and Field, <xref ref-type="bibr" rid="B248">1996</xref>), topographic maps have been used to explain the formation of cortical maps (Obermayer, <xref ref-type="bibr" rid="B246">1990</xref>; Aflalo, <xref ref-type="bibr" rid="B6">2006</xref>), Hebbian learning has been used to explain neural tuning to face orientation (Leibo et al., <xref ref-type="bibr" rid="B191">2017</xref>), and networks trained by backpropagation have been used to model the response properties of posterior parietal neurons (Zipser and Andersen, <xref ref-type="bibr" rid="B384">1988</xref>). Neural networks have also been used to model central pattern generators that drive behavior (Duysens and Van de Crommert, <xref ref-type="bibr" rid="B82">1998</xref>; Ijspeert, <xref ref-type="bibr" rid="B150">2008</xref>) as well as the perception of rhythmic stimuli (Torras i Gen&#x000ED;s, <xref ref-type="bibr" rid="B349">1986</xref>; Gasser, Eck and Port, <xref ref-type="bibr" rid="B110">1999</xref>). Furthermore, reinforcement learning algorithms used to train neural networks for action selection have strong ties with the brain&#x00027;s reward system (Schultz et al., <xref ref-type="bibr" rid="B304">1997</xref>; Sutton and Barto, <xref ref-type="bibr" rid="B333">1998</xref>). It has been shown that RNNs trained to solve a variety of cognitive tasks using reinforcement learning replicate various phenomena observed in biological systems (Song et al., <xref ref-type="bibr" rid="B319">2016</xref>; Miconi, <xref ref-type="bibr" rid="B222">2017</xref>). Crucially, these efforts go beyond descriptive approaches in that they may explain <italic>why</italic> the human brain is organized in a certain manner (Barak, <xref ref-type="bibr" rid="B19">2017</xref>).</p>
<p>Rather than using neural networks to explain certain observed neural or behavioral phenomena, one can also directly fit neural networks to neurobehavioral data. This can be achieved via an indirect approach or via a direct approach. In the <italic>indirect</italic> approach, neural networks are first trained to solve a task of interest. Subsequently, the trained network&#x00027;s responses are fitted to neurobehavioral data obtained as participants engage in the same task. Using this approach, deep convolutional neural networks trained on object recognition, action recognition and music tagging have been used to explain the functional organization of visual as well as auditory cortex (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B128">2015</xref>, <xref ref-type="bibr" rid="B127">2017a</xref>; G&#x000FC;&#x000E7;l&#x000FC; et al., <xref ref-type="bibr" rid="B126">2016</xref>). The indirect approach has also been used to train RNNs via reinforcement learning on a probabilistic categorization task. These networks have been used to fit the learning trajectories and behavioral responses of humans engaged in the same task (Bosch et al., <xref ref-type="bibr" rid="B38">2016</xref>). Mante et al. (<xref ref-type="bibr" rid="B203">2013</xref>) used RNNs to model the population dynamics of single neurons in prefrontal cortex during a context-dependent choice task. In the <italic>direct</italic> approach, neural networks are trained to directly predict neural responses. For example, Mcintosh et al. (<xref ref-type="bibr" rid="B219">2016</xref>) trained convolutional neural networks to predict retinal responses to natural scenes, Joukes et al. (<xref ref-type="bibr" rid="B159">2014</xref>) trained RNNs to predict neural responses to motion stimuli, and G&#x000FC;&#x000E7;l&#x000FC; and van Gerven (<xref ref-type="bibr" rid="B129">2017b</xref>) used RNNs to predict cortical responses to naturalistic video clips. This ability of neural networks to explain neural recordings is expected to become increasingly important (Sompolinsky, <xref ref-type="bibr" rid="B318">2014</xref>; Marder, <xref ref-type="bibr" rid="B206">2015</xref>), given the emergence of new imaging technology where the activity of thousands of neurons can be measured in parallel (Ahrens et al., <xref ref-type="bibr" rid="B8">2013</xref>; Churchland and Sejnowski, <xref ref-type="bibr" rid="B55">2016</xref>; Lopez et al., <xref ref-type="bibr" rid="B196">2016</xref>; Pachitariu et al., <xref ref-type="bibr" rid="B254">2016</xref>; Yang and Yuste, <xref ref-type="bibr" rid="B376">2017</xref>). Better understanding will also be facilitated by the development of new data analysis techniques to elucidate human brain function (Kass et al., <xref ref-type="bibr" rid="B164">2014</xref>)<xref ref-type="fn" rid="fn0012"><sup>12</sup></xref>, the use of ANNs to decode neural representations (Schoenmakers et al., <xref ref-type="bibr" rid="B300">2013</xref>; G&#x000FC;&#x000E7;l&#x000FC;t&#x000FC;rk et al., <xref ref-type="bibr" rid="B130">2017</xref>), as well as the development of approaches that elucidate the functioning of ANNs (e.g., Nguyen et al., <xref ref-type="bibr" rid="B244">2016</xref>; Kindermans et al., <xref ref-type="bibr" rid="B170">2017</xref>; Miller, <xref ref-type="bibr" rid="B225">2017</xref>)<xref ref-type="fn" rid="fn0013"><sup>13</sup></xref>.</p>
</sec>
<sec>
<title>5.3. Next-generation artificial neural networks</title>
<p>The previous sections outlined how neural networks can be made to solve challenging tasks and provide explanations of neural and behavioral responses in biological agents. In this final section, we consider some developments that are expected to fuel the next generation of ANNs.</p>
<p>First, a major driving force in neural network research will be theoretical and algorithmic developments that inform why ANNs work so well in practice, what their fundamental limitations are, as well as how to overcome these. From a theoretical point of view, substantial advances have already been made pertaining to, for example, understanding the nature of representations (Anselmi and Poggio, <xref ref-type="bibr" rid="B14">2014</xref>; Lin and Tegmark, <xref ref-type="bibr" rid="B195">2016</xref>; Shwartz-Ziv and Tishby, <xref ref-type="bibr" rid="B310">2017</xref>), the statistical mechanics of neural networks (Sompolinsky, <xref ref-type="bibr" rid="B317">1988</xref>; Advani et al., <xref ref-type="bibr" rid="B5">2013</xref>), as well as the expressiveness (Pascanu et al., <xref ref-type="bibr" rid="B260">2013</xref>; Bianchini and Scarselli, <xref ref-type="bibr" rid="B31">2014</xref>; Kadmon and Sompolinsky, <xref ref-type="bibr" rid="B160">2016</xref>; Mhaskar et al., <xref ref-type="bibr" rid="B221">2016</xref>; Poole et al., <xref ref-type="bibr" rid="B267">2016</xref>; Raghu et al., <xref ref-type="bibr" rid="B272">2016</xref>; Weichwald et al., <xref ref-type="bibr" rid="B361">2016</xref>), generalizability (Kawaguchi et al., <xref ref-type="bibr" rid="B165">2017</xref>) and learnability (Dauphin et al., <xref ref-type="bibr" rid="B64">2014</xref>; Saxe et al., <xref ref-type="bibr" rid="B291">2014</xref>; Schoenholz et al., <xref ref-type="bibr" rid="B299">2017</xref>) of DNNs.</p>
<p>From an algorithmic point of view, great strides have been made in improving training of deep (Srivastava et al., <xref ref-type="bibr" rid="B321">2014</xref>; He et al., <xref ref-type="bibr" rid="B136">2015</xref>; Ioffe and Szegedy, <xref ref-type="bibr" rid="B151">2015</xref>) and recurrent neural networks (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B144">1997</xref>; Pascanu et al., <xref ref-type="bibr" rid="B259">2012</xref>), overcoming the reality gap (Tobin et al., <xref ref-type="bibr" rid="B345">2017</xref>), adding modularity to neural networks (Fernando et al., <xref ref-type="bibr" rid="B90">2017</xref>), as well as on improving the efficacy of reinforcement learning algorithms (Schulman et al., <xref ref-type="bibr" rid="B303">2015</xref>; Mnih et al., <xref ref-type="bibr" rid="B227">2016</xref>; Pritzel et al., <xref ref-type="bibr" rid="B269">2017</xref>).</p>
<p>Second, it is expected that as neural network models become more plausible from a biological point of view, model fit and task performance will further improve (Cox and Dean, <xref ref-type="bibr" rid="B60">2014</xref>). This is important in driving new developments in model-based cognitive neuroscience but also in developing intelligent machines that show human-like behavior. One example is to match the object recognition capabilities of extremely deep neural networks with more biologically plausible RNNs of limited depth (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B251">2013</xref>; Liao and Poggio, <xref ref-type="bibr" rid="B193">2016</xref>) and achieving category selectivity in a more realistic manner (Peelen and Downing, <xref ref-type="bibr" rid="B262">2017</xref>; Scholte et al., <xref ref-type="bibr" rid="B301">2017</xref>). Another example is to incorporate predictive coding principles in neural network architectures (Lotter et al., <xref ref-type="bibr" rid="B197">2016</xref>). Furthermore, more human-like perceptual systems can be arrived at by including attentional mechanisms (Mnih et al., <xref ref-type="bibr" rid="B228">2014</xref>) as well as mechanisms for saccade planning (Najemnik and Geisler, <xref ref-type="bibr" rid="B237">2005</xref>; Larochelle and Hinton, <xref ref-type="bibr" rid="B183">2010</xref>; Gregor et al., <xref ref-type="bibr" rid="B122">2014</xref>).</p>
<p>In general, ANN research can benefit from a close interaction between the AI and neuroscience communities (Yuste, <xref ref-type="bibr" rid="B379">2015</xref>; Hassabis et al., <xref ref-type="bibr" rid="B134">2017</xref>). For example, neural network research may be shaped by general guiding principles of brain function at different levels of analysis (O&#x00027;Reilly, <xref ref-type="bibr" rid="B249">1998</xref>; Maass, <xref ref-type="bibr" rid="B200">2016</xref>; Sterling and Laughlin, <xref ref-type="bibr" rid="B326">2016</xref>). We may also strive to incorporate more biological detail. For example, to obtain accurate models of neural information processing we may need to embrace spike-based rather than rate-based neural networks (Brette, <xref ref-type="bibr" rid="B43">2015</xref>)<xref ref-type="fn" rid="fn0014"><sup>14</sup></xref>. Efforts are underway to effectively train spiking neural networks (Maass, <xref ref-type="bibr" rid="B199">1997</xref>; Gerstner and Kistler, <xref ref-type="bibr" rid="B114">2002</xref>; Gerstner et al., <xref ref-type="bibr" rid="B115">2014</xref>; O&#x00027;Connor and Welling, <xref ref-type="bibr" rid="B247">2016</xref>; Huh and Sejnowski, <xref ref-type="bibr" rid="B148">2017</xref>) and endow them with the same cognitive capabilities as their rate-based cousins (Thalmeier et al., <xref ref-type="bibr" rid="B340">2015</xref>; Abbott et al., <xref ref-type="bibr" rid="B1">2016</xref>; Kheradpisheh et al., <xref ref-type="bibr" rid="B168">2016</xref>; Lee et al., <xref ref-type="bibr" rid="B188">2016</xref>; Zambrano and Bohte, <xref ref-type="bibr" rid="B381">2016</xref>).</p>
<p>In the same vein, researchers are exploring how probabilistic computations can be performed in neural networks (Nessler et al., <xref ref-type="bibr" rid="B241">2013</xref>; Pouget et al., <xref ref-type="bibr" rid="B268">2013</xref>; Gal, <xref ref-type="bibr" rid="B107">2016</xref>; Orhan and Ma, <xref ref-type="bibr" rid="B252">2016</xref>; Ambrogioni et al., <xref ref-type="bibr" rid="B9">2017</xref>; Heeger, <xref ref-type="bibr" rid="B137">2017</xref>; Mandt et al., <xref ref-type="bibr" rid="B202">2017</xref>) and deriving new biologically plausible synaptic plasticity rules (Brea and Gerstner, <xref ref-type="bibr" rid="B42">2016</xref>; Brea et al., <xref ref-type="bibr" rid="B41">2016</xref>; Schiess et al., <xref ref-type="bibr" rid="B295">2016</xref>). Biologically-inspired principles may also be incorporated at a more conceptual level. For instance, researchers have shown that neural networks can be protected from adversarial attacks (i.e., the construction of stimuli that cause networks to make mistakes) by integrating the notion of nonlinear computations encountered in the branched dendritic structures of real neurons (Nayebi and Ganguli, <xref ref-type="bibr" rid="B238">2016</xref>).</p>
<p>Finally, research is invested in implementing ANNs in hardware, also referred to as neuromorphic computing (Mead, <xref ref-type="bibr" rid="B220">1990</xref>). These brain-based parallel chip architectures hold the promise of devices that operate in real time and with very low power consumption (Schuman et al., <xref ref-type="bibr" rid="B305">2017</xref>), driving new advances in cognitive computing (Modha et al., <xref ref-type="bibr" rid="B230">2011</xref>; Neftci et al., <xref ref-type="bibr" rid="B239">2013</xref>; Van de Burgt et al., <xref ref-type="bibr" rid="B351">2017</xref>). On a related note, nanotechnology may 1 day drive the development of new neural network architectures whose operation is closer to the molecular machines that mediate the operation of biological neural networks (Drexler, <xref ref-type="bibr" rid="B80">1992</xref>; Strukov, <xref ref-type="bibr" rid="B327">2011</xref>). In the words of Feynman (<xref ref-type="bibr" rid="B93">1992</xref>): &#x0201C;There&#x00027;s plenty of room at the bottom.&#x0201D;</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusion</title>
<p>As cognitive scientists, we live in exciting times. Cognitivism offers an interpretation of agents as information processing systems that are engaged in formal symbol manipulation. The probabilistic approach to cognition extends this interpretation by viewing organisms as rational agents that need to act in the face of uncertainty under limited resources. Finally, emergentist approaches such as artificial life and connectionism indicate that concerted interactions between simple processing elements can achieve human-level performance at certain cognitive tasks. While these different views have stirred substantial debate in the past, they need not be irreconcilable. Surely we are capable of formal symbol manipulation and decision making under uncertainty in real-life settings. At the same time, these capabilities must be implemented by the neural circuits that make up our own brains, which themselves rely on noisy long-range communication between neuronal populations.</p>
<p>The thesis of this paper is that natural intelligence can be modeled and understood by constructing artificial agents whose synthetic brains are composed of (rate-based) neural networks. To act as explanations of natural intelligence, these synthetic brains should show a functional correspondence with their biological counterparts. To identify such correspondence we can embrace the rich sources of data provided by biology, neuroscience and psychology, providing a link to Marr&#x00027;s implementational level. At the same time, we can use sophisticated machinery developed in mathematics, computer science and physics to gain a better understanding of these systems. Ultimately, these synthetic brains should be able to show the capabilities that are prescribed by normative theories of intelligent behavior, providing a link to Marr&#x00027;s computational level.</p>
<p>The supposition that artificial neural networks are sufficient for modeling all of cognition may seem premature. For example, state-of-the-art question-answering systems such as IBM&#x00027;s Watson (Ferrucci et al., <xref ref-type="bibr" rid="B92">2010</xref>) use ANN technology as a minor component within a larger (symbolic) framework and the AlphaGo system (Silver et al., <xref ref-type="bibr" rid="B312">2017</xref>), which learns to play the game of Go beyond grandmaster level without any human intervention, combines neural networks with Monte Carlo tree search. While it is true that ANNs remain wanting when it comes to logical reasoning, inferring causal relationships or planning, the pace of current research may very well bring these capabilities within reach in the foreseeable future. Such neural networks may turn out to be quite different from current neural network architectures and their operation may be guided by complementary yet-to-be-discovered learning rules.</p>
<p>The quest for natural intelligence can be contrasted with a pure engineering approach. From an engineering perspective, understanding natural intelligence may be considered irrelevant since the main interest is in building devices that do the job. To quote Edsger Dijkstra, &#x0201C;the question whether machines can think [is] as relevant as the question whether submarines can swim.&#x0201D; At the same time, our quest for natural intelligence may facilitate the development of strong AI given the proven ability of our own brains to generate intelligent behavior. Hence, biologically inspired architectures may not only provide new insights into human brain function but could also in the long run yield superior curious and perhaps even conscious machines that surpass humans in terms of intelligence, creativity, playfulness, and empathy (Boden, <xref ref-type="bibr" rid="B35">1998</xref>; Moravec, <xref ref-type="bibr" rid="B231">2000</xref>; Der and Martius, <xref ref-type="bibr" rid="B73">2011</xref>; Modha et al., <xref ref-type="bibr" rid="B230">2011</xref>; Harari, <xref ref-type="bibr" rid="B132">2017</xref>).</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>The author confirms being the sole contributor of this work and approved it for publication.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>This work is supported by a VIDI grant (639.072.513) from the Netherlands Organization for Scientific Research. I would like to thank Nadine Dijkstra, Gabri&#x000EB;lle Ras, Andrew Reid, Katja Seeliger and the reviewers for their useful comments.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abbott</surname> <given-names>L. F.</given-names></name> <name><surname>Depasquale</surname> <given-names>B.</given-names></name> <name><surname>Memmesheimer</surname> <given-names>R.-M.</given-names></name></person-group> (<year>2016</year>). <article-title>Building functional networks of spiking model neurons</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>350</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4241</pub-id><pub-id pub-id-type="pmid">26906501</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abraham</surname> <given-names>W. C.</given-names></name> <name><surname>Robins</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Memory retention - the synaptic stability versus plasticity dilemma</article-title>. <source>Trends Neurosci.</source> <volume>28</volume>, <fpage>73</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2004.12.003</pub-id><pub-id pub-id-type="pmid">15667929</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ackley</surname> <given-names>D.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T.</given-names></name></person-group> (<year>1985</year>). <article-title>A learning algorithm for Boltzmann machines</article-title>. <source>Cogn. Sci.</source> <volume>9</volume>, <fpage>147</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1016/S0364-0213(85)80012-4</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adams</surname> <given-names>S. S.</given-names></name> <name><surname>Arel</surname> <given-names>I.</given-names></name> <name><surname>Bach</surname> <given-names>J.</given-names></name> <name><surname>Coop</surname> <given-names>R.</given-names></name> <name><surname>Furlan</surname> <given-names>R.</given-names></name> <name><surname>Goertzel</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Mapping the landscape of human-level artificial general intelligence</article-title>. <source>AI Mag.</source> <volume>33</volume>, <fpage>25</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1609/aimag.v33i1.2322</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Advani</surname> <given-names>M.</given-names></name> <name><surname>Lahiri</surname> <given-names>S.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Statistical mechanics of complex neural systems and high dimensional data</article-title>. <source>J. Stat. Mech. Theory Exp.</source> <volume>2013</volume>:<fpage>P03014</fpage>. <pub-id pub-id-type="doi">10.1088/1742-5468/2013/03/P03014</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aflalo</surname> <given-names>T. N.</given-names></name></person-group> (<year>2006</year>). <article-title>Possible origins of the complex topographic organization of motor cortex: reduction of a multidimensional space onto a two-dimensional array</article-title> <source>J. Neurosci.</source> <volume>26</volume>, <fpage>6288</fpage>&#x02013;<lpage>6297</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0768-06.2006</pub-id><pub-id pub-id-type="pmid">16763036</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Agrawal</surname> <given-names>A.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Antol</surname> <given-names>S.</given-names></name> <name><surname>Mitchell</surname> <given-names>M.</given-names></name> <name><surname>Zitnick</surname> <given-names>C. L.</given-names></name> <name><surname>Batra</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>VQA: visual question answering</article-title>. ArXiv:1505.00468, <fpage>1</fpage>&#x02013;<lpage>25</lpage>.</citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahrens</surname> <given-names>M. B.</given-names></name> <name><surname>Orger</surname> <given-names>M. B.</given-names></name> <name><surname>Robson</surname> <given-names>D. N.</given-names></name> <name><surname>Li</surname> <given-names>J. M.</given-names></name> <name><surname>Keller</surname> <given-names>P. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Whole-brain functional imaging at cellular resolution using light-sheet microscopy</article-title>. <source>Nat. Methods</source> <volume>10</volume>, <fpage>413</fpage>&#x02013;<lpage>420</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2434</pub-id><pub-id pub-id-type="pmid">23524393</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ambrogioni</surname> <given-names>L.</given-names></name> <name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>Maris</surname> <given-names>E.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Estimating nonlinear dynamics with the ConvNet smoother</article-title>. ArXiv:1702.05243, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amunts</surname> <given-names>K.</given-names></name> <name><surname>Ebell</surname> <given-names>C.</given-names></name> <name><surname>Muller</surname> <given-names>J.</given-names></name> <name><surname>Telefont</surname> <given-names>M.</given-names></name> <name><surname>Knoll</surname> <given-names>A.</given-names></name> <name><surname>Lippert</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>The Human Brain Project: creating a European research infrastructure to decode the human brain</article-title>. <source>Neuron</source> <volume>92</volume>, <fpage>574</fpage>&#x02013;<lpage>581</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.10.046</pub-id><pub-id pub-id-type="pmid">27809997</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name> <name><surname>Bothell</surname> <given-names>D.</given-names></name> <name><surname>Byrne</surname> <given-names>M. D.</given-names></name> <name><surname>Douglass</surname> <given-names>S.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Qin</surname> <given-names>Y.</given-names></name></person-group> (<year>2004</year>). <article-title>An integrated theory of the mind</article-title>. <source>Psychol. Rev.</source> <volume>111</volume>, <fpage>1036</fpage>&#x02013;<lpage>1060</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.111.4.1036</pub-id><pub-id pub-id-type="pmid">15482072</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>M. L.</given-names></name></person-group> (<year>2003</year>). <article-title>Embodied cognition: a field guide</article-title>. <source>Artif. Intell.</source> <volume>149</volume>, <fpage>91</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1016/S0004-3702(03)00054-7</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andrieu</surname> <given-names>C.</given-names></name> <name><surname>De Freitas</surname> <given-names>N.</given-names></name> <name><surname>Doucet</surname> <given-names>A.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2003</year>). <article-title>An introduction to MCMC for machine learning</article-title>. <source>Mach. Learn.</source> <volume>50</volume>, <fpage>5</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1023/A:1020281327116</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Anselmi</surname> <given-names>F.</given-names></name> <name><surname>Poggio</surname> <given-names>T. A.</given-names></name></person-group> (<year>2014</year>). <source>Representation Learning in Sensory Cortex: A Theory</source>. Tech. Rep. CBMM Memo 026, <publisher-name>MIT</publisher-name>.</citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ashby</surname> <given-names>W.</given-names></name></person-group> (<year>1952</year>). <source>Design for a Brain</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Chapman &#x00026; Hall</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-94-015-1320-3</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ay</surname> <given-names>N.</given-names></name> <name><surname>Bertschinger</surname> <given-names>N.</given-names></name> <name><surname>Der</surname> <given-names>R.</given-names></name> <name><surname>G&#x000FC;ttler</surname> <given-names>F.</given-names></name> <name><surname>Olbrich</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>Predictive information and explorative behavior of autonomous robots</article-title>. <source>Eur. Phys. J. B</source> <volume>63</volume>, <fpage>329</fpage>&#x02013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1140/epjb/e2008-00175-0</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bachman</surname> <given-names>P.</given-names></name> <name><surname>Sordoni</surname> <given-names>A.</given-names></name> <name><surname>Trischler</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Towards information-seeking agents</article-title>. ArXiv:1612.02605v1, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Badre</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Cognitive control, hierarchy, and the rostro-caudal organization of the frontal lobes</article-title>. <source>Trends Cogn. Sci.</source> <volume>12</volume>, <fpage>193</fpage>&#x02013;<lpage>200</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2008.02.004</pub-id><pub-id pub-id-type="pmid">18403252</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barak</surname> <given-names>O.</given-names></name></person-group> (<year>2017</year>). <article-title>Recurrent neural networks as versatile tools of neuroscience research</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>46</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2017.06.003</pub-id><pub-id pub-id-type="pmid">28668365</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Barkow</surname> <given-names>J. H.</given-names></name> <name><surname>Cosmides</surname> <given-names>L.</given-names></name> <name><surname>Tooby</surname> <given-names>J.</given-names></name></person-group> (eds.). (<year>1992</year>). <source>The Adapted Mind: Evolutionary Psychology and the Generation of Culture</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/s0730938400018700</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Grandmother cells, symmetry, and invariance: how the term arose and what the facts suggest,</article-title> in <source>Cognitive Neurosciences</source>, ed <person-group person-group-type="editor"><name><surname>Gazzaniga</surname> <given-names>M. S.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>309</fpage>&#x02013;<lpage>320</lpage>.</citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barrio</surname> <given-names>L. C.</given-names></name> <name><surname>Buno</surname> <given-names>W.</given-names></name></person-group> (<year>1990</year>). <article-title>Temporal correlations in sensory-synaptic interactions: example in crayfish stretch receptors</article-title>. <source>J. Neurophys.</source> <volume>63</volume>, <fpage>1520</fpage>&#x02013;<lpage>1528</lpage>. <pub-id pub-id-type="pmid">2358890</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baxter</surname> <given-names>J.</given-names></name></person-group> (<year>1998</year>). <article-title>Theoretical models of learning to learn,</article-title> in <source>Learning to Learn</source>, eds <person-group person-group-type="editor"><name><surname>Thrun</surname> <given-names>S.</given-names></name> <name><surname>Pratt</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Norwell, MA</publisher-loc>: <publisher-name>Kluwer Academic Publishers</publisher-name>), <fpage>71</fpage>&#x02013;<lpage>94</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4615-5529-2_4</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Beattie</surname> <given-names>C.</given-names></name> <name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Teplyashin</surname> <given-names>D.</given-names></name> <name><surname>Ward</surname> <given-names>T.</given-names></name> <name><surname>Wainwright</surname> <given-names>M.</given-names></name> <name><surname>Lefrancq</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>DeepMind lab</article-title>. ArXiv:1612.03801v2, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bechtel</surname> <given-names>W.</given-names></name></person-group> (<year>1993</year>). <article-title>The case for connectionism</article-title>. <source>Philos. Stud.</source> <volume>71</volume>, <fpage>119</fpage>&#x02013;<lpage>154</lpage>. <pub-id pub-id-type="doi">10.1007/bf00989853</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bedau</surname> <given-names>M. A.</given-names></name></person-group> (<year>2003</year>). <article-title>Artificial life: organization, adaptation and complexity from the bottom up</article-title>. <source>Trends Cogn. Sci.</source> <volume>7</volume>, <fpage>505</fpage>&#x02013;<lpage>512</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2003.09.012</pub-id><pub-id pub-id-type="pmid">14585448</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2009</year>). <article-title>Learning deep architectures for AI</article-title>. <source>Found. Trends Mach. Learn.</source> <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1561/2200000006</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Evolving culture vs local minima,</article-title> in <source>Growing Adaptive Machine</source>, eds <person-group person-group-type="editor"><name><surname>Kowaliw</surname> <given-names>T.</given-names></name> <name><surname>Bredeche</surname> <given-names>N.</given-names></name> <name><surname>Doursat</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>), <fpage>109</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-55337-0_3</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name></person-group> (<year>2007</year>). <article-title>Scaling learning algorithms towards AI,</article-title> in <source>Large Scale Kernel Machines</source>, eds <person-group person-group-type="editor"><name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Chapelle</surname> <given-names>O.</given-names></name> <name><surname>DeCoste</surname> <given-names>D.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>), <fpage>321</fpage>&#x02013;<lpage>360</lpage>.</citation></ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Louradour</surname> <given-names>J.</given-names></name> <name><surname>Collobert</surname> <given-names>R.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>Curriculum learning,</article-title> in <source>Proceedings of the 26th Annual International Conference on Machine Learning</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1145/1553374.1553380</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bianchini</surname> <given-names>M.</given-names></name> <name><surname>Scarselli</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>On the complexity of neural network classifiers: a comparison between shallow and deep architectures</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>25</volume>, <fpage>1553</fpage>&#x02013;<lpage>1565</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2013.2293637</pub-id><pub-id pub-id-type="pmid">25050951</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bishop</surname> <given-names>C. M.</given-names></name></person-group> (<year>1995</year>). <source>Neural Networks for Pattern Recognition</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B33">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Blei</surname> <given-names>D. M.</given-names></name> <name><surname>Kucukelbir</surname> <given-names>A.</given-names></name> <name><surname>McAuliffe</surname> <given-names>J. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Variational inference: a review for statisticians</article-title>. ArXiv:1601.00670v5, <fpage>1</fpage>&#x02013;<lpage>33</lpage>.</citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blei</surname> <given-names>D. M.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2003</year>). <article-title>Latent dirichlet allocation</article-title>. <source>J. Mach. Learn. Res.</source> <volume>3</volume>, <fpage>993</fpage>&#x02013;<lpage>1022</lpage>. <pub-id pub-id-type="doi">10.1162/jmlr.2003.3.4-5.993</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boden</surname> <given-names>M. A.</given-names></name></person-group> (<year>1998</year>). <article-title>Creativity and artificial intelligence</article-title>. <source>Artif. Intell.</source> <volume>103</volume>, <fpage>347</fpage>&#x02013;<lpage>356</lpage>. <pub-id pub-id-type="doi">10.1016/S0004-3702(98)00055-1</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bohte</surname> <given-names>S. M.</given-names></name></person-group> (<year>2004</year>). <article-title>The evidence for neural information processing with precise spike-times: a survey</article-title>. <source>Nat. Comput.</source> <volume>3</volume>, <fpage>195</fpage>&#x02013;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1023/b:naco.0000027755.02868.60</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bordes</surname> <given-names>A.</given-names></name> <name><surname>Chopra</surname> <given-names>S.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Large-scale simple question answering with memory networks</article-title>. ArXiv:1506.02075v1, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B38">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bosch</surname> <given-names>S. E.</given-names></name> <name><surname>Seeliger</surname> <given-names>K.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Modeling cognitive processes with neural reinforcement learning</article-title>. BioArxiv, <fpage>1</fpage>&#x02013;<lpage>19</lpage>.</citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brachman</surname> <given-names>R. J.</given-names></name></person-group> (<year>2002</year>). <article-title>Systems that know what they&#x00027;re doing</article-title>. <source>IEEE Intell. Syst.</source> <volume>17</volume>, <fpage>67</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1109/mis.2002.1134363</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Braitenberg</surname> <given-names>V.</given-names></name></person-group> (<year>1986</year>). <source>Vehicles: Experiments in Synthetic Psychology</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.1016/0004-3702(85)90057-8</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brea</surname> <given-names>J.</given-names></name> <name><surname>Ga&#x000E1;l</surname> <given-names>A. T.</given-names></name> <name><surname>Urbanczik</surname> <given-names>R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Prospective coding by spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1005003</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1005003</pub-id><pub-id pub-id-type="pmid">27341100</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brea</surname> <given-names>J.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Does computational neuroscience need new synaptic learning paradigms?</article-title> <source>Curr. Opin. Behav. Sci.</source> <volume>11</volume>, <fpage>61</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1016/j.cobeha.2016.05.012</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brette</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Philosophy of the spike: rate-based vs spike-based theories of the brain</article-title>. <source>Front. Syst. Neurosci.</source> <volume>9</volume>:<fpage>151</fpage>. <pub-id pub-id-type="doi">10.3389/fnsys.2015.00151</pub-id><pub-id pub-id-type="pmid">26617496</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Brockman</surname> <given-names>G.</given-names></name> <name><surname>Cheung</surname> <given-names>V.</given-names></name> <name><surname>Pettersson</surname> <given-names>L.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Schulman</surname> <given-names>J.</given-names></name> <name><surname>Tang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>OpenAI gym</article-title>. ArXiv:1606.01540v1, <fpage>1</fpage>&#x02013;<lpage>4</lpage>.</citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brooks</surname> <given-names>R. A.</given-names></name></person-group> (<year>1992</year>). <article-title>Artificial life and real robots,</article-title> in <source>Toward a Practice of Autonomous Systems, Proceedings of First European Conference on Artificial Life</source>, eds <person-group person-group-type="editor"><name><surname>Varela</surname> <given-names>F. J.</given-names></name> <name><surname>Bourgine</surname> <given-names>P.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press, Bradford Books</publisher-name>).</citation></ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brooks</surname> <given-names>R. A.</given-names></name></person-group> (<year>1996</year>). <article-title>Prospects for human level intelligence for humanoid robots,</article-title> in <source>Proceedings of the First International Symposium on Humanoid Robots</source> (<publisher-loc>Tokyo</publisher-loc>), <fpage>17</fpage>&#x02013;<lpage>24</lpage>.</citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>L. V.</given-names></name></person-group> (<year>2007</year>). <source>Psychology of Motivation</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Nova Publishers</publisher-name>.</citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buschman</surname> <given-names>T. J.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2014</year>). <article-title>Goal-direction and top-down control</article-title>. <source>Philos. Trans. R. Soc. B</source> <volume>369</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2013.0471</pub-id><pub-id pub-id-type="pmid">25267814</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cannon</surname> <given-names>W. B.</given-names></name></person-group> (<year>1929</year>). <article-title>Organization for physiological homeostasis</article-title>. <source>Physiol. Rev.</source> <volume>9</volume>, <fpage>399</fpage>&#x02013;<lpage>431</lpage>.</citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carnevale</surname> <given-names>F.</given-names></name> <name><surname>de Lafuente</surname> <given-names>V.</given-names></name> <name><surname>Romo</surname> <given-names>R.</given-names></name> <name><surname>Barak</surname> <given-names>O.</given-names></name> <name><surname>Parga</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Dynamic control of response criterion in premotor cortex during perceptual detection under temporal uncertainty</article-title>. <source>Neuron</source> <volume>86</volume>, <fpage>1067</fpage>&#x02013;<lpage>1077</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.04.014</pub-id><pub-id pub-id-type="pmid">25959731</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carr</surname> <given-names>C. E.</given-names></name> <name><surname>Konishi</surname> <given-names>M.</given-names></name></person-group> (<year>1990</year>). <article-title>A circuit for detection of interaural time differences in the brain stem of the barn owl</article-title>. <source>J. Neurosci.</source> <volume>10</volume>, <fpage>3227</fpage>&#x02013;<lpage>3246</lpage>. <pub-id pub-id-type="pmid">2213141</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caruana</surname> <given-names>R.</given-names></name></person-group> (<year>1997</year>). <article-title>Multitask learning</article-title>. <source>Mach. Learn.</source> <volume>28</volume>, <fpage>41</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2010.22</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>E. F.</given-names></name></person-group> (<year>2015</year>). <article-title>Towards large-scale, human-based, mesoscopic neurotechnologies</article-title>. <source>Neuron</source> <volume>86</volume>, <fpage>68</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.03.037</pub-id><pub-id pub-id-type="pmid">25856487</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>van Merrienboer</surname> <given-names>B.</given-names></name> <name><surname>Bahdanau</surname> <given-names>D.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>On the properties of neural machine translation: encoder-decoder approaches,</article-title> in <source>Proceedings of the SSST-8, Eighth Work Syntax Semantics and Structure in Statistical Translation</source> (<publisher-loc>Doha</publisher-loc>), <fpage>103</fpage>&#x02013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.3115/v1/w14-4012</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Churchland</surname> <given-names>P. S.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Blending computational and experimental neuroscience</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>17</volume>, <fpage>667</fpage>&#x02013;<lpage>668</lpage>. <pub-id pub-id-type="doi">10.1038/nrn.2016.114</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Whatever next? Predictive brains, situated agents, and the future of cognitive science</article-title>. <source>Behav. Brain Sci.</source> <volume>36</volume>, <fpage>181</fpage>&#x02013;<lpage>204</lpage>. <pub-id pub-id-type="doi">10.1017/s0140525x12000477</pub-id><pub-id pub-id-type="pmid">23663408</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J. D.</given-names></name> <name><surname>McClure</surname> <given-names>S. M.</given-names></name> <name><surname>Yu</surname> <given-names>A. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration</article-title>. <source>Philos. Trans. R. Soc. B</source> <volume>362</volume>, <fpage>933</fpage>&#x02013;<lpage>942</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2007.2098</pub-id><pub-id pub-id-type="pmid">17395573</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Copeland</surname> <given-names>B. J.</given-names></name> <name><surname>Proudfoot</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <article-title>On Alan Turing&#x00027;s anticipation of connectionism</article-title>. <source>Synthese</source> <volume>108</volume>, <fpage>361</fpage>&#x02013;<lpage>377</lpage>. <pub-id pub-id-type="doi">10.1007/bf00413694</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Corneil</surname> <given-names>D.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2015</year>). <article-title>Attractor network dynamics enable preplay and rapid path planning in maze-like environments,</article-title> in <source>Advances in Neural Information Processing Systems 28</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. D.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural networks and neuroscience-inspired computer vision</article-title>. <source>Curr. Biol.</source> <volume>24</volume>, <fpage>R921</fpage>&#x02013;<lpage>R929</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2014.08.026</pub-id><pub-id pub-id-type="pmid">25247371</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crick</surname> <given-names>F.</given-names></name> <name><surname>Mitchison</surname> <given-names>G.</given-names></name></person-group> (<year>1983</year>). <article-title>The function of dream sleep</article-title>. <source>Nature</source> <volume>304</volume>, <fpage>111</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1038/304111a0</pub-id><pub-id pub-id-type="pmid">6866101</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Csikszentmihalyi</surname> <given-names>M.</given-names></name></person-group> (<year>1975</year>). <source>Beyond Boredom and Anxiety: Experiencing Flow in Work and Play</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons Inc</publisher-name>. <pub-id pub-id-type="doi">10.2307/2065805</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cybenko</surname> <given-names>G.</given-names></name></person-group> (<year>1989</year>). <article-title>Approximation by superpositions of a sigmoidal function</article-title>. <source>Math. Control Signals Syst.</source> <volume>2</volume>, <fpage>303</fpage>&#x02013;<lpage>314</lpage>. <pub-id pub-id-type="doi">10.1007/BF02134016</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Dauphin</surname> <given-names>Y.</given-names></name> <name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Gulcehre</surname> <given-names>C.</given-names></name> <name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Identifying and attacking the saddle point problem in high-dimensional non-convex optimization</article-title>. ArXiv:1406.2572, <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B65">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Davies</surname> <given-names>N. B.</given-names></name> <name><surname>Krebs</surname> <given-names>J. R.</given-names></name> <name><surname>West</surname> <given-names>S. A.</given-names></name></person-group> (<year>2012</year>). <source>An Introduction to Behavioral Ecology, 4th Edn</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons</publisher-name>. <pub-id pub-id-type="doi">10.1037/026600</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Daw</surname> <given-names>N. D.</given-names></name></person-group> (<year>2012</year>). <article-title>Model-based reinforcement learning as cognitive search: neurocomputational theories,</article-title> in <source>Cognitive Search: Evolution, Algorithms, and the Brain</source>, eds <person-group person-group-type="editor"><name><surname>Todd</surname> <given-names>P. M.</given-names></name> <name><surname>Hills</surname> <given-names>T. T.</given-names></name> <name><surname>Robbins</surname> <given-names>T. W.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>), <fpage>195</fpage>&#x02013;<lpage>208</lpage>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262018098.001.0001</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dawkins</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <source>The Selfish Gene, 4th Edn</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>. <pub-id pub-id-type="doi">10.4324/9781912281251</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dawson</surname> <given-names>M. R. W.</given-names></name> <name><surname>Shamanski</surname> <given-names>K. S.</given-names></name></person-group> (<year>1994</year>). <article-title>Connectionism, confusion, and cognitive science</article-title>. <source>J. Intell. Syst.</source> <volume>4</volume>, <fpage>215</fpage>&#x02013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1515/jisys.1994.4.3-4.215</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Abbott</surname> <given-names>L. F.</given-names></name></person-group> (<year>2005</year>). <source>Theoretical Neuroscience</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Neal</surname> <given-names>R.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name></person-group> (<year>1995</year>). <article-title>The Helmholtz machine</article-title>. <source>Neural Comput.</source> <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1995.7.5.889</pub-id><pub-id pub-id-type="pmid">7584891</pub-id></citation></ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Garis</surname> <given-names>H.</given-names></name> <name><surname>Shuo</surname> <given-names>C.</given-names></name> <name><surname>Goertzel</surname> <given-names>B.</given-names></name> <name><surname>Ruiting</surname> <given-names>L.</given-names></name></person-group> (<year>2010</year>). <article-title>A world survey of artificial brain projects, Part I Large-scale brain simulations</article-title>. <source>Neurocomputing</source> <volume>74</volume>, <fpage>3</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2010.08.004</pub-id></citation></ref>
<ref id="B72">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Delalleau</surname> <given-names>O.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). <article-title>Shallow vs. deep sum-product networks,</article-title> in <source>Advances in Neural Information Processing Systems 24</source> (<publisher-loc>Granada</publisher-loc>), <fpage>666</fpage>&#x02013;<lpage>674</lpage>.</citation></ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Der</surname> <given-names>R.</given-names></name> <name><surname>Martius</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <source>The Playful Machine: Theoretical Foundation and Practical Realization of Self-Organizing Robots</source>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer Verlag</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-3-642-20253-7</pub-id></citation></ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Der</surname> <given-names>R.</given-names></name> <name><surname>Steinmetz</surname> <given-names>U.</given-names></name> <name><surname>Pasemann</surname> <given-names>F.</given-names></name></person-group> (<year>1999</year>). <article-title>Homeokinesis - a new principle to back up evolution with learning</article-title>. <source>Comput. Intell. Model. Control. Autom.</source> <volume>55</volume>, <fpage>43</fpage>&#x02013;<lpage>47</lpage>.</citation></ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dewar</surname> <given-names>R. C.</given-names></name></person-group> (<year>2003</year>). <article-title>Information theory explanation of the fluctuation theorem, maximum entropy production and self-organized criticality in non-equilibrium stationary states</article-title>. <source>J. Phys. A Math. Gen.</source> <volume>36</volume>, <fpage>631</fpage>&#x02013;<lpage>641</lpage>. <pub-id pub-id-type="doi">10.1088/0305-4470/36/3/303</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dewar</surname> <given-names>R. C.</given-names></name></person-group> (<year>2005</year>). <article-title>Maximum entropy production and the fluctuation theorem</article-title>. <source>J. Phys. A Math. Gen.</source> <volume>38</volume>, <fpage>L371</fpage>&#x02013;<lpage>L381</lpage>. <pub-id pub-id-type="doi">10.1088/0305-4470/38/21/L01</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dewey</surname> <given-names>J.</given-names></name></person-group> (<year>1896</year>). <article-title>The reflex arc concept in psychology</article-title>. <source>Psychol. Rev.</source> <volume>3</volume>, <fpage>357</fpage>&#x02013;<lpage>370</lpage>. <pub-id pub-id-type="doi">10.1037/11304-041</pub-id></citation></ref>
<ref id="B78">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Doya</surname> <given-names>K.</given-names></name> <name><surname>Ishii</surname> <given-names>S.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name> <name><surname>Rao</surname> <given-names>R. P. N.</given-names></name></person-group> (eds.). (<year>2006</year>). <source>Bayesian Brain: Probabilistic Approaches to Neural Coding</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation>
</ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dragoi</surname> <given-names>G.</given-names></name> <name><surname>Tonegawa</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Hippocampal cellular assemblies</article-title>. <source>Nature</source> <volume>469</volume>, <fpage>397</fpage>&#x02013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1038/nature09633</pub-id><pub-id pub-id-type="pmid">21179088</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Drexler</surname> <given-names>K. E.</given-names></name></person-group> (<year>1992</year>). <source>Nanosystems: Molecular Machinery, Manufacturing, and Computation</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley Interscience</publisher-name>. <pub-id pub-id-type="doi">10.1016/S0010-8545(96)90165-4</pub-id></citation></ref>
<ref id="B81">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>Y.</given-names></name> <name><surname>Andrychowicz</surname> <given-names>M.</given-names></name> <name><surname>Stadie</surname> <given-names>B. C.</given-names></name> <name><surname>Ho</surname> <given-names>J.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>One-shot imitation learning</article-title>. ArXiv:1703.07326v2, <fpage>1</fpage>&#x02013;<lpage>23</lpage>.</citation></ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duysens</surname> <given-names>J.</given-names></name> <name><surname>Van de Crommert</surname> <given-names>H. W. A. A.</given-names></name></person-group> (<year>1998</year>). <article-title>Neural control of locomotion; The central pattern generator from cats to humans</article-title>. <source>Gait Posture</source> <volume>7</volume>, <fpage>131</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1016/S0966-6362(97)00042-8</pub-id><pub-id pub-id-type="pmid">10200383</pub-id></citation></ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edelman</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>The minority report: some common assumptions to reconsider in the modelling of the brain and behavior</article-title>. <source>J. Exp. Theor. Artif. Intell.</source> <volume>3079</volume>, <fpage>1</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1080/0952813X.2015.1042534</pub-id></citation></ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name></person-group> (<year>1990</year>). <article-title>Finding structure in time</article-title>. <source>Cogn. Sci.</source> <volume>14</volume>, <fpage>179</fpage>&#x02013;<lpage>211</lpage>. <pub-id pub-id-type="doi">10.1016/0364-0213(90)90002-E</pub-id></citation></ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name></person-group> (<year>1991</year>). <article-title>Distributed representations, simple recurrent networks, and grammatical structure</article-title>. <source>Mach. Learn.</source> <volume>7</volume>, <fpage>195</fpage>&#x02013;<lpage>225</lpage>. <pub-id pub-id-type="doi">10.1023/A:1022699029236</pub-id></citation></ref>
<ref id="B86">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name></person-group> (<year>1993</year>). <article-title>Learning and development in neural networks - The importance of starting small</article-title>. <source>Cognition</source> <volume>48</volume>, <fpage>71</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/S0010-0277(02)00106-3</pub-id><pub-id pub-id-type="pmid">8403835</pub-id></citation></ref>
<ref id="B87">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name> <name><surname>Bates</surname> <given-names>E. A.</given-names></name> <name><surname>Johnson</surname> <given-names>M. H.</given-names></name> <name><surname>Karmiloff-Smith</surname> <given-names>A.</given-names></name> <name><surname>Parisi</surname> <given-names>D.</given-names></name> <name><surname>Plunkett</surname> <given-names>K.</given-names></name></person-group> (<year>1996</year>). <source>Rethinking Innateness: A Connectionist Perspective on Development</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/s0272263198333070</pub-id></citation></ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fei-Fei</surname> <given-names>L.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>Member</surname> <given-names>S.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>One-shot learning of object categories</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>28</volume>, <fpage>594</fpage>&#x02013;<lpage>611</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2006.79</pub-id><pub-id pub-id-type="pmid">16566508</pub-id></citation></ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Felleman</surname> <given-names>D. J.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name></person-group> (<year>1991</year>). <article-title>Distributed hierarchical processing in the primate cerebral cortex</article-title>. <source>Cereb. Cortex</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/1.1.1</pub-id><pub-id pub-id-type="pmid">1822724</pub-id></citation></ref>
<ref id="B90">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Fernando</surname> <given-names>C.</given-names></name> <name><surname>Banarse</surname> <given-names>D.</given-names></name> <name><surname>Blundell</surname> <given-names>C.</given-names></name> <name><surname>Zwols</surname> <given-names>Y.</given-names></name> <name><surname>Ha</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>PathNet: evolution channels gradient descent in super neural networks</article-title>. ArXiv:1701.08734v1.</citation></ref>
<ref id="B91">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ferrone</surname> <given-names>L.</given-names></name> <name><surname>Zanzotto</surname> <given-names>F. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Symbolic, distributed and distributional representations for natural language processing in the era of deep learning: a survey</article-title>. ArXiv:1702.00764, <fpage>1</fpage>&#x02013;<lpage>25</lpage>.</citation></ref>
<ref id="B92">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrucci</surname> <given-names>D.</given-names></name> <name><surname>Brown</surname> <given-names>E.</given-names></name> <name><surname>Chu-carroll</surname> <given-names>J.</given-names></name> <name><surname>Fan</surname> <given-names>J.</given-names></name> <name><surname>Gondek</surname> <given-names>D.</given-names></name> <name><surname>Kalyanpur</surname> <given-names>A. A.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Building Watson: an overview of the DeepQA project</article-title>. <source>AI Mag.</source> <volume>31</volume>, <fpage>59</fpage>&#x02013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1609/aimag.v31i3.2303</pub-id></citation></ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feynman</surname> <given-names>R.</given-names></name></person-group> (<year>1992</year>). <article-title>There&#x00027;s plenty of room at the bottom</article-title>. <source>J. Microelectromech. Syst.</source> <volume>1</volume>, <fpage>60</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1109/84.128057</pub-id></citation></ref>
<ref id="B94">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Floreano</surname> <given-names>D.</given-names></name> <name><surname>D&#x000FC;rr</surname> <given-names>P.</given-names></name> <name><surname>Mattiussi</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>Neuroevolution: from architectures to learning</article-title>. <source>Evol. Intell.</source> <volume>1</volume>, <fpage>47</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1007/s12065-007-0002-4</pub-id></citation></ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fodor</surname> <given-names>J. A.</given-names></name> <name><surname>Pylyshyn</surname> <given-names>Z. W.</given-names></name></person-group> (<year>1988</year>). <article-title>Connectionism and cognitive architecture: a critical analysis</article-title>. <source>Cognition</source> <volume>28</volume>, <fpage>3</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(88)90031-5</pub-id><pub-id pub-id-type="pmid">2450716</pub-id></citation></ref>
<ref id="B96">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Forstmann</surname> <given-names>B. U.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2015</year>). <source>Model-Based Cognitive Neuroscience: A Conceptual Introduction</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-1-4939-2236-9_7</pub-id></citation></ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>French</surname> <given-names>R. M.</given-names></name></person-group> (<year>1999</year>). <article-title>Catastrophic forgetting in connectionist networks</article-title>. <source>Trends Cogn. Sci.</source> <volume>3</volume>, <fpage>128</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1016/s1364-6613(99)01294-2</pub-id><pub-id pub-id-type="pmid">10322466</pub-id></citation></ref>
<ref id="B98">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2009</year>). <article-title>The free-energy principle: a rough guide to the brain?</article-title> <source>Trends Cogn. Sci.</source> <volume>13</volume>, <fpage>293</fpage>&#x02013;<lpage>301</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2009.04.005</pub-id><pub-id pub-id-type="pmid">19559644</pub-id></citation></ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>11</volume>, <fpage>127</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id><pub-id pub-id-type="pmid">20068583</pub-id></citation></ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Daunizeau</surname> <given-names>J.</given-names></name> <name><surname>Kilner</surname> <given-names>J.</given-names></name> <name><surname>Kiebel</surname> <given-names>S. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Action and behavior: a free-energy formulation</article-title>. <source>Biol. Cybern.</source> <volume>102</volume>, <fpage>227</fpage>&#x02013;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-010-0364-z</pub-id><pub-id pub-id-type="pmid">20148260</pub-id></citation></ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fry</surname> <given-names>R. L.</given-names></name></person-group> (<year>2017</year>). <article-title>Physical intelligence and thermodynamic computing</article-title>. <source>Entropy</source> <volume>19</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.20944/PREPRINTS201701.0097.V1</pub-id></citation></ref>
<ref id="B102">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fukushima</surname> <given-names>K.</given-names></name></person-group> (<year>1980</year>). <article-title>Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position</article-title>. <source>Biol. Cybern.</source> <volume>36</volume>, <fpage>193</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1007/bf00344251</pub-id><pub-id pub-id-type="pmid">7370364</pub-id></citation></ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fukushima</surname> <given-names>K.</given-names></name></person-group> (<year>2013</year>). <article-title>Artificial vision by multi-layered neural networks: neocognitron and its advances</article-title>. <source>Neural Netw.</source> <volume>37</volume>, <fpage>103</fpage>&#x02013;<lpage>119</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2012.09.016</pub-id><pub-id pub-id-type="pmid">23098752</pub-id></citation></ref>
<ref id="B104">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funahashi</surname> <given-names>K.-I.</given-names></name> <name><surname>Nakamura</surname> <given-names>Y.</given-names></name></person-group> (<year>1993</year>). <article-title>Approximation of dynamical systems by continuous time recurrent neural networks</article-title>. <source>Neural Netw.</source> <volume>6</volume>, <fpage>801</fpage>&#x02013;<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1016/s0893-6080(05)80125-x</pub-id></citation></ref>
<ref id="B105">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuster</surname> <given-names>J. M.</given-names></name></person-group> (<year>2001</year>). <article-title>The prefrontal cortex - An update: time is of the essence</article-title>. <source>Neuron</source> <volume>30</volume>, <fpage>319</fpage>&#x02013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1016/S0896-6273(01)00285-9</pub-id><pub-id pub-id-type="pmid">11394996</pub-id></citation></ref>
<ref id="B106">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuster</surname> <given-names>J. M.</given-names></name></person-group> (<year>2004</year>). <article-title>Upper processing stages of the perception-action cycle</article-title>. <source>Trends Cogn. Sci.</source> <volume>8</volume>, <fpage>143</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2004.02.004</pub-id><pub-id pub-id-type="pmid">15551481</pub-id></citation></ref>
<ref id="B107">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Gal</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Dropout as a Bayesian approximation: representing model uncertainty in deep learning</article-title>. ArXiv:1506.02142v6, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gardner</surname> <given-names>E.</given-names></name></person-group> (<year>1988</year>). <article-title>The space of interactions in neural network models</article-title>. <source>J. Phys. A. Math. Gen.</source> <volume>21</volume>, <fpage>257</fpage>&#x02013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1088/0305-4470/21/1/030</pub-id></citation></ref>
<ref id="B109">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gardner</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <source>The Colossal Book of Mathematics: Classic Puzzles, Paradoxes, and Problems</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>W. W. Norton &#x00026; Company</publisher-name>.</citation></ref>
<ref id="B110">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gasser</surname> <given-names>M.</given-names></name> <name><surname>Eck</surname> <given-names>D.</given-names></name> <name><surname>Port</surname> <given-names>R.</given-names></name></person-group> (<year>1999</year>). <article-title>Meter as mechanism: a neural network model that learns metrical patterns</article-title>. <source>Conn. Sci.</source> <volume>11</volume>, <fpage>187</fpage>&#x02013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1080/095400999116331</pub-id></citation></ref>
<ref id="B111">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gauci</surname> <given-names>J.</given-names></name> <name><surname>Stanley</surname> <given-names>K. O.</given-names></name></person-group> (<year>2010</year>). <article-title>Autonomous evolution of topographic regularities in artificial neural networks</article-title>. <source>Neural Comput.</source> <volume>22</volume>, <fpage>1860</fpage>&#x02013;<lpage>1898</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2010.06-09-1042</pub-id><pub-id pub-id-type="pmid">20235822</pub-id></citation></ref>
<ref id="B112">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <name><surname>Beck</surname> <given-names>J. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Complex probabilistic inference: from cognition to neural computation,</article-title> in <source>Computational Models of Brain and Behavior</source>, ed <person-group person-group-type="editor"><name><surname>Moustafa</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Chichester, UK</publisher-loc>: <publisher-name>Wiley-Blackwell</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1002/9781119159193.ch33</pub-id></citation></ref>
<ref id="B113">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <name><surname>Horvitz</surname> <given-names>E. J.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <article-title>Computational rationality: a converging paradigm for intelligence in brains, minds, and machines</article-title>. <source>Science</source> <volume>349</volume>, <fpage>273</fpage>&#x02013;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1126/science.aac6076</pub-id><pub-id pub-id-type="pmid">26185246</pub-id></citation></ref>
<ref id="B114">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Kistler</surname> <given-names>W. M.</given-names></name></person-group> (<year>2002</year>). <source>Spiking Neuron Models</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="pmid">16802968</pub-id></citation></ref>
<ref id="B115">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Kistler</surname> <given-names>W. M.</given-names></name> <name><surname>Naud</surname> <given-names>R.</given-names></name> <name><surname>Paninski</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <source>Neuronal Dynamics: From Single Neurons to Networks and Models of Cognition</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9781107447615</pub-id></citation></ref>
<ref id="B116">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gibson</surname> <given-names>J.</given-names></name></person-group> (<year>1979</year>). <source>The Ecological Approach to Visual Perception</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Houghton Mifflin</publisher-name>. <pub-id pub-id-type="doi">10.1002/bs.3830260313</pub-id></citation></ref>
<ref id="B117">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gigerenzer</surname> <given-names>G.</given-names></name> <name><surname>Goldstein</surname> <given-names>D. G.</given-names></name></person-group> (<year>1996</year>). <article-title>Reasoning the fast and frugal way: models of bounded rationality</article-title>. <source>Psychol. Rev.</source> <volume>103</volume>, <fpage>650</fpage>&#x02013;<lpage>669</lpage>. <pub-id pub-id-type="doi">10.1037//0033-295x.103.4.650</pub-id><pub-id pub-id-type="pmid">8888650</pub-id></citation></ref>
<ref id="B118">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilbert</surname> <given-names>C. D.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Top-down influences on visual processing</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>14</volume>, <fpage>350</fpage>&#x02013;<lpage>363</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3476</pub-id><pub-id pub-id-type="pmid">23595013</pub-id></citation></ref>
<ref id="B119">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I.</given-names></name> <name><surname>Pouget-Abadie</surname> <given-names>J.</given-names></name> <name><surname>Mirza</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Warde-Farley</surname> <given-names>D.</given-names></name> <name><surname>Ozair</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Generative adversarial nets</article-title>. ArXiv:1406.2661v1, <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B120">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gordon</surname> <given-names>G.</given-names></name> <name><surname>Ahissar</surname> <given-names>E.</given-names></name></person-group> (<year>2012</year>). <article-title>Hierarchical curiosity loops and active sensing</article-title>. <source>Neural Netw.</source> <volume>32</volume>, <fpage>119</fpage>&#x02013;<lpage>129</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2012.02.024</pub-id><pub-id pub-id-type="pmid">22386787</pub-id></citation></ref>
<ref id="B121">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Wayne</surname> <given-names>G.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural turing machines</article-title>. ArXiv:1410.5401, <fpage>1</fpage>&#x02013;<lpage>26</lpage>.</citation></ref>
<ref id="B122">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Gregor</surname> <given-names>K.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>DRAW: a recurrent neural network for image generation</article-title>. ArXiv:1502.04623v1, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B123">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Griffiths</surname> <given-names>T.</given-names></name> <name><surname>Chater</surname> <given-names>N.</given-names></name> <name><surname>Kemp</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Probabilistic models of cognition: exploring representations and inductive biases</article-title>. <source>Trends Cogn. Sci.</source> <volume>14</volume>, <fpage>357</fpage>&#x02013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.05.004</pub-id><pub-id pub-id-type="pmid">20576465</pub-id></citation></ref>
<ref id="B124">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grinstein</surname> <given-names>G.</given-names></name> <name><surname>Linsker</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <article-title>Comments on a derivation and application of the &#x02018;maximum entropy production&#x02019; principle</article-title>. <source>J. Phys. A Math. Theor.</source> <volume>40</volume>, <fpage>9717</fpage>&#x02013;<lpage>9720</lpage>. <pub-id pub-id-type="doi">10.1088/1751-8113/40/31/n01</pub-id></citation></ref>
<ref id="B125">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grothe</surname> <given-names>B.</given-names></name></person-group> (<year>2003</year>). <article-title>New roles for synaptic inhibition in sound localization</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>4</volume>, <fpage>540</fpage>&#x02013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1136</pub-id><pub-id pub-id-type="pmid">12838329</pub-id></citation></ref>
<ref id="B126">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>Thielen</surname> <given-names>J.</given-names></name> <name><surname>Hanke</surname> <given-names>M.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Brains on beats,</article-title> in <source>Advances in Neural Information Processing Systems 29</source> (<publisher-loc>Barcelona</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B127">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M.</given-names></name></person-group> (<year>2017a</year>). <article-title>Increasingly complex representations of natural movies across the dorsal stream are shared between subjects</article-title>. <source>Neuroimage</source> <volume>145</volume>, <fpage>329</fpage>&#x02013;<lpage>336</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2015.12.036</pub-id><pub-id pub-id-type="pmid">26724778</pub-id></citation></ref>
<ref id="B128">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>10005</fpage>&#x02013;<lpage>10014</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5023-14.2015</pub-id><pub-id pub-id-type="pmid">26157000</pub-id></citation></ref>
<ref id="B129">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2017b</year>). <article-title>Modeling the dynamics of human brain activity with recurrent neural networks</article-title>. <source>Front. Comput. Neurosci.</source> <volume>11</volume>:<fpage>7</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2017.00007</pub-id><pub-id pub-id-type="pmid">28232797</pub-id></citation></ref>
<ref id="B130">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;t&#x000FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>Seeliger</surname> <given-names>K.</given-names></name> <name><surname>Bosch</surname> <given-names>S.</given-names></name> <name><surname>van Lier</surname> <given-names>R.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep adversarial neural decoding,</article-title> in <source>Advances in Neural Information Processing Systems 30</source> (<publisher-loc>Long Beach</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B131">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;t&#x000FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name> <name><surname>van Lier</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep impression: audiovisual deep residual networks for multimodal apparent personality trait recognition,</article-title> in <source>Proceedings of the 14th European Conference on Computer Vision</source> (<publisher-loc>Amsterdam</publisher-loc>).</citation></ref>
<ref id="B132">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Harari</surname> <given-names>Y. N.</given-names></name></person-group> (<year>2017</year>). <source>Homo Deus: A Brief History of Tomorrow, 1st Edn</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Vintage Books</publisher-name>.</citation></ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harnad</surname> <given-names>S.</given-names></name></person-group> (<year>1990</year>). <article-title>The symbol grounding problem</article-title>. <source>Phys. D Nonlin. Phenom.</source> <volume>42</volume>, <fpage>335</fpage>&#x02013;<lpage>346</lpage>. <pub-id pub-id-type="doi">10.1016/0167-2789(90)90087-6</pub-id></citation></ref>
<ref id="B134">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Kumaran</surname> <given-names>D.</given-names></name> <name><surname>Summerfield</surname> <given-names>C.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Neuroscience-inspired artificial intelligence</article-title>. <source>Neuron</source> <volume>95</volume>, <fpage>245</fpage>&#x02013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2017.06.011</pub-id><pub-id pub-id-type="pmid">28728020</pub-id></citation></ref>
<ref id="B135">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hatfield</surname> <given-names>G.</given-names></name></person-group> (<year>2002</year>). <article-title>Perception and the physical world: psychological and philosophical issues in perception,</article-title> in <source>Perception and the Physical World: Psychological and Philosophical Issues in Perception</source>, eds <person-group person-group-type="editor"><name><surname>Heyer</surname> <given-names>D.</given-names></name> <name><surname>Mausfeld</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley and Sons</publisher-name>), <fpage>113</fpage>&#x02013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1002/0470013427.ch5</pub-id></citation></ref>
<ref id="B136">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep residual learning for image recognition</article-title>. ArXiv:1512.03385, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B137">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heeger</surname> <given-names>D. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Theory of cortical function</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>114</volume>, <fpage>1773</fpage>&#x02013;<lpage>1782</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1619788114</pub-id><pub-id pub-id-type="pmid">28167793</pub-id></citation></ref>
<ref id="B138">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herculano-Houzel</surname> <given-names>S.</given-names></name> <name><surname>Lent</surname> <given-names>R.</given-names></name></person-group> (<year>2005</year>). <article-title>Isotropic fractionator: a simple, rapid method for the quantification of total cell and neuron numbers in the brain</article-title>. <source>J. Neurosci.</source> <volume>25</volume>, <fpage>2518</fpage>&#x02013;<lpage>2521</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4526-04.2005</pub-id><pub-id pub-id-type="pmid">15758160</pub-id></citation></ref>
<ref id="B139">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hertz</surname> <given-names>J. A.</given-names></name> <name><surname>Krogh</surname> <given-names>A. S.</given-names></name> <name><surname>Palmer</surname> <given-names>R. G.</given-names></name></person-group> (<year>1991</year>). <source>Introduction to the Theory of Neural Computation</source>. <publisher-loc>Boulder, CO</publisher-loc>: <publisher-name>Westview Press</publisher-name>. <pub-id pub-id-type="doi">10.1063/1.2810360</pub-id></citation></ref>
<ref id="B140">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>Where do features come from?</article-title> <source>Cogn. Sci.</source> <volume>38</volume>, <fpage>1078</fpage>&#x02013;<lpage>1101</lpage>. <pub-id pub-id-type="doi">10.1111/cogs.12049</pub-id><pub-id pub-id-type="pmid">23800216</pub-id></citation></ref>
<ref id="B141">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>McLelland</surname> <given-names>J. L.</given-names></name> <name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name></person-group> (<year>1986</year>). <article-title>Distributed representations,</article-title> in <source>Parallel Distributed Processing Explorations in the Microstructure of Cognition</source>, <volume>vol. 1</volume>, eds <person-group person-group-type="editor"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>77</fpage>&#x02013;<lpage>109</lpage>.</citation></ref>
<ref id="B142">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Osindero</surname> <given-names>S.</given-names></name> <name><surname>Teh</surname> <given-names>Y.</given-names></name></person-group> (<year>2006</year>). <article-title>A fast learning algorithm for deep belief nets</article-title>. <source>Neural Comput.</source> <volume>18</volume>, <fpage>1527</fpage>&#x02013;<lpage>1554</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2006.18.7.1527</pub-id><pub-id pub-id-type="pmid">16764513</pub-id></citation></ref>
<ref id="B143">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>1983</year>). <article-title>Optimal perceptual inference,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Washington, DC</publisher-loc>).</citation></ref>
<ref id="B144">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id></citation></ref>
<ref id="B145">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>1982</year>). <article-title>Neural networks and physical systems with emergent collective computational abilities</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>79</volume>, <fpage>2554</fpage>&#x02013;<lpage>2558</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.79.8.2554</pub-id><pub-id pub-id-type="pmid">6953413</pub-id></citation></ref>
<ref id="B146">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hornik</surname> <given-names>K.</given-names></name></person-group> (<year>1991</year>). <article-title>Approximation capabilities of multilayer feedforward networks</article-title>. <source>Neural Netw.</source> <volume>4</volume>, <fpage>251</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.1016/0893-6080(91)90009-T</pub-id></citation></ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Rao</surname> <given-names>R. P. N.</given-names></name></person-group> (<year>2011</year>). <article-title>Predictive coding</article-title>. <source>WIREs Cogn. Sci.</source> <volume>2</volume>, <fpage>580</fpage>&#x02013;<lpage>593</lpage>. <pub-id pub-id-type="doi">10.1002/wcs.142</pub-id><pub-id pub-id-type="pmid">26302308</pub-id></citation></ref>
<ref id="B148">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Huh</surname> <given-names>D.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Gradient descent for spiking neural networks</article-title>. ArXiv:1706.04698, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B149">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huo</surname> <given-names>J.</given-names></name> <name><surname>Murray</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>The adaptation of visual and auditory integration in the barn owl superior colliculus with spike timing dependent plasticity</article-title>. <source>Neural Netw.</source> <volume>22</volume>, <fpage>913</fpage>&#x02013;<lpage>921</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2008.10.007</pub-id><pub-id pub-id-type="pmid">19084371</pub-id></citation></ref>
<ref id="B150">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ijspeert</surname> <given-names>A. J.</given-names></name></person-group> (<year>2008</year>). <article-title>Central pattern generators for locomotion control in animals and robots: a review</article-title>. <source>Neural Netw.</source> <volume>21</volume>, <fpage>642</fpage>&#x02013;<lpage>653</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2008.03.014</pub-id><pub-id pub-id-type="pmid">18555958</pub-id></citation></ref>
<ref id="B151">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Szegedy</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Batch normalization: accelerating deep network training by reducing internal covariate shift</article-title>. ArXiv:1502.03167, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B152">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Izhikevich</surname> <given-names>E. M.</given-names></name> <name><surname>Edelman</surname> <given-names>G. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Large-scale model of mammalian thalamocortical systems</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>105</volume>, <fpage>3593</fpage>&#x02013;<lpage>3598</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0712231105</pub-id><pub-id pub-id-type="pmid">18292226</pub-id></citation></ref>
<ref id="B153">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaynes</surname> <given-names>E.</given-names></name></person-group> (<year>1988</year>). <article-title>How does the brain do plausible reasoning?</article-title> <source>Maximum Entropy Bayesian Methods Sci. Eng.</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1007/978-94-009-3049-0_1</pub-id></citation></ref>
<ref id="B154">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jeffress</surname> <given-names>L. A.</given-names></name></person-group> (<year>1948</year>). <article-title>A place theory of sound localization</article-title>. <source>J. Comp. Physiol. Psychol.</source> <volume>41</volume>, <fpage>35</fpage>&#x02013;<lpage>39</lpage>. <pub-id pub-id-type="doi">10.1037/h0061495</pub-id><pub-id pub-id-type="pmid">18904764</pub-id></citation></ref>
<ref id="B155">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>J.</given-names></name> <name><surname>Hariharan</surname> <given-names>B.</given-names></name> <name><surname>van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Hoffman</surname> <given-names>J.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name> <name><surname>Zitnick</surname> <given-names>C. L.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Inferring and executing programs for visual reasoning</article-title>. ArXiv:1705.03633.</citation></ref>
<ref id="B156">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jonas</surname> <given-names>E.</given-names></name> <name><surname>Kording</surname> <given-names>K. P.</given-names></name></person-group> (<year>2017</year>). <article-title>Could a neuroscientist understand a microprocessor?</article-title> <source>PloS Comput. Biol.</source> <volume>13</volume>:<fpage>e1005268</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1005268</pub-id><pub-id pub-id-type="pmid">28081141</pub-id></citation></ref>
<ref id="B157">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>1987</year>). <article-title>Attractor dynamics and parallelism in a connectionist sequential machine,</article-title> in <source>Proceedings of the Eighth Annual Conference of the Cognitive Science Society</source>, <fpage>531</fpage>&#x02013;<lpage>546</lpage>.</citation></ref>
<ref id="B158">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jordan</surname> <given-names>M. I.</given-names></name> <name><surname>Mitchell</surname> <given-names>T. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Machine learning: trends, perspectives, and prospects</article-title>. <source>Science</source> <volume>349</volume>, <fpage>255</fpage>&#x02013;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaa8415</pub-id><pub-id pub-id-type="pmid">26185243</pub-id></citation></ref>
<ref id="B159">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joukes</surname> <given-names>J.</given-names></name> <name><surname>Hartmann</surname> <given-names>T. S.</given-names></name> <name><surname>Krekelberg</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>Motion detection based on recurrent network dynamics</article-title>. <source>Front. Syst. Neurosci.</source> <volume>8</volume>:<fpage>239</fpage>. <pub-id pub-id-type="doi">10.3389/fnsys.2014.00239</pub-id><pub-id pub-id-type="pmid">25565992</pub-id></citation></ref>
<ref id="B160">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kadmon</surname> <given-names>J.</given-names></name> <name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>Optimal architectures in a solvable model of deep networks,</article-title> in <source>Advances in Neural Information Processing Systems 29</source> (<publisher-loc>Barcelona</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B161">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kaiser</surname> <given-names>&#x00141;.</given-names></name> <name><surname>Roy</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Learning to remember rare events,</article-title> in <source>5th International Conference on Learning Representations</source> (<publisher-loc>Toulon</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B162">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kanitscheider</surname> <given-names>I.</given-names></name> <name><surname>Fiete</surname> <given-names>I.</given-names></name></person-group> (<year>2016</year>). <article-title>Training recurrent networks to generate hypotheses about how the brain solves hard navigation problems</article-title>. ArXiv:1609.09059, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B163">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaplan</surname> <given-names>F.</given-names></name> <name><surname>Oudeyer</surname> <given-names>P.-Y.</given-names></name></person-group> (<year>2004</year>). <article-title>Maximizing learning progress: an internal reward system for development</article-title>. <source>Embodied Artif. Intell.</source> <volume>3139</volume>, <fpage>259</fpage>&#x02013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1007/b99075</pub-id></citation></ref>
<ref id="B164">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kass</surname> <given-names>R.</given-names></name> <name><surname>Eden</surname> <given-names>U.</given-names></name> <name><surname>Brown</surname> <given-names>E.</given-names></name></person-group> (<year>2014</year>). <source>Analysis of Neural Data</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-1-4614-9602-1</pub-id></citation></ref>
<ref id="B165">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kawaguchi</surname> <given-names>K.</given-names></name> <name><surname>Kaelbling</surname> <given-names>L. P.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Generalization in deep learning</article-title>. ArXiv: 1710.045468v1, <fpage>1</fpage>&#x02013;<lpage>15</lpage>.</citation></ref>
<ref id="B166">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kemp</surname> <given-names>C.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2008</year>). <article-title>The discovery of structural form</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>105</volume>:<fpage>10687</fpage>. <pub-id pub-id-type="doi">10.1073/pnas.0802631105</pub-id><pub-id pub-id-type="pmid">18669663</pub-id></citation></ref>
<ref id="B167">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kempka</surname> <given-names>M.</given-names></name> <name><surname>Wydmuch</surname> <given-names>M.</given-names></name> <name><surname>Runc</surname> <given-names>G.</given-names></name> <name><surname>Toczek</surname> <given-names>J.</given-names></name> <name><surname>Ja</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>ViZDoom: a Doom-based AI research platform for visual reinforcement learning</article-title>. ArXiv:1605.02097v2, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B168">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kheradpisheh</surname> <given-names>S. R.</given-names></name> <name><surname>Ganjtabesh</surname> <given-names>M.</given-names></name> <name><surname>Thorpe</surname> <given-names>S. J.</given-names></name></person-group> (<year>2016</year>). <article-title>STDP-based spiking deep neural networks for object recognition</article-title>. ArXiv:1611.01421v1, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B169">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kietzmann</surname> <given-names>T. C.</given-names></name> <name><surname>McClure</surname> <given-names>P.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep neural networks in computational neuroscience</article-title>. BioRxiv, <fpage>1</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1101/133504</pub-id></citation></ref>
<ref id="B170">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kindermans</surname> <given-names>P.-J.</given-names></name> <name><surname>Sch&#x000FC;tt</surname> <given-names>K. T.</given-names></name> <name><surname>Alber</surname> <given-names>M.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>K.-R.</given-names></name> <name><surname>D&#x000E4;hne</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>PatternNet and PatternLRP &#x02013; improving the interpretability of neural networks</article-title>. ArXiv:1705.05598, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B171">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Auto-encoding variational Bayes</article-title>. ArXiv:1312.6114, <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B172">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kirkpatrick</surname> <given-names>J.</given-names></name> <name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Rabinowitz</surname> <given-names>N.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Desjardins</surname> <given-names>G.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name></person-group> (<year>2015</year>). <article-title>Overcoming catastrophic forgetting in neural networks</article-title>. ArXiv:1612.00796v1, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="pmid">28292907</pub-id></citation></ref>
<ref id="B173">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klyubin</surname> <given-names>A.</given-names></name> <name><surname>Polani</surname> <given-names>D.</given-names></name> <name><surname>Nehaniv</surname> <given-names>C.</given-names></name></person-group> (<year>2005a</year>). <article-title>Empowerment: a universal agent-centric measure of control</article-title>. <source>IEEE Congr. Evol. Comput.</source> <volume>1</volume>, <fpage>128</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1109/CEC.2005.1554676</pub-id></citation></ref>
<ref id="B174">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Klyubin</surname> <given-names>A. S.</given-names></name> <name><surname>Polani</surname> <given-names>D.</given-names></name> <name><surname>Nehaniv</surname> <given-names>C. L.</given-names></name></person-group> (<year>2005b</year>). <article-title>All else being equal be empowered,</article-title> in <source>Lecture Notes in Computer Science</source>, <volume>Vol. 3630</volume> (<publisher-loc>Canterbury</publisher-loc>), <fpage>744</fpage>&#x02013;<lpage>753</lpage>. <pub-id pub-id-type="doi">10.1007/11553090_75</pub-id></citation></ref>
<ref id="B175">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Koller</surname> <given-names>D.</given-names></name> <name><surname>Friedman</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <source>Probabilistic Graphical Models: Principles and Techniques</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation></ref>
<ref id="B176">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks: a new framework for modeling biological vision and brain information processing</article-title>. <source>Annu. Rev. Vis. Sci.</source> <volume>1</volume>, <fpage>417</fpage>&#x02013;<lpage>446</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-vision-082114-035447</pub-id><pub-id pub-id-type="pmid">28532370</pub-id></citation></ref>
<ref id="B177">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>ImageNet classification with deep convolutional neural networks,</article-title> in <source>Advances in Neural Information Processing Systems 25</source> (<publisher-loc>Lake Tahoe</publisher-loc>), <fpage>1106</fpage>&#x02013;<lpage>1114</lpage>.</citation></ref>
<ref id="B178">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kruschke</surname> <given-names>J. K.</given-names></name></person-group> (<year>1992</year>). <article-title>ALCOVE: an exemplar-based connectionist model of category learning</article-title>. <source>Psychol. Rev.</source> <volume>99</volume>, <fpage>22</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.99.1.22</pub-id><pub-id pub-id-type="pmid">1546117</pub-id></citation></ref>
<ref id="B179">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumaran</surname> <given-names>D.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<year>2016</year>). <article-title>What learning systems do intelligent agents need? Complementary learning systems theory updated</article-title>. <source>Trends Cogn. Sci.</source> <volume>20</volume>, <fpage>512</fpage>&#x02013;<lpage>534</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2016.05.004</pub-id><pub-id pub-id-type="pmid">27315762</pub-id></citation></ref>
<ref id="B180">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Laird</surname> <given-names>J. E.</given-names></name></person-group> (<year>2012</year>). <source>The SOAR Cognitive Architecture</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation></ref>
<ref id="B181">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laje</surname> <given-names>R.</given-names></name> <name><surname>Buonomano</surname> <given-names>D. V.</given-names></name></person-group> (<year>2013</year>). <article-title>Robust timing and motor patterns by taming chaos in recurrent neural networks</article-title>. <source>Nat. Neurosci.</source> <volume>16</volume>, <fpage>925</fpage>&#x02013;<lpage>933</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3405</pub-id><pub-id pub-id-type="pmid">23708144</pub-id></citation></ref>
<ref id="B182">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lake</surname> <given-names>B. M.</given-names></name> <name><surname>Ullman</surname> <given-names>T. D.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Building machines that learn and think like people</article-title>. <source>Behav. Brain Sci</source>. <pub-id pub-id-type="doi">10.1017/s0140525x16001837</pub-id><pub-id pub-id-type="pmid">27881212</pub-id></citation></ref>
<ref id="B183">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2010</year>). <article-title>Learning to combine foveal glimpses with a third-order Boltzmann machine,</article-title> in <source>Advances in Neural Information Processing Systems 23</source>, <volume>Vol. 40</volume> (<publisher-loc>Vancouver</publisher-loc>), <fpage>1243</fpage>&#x02013;<lpage>1251</lpage>.</citation></ref>
<ref id="B184">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laughlin</surname> <given-names>S. B.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Communication in neuronal networks</article-title>. <source>Science</source> <volume>301</volume>, <fpage>1870</fpage>&#x02013;<lpage>1874</lpage>. <pub-id pub-id-type="doi">10.1126/science.1089662</pub-id><pub-id pub-id-type="pmid">14512617</pub-id></citation></ref>
<ref id="B185">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le Roux</surname> <given-names>N.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <article-title>Deep belief networks are compact universal approximators</article-title>. <source>Neural Comput.</source> <volume>22</volume>, <fpage>2192</fpage>&#x02013;<lpage>2207</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2010.08-09-1081</pub-id></citation></ref>
<ref id="B186">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B187">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Haffner</surname> <given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x02013;<lpage>2324</lpage>. <pub-id pub-id-type="doi">10.1109/5.726791</pub-id></citation></ref>
<ref id="B188">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>J. H.</given-names></name> <name><surname>Delbruck</surname> <given-names>T.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Training deep spiking neural networks using backpropagation</article-title>. ArXiv:1608.08782, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="pmid">27877107</pub-id></citation></ref>
<ref id="B189">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>T.</given-names></name> <name><surname>Mumford</surname> <given-names>D.</given-names></name></person-group> (<year>2003</year>). <article-title>Hierarchical Bayesian inference in the visual cortex</article-title>. <source>J. Opt. Soc. Am. A</source> <volume>20</volume>, <fpage>1434</fpage>&#x02013;<lpage>1448</lpage>. <pub-id pub-id-type="doi">10.1364/josaa.20.001434</pub-id><pub-id pub-id-type="pmid">12868647</pub-id></citation></ref>
<ref id="B190">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lehky</surname> <given-names>S. R.</given-names></name> <name><surname>Tanaka</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Neural representation for object recognition in inferotemporal cortex</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>37</volume>, <fpage>23</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2015.12.001</pub-id><pub-id pub-id-type="pmid">26771242</pub-id></citation></ref>
<ref id="B191">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Anselmi</surname> <given-names>F.</given-names></name> <name><surname>Freiwald</surname> <given-names>W. A.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>View-tolerant face recognition and Hebbian learning imply mirror-symmetric neural tuning to head orientation</article-title>. <source>Curr. Biol.</source> <volume>27</volume>, <fpage>62</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2016.10.015</pub-id><pub-id pub-id-type="pmid">27916522</pub-id></citation></ref>
<ref id="B192">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Finn</surname> <given-names>C.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>End-to-end training of deep visuomotor policies</article-title>. ArXiv:1504.00702v1, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B193">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Bridging the gaps between residual learning, recurrent neural networks and visual cortex</article-title>. ArXiv:1604.03640, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B194">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lillicrap</surname> <given-names>T. P.</given-names></name> <name><surname>Cownden</surname> <given-names>D.</given-names></name> <name><surname>Tweed</surname> <given-names>D. B.</given-names></name> <name><surname>Akerman</surname> <given-names>C. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Random feedback weights support learning in deep neural networks</article-title>. <source>Nat. Commun.</source> <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1038/ncomms13276</pub-id></citation></ref>
<ref id="B195">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>H. W.</given-names></name> <name><surname>Tegmark</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Why does deep and cheap learning work so well?</article-title> ArXiv:1608.08225, <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B196">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lopez</surname> <given-names>C. M.</given-names></name> <name><surname>Mitra</surname> <given-names>S.</given-names></name> <name><surname>Putzeys</surname> <given-names>J.</given-names></name> <name><surname>Raducanu</surname> <given-names>B.</given-names></name> <name><surname>Ballini</surname> <given-names>M.</given-names></name> <name><surname>Andrei</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>A 966-electrode neural probe with 384 configurable channels in 0.13&#x003BC;m SOI CMOS,</article-title> in <source>Solid State Circuits Conference Dig Technical Papers</source> (<publisher-loc>San Francisco, CA</publisher-loc>), <fpage>21</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/ISSCC.2016.7418072</pub-id></citation></ref>
<ref id="B197">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lotter</surname> <given-names>W.</given-names></name> <name><surname>Kreiman</surname> <given-names>G.</given-names></name> <name><surname>Cox</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep predictive coding networks for video prediction and unsupervised learning</article-title>. ArXiv:1605.08104, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B198">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Louizos</surname> <given-names>C.</given-names></name> <name><surname>Shalit</surname> <given-names>U.</given-names></name> <name><surname>Mooij</surname> <given-names>J.</given-names></name> <name><surname>Sontag</surname> <given-names>D.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Causal effect inference with deep latent-variable models</article-title>. ArXiv:1705.08821, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B199">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>1997</year>). <article-title>Networks of spiking neurons: the third generation of neural network models</article-title>. <source>Neural Netw.</source> <volume>10</volume>, <fpage>1659</fpage>&#x02013;<lpage>1671</lpage>.</citation></ref>
<ref id="B200">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Searching for principles of brain computation</article-title>. BioRxiv, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B201">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>MacKay</surname> <given-names>D. J. C.</given-names></name></person-group> (<year>2003</year>). <source>Information Theory, Inference and Learning Algorithms</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1108/03684920410534506</pub-id></citation></ref>
<ref id="B202">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mandt</surname> <given-names>S.</given-names></name> <name><surname>Hoffman</surname> <given-names>M. D.</given-names></name> <name><surname>Blei</surname> <given-names>D. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Stochastic gradient descent as approximate Bayesian inference</article-title>. ArXiv:1704.04289v1, <fpage>1</fpage>&#x02013;<lpage>30</lpage>.</citation></ref>
<ref id="B203">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mante</surname> <given-names>V.</given-names></name> <name><surname>Sussillo</surname> <given-names>D.</given-names></name> <name><surname>Shenoy</surname> <given-names>K. V.</given-names></name> <name><surname>Newsome</surname> <given-names>W. T.</given-names></name></person-group> (<year>2013</year>). <article-title>Context-dependent computation by recurrent dynamics in prefrontal cortex</article-title>. <source>Nature</source> <volume>503</volume>, <fpage>78</fpage>&#x02013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1038/nature12742</pub-id><pub-id pub-id-type="pmid">24201281</pub-id></citation></ref>
<ref id="B204">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marblestone</surname> <given-names>A. H.</given-names></name> <name><surname>Wayne</surname> <given-names>G.</given-names></name> <name><surname>Kording</surname> <given-names>K. P.</given-names></name></person-group> (<year>2016</year>). <article-title>Towards an integration of deep learning and neuroscience</article-title>. <source>Front. Comput. Neurosci.</source> <volume>10</volume>:<fpage>94</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2016.00094</pub-id></citation></ref>
<ref id="B205">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcus</surname> <given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>How does the mind work? Insights from biology</article-title>. <source>Top. Cogn. Sci.</source> <volume>1</volume>, <fpage>145</fpage>&#x02013;<lpage>172</lpage>. <pub-id pub-id-type="doi">10.1111/j.1756-8765.2008.01007.x</pub-id><pub-id pub-id-type="pmid">19890489</pub-id></citation></ref>
<ref id="B206">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marder</surname> <given-names>E.</given-names></name></person-group> (<year>2015</year>). <article-title>Understanding brains: details, intuition, and big data</article-title>. <source>PLoS Biol.</source> <volume>13</volume>:<fpage>e1002147</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pbio.1002147</pub-id><pub-id pub-id-type="pmid">25965068</pub-id></citation></ref>
<ref id="B207">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markram</surname> <given-names>H.</given-names></name></person-group> (<year>2006</year>). <article-title>The blue brain project</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>7</volume>, <fpage>153</fpage>&#x02013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1848</pub-id><pub-id pub-id-type="pmid">16429124</pub-id></citation></ref>
<ref id="B208">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markram</surname> <given-names>H.</given-names></name> <name><surname>Meier</surname> <given-names>K.</given-names></name> <name><surname>Lippert</surname> <given-names>T.</given-names></name> <name><surname>Grillner</surname> <given-names>S.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Introducing the human brain project</article-title>. <source>Proc. Comput. Sci.</source> <volume>7</volume>, <fpage>39</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2011.12.015</pub-id></citation></ref>
<ref id="B209">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name></person-group> (<year>1969</year>). <article-title>A theory of cerebellar cortex</article-title>. <source>J. Physiol.</source> <volume>202</volume>, <fpage>437</fpage>&#x02013;<lpage>470</lpage>. <pub-id pub-id-type="doi">10.2307/1776957</pub-id><pub-id pub-id-type="pmid">5784296</pub-id></citation></ref>
<ref id="B210">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name></person-group> (<year>1982</year>). <source>Vision: A Computational Investigation into the Human Representation and Processing of Visual Information</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262514620.001.0001</pub-id></citation></ref>
<ref id="B211">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>1976</year>). <source>From Understanding Computation to Understanding Neural Circuitry</source>. Tech. Rep. <publisher-name>MIT</publisher-name>.</citation></ref>
<ref id="B212">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mathieu</surname> <given-names>M.</given-names></name> <name><surname>Couprie</surname> <given-names>C.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep multi-scale video prediction beyond mean square error,</article-title> in <source>4th International Conference on Learning Representations</source> (<publisher-loc>San Juan</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B213">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Maturana</surname> <given-names>H.</given-names></name> <name><surname>Varela</surname> <given-names>F.</given-names></name></person-group> (<year>1980</year>). <source>Autopoiesis and Cognition: The Realization of the Living, 1st Edn</source>. <publisher-loc>Dordrecht</publisher-loc>: <publisher-name>D. Reidel Publishing Company</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-94-009-8947-4</pub-id></citation></ref>
<ref id="B214">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Maturana</surname> <given-names>H.</given-names></name> <name><surname>Varela</surname> <given-names>F.</given-names></name></person-group> (<year>1987</year>). <source>The Tree of Knowledge - The Biological Roots of Human Understanding</source>. <publisher-loc>London</publisher-loc>: <publisher-name>New Science Library</publisher-name>.</citation></ref>
<ref id="B215">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<year>2003</year>). <article-title>The parallel distributed processing approach to semantic cognition</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>4</volume>, <fpage>310</fpage>&#x02013;<lpage>322</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1076</pub-id><pub-id pub-id-type="pmid">12671647</pub-id></citation></ref>
<ref id="B216">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McClelland</surname> <given-names>J. L.</given-names></name> <name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Noelle</surname> <given-names>D. C.</given-names></name> <name><surname>Plaut</surname> <given-names>D. C.</given-names></name> <name><surname>Rogers</surname> <given-names>T. T.</given-names></name> <name><surname>Seidenberg</surname> <given-names>M. S.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Letting structure emerge: connectionist and dynamical systems approaches to cognition</article-title>. <source>Trends Cogn. Sci.</source> <volume>14</volume>, <fpage>348</fpage>&#x02013;<lpage>356</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.06.002</pub-id><pub-id pub-id-type="pmid">20598626</pub-id></citation></ref>
<ref id="B217">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCloskey</surname> <given-names>M.</given-names></name> <name><surname>Cohen</surname> <given-names>N. J.</given-names></name></person-group> (<year>1989</year>). <article-title>Catastrophic inference in connectionist networks: the sequential learning problem</article-title>. <source>Psychol. Learn. Motiv.</source> <volume>24</volume>, <fpage>109</fpage>&#x02013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/s0079-7421(08)60536-8</pub-id></citation></ref>
<ref id="B218">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McCorduck</surname> <given-names>P.</given-names></name></person-group> (<year>2004</year>). <source>Machines Who Think, 2nd Edn</source>. <publisher-loc>Natick, MA</publisher-loc>: <publisher-name>A. K. Peters, Ltd</publisher-name>.</citation></ref>
<ref id="B219">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mcintosh</surname> <given-names>L. T.</given-names></name> <name><surname>Maheswaranathan</surname> <given-names>N.</given-names></name> <name><surname>Nayebi</surname> <given-names>A.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Baccus</surname> <given-names>S. A.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning models of the retinal response to natural scenes,</article-title> in <source>Advances in Neural Information Processing Systems 29</source> (<publisher-loc>Barcelona</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B220">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mead</surname> <given-names>C.</given-names></name></person-group> (<year>1990</year>). <article-title>Neuromorphic electronic systems</article-title>. <source>Proc. IEEE</source> <volume>78</volume>, <fpage>1629</fpage>&#x02013;<lpage>1636</lpage> <pub-id pub-id-type="doi">10.1109/5.58356</pub-id></citation></ref>
<ref id="B221">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mhaskar</surname> <given-names>H.</given-names></name> <name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Learning functions: when is deep better than shallow</article-title>. ArXiv:1603.00988v4, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B222">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miconi</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Biologically plausible learning in recurrent neural networks for flexible decision tasks</article-title>. <source>Elife</source> <volume>6</volume>:<fpage>e20899</fpage>. <pub-id pub-id-type="doi">10.16373/j.cnki.ahr.150049</pub-id></citation></ref>
<ref id="B223">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Efficient estimation of word representations in vector space,</article-title> in <source>1st International Conference on Learning Representations</source> (<publisher-loc>Scottsdale</publisher-loc>).</citation></ref>
<ref id="B224">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>E. K.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name></person-group> (<year>2001</year>). <article-title>An integrative theory of prefrontal cortex function</article-title>. <source>Annu. Rev. Neurosci.</source> <volume>24</volume>, <fpage>167</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.neuro.24.1.167</pub-id><pub-id pub-id-type="pmid">11283309</pub-id></citation></ref>
<ref id="B225">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Explanation in artificial intelligence: insights from the social sciences</article-title>. ArXiv:1706.07269v1, <fpage>1</fpage>&#x02013;<lpage>57</lpage>.</citation></ref>
<ref id="B226">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M.</given-names></name> <name><surname>Papert</surname> <given-names>S.</given-names></name></person-group> (<year>1969</year>). <source>Perceptrons. An Introduction to Computational Geometry</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B227">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Badia</surname> <given-names>A. P.</given-names></name> <name><surname>Mirza</surname> <given-names>M.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Lillicrap</surname> <given-names>T. P.</given-names></name> <name><surname>Harley</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Asynchronous methods for deep reinforcement learning</article-title>. ArXiv:1602.01783, <fpage>1</fpage>&#x02013;<lpage>28</lpage>.</citation></ref>
<ref id="B228">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Heess</surname> <given-names>N.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <article-title>Recurrent models of visual attention,</article-title> <source>Advances in Neural Information Processing Systems 27</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B229">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source> <volume>518</volume>, <fpage>529</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id></citation></ref>
<ref id="B230">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Modha</surname> <given-names>D. S.</given-names></name> <name><surname>Ananthanarayanan</surname> <given-names>R.</given-names></name> <name><surname>Esser</surname> <given-names>S. K.</given-names></name> <name><surname>Ndirango</surname> <given-names>A.</given-names></name> <name><surname>Sherbondy</surname> <given-names>A. J.</given-names></name> <name><surname>Singh</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Cognitive computing</article-title>. <source>Commun. ACM</source> <volume>54</volume>, <fpage>62</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1145/1978542.1978559</pub-id></citation></ref>
<ref id="B231">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moravec</surname> <given-names>H. P.</given-names></name></person-group> (<year>2000</year>). <source>Robot: Mere Machine to Transcendent Mind</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B232">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moser</surname> <given-names>M.-B.</given-names></name> <name><surname>Rowland</surname> <given-names>D. C.</given-names></name> <name><surname>Moser</surname> <given-names>E. I.</given-names></name></person-group> (<year>2015</year>). <article-title>Place cells, grid cells, and memory</article-title>. <source>Cold Spring Harb. Perspect. Biol.</source> <volume>7</volume>:<fpage>a021808</fpage> <pub-id pub-id-type="doi">10.1101/cshperspect.a021808</pub-id><pub-id pub-id-type="pmid">25646382</pub-id></citation></ref>
<ref id="B233">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moulton</surname> <given-names>S. T.</given-names></name> <name><surname>Kosslyn</surname> <given-names>S. M.</given-names></name></person-group> (<year>2009</year>). <article-title>Imagining predictions: mental imagery as mental emulation</article-title>. <source>Philos. Trans. R. Soc. B</source> <volume>364</volume>, <fpage>1273</fpage>&#x02013;<lpage>1280</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2008.0314</pub-id><pub-id pub-id-type="pmid">19528008</pub-id></citation></ref>
<ref id="B234">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mozer</surname> <given-names>M. C.</given-names></name></person-group> (<year>1989</year>). <article-title>A focused back-propagation algorithm for temporal pattern recognition</article-title>. <source>Complex Syst.</source> <volume>3</volume>, <fpage>349</fpage>&#x02013;<lpage>381</lpage>.</citation></ref>
<ref id="B235">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mozer</surname> <given-names>M. C.</given-names></name> <name><surname>Smolensky</surname> <given-names>P.</given-names></name></person-group> (<year>1989</year>). <article-title>Using relevance to reduce network size automatically</article-title>. <source>Conn. Sci.</source> <volume>1</volume>, <fpage>3</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1080/09540098908915626</pub-id></citation></ref>
<ref id="B236">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mujika</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Multi-task learning with deep model based reinforcement learning</article-title>. ArXiv:1611.01457, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B237">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Najemnik</surname> <given-names>J.</given-names></name> <name><surname>Geisler</surname> <given-names>W. S.</given-names></name></person-group> (<year>2005</year>). <article-title>Optimal eye movement strategies in visual search</article-title>. <source>Nature</source> <volume>434</volume>, <fpage>387</fpage>&#x02013;<lpage>391</lpage>. <pub-id pub-id-type="doi">10.1038/nature03390</pub-id><pub-id pub-id-type="pmid">15772663</pub-id></citation></ref>
<ref id="B238">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Nayebi</surname> <given-names>A.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Biologically inspired protection of deep networks from adversarial attacks</article-title>. ArXiv:1703.09202v1, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B239">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neftci</surname> <given-names>E.</given-names></name> <name><surname>Binas</surname> <given-names>J.</given-names></name> <name><surname>Rutishauser</surname> <given-names>U.</given-names></name> <name><surname>Chicca</surname> <given-names>E.</given-names></name> <name><surname>Indiveri</surname> <given-names>G.</given-names></name> <name><surname>Douglas</surname> <given-names>R. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Synthesizing cognition in neuromorphic electronic systems</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>110</volume>, <fpage>E3468</fpage>&#x02013;<lpage>E3476</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1212083110</pub-id><pub-id pub-id-type="pmid">23878215</pub-id></citation></ref>
<ref id="B240">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Neil</surname> <given-names>D.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>S.-C.</given-names></name></person-group> (<year>2016</year>). <article-title>Phased LSTM: accelerating recurrent network training for long or event-based sequences</article-title>. ArXiv:1610.09513v1, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="pmid">19834158</pub-id></citation></ref>
<ref id="B241">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name> <name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Bayesian computation emerges in generic cortical microcircuits through spike-timing-dependent plasticity</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003037</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003037</pub-id><pub-id pub-id-type="pmid">23633941</pub-id></citation></ref>
<ref id="B242">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Newell</surname> <given-names>A.</given-names></name></person-group> (<year>1991</year>). <source>Unified Theories of Cognition</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Harvard University Press</publisher-name>. <pub-id pub-id-type="pmid">24924001</pub-id></citation></ref>
<ref id="B243">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newell</surname> <given-names>A.</given-names></name> <name><surname>Simon</surname> <given-names>H. A.</given-names></name></person-group> (<year>1976</year>). <article-title>Computer science as empirical inquiry: symbols and search</article-title>. <source>Commun. ACM</source> <volume>19</volume>, <fpage>113</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1145/360018.360022</pub-id></citation></ref>
<ref id="B244">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>A.</given-names></name> <name><surname>Dosovitskiy</surname> <given-names>A.</given-names></name> <name><surname>Yosinski</surname> <given-names>J.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name> <name><surname>Clune</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Synthesizing the preferred inputs for neurons in neural networks via deep generator networks</article-title>. ArXiv:1605.09304, <fpage>1</fpage>&#x02013;<lpage>29</lpage>.</citation></ref>
<ref id="B245">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nilsson</surname> <given-names>N.</given-names></name></person-group> (<year>2005</year>). <article-title>Human-level artificial intelligence? Be serious!</article-title> <source>AI Mag.</source> <volume>26</volume>, <fpage>68</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1609/aimag.v26i4.1850</pub-id></citation></ref>
<ref id="B246">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Obermayer</surname> <given-names>K.</given-names></name></person-group> (<year>1990</year>). <article-title>A principle for the formation of the spatial structure of cortical feature maps</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>87</volume>, <fpage>8345</fpage>&#x02013;<lpage>8349</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.87.21.8345</pub-id><pub-id pub-id-type="pmid">2236045</pub-id></citation></ref>
<ref id="B247">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>O&#x00027;Connor</surname> <given-names>P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep spiking networks</article-title>. ArXiv:1602.08323, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B248">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B. A.</given-names></name> <name><surname>Field</surname> <given-names>D. J.</given-names></name></person-group> (<year>1996</year>). <article-title>Emergence of simple-cell receptive field properties by learning a sparse code for natural images</article-title>. <source>Nature</source> <volume>381</volume>, <fpage>607</fpage>&#x02013;<lpage>609</lpage>. <pub-id pub-id-type="doi">10.1038/381607a0</pub-id><pub-id pub-id-type="pmid">8637596</pub-id></citation></ref>
<ref id="B249">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R.</given-names></name></person-group> (<year>1998</year>). <article-title>Six principles for biologically based computational models of cortical cognition</article-title>. <source>Trends Cogn. Sci.</source> <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1016/s1364-6613(98)01241-8</pub-id><pub-id pub-id-type="pmid">21227277</pub-id></citation></ref>
<ref id="B250">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R.</given-names></name> <name><surname>Hazy</surname> <given-names>T.</given-names></name> <name><surname>Herd</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>The Leabra cognitive architecture: how to play 20 principles with nature and win!,</article-title> in <source>The Oxford Handbook of Cognitive Science</source>, ed <person-group person-group-type="editor"><name><surname>Chipman</surname> <given-names>S. E. F.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordhb/9780199842193.013.8</pub-id></citation></ref>
<ref id="B251">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Wyatte</surname> <given-names>D.</given-names></name> <name><surname>Herd</surname> <given-names>S.</given-names></name> <name><surname>Mingus</surname> <given-names>B.</given-names></name> <name><surname>Jilk</surname> <given-names>D. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Recurrent processing during object recognition</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>124</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00124</pub-id><pub-id pub-id-type="pmid">23554596</pub-id></citation></ref>
<ref id="B252">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Orhan</surname> <given-names>A. E.</given-names></name> <name><surname>Ma</surname> <given-names>W. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Probabilistic inference in generic neural networks trained with non-probabilistic feedback</article-title>. ArXiv:1601.03060v4, <fpage>1</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="pmid">28743932</pub-id></citation></ref>
<ref id="B253">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oudeyer</surname> <given-names>P.-Y.</given-names></name></person-group> (<year>2007</year>). <article-title>Intrinsically motivated machines,</article-title> in <source>Lecture Notes Computer Science</source>, <volume>Vol. 4850</volume>, eds <person-group person-group-type="editor"><name><surname>Lungarella</surname> <given-names>M.</given-names></name> <name><surname>Iida</surname> <given-names>F.</given-names></name> <name><surname>Bongard</surname> <given-names>J.</given-names></name> <name><surname>Pfeifer</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>304</fpage>&#x02013;<lpage>314</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-77296-5_27</pub-id></citation></ref>
<ref id="B254">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pachitariu</surname> <given-names>M.</given-names></name> <name><surname>Stringer</surname> <given-names>C.</given-names></name> <name><surname>Schr&#x000F6;der</surname> <given-names>S.</given-names></name> <name><surname>Dipoppa</surname> <given-names>M.</given-names></name> <name><surname>Rossi</surname> <given-names>L. F.</given-names></name> <name><surname>Carandini</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Suite2p: beyond 10,000 neurons with standard two-photon microscopy</article-title>. BioRxiv, <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B255">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pakkenberg</surname> <given-names>B.</given-names></name> <name><surname>Gundersen</surname> <given-names>H.</given-names></name></person-group> (<year>1997</year>). <article-title>Neocortical neuron number in humans: effect of sex and age</article-title>. <source>J. Comp. Neurol.</source> <volume>384</volume>, <fpage>312</fpage>&#x02013;<lpage>320</lpage>. <pub-id pub-id-type="pmid">9215725</pub-id></citation></ref>
<ref id="B256">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pakkenberg</surname> <given-names>B.</given-names></name> <name><surname>Pelvig</surname> <given-names>D.</given-names></name> <name><surname>Marner</surname> <given-names>L.</given-names></name> <name><surname>Bundgaard</surname> <given-names>M.</given-names></name> <name><surname>Gundersen</surname> <given-names>H.</given-names></name> <name><surname>Nyengaard</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Aging and the human neocortex</article-title>. <source>Exp. Gerontol.</source> <volume>38</volume>, <fpage>95</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/s0531-5565(02)00151-1</pub-id><pub-id pub-id-type="pmid">12543266</pub-id></citation></ref>
<ref id="B257">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Palatucci</surname> <given-names>M.</given-names></name> <name><surname>Pomerleau</surname> <given-names>D.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Mitchell</surname> <given-names>T.</given-names></name></person-group> (<year>2009</year>). <article-title>Zero-shot learning with semantic output codes,</article-title> in <source>Advances in Neural Information Processing Systems 22</source>, eds <person-group person-group-type="editor"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Schuurmans</surname> <given-names>D.</given-names></name> <name><surname>Lafferty</surname> <given-names>J.</given-names></name> <name><surname>Williams</surname> <given-names>C. K. I.</given-names></name> <name><surname>Culotta</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Vancouver</publisher-loc>), <fpage>1410</fpage>&#x02013;<lpage>1418</lpage>.</citation></ref>
<ref id="B258">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Fellow</surname> <given-names>Q. Y.</given-names></name></person-group> (<year>2009</year>). <article-title>A survey on transfer learning</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>22</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id></citation></ref>
<ref id="B259">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>On the difficulty of training recurrent neural networks,</article-title> in <source>Proceedings of the 30th International Conference on Machine Learning</source> (<publisher-loc>Atlanta</publisher-loc>), <fpage>1310</fpage>&#x02013;<lpage>1318</lpage>.</citation></ref>
<ref id="B260">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Montufar</surname> <given-names>G.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>On the number of response regions of deep feed forward networks with piece-wise linear activations</article-title>. ArXiv:1312.6098, <fpage>1</fpage>&#x02013;<lpage>17</lpage>.</citation></ref>
<ref id="B261">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Pathak</surname> <given-names>D.</given-names></name> <name><surname>Agrawal</surname> <given-names>P.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Curiosity-driven exploration by self-supervised prediction</article-title>. ArXiv:1705.05363, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B262">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peelen</surname> <given-names>M. V.</given-names></name> <name><surname>Downing</surname> <given-names>P. E.</given-names></name></person-group> (<year>2017</year>). <article-title>Category selectivity in human visual cortex: beyond visual object recognition</article-title>. <source>Neuropsychologia</source> <volume>105</volume>, <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuropsychologia.2017.03.033</pub-id><pub-id pub-id-type="pmid">28377161</pub-id></citation></ref>
<ref id="B263">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Perunov</surname> <given-names>N.</given-names></name> <name><surname>Marsland</surname> <given-names>R.</given-names></name> <name><surname>England</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Statistical physics of adaptation</article-title>. ArXiv:1412.1875, <fpage>1</fpage>&#x02013;<lpage>24</lpage>.</citation></ref>
<ref id="B264">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Peterson</surname> <given-names>J. C.</given-names></name> <name><surname>Abbott</surname> <given-names>J. T.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. L.</given-names></name></person-group> (<year>2016</year>). <article-title>Adapting deep network features to capture psychological representations</article-title>. ArXiv:1608.02164, <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation></ref>
<ref id="B265">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Pinker</surname> <given-names>S.</given-names></name> <name><surname>Mehler</surname> <given-names>J.</given-names></name></person-group> (eds.). (<year>1988</year>). <source>Connections and Symbols</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation>
</ref>
<ref id="B266">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>The levels of understanding framework, revised <italic>Perception</italic></article-title> <volume>41</volume>, <fpage>1017</fpage>&#x02013;<lpage>1023</lpage>. <pub-id pub-id-type="doi">10.1068/p7299</pub-id></citation></ref>
<ref id="B267">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Poole</surname> <given-names>B.</given-names></name> <name><surname>Lahiri</surname> <given-names>S.</given-names></name> <name><surname>Raghu</surname> <given-names>M.</given-names></name> <name><surname>Sohl-Dickstein</surname> <given-names>J.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Exponential expressivity in deep neural networks through transient chaos</article-title>. ArXiv:1606.05340, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B268">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pouget</surname> <given-names>A.</given-names></name> <name><surname>Beck</surname> <given-names>J. M.</given-names></name> <name><surname>Ma</surname> <given-names>W. J.</given-names></name> <name><surname>Latham</surname> <given-names>P. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Probabilistic brains: knowns and unknowns</article-title>. <source>Nat. Neurosci.</source> <volume>16</volume>, <fpage>1170</fpage>&#x02013;<lpage>1178</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3495</pub-id><pub-id pub-id-type="pmid">23955561</pub-id></citation></ref>
<ref id="B269">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Pritzel</surname> <given-names>A.</given-names></name> <name><surname>Uria</surname> <given-names>B.</given-names></name> <name><surname>Srinivasan</surname> <given-names>S.</given-names></name> <name><surname>Puigdom&#x000E8;nech</surname> <given-names>A.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Neural episodic control</article-title>. ArXiv:1703.01988, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B270">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quian Quiroga</surname> <given-names>R.</given-names></name> <name><surname>Reddy</surname> <given-names>L.</given-names></name> <name><surname>Kreiman</surname> <given-names>G.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name> <name><surname>Fried</surname> <given-names>I.</given-names></name></person-group> (<year>2005</year>). <article-title>Invariant visual representation by single neurons in the human brain</article-title>. <source>Nature</source> <volume>435</volume>, <fpage>1102</fpage>&#x02013;<lpage>1107</lpage>. <pub-id pub-id-type="doi">10.1038/nature03687</pub-id><pub-id pub-id-type="pmid">15973409</pub-id></citation></ref>
<ref id="B271">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Rafler</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Generalization of Conway&#x00027;s &#x0201C;Game of Life&#x0201D; to a continuous domain - SmoothLife</article-title>. ArXiv:1111.1567v2, <fpage>1</fpage>&#x02013;<lpage>4</lpage>.</citation></ref>
<ref id="B272">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Raghu</surname> <given-names>M.</given-names></name> <name><surname>Kleinberg</surname> <given-names>J.</given-names></name> <name><surname>Poole</surname> <given-names>B.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Sohl-Dickstein</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Survey of expressivity in deep neural networks</article-title>. ArXiv:1611.08083v1, <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation></ref>
<ref id="B273">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Raina</surname> <given-names>R.</given-names></name> <name><surname>Madhavan</surname> <given-names>A.</given-names></name> <name><surname>Ng</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Large-scale deep unsupervised learning using graphics processors,</article-title> in <source>Proceedings of the 26th Annual International Conference on Machine Learning</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1145/1553374.1553486</pub-id></citation></ref>
<ref id="B274">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajan</surname> <given-names>K.</given-names></name> <name><surname>Harvey</surname> <given-names>C. D.</given-names></name> <name><surname>Tank</surname> <given-names>D. W.</given-names></name></person-group> (<year>2015</year>). <article-title>Recurrent network models of sequence generation and memory</article-title>. <source>Neuron</source> <volume>90</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.02.009</pub-id><pub-id pub-id-type="pmid">26971945</pub-id></citation></ref>
<ref id="B275">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ramsey</surname> <given-names>F. P.</given-names></name></person-group> (<year>1926</year>). <article-title>Truth and probability,</article-title> in <source>The Foundations of Mathematics and other Logical Essays</source>, ed <person-group person-group-type="editor"><name><surname>Braithwaite</surname> <given-names>R. B.</given-names></name></person-group> (<publisher-loc>Abingdon</publisher-loc>: <publisher-name>Routledge</publisher-name>), <fpage>156</fpage>&#x02013;<lpage>198</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-20451-2_3</pub-id></citation></ref>
<ref id="B276">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>R. P.</given-names></name> <name><surname>Ballard</surname> <given-names>D. H.</given-names></name></person-group> (<year>1999</year>). <article-title>Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects</article-title>. <source>Nat. Neurosci.</source> <volume>2</volume>, <fpage>79</fpage>&#x02013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1038/4580</pub-id><pub-id pub-id-type="pmid">10195184</pub-id></citation></ref>
<ref id="B277">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Real</surname> <given-names>E.</given-names></name> <name><surname>Moore</surname> <given-names>S.</given-names></name> <name><surname>Selle</surname> <given-names>A.</given-names></name> <name><surname>Saxena</surname> <given-names>S.</given-names></name> <name><surname>Suematsu</surname> <given-names>Y. L.</given-names></name> <name><surname>Le</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Large-scale evolution of image classifiers</article-title>. ArXiv:1703.01041v1, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B278">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Regan</surname> <given-names>J. K. O.</given-names></name> <name><surname>No&#x000EB;</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>A sensorimotor account of vision and visual consciousness</article-title>. <source>Behav. Brain Sci.</source> <volume>24</volume>, <fpage>939</fpage>&#x02013;<lpage>1031</lpage>. <pub-id pub-id-type="doi">10.1017/s0140525x01000115</pub-id></citation></ref>
<ref id="B279">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rid</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>Rise of the Machines: A Cybernetic History</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>W. W. Norton &#x00026; Company</publisher-name>.</citation></ref>
<ref id="B280">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riesenhuber</surname> <given-names>M.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>1999</year>). <article-title>Hierarchical models of object recognition in cortex</article-title>. <source>Nat. Neurosci.</source> <volume>2</volume>, <fpage>1019</fpage>&#x02013;<lpage>1025</lpage>. <pub-id pub-id-type="pmid">10526343</pub-id></citation></ref>
<ref id="B281">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ritter</surname> <given-names>H.</given-names></name> <name><surname>Kohonen</surname> <given-names>T.</given-names></name></person-group> (<year>1989</year>). <article-title>Self-organizing semantic maps</article-title>. <source>Biol. Cybern.</source> <volume>61</volume>, <fpage>241</fpage>&#x02013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1007/bf00203171</pub-id></citation></ref>
<ref id="B282">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robinson</surname> <given-names>L.</given-names></name> <name><surname>Rolls</surname> <given-names>E. T.</given-names></name></person-group> (<year>2015</year>). <article-title>Invariant visual object recognition: biologically plausible approaches</article-title>. <source>Biol. Cybern.</source> <volume>209</volume>, <fpage>505</fpage>&#x02013;<lpage>535</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-015-0658-2</pub-id></citation></ref>
<ref id="B283">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name> <name><surname>van Ooyen</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Attention-gated reinforcement learning of internal representations for classification</article-title>. <source>Neural Comput.</source> <volume>17</volume>, <fpage>2176</fpage>&#x02013;<lpage>2214</lpage>. <pub-id pub-id-type="doi">10.1162/0899766054615699</pub-id><pub-id pub-id-type="pmid">16105222</pub-id></citation></ref>
<ref id="B284">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosenblatt</surname> <given-names>F.</given-names></name></person-group> (<year>1958</year>). <article-title>The perceptron: a probabilistic model for information storage and organization in the brain</article-title>. <source>Psychol. Rev.</source> <volume>65</volume>, <fpage>386</fpage>&#x02013;<lpage>408</lpage>. <pub-id pub-id-type="doi">10.1037/h0042519</pub-id><pub-id pub-id-type="pmid">13602029</pub-id></citation></ref>
<ref id="B285">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Williams</surname> <given-names>R.</given-names></name></person-group> (<year>1986</year>). <article-title>Learning internal representations by error propagation,</article-title> in <source>Parallel Distributed Processing, Explorations in the Microstructure of Cognition</source>, eds <person-group person-group-type="editor"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>318</fpage>&#x02013;<lpage>362</lpage>.</citation></ref>
<ref id="B286">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Salge</surname> <given-names>C.</given-names></name> <name><surname>Glackin</surname> <given-names>C.</given-names></name> <name><surname>Polani</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>Empowerment - An introduction</article-title>. ArXiv:1310.1863, <fpage>1</fpage>&#x02013;<lpage>46</lpage>.</citation></ref>
<ref id="B287">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Salimans</surname> <given-names>T.</given-names></name> <name><surname>Ho</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). <article-title>Evolution strategies as a scalable alternative to reinforcement learning</article-title>. ArXiv:1703.03864v2, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B288">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Santana</surname> <given-names>E.</given-names></name> <name><surname>Hotz</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Learning a driving simulator</article-title>. ArXiv:1608.01230, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B289">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Santoro</surname> <given-names>A.</given-names></name> <name><surname>Bartunov</surname> <given-names>S.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name> <name><surname>Lillicrap</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>One-shot learning with memory-augmented neural networks</article-title>. ArXiv:1605.06065v1, <fpage>1</fpage>&#x02013;<lpage>13</lpage>.</citation></ref>
<ref id="B290">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Santoro</surname> <given-names>A.</given-names></name> <name><surname>Raposo</surname> <given-names>D.</given-names></name> <name><surname>Barrett</surname> <given-names>D. G. T.</given-names></name> <name><surname>Malinowski</surname> <given-names>M.</given-names></name> <name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Battaglia</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A simple neural network module for relational reasoning</article-title>. ArXiv:1706.01427v1, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B291">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Saxe</surname> <given-names>A.</given-names></name> <name><surname>McClelland</surname> <given-names>J.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Exact solutions to the nonlinear dynamics of learning in deep linear neural networks Andrew,</article-title> in <source>2nd International Conference on Learning Representations</source> (<publisher-loc>Banff</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>22</lpage>.</citation></ref>
<ref id="B292">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scellier</surname> <given-names>B.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Equilibrium propagation: bridging the gap between energy-based models and backpropagation</article-title>. <source>Front. Comput. Neurosci.</source> <volume>11</volume>:<fpage>24</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2017.00024</pub-id><pub-id pub-id-type="pmid">28522969</pub-id></citation></ref>
<ref id="B293">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schaal</surname> <given-names>S.</given-names></name></person-group> (<year>1999</year>). <article-title>Is imitation learning the route to humanoid robots?</article-title> <source>Trends Cogn. Sci.</source> <volume>3</volume>, <fpage>233</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1016/s1364-6613(99)01327-3</pub-id><pub-id pub-id-type="pmid">10354577</pub-id></citation></ref>
<ref id="B294">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schacter</surname> <given-names>D. L.</given-names></name> <name><surname>Addis</surname> <given-names>D. R.</given-names></name> <name><surname>Buckner</surname> <given-names>R. L.</given-names></name></person-group> (<year>2007</year>). <article-title>Remembering the past to imagine the future: the prospective brain</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>8</volume>, <fpage>657</fpage>&#x02013;<lpage>661</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2213</pub-id><pub-id pub-id-type="pmid">17700624</pub-id></citation></ref>
<ref id="B295">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schiess</surname> <given-names>M.</given-names></name> <name><surname>Urbanczik</surname> <given-names>R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Somato-dendritic synaptic plasticity and error-backpropagation in active dendrites</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1004638</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004638</pub-id><pub-id pub-id-type="pmid">26841235</pub-id></citation></ref>
<ref id="B296">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1991</year>). <article-title>Curious model-building control systems,</article-title> in <source>Proceedings of International Joint Conference on Neural Networks</source>, <volume>Vol. 2</volume> (<publisher-loc>Singapore</publisher-loc>), <fpage>1458</fpage>&#x02013;<lpage>1463</lpage>. <pub-id pub-id-type="doi">10.1109/IJCNN.1991.170605</pub-id></citation></ref>
<ref id="B297">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>Exploring the predictable,</article-title> in <source>Advances in Evolutionary Computing</source>, eds <person-group person-group-type="editor"><name><surname>Ghosh</surname> <given-names>A.</given-names></name> <name><surname>Tsutsui</surname> <given-names>S.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>579</fpage>&#x02013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1017/CBO9781107415324.004</pub-id></citation></ref>
<ref id="B298">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>On learning to think: algorithmic information theory for novel combinations of reinforcement learning controllers and recurrent neural world models</article-title>. ArXiv:1511.09249, <fpage>1</fpage>&#x02013;<lpage>36</lpage>.</citation></ref>
<ref id="B299">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schoenholz</surname> <given-names>S. S.</given-names></name> <name><surname>Gilmer</surname> <given-names>J.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Sohl-Dickstein</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep information propagation,</article-title> in <source>5th International Conference on Learning Representations</source> (<publisher-loc>Toulon</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>18</lpage>.</citation></ref>
<ref id="B300">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schoenmakers</surname> <given-names>S.</given-names></name> <name><surname>Barth</surname> <given-names>M.</given-names></name> <name><surname>Heskes</surname> <given-names>T.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Linear reconstruction of perceived images from human brain activity</article-title>. <source>Neuroimage</source> <volume>83</volume>, <fpage>951</fpage>&#x02013;<lpage>961</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.07.043</pub-id><pub-id pub-id-type="pmid">23886984</pub-id></citation></ref>
<ref id="B301">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Scholte</surname> <given-names>H. S.</given-names></name> <name><surname>Losch</surname> <given-names>M. M.</given-names></name> <name><surname>Ramakrishnan</surname> <given-names>K.</given-names></name> <name><surname>de Haan</surname> <given-names>E. H. F.</given-names></name> <name><surname>Bohte</surname> <given-names>S. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Visual pathways from the perspective of cost functions and deep learning</article-title>. BioRxiv, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B302">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schroeder</surname> <given-names>C. E.</given-names></name> <name><surname>Wilson</surname> <given-names>D. A.</given-names></name> <name><surname>Radman</surname> <given-names>T.</given-names></name> <name><surname>Scharfman</surname> <given-names>H.</given-names></name> <name><surname>Lakatos</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Dynamics of active sensing and perceptual selection</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>20</volume>, <fpage>172</fpage>&#x02013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2010.02.010</pub-id><pub-id pub-id-type="pmid">20307966</pub-id></citation></ref>
<ref id="B303">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Schulman</surname> <given-names>J.</given-names></name> <name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Moritz</surname> <given-names>P.</given-names></name> <name><surname>Jordan</surname> <given-names>M.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Trust region policy optimization</article-title>. ArXiv:1502.05477v4, <fpage>1</fpage>&#x02013;<lpage>16</lpage>.</citation></ref>
<ref id="B304">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schultz</surname> <given-names>W.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Montague</surname> <given-names>P. R.</given-names></name></person-group> (<year>1997</year>). <article-title>A neural substrate of prediction and reward</article-title>. <source>Science</source> <volume>275</volume>, <fpage>1593</fpage>&#x02013;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1126/science.275.5306.1593</pub-id><pub-id pub-id-type="pmid">9054347</pub-id></citation></ref>
<ref id="B305">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Schuman</surname> <given-names>C. D.</given-names></name> <name><surname>Potok</surname> <given-names>T. E.</given-names></name> <name><surname>Patton</surname> <given-names>R. M.</given-names></name> <name><surname>Birdwell</surname> <given-names>J. D.</given-names></name> <name><surname>Dean</surname> <given-names>M. E.</given-names></name> <name><surname>Rose</surname> <given-names>G. S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A survey of neuromorphic computing and neural networks in hardware</article-title>. ArXiv:1705.06963, <fpage>1</fpage>&#x02013;<lpage>88</lpage>.</citation></ref>
<ref id="B306">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Searle</surname> <given-names>J. R.</given-names></name></person-group> (<year>1980</year>). <article-title>Minds, brains and Programs</article-title>. <source>Behav. Brain Sci.</source> <volume>3</volume>, <fpage>417</fpage>&#x02013;<lpage>424</lpage>. <pub-id pub-id-type="doi">10.1017/s0140525x00005756</pub-id></citation></ref>
<ref id="B307">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Segundo</surname> <given-names>J. P.</given-names></name> <name><surname>Perkel</surname> <given-names>D. H.</given-names></name> <name><surname>Moore</surname> <given-names>G. P.</given-names></name></person-group> (<year>1966</year>). <article-title>Spike probability in neurones: influence of temporal structure in the train of synaptic events</article-title>. <source>Kybernetik</source> <volume>3</volume>, <fpage>67</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1007/BF00299899</pub-id><pub-id pub-id-type="pmid">6003993</pub-id></citation></ref>
<ref id="B308">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seising</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Marvin Lee Minsky (1927-2016)</article-title>. <source>Artif. Intell. Med.</source> <volume>75</volume>, <fpage>24</fpage>&#x02013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2016.12.001</pub-id></citation></ref>
<ref id="B309">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Selfridge</surname> <given-names>O.</given-names></name></person-group> (<year>1959</year>). <article-title>Pandemonium: a paradigm for learning,</article-title> in <source>Symposium on the Mechanization of Thought Processes</source> (<publisher-loc>Teddington</publisher-loc>), <fpage>513</fpage>&#x02013;<lpage>526</lpage>.</citation></ref>
<ref id="B310">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Shwartz-Ziv</surname> <given-names>R.</given-names></name> <name><surname>Tishby</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>Opening the black box of deep neural networks via information</article-title>. ArXiv:1703.00810, <fpage>1</fpage>&#x02013;<lpage>19</lpage>.</citation></ref>
<ref id="B311">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Lever</surname> <given-names>G.</given-names></name> <name><surname>Heess</surname> <given-names>N.</given-names></name> <name><surname>Degris</surname> <given-names>T.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name> <name><surname>Riedmiller</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Deterministic policy gradient algorithms,</article-title> in <source>2nd International Conference on Learning Representations</source> (<publisher-loc>Banff</publisher-loc>), <fpage>387</fpage>&#x02013;<lpage>395</lpage>.</citation></ref>
<ref id="B312">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Schrittwieser</surname> <given-names>J.</given-names></name> <name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Antonoglou</surname> <given-names>I.</given-names></name> <name><surname>Huang</surname> <given-names>A.</given-names></name> <name><surname>Guez</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Mastering the game of Go without human knowledge</article-title>. <source>Nature</source> <volume>550</volume>, <fpage>354</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1038/nature24270</pub-id><pub-id pub-id-type="pmid">29052630</pub-id></citation></ref>
<ref id="B313">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simon</surname> <given-names>H. A.</given-names></name></person-group> (<year>1962</year>). <article-title>The architecture of complexity</article-title>. <source>Proc. Am. Philos. Soc.</source> <volume>106</volume>, <fpage>467</fpage>&#x02013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4899-0718-9_31</pub-id></citation></ref>
<ref id="B314">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Simon</surname> <given-names>H. A.</given-names></name></person-group> (<year>1996</year>). <source>The Sciences of the Artificial, 3rd Edn</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation></ref>
<ref id="B315">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singer</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Cortical dynamics revisited</article-title>. <source>Trends Cogn. Sci.</source> <volume>17</volume>, <fpage>616</fpage>&#x02013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2013.09.006</pub-id><pub-id pub-id-type="pmid">24139950</pub-id></citation></ref>
<ref id="B316">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smolensky</surname> <given-names>P.</given-names></name></person-group> (<year>1987</year>). <article-title>Connectionist AI, symbolic AI, and the brain</article-title>. <source>Artif. Intell. Rev.</source> <volume>1</volume>, <fpage>95</fpage>&#x02013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1007/BF00130011</pub-id></citation></ref>
<ref id="B317">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>1988</year>). <article-title>Statistical mechanics of neural networks</article-title>. <source>Phys. Today</source> <volume>40</volume>, <fpage>70</fpage>&#x02013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1063/1.881142</pub-id></citation></ref>
<ref id="B318">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Computational neuroscience: beyond the local circuit</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>25</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2014.02.002</pub-id><pub-id pub-id-type="pmid">24602868</pub-id></citation></ref>
<ref id="B319">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>H. F.</given-names></name> <name><surname>Yang</surname> <given-names>G. R.</given-names></name> <name><surname>Wang</surname> <given-names>X.-J.</given-names></name></person-group> (<year>2016</year>). <article-title>Reward-based training of recurrent neural networks for diverse cognitive and value-based tasks</article-title>. <source>Elife</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1101/070375</pub-id></citation></ref>
<ref id="B320">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sperry</surname> <given-names>R. W.</given-names></name></person-group> (<year>1952</year>). <article-title>Neurology and the mind-brain problem</article-title>. <source>Am. Sci.</source> <volume>40</volume>, <fpage>291</fpage>&#x02013;<lpage>312</lpage>.</citation></ref>
<ref id="B321">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srivastava</surname> <given-names>N.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>. <source>J. Mach. Learn. Res.</source> <volume>15</volume>, <fpage>1929</fpage>&#x02013;<lpage>1958</lpage>. <pub-id pub-id-type="doi">10.1214/12-AOS1000</pub-id></citation></ref>
<ref id="B322">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stanley</surname> <given-names>K.</given-names></name> <name><surname>Miikkulainen</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>Evolving neural networks through augmenting topologies</article-title>. <source>Evol. Comput.</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1162/106365602320169811</pub-id><pub-id pub-id-type="pmid">12180173</pub-id></citation></ref>
<ref id="B323">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steels</surname> <given-names>L.</given-names></name></person-group> (<year>1993</year>). <article-title>The artificial life roots of artificial intelligence</article-title>. <source>Artif. Life</source> <volume>1</volume>, <fpage>75</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1162/artl.1993.1.1_2.75</pub-id></citation></ref>
<ref id="B324">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Steels</surname> <given-names>L.</given-names></name></person-group> (<year>2004</year>). <article-title>The autotelic principle,</article-title> in <source>Embodied Artificial Intelligence. Lecture Notes in Computer Science</source>, eds <person-group person-group-type="editor"><name><surname>Iida</surname> <given-names>F.</given-names></name> <name><surname>Pfeifer</surname> <given-names>R.</given-names></name> <name><surname>Steels</surname> <given-names>L.</given-names></name> <name><surname>Kuniyoshi</surname> <given-names>Y.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>231</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-27833-7_17</pub-id></citation></ref>
<ref id="B325">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sterling</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <article-title>Allostasis: a model of predictive regulation</article-title>. <source>Physiol. Behav.</source> <volume>106</volume>, <fpage>5</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1016/j.physbeh.2011.06.004</pub-id><pub-id pub-id-type="pmid">21684297</pub-id></citation></ref>
<ref id="B326">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sterling</surname> <given-names>P.</given-names></name> <name><surname>Laughlin</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <source>Principles of Neural Design</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262028707.001.0001</pub-id></citation></ref>
<ref id="B327">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Strukov</surname> <given-names>D. B.</given-names></name></person-group> (<year>2011</year>). <article-title>Smart connections</article-title>. <source>Nature</source> <volume>476</volume>, <fpage>403</fpage>&#x02013;<lpage>405</lpage>. <pub-id pub-id-type="doi">10.1038/476403a</pub-id><pub-id pub-id-type="pmid">21866148</pub-id></citation></ref>
<ref id="B328">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Summerfield</surname> <given-names>C.</given-names></name> <name><surname>de Lange</surname> <given-names>F. P.</given-names></name></person-group> (<year>2014</year>). <article-title>Expectation in perceptual decision making: neural and computational mechanisms</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>15</volume>, <fpage>745</fpage>&#x02013;<lpage>756</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3838</pub-id><pub-id pub-id-type="pmid">25315388</pub-id></citation></ref>
<ref id="B329">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>R.</given-names></name></person-group> (<year>2004</year>). <article-title>Desiderata for cognitive architectures</article-title>. <source>Philos. Psychol.</source> <volume>17</volume>, <fpage>341</fpage>&#x02013;<lpage>373</lpage>. <pub-id pub-id-type="doi">10.1080/0951508042000286721</pub-id></citation></ref>
<ref id="B330">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>R.</given-names></name> <name><surname>Coward</surname> <given-names>L. A.</given-names></name> <name><surname>Zenzen</surname> <given-names>M. J.</given-names></name></person-group> (<year>2005</year>). <article-title>On levels of cognitive modeling</article-title>. <source>Philos. Psychol.</source> <volume>18</volume>, <fpage>613</fpage>&#x02013;<lpage>637</lpage>. <pub-id pub-id-type="doi">10.1080/09515080500264248</pub-id></citation></ref>
<ref id="B331">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussillo</surname> <given-names>D.</given-names></name> <name><surname>Churchland</surname> <given-names>M. M.</given-names></name> <name><surname>Kaufman</surname> <given-names>M. T.</given-names></name> <name><surname>Shenoy</surname> <given-names>K. V.</given-names></name></person-group> (<year>2015</year>). <article-title>A neural network that finds a naturalistic solution for the production of muscle activity</article-title>. <source>Nat. Neurosci.</source> <volume>18</volume>, <fpage>1025</fpage>&#x02013;<lpage>1033</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4042</pub-id><pub-id pub-id-type="pmid">26075643</pub-id></citation></ref>
<ref id="B332">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2014</year>). <article-title>Sequence to sequence learning with neural networks,</article-title> in <source>Advances in Neural Information Processing Systems 27</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>3104</fpage>&#x02013;<lpage>3112</lpage>.</citation></ref>
<ref id="B333">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source>Reinforcement Learning: An Introduction</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.1016/j.brainres.2010.09.091</pub-id></citation></ref>
<ref id="B334">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swanson</surname> <given-names>L. W.</given-names></name></person-group> (<year>2000</year>). <article-title>Cerebral hemisphere regulation of motivated behavior</article-title>. <source>Brain Res.</source> <volume>886</volume>, <fpage>113</fpage>&#x02013;<lpage>164</lpage>. <pub-id pub-id-type="doi">10.1016/s0006-8993(00)02905-x</pub-id><pub-id pub-id-type="pmid">11119693</pub-id></citation></ref>
<ref id="B335">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Swanson</surname> <given-names>L. W.</given-names></name></person-group> (<year>2012</year>). <source>Brain Architecture: Understanding the Basic Plan, 2nd Edn</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B336">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Synnaeve</surname> <given-names>G.</given-names></name> <name><surname>Nardelli</surname> <given-names>N.</given-names></name> <name><surname>Auvolat</surname> <given-names>A.</given-names></name> <name><surname>Chintala</surname> <given-names>S.</given-names></name> <name><surname>Lacroix</surname> <given-names>T.</given-names></name> <name><surname>Lin</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>TorchCraft: a library for machine learning research on real-time strategy games</article-title>. ArXiv:1611.00625v2, <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation></ref>
<ref id="B337">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szigeti</surname> <given-names>B.</given-names></name> <name><surname>Gleeson</surname> <given-names>P.</given-names></name> <name><surname>Vella</surname> <given-names>M.</given-names></name> <name><surname>Khayrulin</surname> <given-names>S.</given-names></name> <name><surname>Palyanov</surname> <given-names>A.</given-names></name> <name><surname>Hokanson</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>OpenWorm: an open-science approach to modeling Caenorhabditis elegans</article-title>. <source>Front. Comput. Neurosci.</source> <volume>8</volume>:<fpage>137</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2014.00137</pub-id><pub-id pub-id-type="pmid">25404913</pub-id></citation></ref>
<ref id="B338">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tapaswi</surname> <given-names>M.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Stiefelhagen</surname> <given-names>R.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Urtasun</surname> <given-names>R.</given-names></name> <name><surname>Fidler</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>MovieQA: understanding stories in movies through question-answering</article-title>. ArXiv:1512.02902, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B339">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Kemp</surname> <given-names>C.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. L.</given-names></name> <name><surname>Goodman</surname> <given-names>N. D.</given-names></name></person-group> (<year>2011</year>). <article-title>How to grow a mind: statistics, structure, and abstraction</article-title>. <source>Science</source> <volume>331</volume>, <fpage>1279</fpage>&#x02013;<lpage>1285</lpage>. <pub-id pub-id-type="doi">10.1126/science.1192788</pub-id><pub-id pub-id-type="pmid">21393536</pub-id></citation></ref>
<ref id="B340">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Thalmeier</surname> <given-names>D.</given-names></name> <name><surname>Uhlmann</surname> <given-names>M.</given-names></name> <name><surname>Kappen</surname> <given-names>H. J.</given-names></name> <name><surname>Memmesheimer</surname> <given-names>R.-M.</given-names></name> <name><surname>May</surname> <given-names>N. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning universal computations with spikes</article-title>. ArXiv:1505.07866v1, <fpage>1</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="pmid">27309381</pub-id></citation></ref>
<ref id="B341">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thorpe</surname> <given-names>S. J.</given-names></name> <name><surname>Fabre-Thorpe</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>Seeking categories in the brain</article-title>. <source>Science</source> <volume>291</volume>, <fpage>260</fpage>&#x02013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1126/science.1058249</pub-id><pub-id pub-id-type="pmid">11253215</pub-id></citation></ref>
<ref id="B342">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thrun</surname> <given-names>S.</given-names></name> <name><surname>Mitchell</surname> <given-names>T. M.</given-names></name></person-group> (<year>1995</year>). <article-title>Lifelong robot learning</article-title>. <source>Robot. Auton. Syst.</source> <volume>15</volume>, <fpage>25</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/0921-8890(95)00004-y</pub-id></citation></ref>
<ref id="B343">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thurstone</surname> <given-names>L.</given-names></name></person-group> (<year>1923</year>). <article-title>The stimulus-response fallacy in psychology</article-title>. <source>Psychol. Rev.</source> <volume>30</volume>:<fpage>354369</fpage>. <pub-id pub-id-type="doi">10.1037/h0074251</pub-id></citation></ref>
<ref id="B344">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tinbergen</surname> <given-names>N.</given-names></name></person-group> (<year>1951</year>). <source>The Study of Instinct</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B345">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tobin</surname> <given-names>J.</given-names></name> <name><surname>Fong</surname> <given-names>R.</given-names></name> <name><surname>Ray</surname> <given-names>A.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Zaremba</surname> <given-names>W.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Domain randomization for transferring deep neural networks from simulation to the real world</article-title>. ArXiv:1703.06907, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B346">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name> <name><surname>Erez</surname> <given-names>T.</given-names></name> <name><surname>Tassa</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>MuJoCo: a physics engine for model-based control,</article-title> in <source>International Conference on Intelligent Robots and Systems</source> (<publisher-loc>Vilamoura</publisher-loc>), <fpage>5026</fpage>&#x02013;<lpage>5033</lpage>. <pub-id pub-id-type="doi">10.1109/iros.2012.6386109</pub-id></citation></ref>
<ref id="B347">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2002</year>). <article-title>Optimal feedback control as a theory of motor coordination</article-title>. <source>Nat. Neurosci.</source> <volume>5</volume>, <fpage>1226</fpage>&#x02013;<lpage>1235</lpage>. <pub-id pub-id-type="doi">10.1038/nn963</pub-id><pub-id pub-id-type="pmid">12404008</pub-id></citation></ref>
<ref id="B348">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tolman</surname> <given-names>E.</given-names></name></person-group> (<year>1932</year>). <source>Purposive Behavior in Animals and Men</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Century</publisher-name>.</citation></ref>
<ref id="B349">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Torras i Gen&#x000ED;s</surname> <given-names>C.</given-names></name></person-group> (<year>1986</year>). <article-title>Neural network model with rhythm-assimilation capacity</article-title>. <source>IEEE Trans. Syst. Man Cybern.</source> <volume>16</volume>, <fpage>680</fpage>&#x02013;<lpage>693</lpage>. <pub-id pub-id-type="doi">10.1109/TSMC.1986.289312</pub-id></citation></ref>
<ref id="B350">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turing</surname> <given-names>A. M.</given-names></name></person-group> (<year>1950</year>). <article-title>Computing machinery and intelligence</article-title>. <source>Mind</source> <volume>49</volume>, <fpage>433</fpage>&#x02013;<lpage>460</lpage>. <pub-id pub-id-type="doi">10.1093/mind/LIX.236.433</pub-id></citation></ref>
<ref id="B351">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van de Burgt</surname> <given-names>Y.</given-names></name> <name><surname>Lubberman</surname> <given-names>E.</given-names></name> <name><surname>Fuller</surname> <given-names>E. J.</given-names></name> <name><surname>Keene</surname> <given-names>S. T.</given-names></name> <name><surname>Faria</surname> <given-names>G. C.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A non-volatile organic electrochemical device as a low-voltage artificial synapse for neuromorphic computing</article-title>. <source>Nat. Mater.</source> <volume>16</volume>, <fpage>414</fpage>&#x02013;<lpage>419</lpage>. <pub-id pub-id-type="doi">10.1038/NMAT4856</pub-id><pub-id pub-id-type="pmid">28218920</pub-id></citation></ref>
<ref id="B352">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2017</year>). <article-title>A primer on encoding models in sensory neuroscience</article-title>. <source>J. Math. Psychol.</source> <volume>76</volume>, <fpage>172</fpage>&#x02013;<lpage>183</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2016.06.009</pub-id></citation></ref>
<ref id="B353">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vanrullen</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <article-title>The power of the feed-forward sweep</article-title>. <source>Adv. Cogn. Psychol.</source> <volume>3</volume>, <fpage>167</fpage>&#x02013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.2478/v10053-008-0022-3</pub-id><pub-id pub-id-type="pmid">20517506</pub-id></citation></ref>
<ref id="B354">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vanrullen</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Perception science in the age of deep neural networks</article-title>. <source>Front. Psychol.</source> <volume>8</volume>:<fpage>142</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2017.00142</pub-id><pub-id pub-id-type="pmid">28210237</pub-id></citation></ref>
<ref id="B355">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varshney</surname> <given-names>L. R.</given-names></name> <name><surname>Chen</surname> <given-names>B. L.</given-names></name> <name><surname>Paniagua</surname> <given-names>E.</given-names></name> <name><surname>Hall</surname> <given-names>D. H.</given-names></name> <name><surname>Chklovskii</surname> <given-names>D. B.</given-names></name></person-group> (<year>2011</year>). <article-title>Structural properties of the Caenorhabditis elegans neuronal network</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1001066</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1001066</pub-id><pub-id pub-id-type="pmid">21304930</pub-id></citation></ref>
<ref id="B356">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vernon</surname> <given-names>D.</given-names></name> <name><surname>Metta</surname> <given-names>G.</given-names></name> <name><surname>Sandini</surname> <given-names>G.</given-names></name></person-group> (<year>2007</year>). <article-title>A survey of artificial cognitive systems: implications for the autonomous development of mental capbilities in computational agents</article-title>. <source>IEEE Trans. Evol. Comput.</source> <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1109/TEVC.2006.890274</pub-id></citation></ref>
<ref id="B357">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Blundell</surname> <given-names>C.</given-names></name> <name><surname>Lillicrap</surname> <given-names>T.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Matching networks for one shot learning</article-title>. ArXiv:1606.04080v1, <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B358">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Brain</surname> <given-names>G.</given-names></name> <name><surname>Fortunato</surname> <given-names>M.</given-names></name> <name><surname>Jaitly</surname> <given-names>N.</given-names></name> <name><surname>Brain</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Pointer networks</article-title>. ArXiv:1506.03134v2, <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B359">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>von Neumann</surname> <given-names>J.</given-names></name></person-group> (<year>1966</year>). <source>Theory of Self-Reproducing Automata</source>. <publisher-loc>Champaign, IL</publisher-loc>: <publisher-name>University of Illinois Press</publisher-name>.</citation></ref>
<ref id="B360">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>von Neumann</surname> <given-names>J.</given-names></name> <name><surname>Morgenstern</surname> <given-names>O.</given-names></name></person-group> (<year>1953</year>). <source>Theory of Games and Economic Behavior, 3rd Edn</source>. <publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>Princeton University Press</publisher-name>.</citation></ref>
<ref id="B361">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Weichwald</surname> <given-names>S.</given-names></name> <name><surname>Fomina</surname> <given-names>T.</given-names></name> <name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name> <name><surname>Grosse-Wentrup</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Optimal coding in biological and artificial neural networks</article-title>. ArXiv:1605.07094v2, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B362">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Weston</surname> <given-names>J.</given-names></name> <name><surname>Chopra</surname> <given-names>S.</given-names></name> <name><surname>Bordes</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Memory networks,</article-title> in <source>3rd International Conference on Learning Representations</source> (<publisher-loc>San Diego</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B363">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>White</surname> <given-names>R. W.</given-names></name></person-group> (<year>1959</year>). <article-title>Motivation reconsidered: the concept of competence</article-title>. <source>Psychol. Rev.</source> <volume>66</volume>, <fpage>297</fpage>&#x02013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1037/h0040934</pub-id><pub-id pub-id-type="pmid">13844397</pub-id></citation></ref>
<ref id="B364">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>White</surname> <given-names>S. G.</given-names></name> <name><surname>Southgate</surname> <given-names>E.</given-names></name> <name><surname>Thomson</surname> <given-names>J.</given-names></name> <name><surname>Brenner</surname> <given-names>S.</given-names></name></person-group> (<year>1986</year>). <article-title>The structure of the nervous system of the nematode <italic>C. elegans</italic></article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>314</volume>, <fpage>1</fpage>&#x02013;<lpage>340</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.1986.0056</pub-id></citation></ref>
<ref id="B365">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Whitehead</surname> <given-names>S. D.</given-names></name> <name><surname>Ballard</surname> <given-names>D. H.</given-names></name></person-group> (<year>1991</year>). <article-title>Learning to perceive and act by trial and error</article-title>. <source>Mach. Learn.</source> <volume>7</volume>, <fpage>45</fpage>&#x02013;<lpage>83</lpage>. <pub-id pub-id-type="doi">10.1007/bf00058926</pub-id></citation></ref>
<ref id="B366">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Widrow</surname> <given-names>B.</given-names></name> <name><surname>Lehr</surname> <given-names>M. A.</given-names></name></person-group> (<year>1990</year>). <article-title>30 Years of adaptive neural networks: perceptron, madaline, and backpropagation</article-title>. <source>Proc. IEEE</source> <volume>78</volume>, <fpage>1415</fpage>&#x02013;<lpage>1442</lpage>. <pub-id pub-id-type="doi">10.1109/5.58323</pub-id></citation></ref>
<ref id="B367">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wills</surname> <given-names>T. J.</given-names></name> <name><surname>Lever</surname> <given-names>C.</given-names></name> <name><surname>Cacucci</surname> <given-names>F.</given-names></name> <name><surname>Burgess</surname> <given-names>N.</given-names></name> <name><surname>Keefe</surname> <given-names>J. O.</given-names></name></person-group> (<year>2005</year>). <article-title>Attractor ddynamics in the hippocampal representation of the local environment</article-title>. <source>Science</source> <volume>308</volume>, <fpage>873</fpage>&#x02013;<lpage>876</lpage>. <pub-id pub-id-type="doi">10.1126/science.1108905.Attractor</pub-id><pub-id pub-id-type="pmid">15879220</pub-id></citation></ref>
<ref id="B368">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Willshaw</surname> <given-names>D. J.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Morris</surname> <given-names>R. G. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Memory, modelling and Marr: a commentary on Marr (1971) &#x02018;Simple memory: A theory of archicortex&#x02019;</article-title>. <source>Philos. Trans. R. Soc. B</source> <volume>370</volume>:<fpage>20140383</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2014.0383</pub-id><pub-id pub-id-type="pmid">25750246</pub-id></citation></ref>
<ref id="B369">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Winograd</surname> <given-names>T.</given-names></name></person-group> (<year>1972</year>). <article-title>Understanding natural language</article-title>. <source>Cogn. Psychol.</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0285(72)90002-3</pub-id></citation></ref>
<ref id="B370">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wissner-Gross</surname> <given-names>A. D.</given-names></name> <name><surname>Freer</surname> <given-names>C. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Causal entropic forces</article-title>. <source>Phys. Rev. Lett.</source> <volume>110</volume>:<fpage>168702</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.110.168702</pub-id><pub-id pub-id-type="pmid">23679649</pub-id></citation></ref>
<ref id="B371">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wolfram</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <source>A New Kind of Science</source>. <publisher-loc>Champaign, IL</publisher-loc>: <publisher-name>Wolfram Media</publisher-name>.</citation></ref>
<ref id="B372">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>On multiplicative integration with recurrent neural networks</article-title>. ArXiv:1606.06630v2, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B373">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Xue</surname> <given-names>T.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Bouman</surname> <given-names>K. L.</given-names></name> <name><surname>Freeman</surname> <given-names>W. T.</given-names></name></person-group> (<year>2016</year>). <article-title>Visual dynamics: probabilistic future frame synthesis via cross convolutional networks</article-title>. ArXiv:1607.02586, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B374">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L. K.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Using goal-driven deep learning models to understand sensory cortex</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>356</fpage>&#x02013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4244</pub-id><pub-id pub-id-type="pmid">26906502</pub-id></citation></ref>
<ref id="B375">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>G. R.</given-names></name> <name><surname>Song</surname> <given-names>H. F.</given-names></name> <name><surname>Newsome</surname> <given-names>W. T.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Clustering and compositionality of task representations in a neural network trained to perform many cognitive tasks</article-title>. BioRXiv, <fpage>1</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1101/183632</pub-id></citation></ref>
<ref id="B376">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>W.</given-names></name> <name><surname>Yuste</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title><italic>In vivo</italic> imaging of neural activity</article-title>. <source>Nat. Methods</source> <volume>14</volume>, <fpage>349</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.4230</pub-id><pub-id pub-id-type="pmid">28362436</pub-id></citation></ref>
<ref id="B377">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yarbus</surname> <given-names>A. L.</given-names></name></person-group> (<year>1967</year>). <source>Eye Movements and Vision</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Plenum</publisher-name>.</citation></ref>
<ref id="B378">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuille</surname> <given-names>A.</given-names></name> <name><surname>Kersten</surname> <given-names>D.</given-names></name></person-group> (<year>2006</year>). <article-title>Vision as Bayesian inference: analysis by synthesis?</article-title> <source>Trends Cogn. Sci.</source> <volume>10</volume>, <fpage>301</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2006.05.002</pub-id><pub-id pub-id-type="pmid">16784882</pub-id></citation></ref>
<ref id="B379">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuste</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>From the neuron doctrine to neural networks</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>16</volume>, <fpage>487</fpage>&#x02013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3962</pub-id><pub-id pub-id-type="pmid">26152865</pub-id></citation></ref>
<ref id="B380">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Zagoruyko</surname> <given-names>S.</given-names></name> <name><surname>Komodakis</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>DiracNets: training very deep neural networks without skip-connections</article-title>. ArXiv:1706.00388, <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B381">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Zambrano</surname> <given-names>D.</given-names></name> <name><surname>Bohte</surname> <given-names>S. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Fast and efficient asynchronous neural computation with adapting spiking neural networks</article-title>. ArXiv:1609.02053, <fpage>1</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B382">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Zenke</surname> <given-names>F.</given-names></name> <name><surname>Poole</surname> <given-names>B.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Improved multitask learning through synaptic intelligence</article-title>. ArXiv:1703.04200v2, <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B383">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Gordon</surname> <given-names>D.</given-names></name> <name><surname>Kolve</surname> <given-names>E.</given-names></name> <name><surname>Fox</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Visual semantic planning using deep successor representations</article-title>. ArXiv:1705.08080v1, <fpage>1</fpage>&#x02013;<lpage>13</lpage>.</citation></ref>
<ref id="B384">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zipser</surname> <given-names>D.</given-names></name> <name><surname>Andersen</surname> <given-names>R.</given-names></name></person-group> (<year>1988</year>). <article-title>A back-propagation programmed network that simulates response properties of a subset of posterior parietal neurons</article-title>. <source>Nature</source> <volume>331</volume>, <fpage>679</fpage>&#x02013;<lpage>684</lpage>. <pub-id pub-id-type="doi">10.1038/331679a0</pub-id><pub-id pub-id-type="pmid">3344044</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>Figure modified from <ext-link ext-link-type="uri" xlink:href="http://larrywswanson.com/?page_id=1523">http://larrywswanson.com/?page_id=1523</ext-link> with permission.</p></fn>
<fn id="fn0002"><p><sup>2</sup>Figure by K. D. Schroeder, CC BY-SA 3.0, <ext-link ext-link-type="uri" xlink:href="https://commons.wikimedia.org/w/index.php?curid=26958836">https://commons.wikimedia.org/w/index.php?curid=26958836</ext-link>. Used with permission.</p></fn>
<fn id="fn0003"><p><sup>3</sup>In fact, ACT-R also uses some subsymbolic elements and can therefore be considered a <italic>hybrid</italic> architecture.</p></fn>
<fn id="fn0004"><p><sup>4</sup>Figure modified from <ext-link ext-link-type="uri" xlink:href="http://act-r.psy.cmu.edu/about">http://act-r.psy.cmu.edu/about</ext-link> with permission.</p></fn>
<fn id="fn0005"><p><sup>5</sup>Beliefs over continuous quantities can be expressed by replacing summation with integration.</p></fn>
<fn id="fn0006"><p><sup>6</sup>From: <ext-link ext-link-type="uri" xlink:href="https://www.izhikevich.org/human_brain_simulation/why.htm">https://www.izhikevich.org/human_brain_simulation/why.htm</ext-link></p></fn>
<fn id="fn0007"><p><sup>7</sup>Intentionality or &#x0201C;aboutness&#x0201D; refers to the quality of mental states as being directed toward an object or state of affairs.</p></fn>
<fn id="fn0008"><p><sup>8</sup>In practice, it is more efficient to iterate over subsets of datapoints, known as mini-batches, in sequence. That is, training is organized in terms of epochs in which all datapoints are processed by iterating over mini-batches. Note that, whenever we are not processing all data points in parallel, we are not exactly following the gradient. Therefore, any such procedure is known as <italic>stochastic gradient descent</italic>.</p></fn>
<fn id="fn0009"><p><sup>9</sup>The ability of simple RNNs to integrate information over time remains limited, which led to the introduction of various extensions that perform more favorably in this regard (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B144">1997</xref>; Cho et al., <xref ref-type="bibr" rid="B54">2014</xref>; Neil et al., <xref ref-type="bibr" rid="B240">2016</xref>; Wu et al., <xref ref-type="bibr" rid="B372">2016</xref>).</p></fn>
<fn id="fn0010"><p><sup>10</sup>See SHRDLU for an early example of such a virtual world (Winograd, <xref ref-type="bibr" rid="B369">1972</xref>).</p></fn>
<fn id="fn0011"><p><sup>11</sup>The notion of <italic>wanting</italic> agents was already present in the writings of Thurstone (<xref ref-type="bibr" rid="B343">1923</xref>), who wrote: &#x0201C;My main thesis is that conduct originates in the organism itself and not in the environment in the form of a stimulus. [&#x02026;] All mental life may be looked upon as incomplete behavior which is in the process of being formed. [&#x02026;] Perception is the discovery of the suitable stimulus which is often anticipated imaginally. The appearance of the stimulus is one of the last events in the expression of impulses in conduct. The stimulus is not the starting point for behavior.&#x0201D;</p></fn>
<fn id="fn0012"><p><sup>12</sup>But see Jonas and Kording (<xref ref-type="bibr" rid="B156">2017</xref>) for a critical appraisal of the informativeness of such techniques.</p></fn>
<fn id="fn0013"><p><sup>13</sup>These techniques aim to overcome the interpretability problem raised by Mozer and Smolensky (<xref ref-type="bibr" rid="B235">1989</xref>), who state: &#x0201D;One thing that connectionist networks have in common with brains is that if you open them up and peer inside, all you can see is a big pile of goo.&#x0201D;</p></fn>
<fn id="fn0014"><p><sup>14</sup>While there surely exists neurobiological evidence for temporal coding with spikes (Segundo et al., <xref ref-type="bibr" rid="B307">1966</xref>; Barrio and Buno, <xref ref-type="bibr" rid="B22">1990</xref>; Bohte, <xref ref-type="bibr" rid="B36">2004</xref>), it remains an open question if temporal coding is absolutely necessary for the generation of adaptive behavior. In the end, computing with spikes may have emerged chiefly to promote efficiency and allow long-distance neuronal communication (Laughlin and Sejnowski, <xref ref-type="bibr" rid="B184">2003</xref>).</p></fn>
</fn-group>
</back>
</article>
