<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2016.00094</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Hypothesis and Theory</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Toward an Integration of Deep Learning and Neuroscience</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Marblestone</surname> <given-names>Adam H.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/101076/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wayne</surname> <given-names>Greg</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/375596/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kording</surname> <given-names>Konrad P.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/231/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Synthetic Neurobiology Group, Massachusetts Institute of Technology, Media Lab</institution> <country>Cambridge, MA, USA</country></aff>
<aff id="aff2"><sup>2</sup><institution>Google Deepmind</institution> <country>London, UK</country></aff>
<aff id="aff3"><sup>3</sup><institution>Rehabilitation Institute of Chicago, Northwestern University</institution> <country>Chicago, IL, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Sander Bohte, Centrum Wiskunde &#x00026; Informatica, Netherlands</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Mattia Rigotti, IBM T.J. Watson Research Center, USA; H. Steven Scholte, University of Amsterdam, Netherlands; Petia D. Koprinkova-Hristova, Bulgarian Academy of Sciences, Bulgaria</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Adam H. Marblestone <email>adam.h.marblestone&#x00040;gmail.com</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>09</month>
<year>2016</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>10</volume>
<elocation-id>94</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>06</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>08</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2016 Marblestone, Wayne and Kording.</copyright-statement>
<copyright-year>2016</copyright-year>
<copyright-holder>Marblestone, Wayne and Kording</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Neuroscience has focused on the detailed implementation of computation, studying neural codes, dynamics and circuits. In machine learning, however, artificial neural networks tend to eschew precisely designed codes, dynamics or circuits in favor of brute force optimization of a cost function, often using simple and relatively uniform initial architectures. Two recent developments have emerged within machine learning that create an opportunity to connect these seemingly divergent perspectives. First, structured architectures are used, including dedicated systems for attention, recursion and various forms of short- and long-term memory storage. Second, cost functions and training procedures have become more complex and are varied across layers and over time. Here we think about the brain in terms of these ideas. We hypothesize that (1) the brain optimizes cost functions, (2) the cost functions are diverse and differ across brain locations and over development, and (3) optimization operates within a pre-structured architecture matched to the computational problems posed by behavior. In support of these hypotheses, we argue that a range of implementations of credit assignment through multiple layers of neurons are compatible with our current knowledge of neural circuitry, and that the brain&#x00027;s specialized systems can be interpreted as enabling efficient optimization for specific problem classes. Such a heterogeneously optimized system, enabled by a series of interacting cost functions, serves to make learning data-efficient and precisely targeted to the needs of the organism. We suggest directions by which neuroscience could seek to refine and test these hypotheses.</p></abstract>
<kwd-group><kwd>cost functions</kwd>
<kwd>neural networks</kwd>
<kwd>neuroscience</kwd>
<kwd>cognitive architecture</kwd>
</kwd-group>
<counts>
<fig-count count="1"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="488"/>
<page-count count="41"/>
<word-count count="41961"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Machine learning and neuroscience speak different languages today. Brain science has discovered a dazzling array of brain areas (Solari and Stoner, <xref ref-type="bibr" rid="B406">2011</xref>), cell types, molecules, cellular states, and mechanisms for computation and information storage. Machine learning, in contrast, has largely focused on instantiations of a single principle: function optimization. It has found that simple optimization objectives, like minimizing classification error, can lead to the formation of rich internal representations and powerful algorithmic capabilities in multilayer and recurrent networks (LeCun et al., <xref ref-type="bibr" rid="B247">2015</xref>; Schmidhuber, <xref ref-type="bibr" rid="B386">2015</xref>). Here we seek to connect these perspectives.</p>
<p>The artificial neural networks now prominent in machine learning were, of course, originally inspired by neuroscience (McCulloch and Pitts, <xref ref-type="bibr" rid="B293">1943</xref>). While neuroscience has continued to play a role (Cox and Dean, <xref ref-type="bibr" rid="B85">2014</xref>), many of the major developments were guided by insights into the mathematics of efficient optimization, rather than neuroscientific findings (Sutskever and Martens, <xref ref-type="bibr" rid="B421">2013</xref>). The field has advanced from simple linear systems (Minsky and Papert, <xref ref-type="bibr" rid="B308">1972</xref>), to nonlinear networks (Haykin, <xref ref-type="bibr" rid="B177">1994</xref>), to deep and recurrent networks (LeCun et al., <xref ref-type="bibr" rid="B247">2015</xref>; Schmidhuber, <xref ref-type="bibr" rid="B386">2015</xref>). Backpropagation of error (Werbos, <xref ref-type="bibr" rid="B459">1974</xref>, <xref ref-type="bibr" rid="B460">1982</xref>; Rumelhart et al., <xref ref-type="bibr" rid="B376">1986</xref>) enabled neural networks to be trained efficiently, by providing an efficient means to compute the gradient with respect to the weights of a multi-layer network. Methods of training have improved to include momentum terms, better weight initializations, conjugate gradients and so forth, evolving to the current breed of networks optimized using batch-wise stochastic gradient descent. These developments have little obvious connection to neuroscience.</p>
<p>We will argue here, however, that neuroscience and machine learning are again ripe for convergence. Three aspects of machine learning are particularly important in the context of this paper. First, machine learning has focused on the optimization of cost functions (Figure <xref ref-type="fig" rid="F1">1A</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Putative differences between conventional and brain-like neural network designs</bold>. <bold>(A)</bold> In conventional deep learning, supervised training is based on externally-supplied, labeled data. <bold>(B)</bold> In the brain, supervised training of networks can still occur via gradient descent on an error signal, but this error signal must arise from internally generated cost functions. These cost functions are themselves computed by neural modules specified by both genetics and learning. Internally generated cost functions create heuristics that are used to bootstrap more complex learning. For example, an area which recognizes faces might first be trained to detect faces using simple heuristics, like the presence of two dots above a line, and then further trained to discriminate salient facial expressions using representations arising from unsupervised learning and error signals from other brain areas related to social reward processing. <bold>(C)</bold> Internally generated cost functions and error-driven training of cortical deep networks form part of a larger architecture containing several specialized systems. Although the trainable cortical areas are schematized as feedforward neural networks here, LSTMs or other types of recurrent networks may be a more accurate analogy, and many neuronal and network properties such as spiking, dendritic computation, neuromodulation, adaptation and homeostatic plasticity, timing-dependent plasticity, direct electrical connections, transient synaptic dynamics, excitatory/inhibitory balance, spontaneous oscillatory activity, axonal conduction delays (Izhikevich, <xref ref-type="bibr" rid="B203">2006</xref>) and others, will influence what and how such networks learn.</p></caption>
<graphic xlink:href="fncom-10-00094-g0001.tif"/>
</fig>
<p>Second, recent work in machine learning has started to introduce complex cost functions, those that are not uniform across layers and time, and those that arise from interactions between different parts of a network. For example, introducing the objective of temporal coherence for lower layers (non-uniform cost function over space) improves feature learning (Sermanet and Kavukcuoglu, <xref ref-type="bibr" rid="B390">2013</xref>), cost function schedules (non-uniform cost function over time) improve<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> generalization (Saxe et al., <xref ref-type="bibr" rid="B382">2013</xref>; Goodfellow et al., <xref ref-type="bibr" rid="B149">2014b</xref>; G&#x000FC;l&#x000E7;ehre and Bengio, <xref ref-type="bibr" rid="B158">2016</xref>) and adversarial networks&#x02014;an example of a cost function arising from internal interactions&#x02014;allow gradient-based training of generative models (Goodfellow et al., <xref ref-type="bibr" rid="B148">2014a</xref>)<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>. Networks that are easier to train are being used to provide &#x0201C;hints&#x0201D; to help bootstrap the training of more powerful networks (Romero et al., <xref ref-type="bibr" rid="B372">2014</xref>).</p>
<p>Third, machine learning has also begun to diversify the architectures that are subject to optimization. It has introduced simple memory cells with multiple persistent states (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B192">1997</xref>; Chung et al., <xref ref-type="bibr" rid="B78">2014</xref>), more complex elementary units such as &#x0201C;capsules&#x0201D; and other structures (Delalleau and Bengio, <xref ref-type="bibr" rid="B97">2011</xref>; Hinton et al., <xref ref-type="bibr" rid="B189">2011</xref>; Tang et al., <xref ref-type="bibr" rid="B427">2012</xref>; Livni et al., <xref ref-type="bibr" rid="B265">2013</xref>), content addressable (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>; Weston et al., <xref ref-type="bibr" rid="B464">2014</xref>) and location addressable memories (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>), as well as pointers (Kurach et al., <xref ref-type="bibr" rid="B242">2015</xref>) and hard-coded arithmetic operations (Neelakantan et al., <xref ref-type="bibr" rid="B321">2015</xref>).</p>
<p>These three ideas have, so far, not received much attention in neuroscience. We thus formulate these ideas as three hypotheses about the brain, examine evidence for them, and sketch how experiments could test them. But first, let us state the hypotheses more precisely.</p>
<sec>
<title>1.1. <italic>Hypothesis 1</italic> &#x02013; the brain optimizes cost functions</title>
<p>The central hypothesis for linking the two fields is that biological systems, like many machine-learning systems, are able to optimize cost functions. The idea of cost functions means that neurons in a brain area can somehow change their properties, e.g., the properties of their synapses, so that they get better at doing whatever the cost function defines as their role. Human behavior sometimes approaches optimality in a domain, e.g., during movement (K&#x000F6;rding, <xref ref-type="bibr" rid="B228">2007</xref>), which suggests that the brain may have learned optimal strategies. Subjects minimize energy consumption of their movement system (Taylor and Faisal, <xref ref-type="bibr" rid="B431">2011</xref>), and minimize risk and damage to their body, while maximizing financial and movement gains. Computationally, we now know that optimization of trajectories gives rise to elegant solutions for very complex motor tasks (Harris and Wolpert, <xref ref-type="bibr" rid="B165">1998</xref>; Todorov and Jordan, <xref ref-type="bibr" rid="B439">2002</xref>; Mordatch et al., <xref ref-type="bibr" rid="B317">2012</xref>). We suggest that cost function optimization occurs much more generally in shaping the internal representations and processes used by the brain. Importantly, we also suggest that this requires the brain to have mechanisms for efficient credit assignment in multilayer and recurrent networks.</p>
</sec>
<sec>
<title>1.2. <italic>Hypothesis 2</italic> &#x02013; cost functions are diverse across areas and change over development</title>
<p>A second realization is that cost functions need not be global. Neurons in different brain areas may optimize different things, e.g., the mean squared error of movements, surprise in a visual stimulus, or the allocation of attention. Importantly, such a cost function could be locally generated. For example, neurons could locally evaluate the quality of their statistical model of their inputs (Figure <xref ref-type="fig" rid="F1">1B</xref>). Alternatively, cost functions for one area could be generated by another area. Moreover, cost functions may change over time, e.g., guiding young humans to understanding simple visual contrasts early on, and faces a bit later<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>. This could allow the developing brain to bootstrap more complex knowledge based on simpler knowledge. Cost functions in the brain are likely to be complex and to be arranged to vary across areas and over development.</p>
</sec>
<sec>
<title>1.3. <italic>Hypothesis 3</italic> &#x02013; specialized systems allow efficient solution of key computational problems</title>
<p>A third realization is that structure matters. The patterns of information flow seem fundamentally different across brain areas, suggesting that they solve distinct computational problems. Some brain areas are highly recurrent, perhaps making them predestined for short-term memory storage (Wang, <xref ref-type="bibr" rid="B454">2012</xref>). Some areas contain cell types that can switch between qualitatively different states of activation, such as a persistent firing mode vs. a transient firing mode, in response to particular neurotransmitters (Hasselmo, <xref ref-type="bibr" rid="B169">2006</xref>). Other areas, like the thalamus appear to have the information from other areas flowing through them, perhaps allowing them to determine information routing (Sherman, <xref ref-type="bibr" rid="B397">2005</xref>). Areas like the basal ganglia are involved in reinforcement learning and gating of discrete decisions (Doya, <xref ref-type="bibr" rid="B102">1999</xref>; Sejnowski and Poizner, <xref ref-type="bibr" rid="B389">2014</xref>). As every programmer knows, specialized algorithms matter for efficient solutions to computational problems, and the brain is likely to make good use of such specialization (Figure <xref ref-type="fig" rid="F1">1C</xref>).</p>
<p>These ideas are inspired by recent advances in machine learning, but we also propose that the brain has major differences from any of today&#x00027;s machine learning techniques. In particular, the world gives us a relatively limited amount of information that we could use for supervised learning (Fodor and Crowther, <xref ref-type="bibr" rid="B123">2002</xref>). There is a huge amount of information available for unsupervised learning, but there is no reason to assume that a <italic>generic</italic> unsupervised algorithm, no matter how powerful, would learn the precise things that humans need to know, in the order that they need to know it. The evolutionary challenge of making unsupervised learning solve the &#x0201C;right&#x0201D; problems is, therefore, to find a sequence of cost functions that will deterministically build circuits and behaviors according to prescribed developmental stages, so that in the end a relatively small amount of information suffices to produce the right behavior. For example, a developing duck imprints (Tinbergen, <xref ref-type="bibr" rid="B436">1965</xref>) a template of its parent, and then uses that template to generate goal-targets that help it develop other skills like foraging.</p>
<p>Generalizing from this and from other studies (Minsky, <xref ref-type="bibr" rid="B304">1977</xref>; Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>), we propose that many of the brain&#x00027;s cost functions arise from such an internal bootstrapping process. Indeed, we propose that biological development and reinforcement learning can, in effect, program the emergence of a sequence of cost functions that precisely anticipates the future needs faced by the brain&#x00027;s internal subsystems, as well as by the organism as a whole. This type of developmentally programmed bootstrapping generates an internal infrastructure of cost functions which is diverse and complex, while simplifying the learning problems faced by the brain&#x00027;s internal processes. Beyond simple tasks like familial imprinting, this type of bootstrapping could extend to higher cognition, e.g., internally generated cost functions could train a developing brain to properly access its memory or to organize its actions in ways that will prove to be useful later on. The potential bootstrapping mechanisms that we will consider operate in the context of unsupervised and reinforcement learning, and go well beyond the types of curriculum learning ideas used in today&#x00027;s machine learning (Bengio et al., <xref ref-type="bibr" rid="B39">2009</xref>).</p>
<p>In the rest of this paper, we will elaborate on these hypotheses. First, we will argue that both local and multi-layer optimization is, perhaps surprisingly, compatible with what we know about the brain. Second, we will argue that cost functions differ across brain areas and change over time and describe how cost functions interacting in an orchestrated way could allow bootstrapping of complex function. Third, we will list a broad set of specialized problems that need to be solved by neural computation, and the brain areas that have structure that seems to be matched to a particular computational problem. We then discuss some implications of the above hypotheses for research approaches in neuroscience and machine learning, and sketch a set of experiments to test these hypotheses. Finally, we discuss this architecture from the perspective of evolution.</p>
</sec>
</sec>
<sec id="s2">
<title>2. The brain can optimize cost functions</title>
<p>Much of machine learning is based on efficiently optimizing functions, and, as we will detail below, the ability to use backpropagation of error (Werbos, <xref ref-type="bibr" rid="B459">1974</xref>; Rumelhart et al., <xref ref-type="bibr" rid="B376">1986</xref>) to calculate gradients of arbitrary parametrized functions has been a key breakthrough. In Hypothesis 1, we claim that the brain is also, at least in part<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref>, an optimization machine. But what exactly does it mean to say that the brain can optimize cost functions? After all, many processes can be viewed as optimizations. For example, the laws of physics are often viewed as minimizing an action functional, while evolution optimizes the fitness of replicators over a long timescale. To be clear, our main claims are: that (a) the brain has powerful mechanisms for credit assignment during learning that allow it to optimize global functions in multi-layer networks by adjusting the properties of each neuron to contribute to the global outcome, and that (b) the brain has mechanisms to specify exactly which cost functions it subjects its networks to, i.e., that the cost functions are highly tunable, shaped by evolution and matched to the animal&#x00027;s ethological needs. Thus, the brain uses cost functions as a key driving force of its development, much as modern machine learning systems do.</p>
<p>To understand the basis of these claims, we must now delve into the details of how the brain might efficiently perform credit assignment throughout large, multi-layered networks, in order to optimize complex functions. We argue that the brain uses several different types of optimization to solve distinct problems. In some structures, it may use genetic pre-specification of circuits for problems that require only limited learning based on data, or it may exploit local optimization to avoid the need to assign credit through many layers of neurons. It may also use a host of proposed circuit structures that would allow it to actually perform, in effect, backpropagation of errors through a multi-layer network, using biologically realistic mechanisms&#x02014;a feat that had once been widely believed to be biologically implausible (Crick, <xref ref-type="bibr" rid="B86">1989</xref>; Stork, <xref ref-type="bibr" rid="B414">1989</xref>). Potential such mechanisms include circuits that literally backpropagate error derivatives in the manner of conventional backpropagation, as well as circuits that provide other efficient means of approximating the effects of backpropagation, i.e., of rapidly computing the approximate gradient of a cost function relative to any given connection weight in the network. Lastly, the brain may use algorithms that exploit specific aspects of neurophysiology&#x02014;such as spike timing dependent plasticity, dendritic computation, local excitatory-inhibitory networks, or other properties&#x02014;as well as the integrated nature of higher-level brain systems. Such mechanisms promise to allow learning capabilities that go even beyond those of current backpropagation networks.</p>
<sec>
<title>2.1. Local self-organization and optimization without multi-layer credit assignment</title>
<p>Not all learning requires a general-purpose optimization mechanism like gradient descent<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref>. Many theories of cortex (George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>; Kappel et al., <xref ref-type="bibr" rid="B221">2014</xref>) emphasize potential self-organizing and unsupervised learning properties that may obviate the need for multi-layer backpropagation as such. Hebbian plasticity, which adjusts weights according to correlations in pre-synaptic and post-synaptic activity, is well established<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>. Various versions of Hebbian plasticity (Miller and MacKay, <xref ref-type="bibr" rid="B303">1994</xref>), e.g., with nonlinearities (Brito and Gerstner, <xref ref-type="bibr" rid="B58">2016</xref>), can give rise to different forms of correlation and competition between neurons, leading to the self-organized formation of ocular dominance columns, self-organizing maps and orientation columns (Miller et al., <xref ref-type="bibr" rid="B302">1989</xref>; Ferster and Miller, <xref ref-type="bibr" rid="B115">2000</xref>). Often these types of local self-organization can also be viewed as optimizing a cost function: for example, certain forms of Hebbian plasticity can be viewed as extracting the principal components of the input, which minimizes a reconstruction error (Pehlevan and Chklovskii, <xref ref-type="bibr" rid="B343">2015</xref>).</p>
<p>To generate complex temporal patterns, the brain may also implement other forms of learning that do not require any equivalent of full backpropagation through a multilayer network. For example, &#x0201C;liquid-&#x0201D; (Maass et al., <xref ref-type="bibr" rid="B274">2002</xref>) or &#x0201C;echo-state machines&#x0201D; (Jaeger and Haas, <xref ref-type="bibr" rid="B209">2004</xref>) are randomly connected recurrent networks that form a basis set (also known as a &#x0201C;reservoir&#x0201D;) of random filters, which can be harnessed for learning with tunable readout weights. Variants exhibiting chaotic, spontaneous dynamics can even be trained by feeding back readouts into the network and suppressing the chaotic activity (Sussillo and Abbott, <xref ref-type="bibr" rid="B419">2009</xref>). Learning only the readout layer makes the optimization problem much simpler (indeed, equivalent to regression for supervised learning). Additionally, echo state networks can be trained by reinforcement learning as well as supervised learning (Bush, <xref ref-type="bibr" rid="B67">2007</xref>; Hoerzer et al., <xref ref-type="bibr" rid="B193">2014</xref>). Reservoirs of random nonlinear filters are one interpretation of the diverse, high-dimensional, mixed-selectivity tuning properties of many neurons, e.g., in the prefrontal cortex (Enel et al., <xref ref-type="bibr" rid="B110">2016</xref>). Other variants of learning rules that modify only a fraction of the synapses inside a random network are being developed as models of biological working memory and sequence generation (Rajan et al., <xref ref-type="bibr" rid="B357">2016</xref>).</p>
</sec>
<sec>
<title>2.2. Biological implementation of optimization</title>
<p>We argue that the above mechanisms of local self-organization are likely insufficient to account for the brain&#x00027;s powerful learning performance (Brea and Gerstner, <xref ref-type="bibr" rid="B56">2016</xref>). To elaborate on the need for an efficient means of gradient computation in the brain, we will first place backpropagation into its computational context (Hinton, <xref ref-type="bibr" rid="B183">1989</xref>; Baldi and Sadowski, <xref ref-type="bibr" rid="B27">2015</xref>). Then we will explain how the brain could plausibly implement approximations of gradient descent.</p>
<sec>
<title>2.2.1. The need for efficient gradient descent in multi-layer networks</title>
<p>The simplest mechanism to perform cost function optimization is sometimes known as the &#x0201C;twiddle&#x0201D; algorithm or, more technically, as &#x0201C;serial perturbation.&#x0201D; This mechanism works by perturbing (i.e., &#x0201C;twiddling&#x0201D;), with a small increment, a single weight in the network, and verifying improvement by measuring whether the cost function has decreased compared to the network&#x00027;s performance with the weight unperturbed. If improvement is noticeable, the perturbation is used as a direction of change to the weight; otherwise, the weight is changed in the opposite direction (or not changed at all). Serial perturbation is therefore a method of &#x0201C;coordinate descent&#x0201D; on the cost, but it is slow and requires global coordination: each synapse in turn is perturbed while others remain fixed.</p>
<p>Weight perturbation (or parallel perturbation) perturbs all of the weights in the network at once. It is able to optimize small networks to perform tasks but generally suffers from high variance. That is, the measurement of the gradient direction is noisy and changes drastically from perturbation to perturbation because a weight&#x00027;s influence on the cost is masked by the changes of all other weights, and there is only one scalar feedback signal indicating the change in the cost<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref>. Weight perturbation is dramatically inefficient for large networks. In fact, parallel and serial perturbation learn at approximately the same rate if the time measure counts the number of times the network propagates information from input to output (Werfel et al., <xref ref-type="bibr" rid="B463">2005</xref>).</p>
<p>Some efficiency gain can be achieved by perturbing neural activities instead of synaptic weights, acknowledging the fact that any long-range effect of a synapse is mediated through a neuron. Like weight perturbation and unlike serial perturbation, minimal global coordination is needed: each neuron only needs to receive a feedback signal indicating the global cost. The variance of node perturbation&#x00027;s gradient estimate is far smaller than that of weight perturbation under the assumptions that either all neurons or all weights, respectively, are perturbed and that they are perturbed at the same frequency. In this case, node perturbation&#x00027;s variance is proportional to the number of cells in the network, not the number of synapses.</p>
<p>All of these approaches are slow either due to the time needed for serial iteration over all weights or the time needed for averaging over low signal-to-noise ratio gradient estimates. To their credit however, none of these approaches requires more than knowledge of local activities and the single global cost signal. Real neural circuits in the brain have mechanisms (e.g., diffusible neuromodulators) that appear to code the signals relevant to implementing those algorithms. In many cases, for example in reinforcement learning, the cost function, which is computed based on interaction with an unknown environment, cannot be differentiated directly, and an agent has no choice but to deploy clever twiddling to explore at some level of the system (Williams, <xref ref-type="bibr" rid="B466">1992</xref>).</p>
<p>Backpropagation, in contrast, works by computing the sensitivity of the cost function to each weight based on the layered structure of the system. The derivatives of the cost function with respect to the last layer can be used to compute the derivatives of the cost function with respect to the penultimate layer, and so on, all the way down to the earliest layers<xref ref-type="fn" rid="fn0008"><sup>8</sup></xref>. Backpropagation can be computed rapidly, and for a single input-output pattern, it exhibits no variance in its gradient estimate. The backpropagated gradient has no more noise for a large system than for a small system, so deep and wide architectures with great computational power can be trained efficiently.</p>
</sec>
<sec>
<title>2.2.2. Biologically plausible approximations of gradient descent</title>
<p>To permit biological learning with efficiency approaching that of machine learning methods, some provision for more sophisticated gradient propagation may be suspected. Contrary to what was once a common assumption, there are now many proposed &#x0201C;biologically plausible&#x0201D; mechanisms by which a neural circuit could implement optimization algorithms that, like backpropagation, can efficiently make use of the gradient. These include Generalized Recirculation (O&#x00027;Reilly, <xref ref-type="bibr" rid="B332">1996</xref>), Contrastive Hebbian Learning (Xie and Seung, <xref ref-type="bibr" rid="B476">2003</xref>), random feedback weights together with synaptic homeostasis (Lillicrap et al., <xref ref-type="bibr" rid="B262">2014</xref>; Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>), spike timing dependent plasticity (STDP) with iterative inference and target propagation (Bengio et al., <xref ref-type="bibr" rid="B38">2015a</xref>; Scellier and Bengio, <xref ref-type="bibr" rid="B383">2016</xref>), complex neurons with backpropagating action-potentials (K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B229">2000</xref>), and others (Balduzzi et al., <xref ref-type="bibr" rid="B29">2014</xref>). While these mechanisms differ in detail, they all invoke feedback connections that carry error phasically. Learning occurs by comparing a prediction with a target, and the prediction error is used to drive top-down changes in bottom-up activity.</p>
<p>As an example, consider O&#x00027;Reilly&#x00027;s temporally eXtended Contrastive Attractor Learning (XCAL) algorithm (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B336">2012</xref>, <xref ref-type="bibr" rid="B337">2014b</xref>). Suppose we have a multilayer neural network with an input layer, an output layer, and a set of hidden layers in between. O&#x00027;Reilly showed that the same functionality as backpropagation can be implemented by a bidirectional network with the same weights but symmetric connections. After computing the outputs using the forward connections only, we set the output neurons to the values they should have. The dynamics of the network then cause the hidden layers&#x00027; activities to evolve toward a stable attractor state linking input to output. The XCAL algorithm performs a type of local modified Hebbian learning at each synapse in the network during this process (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B336">2012</xref>). The XCAL Hebbian learning rule compares the local synaptic activity (pre x post) during the early phase of this settling (before the attractor state is reached) to the final phase (once the attractor state has been reached), and adjusts the weights in a way that should make the early phase reflect the later phase more closely. These contrastive Hebbian learning methods even work when the connection weights are not precisely symmetric (O&#x00027;Reilly, <xref ref-type="bibr" rid="B332">1996</xref>). XCAL has been implemented in biologically plausible conductance-based neurons and basically implements the backpropagation of error approach.</p>
<p>Approximations to backpropagation could also be enabled by the millisecond-scale timing of of neural activities (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B337">2014b</xref>). Spike timing dependent plasticity (STDP) (Markram et al., <xref ref-type="bibr" rid="B288">1997</xref>), for example, is a feature of some neurons in which the sign of the synaptic weight change depends on the precise millisecond-scale relative timing of pre-synaptic and post-synaptic spikes. This is conventionally interpreted as Hebbian plasticity that measures the potential for a causal relationship between the pre-synaptic and post-synaptic spikes: a pre-synaptic spike could have contributed to causing a post-synaptic spike only if it occurs shortly beforehand<xref ref-type="fn" rid="fn0009"><sup>9</sup></xref>. To enable a backpropagation mechanism, Hinton has suggested an alternative interpretation: that neurons could encode the types of error derivatives needed for backpropagation in the temporal derivatives of their firing rates (Hinton, <xref ref-type="bibr" rid="B184">2007</xref>, <xref ref-type="bibr" rid="B185">2016</xref>). STDP then corresponds to a learning rule that is sensitive to these error derivatives (Xie and Seung, <xref ref-type="bibr" rid="B475">2000</xref>; Bengio et al., <xref ref-type="bibr" rid="B40">2015b</xref>). In other words, in an appropriate network context, STDP learning could give rise to a biological implementation of backpropagation<xref ref-type="fn" rid="fn0010"><sup>10</sup></xref>.</p>
<p>Another possible mechanism, by which biological neural networks could approximate backpropagation, is &#x0201C;feedback alignment&#x0201D; (Lillicrap et al., <xref ref-type="bibr" rid="B262">2014</xref>; Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>). There, the feedback pathway in backpropagation, by which error derivatives at a layer are computed from error derivatives at the subsequent layer, is replaced by a set of random feedback connections, with no dependence on the forward weights. Subject to the existence of a synaptic normalization mechanism and approximate sign-concordance between the feedforward and feedback connections (Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>), this mechanism of computing error derivatives works nearly as well as backpropagation on a variety of tasks. In effect, the forward weights are able to adapt to bring the network into a regime in which the random backwards weights actually carry the information that is useful for approximating the gradient. This is a remarkable and surprising finding, and is indicative of the fact that our understanding of gradient descent optimization, and specifically of the mechanisms by which backpropagation itself functions, are still incomplete. In neuroscience, meanwhile, we find feedback connections almost wherever we find feed-forward connections, and their role is the subject of diverse theories (Callaway, <xref ref-type="bibr" rid="B71">2004</xref>; Maass et al., <xref ref-type="bibr" rid="B273">2007</xref>). It should be noted that feedback alignment as such does not specify exactly how neurons represent and make use of the error signals; it only relaxes a constraint on the transport of the error signals. Thus, feedback alignment is more a primitive that can be used in fully biological (approximate) implementations of backpropagation, than a fully biological implementation in its own right. As such, it may be possible to incorporate it into several of the other schemes discussed here.</p>
<p>The above &#x0201C;biological&#x0201D; implementations of backpropagation still lack some key aspects of biological realism. For example, in the brain, neurons tend to be either excitatory or inhibitory but not both, whereas in artificial neural networks a single neuron may send both excitatory and inhibitory signals to its downstream neurons. Fortunately, this constraint is unlikely to limit the functions that can be learned (Parisien et al., <xref ref-type="bibr" rid="B340">2008</xref>; Tripp and Eliasmith, <xref ref-type="bibr" rid="B440">2016</xref>). Other biological considerations, however, need to be looked at in more detail: the highly recurrent nature of biological neural networks, which show rich dynamics in time, and the fact that most neurons in mammalian brains communicate via spikes. We now consider these two issues in turn.</p>
<sec>
<title>2.2.2.1. Temporal credit assignment:</title>
<p>The biological implementations of backpropagation proposed above, while applicable to feedforward networks, do not give a natural implementation of &#x0201C;backpropagation through time&#x0201D; (BPTT) (Werbos, <xref ref-type="bibr" rid="B461">1990</xref>) for recurrent networks, which is widely used in machine learning for training recurrent networks on sequential processing tasks. BPTT &#x0201C;unfolds&#x0201D; a recurrent network across multiple discrete time steps and then runs backpropagation on the unfolded network to assign credit to particular units at particular time steps<xref ref-type="fn" rid="fn0011"><sup>11</sup></xref>. While the network unfolding procedure of BPTT itself does not seem biologically plausible, to our intuition, it is unclear to what extent temporal credit assignment is truly needed (Ollivier and Charpiat, <xref ref-type="bibr" rid="B326">2015</xref>) for learning particular temporally extended tasks.</p>
<p>If the system is given access to appropriate memory stores and representations (Buonomano and Merzenich, <xref ref-type="bibr" rid="B63">1995</xref>; Gershman et al., <xref ref-type="bibr" rid="B139">2012</xref>, <xref ref-type="bibr" rid="B140">2014</xref>) of temporal context, this could potentially mitigate the need for temporal credit assignment as such&#x02014;in effect, memory systems could &#x0201C;spatialize&#x0201D; the problem of temporal credit assignment<xref ref-type="fn" rid="fn0012"><sup>12</sup></xref>. For example, memory networks (Weston et al., <xref ref-type="bibr" rid="B464">2014</xref>) store everything by default up to a certain buffer size, eliminating the need to perform credit assignment over the write-to-memory events, such that the network only needs to perform credit assignment over the read-from-memory events. In another example, certain network architectures that are superficially very deep, but which possess particular types of &#x0201C;skip connections,&#x0201D; can actually be seen as ensembles of comparatively shallow networks (Veit et al., <xref ref-type="bibr" rid="B449">2016</xref>); applied in the time domain, this could limit the need to propagate errors far backwards in time. Other, similar specializations or higher-levels of structure could, potentially, further ease the burden on credit assignment.</p>
<p>Can generic recurrent networks perform temporal credit assignment in in a way that is more biologically plausible than BPTT? Indeed, new discoveries are being made about the capacity for supervised learning in continuous-time recurrent networks with more realistic synapses and neural integration properties. In internal FORCE learning (Sussillo and Abbott, <xref ref-type="bibr" rid="B419">2009</xref>), internally generated random fluctuations inside a chaotic recurrent network are adjusted to provide feedback signals that drive weight changes internal to the network while the outputs are clamped to desired patterns. This is made possible by a learning procedure that rapidly adjusts the network output to a state where it is close to the clamped values, and exerts continuous control to keep this difference small throughout the learning process<xref ref-type="fn" rid="fn0013"><sup>13</sup></xref>. This procedure is able to control and exploit the chaotic dynamical patterns that are spontaneously generated by the network.</p>
<p>Werbos has proposed in his &#x0201C;error critic&#x0201D; that an online approximation to BPTT can be achieved by learning to predict the backward-through-time gradient signal (costate) in a manner analogous to the prediction of value functions in reinforcement learning (Werbos and Si, <xref ref-type="bibr" rid="B462">2004</xref>). This kind of idea was recently applied in (Jaderberg et al., <xref ref-type="bibr" rid="B206">2016</xref>) to allow decoupling of different parts of a network during training and to facilitate backpropagation through time. Broadly, we are only beginning to understand how neural activity can itself represent the time variable (Xu et al., <xref ref-type="bibr" rid="B478">2014</xref>; Finnerty et al., <xref ref-type="bibr" rid="B121">2015</xref>)<xref ref-type="fn" rid="fn0014"><sup>14</sup></xref>, and how recurrent networks can learn to generate trajectories of population activity over time (Liu and Buonomano, <xref ref-type="bibr" rid="B264">2009</xref>). Moreover, as we discuss below, a number of cortical models also propose means, other than BPTT, by which networks could be trained on sequential prediction tasks, even in an online fashion (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B337">2014b</xref>; Cui et al., <xref ref-type="bibr" rid="B89">2015</xref>; Brea et al., <xref ref-type="bibr" rid="B55">2016</xref>). A broad range of ideas can be used to approximate BPTT in more realistic ways.</p>
</sec>
<sec>
<title>2.2.2.2. Spiking networks:</title>
<p>It has been difficult to apply gradient descent learning directly to spiking neural networks<xref ref-type="fn" rid="fn0015"><sup>15</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0016"><sup>16</sup></xref>, although there do exist learning rules for doing so in specific representational contexts and network structures (Bekolay et al., <xref ref-type="bibr" rid="B35">2013</xref>). A number of optimization procedures have been used to generate, indirectly, spiking networks which can perform complex tasks, by performing optimization on a continuous representation of the network dynamics and embedding variables into high-dimensional spaces with many spiking neurons representing each variable (Thalmeier et al., <xref ref-type="bibr" rid="B435">2015</xref>; Abbott et al., <xref ref-type="bibr" rid="B1">2016</xref>; DePasquale et al., <xref ref-type="bibr" rid="B98">2016</xref>; Komer and Eliasmith, <xref ref-type="bibr" rid="B227">2016</xref>). The use of recurrent connections with multiple timescales can remove the need for backpropagation in the direct training of spiking recurrent networks (Bourdoukan and Den&#x000E8;ve, <xref ref-type="bibr" rid="B53">2015</xref>). Fast connections maintain the network in a state where slow connections have local access to a global error signal. While the biological realism of these methods is still unknown, they all allow connection weights to be learned in spiking networks.</p>
<p>These and other novel learning procedures illustrate the fact that we are only beginning to understand the connections between the temporal dynamics of biologically realistic networks, and mechanisms of temporal and spatial credit assignment. Nevertheless, we argue here that existing evidence suggests that biologically plausible neural networks can solve these problems&#x02014;in other words, it is possible to efficiently optimize complex functions of temporal history in the context of spiking networks of biologically realistic neurons. In any case, there is little doubt that spiking recurrent networks using realistic population coding schemes can, with an appropriate choice of connection weights, compute complicated, cognitively relevant functions<xref ref-type="fn" rid="fn0017"><sup>17</sup></xref>. The question is how the developing brain efficiently learns such complex functions.</p>
</sec>
</sec>
</sec>
<sec>
<title>2.3. Other principles for biological learning</title>
<p>The brain has mechanisms and structures that could support learning mechanisms different from typical gradient-based optimization algorithms employed in artificial neural networks.</p>
<sec>
<title>2.3.1. Exploiting biological neural mechanisms</title>
<p>The complex physiology of individual biological neurons may not only help explain how some form of efficient gradient descent could be implemented within the brain, but also could provide mechanisms for learning that go beyond backpropagation. This suggests that the brain may have discovered mechanisms of credit assignment quite different from those dreamt up by machine learning.</p>
<p>One such biological primitive is dendritic computation, which could impact prospects for learning algorithms in several ways. First, real neurons are highly nonlinear (Antic et al., <xref ref-type="bibr" rid="B13">2010</xref>), with the dendrites of each <italic>single</italic> neuron implementing<xref ref-type="fn" rid="fn0018"><sup>18</sup></xref> something computationally similar to a three-layer neural network (Mel, <xref ref-type="bibr" rid="B297">1992</xref>)<xref ref-type="fn" rid="fn0019"><sup>19</sup></xref>. Individual neurons thus should not be regarded as single &#x0201C;nodes&#x0201D; but as multi-component sub-networks. Second, when a neuron spikes, its action potential propagates back from the soma into the dendritic tree. However, it propagates more strongly into the branches of the dendritic tree that have been active (Williams and Stuart, <xref ref-type="bibr" rid="B468">2000</xref>), potentially simplifying the problem of credit assignment (K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B229">2000</xref>). Third, neurons can have multiple somewhat independent dendritic compartments, as well as a somewhat independent somatic compartment, which means that the neuron should be thought of as storing more than one variable. Thus, there is the possibility for a neuron to store both its activation itself, and the error derivative of a cost function with respect to its activation, as required in backpropagation, and biological implementations of backpropagation based on this principle have been proposed (K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B230">2001</xref>; Schiess et al., <xref ref-type="bibr" rid="B384">2016</xref>)<xref ref-type="fn" rid="fn0020"><sup>20</sup></xref>. Overall, the implications of dendritic computation for credit assignment in deep networks are only beginning to be considered<xref ref-type="fn" rid="fn0021"><sup>21</sup></xref>. But it is clear that the types of bi-directional, non-linear, multi-variate interactions that are possible <italic>inside</italic> a single neuron could support gradient descent learning or other powerful optimization mechanisms.</p>
<p>Beyond dendritic computation, diverse mechanisms (Marblestone and Boyden, <xref ref-type="bibr" rid="B281">2014</xref>) like retrograde (post-synaptic to pre-synaptic) signals using cannabinoids (Wilson and Nicoll, <xref ref-type="bibr" rid="B469">2001</xref>), or rapidly-diffusing gases such as nitric oxide (Arancio et al., <xref ref-type="bibr" rid="B14">1996</xref>), are among many that could enable learning rules that go beyond conventional conceptions of backpropagation. Harris has suggested (Harris, <xref ref-type="bibr" rid="B166">2008</xref>; Lewis and Harris, <xref ref-type="bibr" rid="B259">2014</xref>) how slow, retroaxonal (i.e., from the outgoing synapses back to the parent cell body) transport of molecules like neurotrophins could allow neural networks to implement an analog of an exchangeable currency in economics, allowing networks to self-organize to efficiently provide information to downstream &#x0201C;consumer&#x0201D; neurons that are trained via faster and more direct error signals. The existence of these diverse mechanisms may call into question traditional, intuitive notions of &#x0201C;biological plausibility&#x0201D; for learning algorithms.</p>
<p>Another potentially important biological primitive is neuromodulation. The same neuron or circuit can exhibit different input-output responses and plasticity depending on a global circuit state, as reflected by the concentrations of various <italic>neuromodulators</italic> like dopamine, serotonin, norepinephrine, acetylcholine, and hundreds of different neuropeptides such as opiods (Bargmann, <xref ref-type="bibr" rid="B30">2012</xref>; Bargmann and Marder, <xref ref-type="bibr" rid="B31">2013</xref>). These modulators interact in complex and cell-type-specific ways to influence circuit function. Interactions with glial cells also play a role in neural signaling and neuromodulation, leading to the concept of &#x0201C;tripartite&#x0201D; synapses that include a glial contribution (Perea et al., <xref ref-type="bibr" rid="B344">2009</xref>). Modulation could have many implications for learning. First, modulators can be used to gate synaptic plasticity on and off selectively in different areas and at different times, allowing precise, rapidly updated orchestration of where and when cost functions are applied. Furthermore, it has been argued that a single neural circuit can be thought of as multiple overlapping circuits with modulation switching between them (Bargmann, <xref ref-type="bibr" rid="B30">2012</xref>; Bargmann and Marder, <xref ref-type="bibr" rid="B31">2013</xref>). In a learning context, this could potentially allow sharing of synaptic weight information between overlapping circuits. Dayan (<xref ref-type="bibr" rid="B93">2012</xref>) discusses further computational aspects of neuromodulation. Overall, neuromodulation seems to expand the range of possible algorithms that could be used for optimization.</p>
</sec>
<sec>
<title>2.3.2. Learning in the cortical sheet</title>
<p>A number of models attempt to explain cortical learning on the basis of specific architectural features of the 6-layered cortical sheet. These models generally agree that a primary function of the cortex is some form of unsupervised learning via prediction (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B337">2014b</xref>; Brea et al., <xref ref-type="bibr" rid="B55">2016</xref>)<xref ref-type="fn" rid="fn0022"><sup>22</sup></xref>. Some cortical learning models are explicit attempts to map cortical structure onto the framework of message-passing algorithms for Bayesian inference (Lee and Mumford, <xref ref-type="bibr" rid="B250">2003</xref>; Dean, <xref ref-type="bibr" rid="B94">2005</xref>; George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>), while others start with particular aspects of cortical neurophysiology and seek to explain those in terms of a learning function, or in terms of a computational function, e.g., hierarchical clustering (Rodriguez et al., <xref ref-type="bibr" rid="B367">2004</xref>). For example, the nonlinear and dynamical properties of cortical pyramidal neurons&#x02014;the principal excitatory neuron type in cortex (Shepherd, <xref ref-type="bibr" rid="B396">2014</xref>)&#x02014;are of particular interest here, especially because these neurons have multiple dendritic zones that are targeted by different kinds of projections, which may allow the pyramidal neuron to make comparisons of top-down and bottom-up inputs<xref ref-type="fn" rid="fn0023"><sup>23</sup></xref>.</p>
<p>Other aspects of the laminar cortical architecture could be crucial to how the brain implements learning. Local inhibitory neurons targeting particular dendritic compartments of the L5 pyramidal could be used to exert precise control over when and how the relevant feedback signals and associative mechanisms are utilized. Notably, local inhibitory networks could also give rise to competition (Petrov et al., <xref ref-type="bibr" rid="B345">2010</xref>) between different representations in the cortex, perhaps allowing one cortical column to suppress others nearby, or perhaps even to send more sophisticated messages to gate the state transitions of its neighbors (Bach and Herger, <xref ref-type="bibr" rid="B22">2015</xref>). Moreover, recurrent connectivity with the thalamus, structured bursts of spiking, and cortical oscillations (not to mention other mechanisms like neuromodulation) could control the storage of information over time, to facilitate learning based on temporal prediction. These concepts begin to suggest preliminary, exploratory models for how the detailed anatomy and physiology of the cortex could be interpreted within a machine-learning framework that goes beyond backpropagation. But these are early days: we still lack detailed structural/molecular and functional maps of even a single local cortical microcircuit.</p>
</sec>
<sec>
<title>2.3.3. One-shot learning</title>
<p>Human learning is often one-shot: it can take just a single exposure to a stimulus to never forget it, as well as to generalize from it to new examples. One way of allowing networks to have such properties is what is described by I-theory, in the context of learning invariant representations for object recognition (Anselmi et al., <xref ref-type="bibr" rid="B12">2015</xref>). Instead of training via gradient descent, image templates are stored in the weights of simple-complex cell networks while objects undergo transformations, similar to the use of stored templates in HMAX (Serre et al., <xref ref-type="bibr" rid="B391">2007</xref>). The theories then aim to show that you can invariantly and discriminatively represent objects using a single sample, even of a new class (Anselmi et al., <xref ref-type="bibr" rid="B12">2015</xref>)<xref ref-type="fn" rid="fn0024"><sup>24</sup></xref>.</p>
<p>Additionally, the nervous system may have a way of quickly storing and replaying sequences of events. This would allow the brain to move an item from episodic memory into a long-term memory stored in the weights of a cortical network (Ji and Wilson, <xref ref-type="bibr" rid="B214">2007</xref>), by replaying the memory over and over. This solution effectively uses many iterations of weight updating to fully learn a single item, even if one has only been exposed to it once. Alternatively, the brain could rapidly store an episodic memory and then retrieve it later without the need to perform slow gradient updates, which has proven to be useful for fast reinforcement learning in scenarios with limited available data (Blundell et al., <xref ref-type="bibr" rid="B47">2016</xref>).</p>
<p>Finally, higher-level systems in the brain may be able to implement Bayesian learning of sequential programs, which is a powerful means of one-shot learning (Lake et al., <xref ref-type="bibr" rid="B243">2015</xref>). This type of cognition likely relies on an interaction between multiple brain areas such as the prefrontal cortex and basal ganglia.</p>
<p>These potential substrates of one-shot learning rely on mechanisms other than simple gradient descent. It should be noted, though, that recent architectural advances, including specialized spatial attention and feedback mechanisms (Rezende et al., <xref ref-type="bibr" rid="B363">2016</xref>), as well as specialized memory mechanisms (Santoro et al., <xref ref-type="bibr" rid="B381">2016</xref>), do allow some types of one-shot generalization to be driven by backpropagation-based learning.</p>
</sec>
<sec>
<title>2.3.4. Active learning</title>
<p>Human learning is often active and deliberate. It seems likely that, in human learning, actions are chosen so as to generate interesting training examples, and sometimes also to test specific hypotheses. Such ideas of active learning and &#x0201C;child as scientist&#x0201D; go back to Piaget and have been elaborated more recently (Gopnik et al., <xref ref-type="bibr" rid="B150">2000</xref>). We want our learning to be based on maximally informative samples, and active querying of the environment (or of internal subsystems) provides a way route to this.</p>
<p>At some level of organization, of course, it would seem useful for a learning system to develop explicit representations of its uncertainty, since this can be used to guide the system to actively seek the information that would reduce its uncertainty most quickly. Moreover, there are population coding mechanisms that could support explicit probabilistic computations (Zemel and Dayan, <xref ref-type="bibr" rid="B487">1997</xref>; Sahani and Dayan, <xref ref-type="bibr" rid="B379">2003</xref>; Rao, <xref ref-type="bibr" rid="B359">2004</xref>; Ma et al., <xref ref-type="bibr" rid="B271">2006</xref>; Eliasmith and Martens, <xref ref-type="bibr" rid="B107">2011</xref>; Gershman and Beck, <xref ref-type="bibr" rid="B138">2016</xref>). Yet it is unclear to what extent and at what levels the brain uses an explicitly probabilistic framework, or to what extent probabilistic computations are emergent from other learning processes (Orhan and Ma, <xref ref-type="bibr" rid="B338">2016</xref>)<xref ref-type="fn" rid="fn0025"><sup>25</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0026"><sup>26</sup></xref>.</p>
<p>Standard gradient descent does not incorporate any such adaptive sampling mechanism, e.g., it does not deliberately sample data so as to maximally reduce its uncertainty. Interestingly, however, stochastic gradient descent can be used to generate a system that samples adaptively (Alain et al., <xref ref-type="bibr" rid="B5">2015</xref>; Bouchard et al., <xref ref-type="bibr" rid="B52">2015</xref>). In other words, a system can learn, by gradient descent, how to choose its own input data samples in order to learn most quickly from them by gradient descent.</p>
<p>Ideally, the learner learns to choose actions that will lead to the largest improvements in its prediction or data compression performance (Schmidhuber, <xref ref-type="bibr" rid="B385">2010</xref>). In Schmidhuber (<xref ref-type="bibr" rid="B385">2010</xref>), this is done in the framework of reinforcement learning, and incorporates a mechanisms for the system to measure its own rate of learning. In other words, it is possible to reinforcement-learn a policy for selecting the most interesting inputs to drive learning. Adaptive sampling methods are also known in reinforcement learning that can achieve optimal Bayesian exploration of Markov Decision Process environments (Sun et al., <xref ref-type="bibr" rid="B417">2011</xref>; Guez et al., <xref ref-type="bibr" rid="B157">2012</xref>).</p>
<p>These approaches achieve optimality in an arbitrary, abstract environment. But of course, evolution may also encode its implicit knowledge of the organism&#x00027;s natural environment, the behavioral goals of the organism, and the developmental stages and processes which occur inside the organism, as priors or heuristics<xref ref-type="fn" rid="fn0027"><sup>27</sup></xref> which would further constrain the types of adaptive sampling that are optimal in practice. For example, simple heuristics like seeking certain perceptual signatures of novelty, or more complex heuristics like monitoring situations that other people seem to find interesting, might be good ways to bias sampling of the environment so as to learn more quickly. Other such heuristics might be used to give internal brain systems the types of training data that will be most useful to those particular systems at any given developmental stage.</p>
<p>We are only beginning to understand how active learning might be implemented in the brain. We speculate that multiple mechanisms, specialized to different brain systems and spatio-temporal scales, could be involved. The above examples suggest that at least some such mechanisms could be understood from the perspective of optimizing cost functions.</p>
</sec>
</sec>
<sec>
<title>2.4. Differing biological requirements for supervised and reinforcement learning</title>
<p>We have suggested ways in which the brain could implement learning mechanisms of comparable power to backpropagation. But in many cases, the system may be more limited by the available training signals than by the optimization process itself. In machine learning, one distinguishes supervised learning, reinforcement learning and unsupervised learning, and the training data limitation manifests differently in each case.</p>
<p>Both supervised and reinforcement learning require some form of teaching signal, but the nature of the teaching signal in supervised learning is different from that in reinforcement learning. In supervised learning, the trainer provides the entire vector of errors for the output layer and these are back-propagated to compute the gradient: a locally optimal direction in which to update all of the weights of a potentially multi-layer and/or recurrent network. In reinforcement learning, however, the trainer provides a scalar evaluation signal, but this is not sufficient to derive a low-variance gradient. Hence, some form of trial and error twiddling must be used to discover how to increase the evaluation signal. Consequently, reinforcement learning is generally much less efficient than supervised learning.</p>
<p>Reinforcement learning in shallow networks is simple to implement biologically. For reinforcement learning of a deep network to be biologically plausible, however, we need a more powerful learning mechanism, since we are learning based on a more limited evaluation signal than in the supervised case: we do not have the full target pattern to train toward. Nevertheless, approximations of gradient descent can be achieved in this case, and there are cases in which the scalar evaluation signal of reinforcement learning can be used to efficiently update a multi-layer network by gradient descent. The &#x0201C;attention-gated reinforcement learning&#x0201D; (AGREL) networks of Stanisor et al. (<xref ref-type="bibr" rid="B411">2013</xref>), Brosch et al. (<xref ref-type="bibr" rid="B59">2015</xref>), and Roelfsema and van Ooyen (<xref ref-type="bibr" rid="B369">2005</xref>), and variants like KickBack (Balduzzi, <xref ref-type="bibr" rid="B28">2014</xref>), give a way to compute an approximation to the full gradient in a reinforcement learning context using a feedback-based attention mechanism for credit assignment within the multi-layer network. The feedback pathway, together with a diffusible reward signal, together gate plasticity. For networks with more than three layers, this gives rise to a model based on columns containing parallel feedforward and feedback pathways (Roelfsema and van Ooyen, <xref ref-type="bibr" rid="B369">2005</xref>), and for recurrent networks that settle into attractor states it gives a reinforcement-trained version (Brosch et al., <xref ref-type="bibr" rid="B59">2015</xref>) of the Almeida/Pineda recurrent backpropagation algorithm (Pineda, <xref ref-type="bibr" rid="B350">1987</xref>). The process is still not as efficient or generic as backpropagation, but it seems that this form of feedback can make reinforcement learning in multi-layer networks more efficient than a naive node perturbation or weight perturbation approach.</p>
<p>The machine-learning field has recently been tackling the question of credit assignment in deep reinforcement learning. Deep Q-learning (Mnih et al., <xref ref-type="bibr" rid="B313">2015</xref>) demonstrates reinforcement learning in a deep network, wherein most of the network is trained via backpropagation. In regular Q learning, we define a function Q, which estimates the best possible sum of future rewards (the return) if we are in a given state and take a given action. In deep Q learning, this function is approximated by a neural network that, in effect, estimates action-dependent returns in a given state. The network is trained using backpropagation of local errors in Q estimation, using the fact that the return decomposes into the current reward plus the discounted estimate of future return at the next moment. During training, as the agent acts in the environment, a series of loss functions is generated at each step, defining target patterns that can be used as the supervision signal for backpropagation. As Q is a highly nonlinear function of the state, tricks are needed to make deep Q learning efficient and stable, including experience replay and a particular type of mini-batch training. It is also necessary to store the outputs from the previous iteration (or clone the entire network) in evaluating the loss function for the subsequent iteration<xref ref-type="fn" rid="fn0028"><sup>28</sup></xref>.</p>
<p>This process for generating learning targets provides a kind of bridge between reinforcement learning and efficient backpropagation-based gradient descent learning<xref ref-type="fn" rid="fn0029"><sup>29</sup></xref>. Importantly, only temporally local information is needed making the approach relatively compatible with what we know about the nervous system.</p>
<p>Even given these advances, a key remaining issue in reinforcement learning is the problem of long timescales, e.g., learning the many small steps needed to navigate from London to Chicago. Many of the formal guarantees of reinforcement learning (Williams and Baird, <xref ref-type="bibr" rid="B467">1993</xref>), for example, suggest that the difference between an optimal policy and the learned policy becomes increasingly loose as the discount factor shifts to take into account reward at longer timescales. Although the degree of optimality of human behavior is unknown, people routinely engage in adaptive behaviors that can take hours or longer to carry out, by using specialized processes like <italic>prospective memory</italic> to &#x0201C;remember to remember&#x0201D; relevant variables at the right times, permitting extremely long timescales of coherent action. Machine learning has not yet developed methods to deal with such a wide range of timescales and scopes of hierarchical action. Below we discuss ideas of hierarchical reinforcement learning that may make use of callable procedures and sub-routines, rather than operating explicitly in a time domain.</p>
<p>As we will discuss below, some form of deep reinforcement learning may be used by the brain for purposes beyond optimizing global rewards, including the training of local networks based on diverse internally generated cost functions. Scalar reinforcement-like signals are easy to compute, and easy to deliver to other areas, making them attractive mechanistically. If the brain does employ internally computed scalar reward-like signals as a basis for cost functions, it seems likely that it will have found an efficient means of reinforcement-based training of deep networks, but it is an open question whether an analog of deep Q networks, AGREL, or some other mechanism entirely, is used in the brain for this purpose. Moreover, as we will discuss further below, it is possible that reinforcement-type learning is made more efficient in the context of specialized brain systems like short term memories, replay mechanisms, and hierarchically organized control systems. These specialized systems could reduce reliance on a need for powerful credit assignment mechanisms for reinforcement learning. Finally, if the brain uses a diversity of scalar reward-like signals to implement different cost functions, then it may need to mediate delivery of those signals via a comparable diversity of molecular substrates. The great diversity of neuromodulatory signals, e.g., neuropeptides, in the brain (Bargmann, <xref ref-type="bibr" rid="B30">2012</xref>; Bargmann and Marder, <xref ref-type="bibr" rid="B31">2013</xref>) makes such diversity quite plausible, and moreover, the brain may have found other, as yet unknown, mechanisms of diversifying reward-like signaling pathways and enabling them to act independently of one another.</p>
</sec>
</sec>
<sec id="s3">
<title>3. The cost functions are diverse across brain areas and time</title>
<p>In the last section, we argued that the brain can optimize functions. This raises the question of what functions it optimizes. Of course, in the brain, a cost function will itself be created (explicitly or implicitly) by a neural network shaped by the genome. Thus, the cost function used to train a given sub-network in the brain is a key innate property that can be built into the system by evolution. It may be much cheaper in biological terms to specify a cost function that allows the rapid learning of the solution to a problem than to specify the solution itself.</p>
<p>In Hypothesis 2, we proposed that the brain optimizes not a single &#x0201C;end-to-end&#x0201D; cost function, but rather a diversity of internally generated cost functions specific to particular brain functions<xref ref-type="fn" rid="fn0030"><sup>30</sup></xref>. To understand how and why the brain may use a diversity of cost functions, it is important to distinguish the differing types of cost functions that would be needed for supervised, unsupervised and reinforcement learning. We can also seek to identify types of cost functions that the brain may need to generate from a functional perspective, and how each may be implemented as supervised, unsupervised, reinforcement-based or hybrid systems.</p>
<sec>
<title>3.1. How cost functions may be represented and applied</title>
<p>What additional circuitry is required to actually impose a cost function on an optimizing network? In the most familiar case, supervised learning may rely on computing a vector of errors at the output of a network, which will rely on some comparator circuitry<xref ref-type="fn" rid="fn0031"><sup>31</sup></xref> to compute the difference between the network outputs and the target values. This difference could then be backpropagated to earlier layers. An alternative way to impose a cost function is to &#x0201C;clamp&#x0201D; the output of the network, forcing it to occupy a desired target state. Such clamping is actually assumed in some of the putative biological implementations of backpropagation described above, such as XCAL and target propagation. Alternatively, as described above, scalar reinforcement signals are attractive as internally-computed cost functions, but using them in deep networks requires special mechanisms for credit assignment.</p>
<p>In unsupervised learning, cost functions may not take the form of externally supplied training or error signals, but rather can be built into the dynamics inherent to the network itself, i.e., there may be no need for a <italic>separate</italic> circuit to compute and impose a cost function on the network. For example, specific spike-timing-dependent and homeostatic plasticity rules have been shown to give rise to gradient descent on a prediction error in recurrent neural networks (Galtier and Wainrib, <xref ref-type="bibr" rid="B134">2013</xref>). Thus, specific unsupervised objectives could be implemented implicitly through specific local network dynamics<xref ref-type="fn" rid="fn0032"><sup>32</sup></xref> and plasticity rules inside a network without explicit computation of cost function, nor explicit propagation of error derivatives.</p>
<p>Alternatively, explicit cost functions could be computed, delivered to an optimizing network, and used for unsupervised learning, following a variety of principles being discovered in machine learning (e.g., Radford et al., <xref ref-type="bibr" rid="B356">2015</xref>; Lotter et al., <xref ref-type="bibr" rid="B266">2015</xref>). These networks rely on backpropagation as the sole learning rule, and typically find a way to encode the desired cost function into the error derivatives which are backpropagated. For example, prediction errors naturally give rise to error signals for unsupervised learning, as do reconstruction errors in autoencoders, and these error signals can also be augmented with additional penalty or regularization terms that enforce objectives like sparsity or continuity, as described below. Then these error derivatives can be propagated throughout the network via standard backpropagation. In such systems, the objective function and the optimization mechanism can thus be mixed and matched modularly. In the next sections, we elaborate on these and other means of specifying and delivering cost functions in different learning contexts.</p>
</sec>
<sec>
<title>3.2. Cost functions for unsupervised learning</title>
<p>There are many objectives that can be optimized in an unsupervised context, to accomplish different kinds of functions or guide a network to form particular kinds of representations.</p>
<sec>
<title>3.2.1. Matching the statistics of the input data using generative models</title>
<p>In one common form of unsupervised learning, higher brain areas attempt to produce samples that are statistically similar to those actually seen in lower layers. For example, the wake-sleep algorithm (Hinton et al., <xref ref-type="bibr" rid="B186">1995</xref>) requires the sleep mode to sample potential data points whose distribution should then match the observed distribution. Unsupervised pre-training of deep networks is an instance of this (Erhan and Manzagol, <xref ref-type="bibr" rid="B111">2009</xref>), typically making use of a stacked auto-encoder framework. Similarly, in target propagation (Bengio, <xref ref-type="bibr" rid="B36">2014</xref>), a top-down circuit, together with lateral information, has to produce data that directs the local learning of a bottom-up circuit and vice-versa. Ladder autoencoders make use of lateral connections and local noise injection to introduce an unsupervised cost function, based on internal reconstructions, that can be readily combined with supervised cost functions defined on the networks top layer outputs (Valpola, <xref ref-type="bibr" rid="B445">2015</xref>). Compositional generative models generate a scene from discrete combinations of template parts and their transformations (Wang and Yuille, <xref ref-type="bibr" rid="B453">2014</xref>), in effect performing a rendering of a scene based on its structural description. Hinton and colleagues have also proposed cortical &#x0201C;capsules&#x0201D; (Hinton et al., <xref ref-type="bibr" rid="B189">2011</xref>; Tang et al., <xref ref-type="bibr" rid="B427">2012</xref>, <xref ref-type="bibr" rid="B428">2013</xref>) for compositional inverse rendering. The network can thus implement a statistical goal that embodies some understanding of the way that the world produces samples<xref ref-type="fn" rid="fn0033"><sup>33</sup></xref>.</p>
<p>Learning rules for generative models have historically involved local message passing of a form quite different from backpropagation, e.g., in a multi-stage process that first learns one layer at a time and then fine-tunes via the wake-sleep algorithm (Hinton et al., <xref ref-type="bibr" rid="B187">2006</xref>). Message-passing implementations of probabilistic inference have also been proposed as an explanation and generalization of deep convolutional networks (Chen et al., <xref ref-type="bibr" rid="B74">2014</xref>; Patel et al., <xref ref-type="bibr" rid="B342">2015</xref>). Various mappings of such processes onto neural circuitry have been attempted (George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>; Lee and Yuille, <xref ref-type="bibr" rid="B249">2011</xref>; Sountsov and Miller, <xref ref-type="bibr" rid="B407">2015</xref>), and related models (Makin et al., <xref ref-type="bibr" rid="B278">2013</xref>, <xref ref-type="bibr" rid="B277">2016</xref>) have been used to account for optimal multi-sensory integration in the brain. Feedback connections tend to terminate in distinct layers of cortex relative to the feedforward ones (Felleman and Van Essen, <xref ref-type="bibr" rid="B114">1991</xref>; Callaway, <xref ref-type="bibr" rid="B71">2004</xref>) making the idea of separate but interacting networks for recognition and generation potentially attractive<xref ref-type="fn" rid="fn0034"><sup>34</sup></xref>. Interestingly, such sub-networks might even be part of the same neuron and map onto &#x0201C;apical&#x0201D; vs. &#x0201C;basal&#x0201D; parts of the dendritic tree (K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B230">2001</xref>; Urbanczik and Senn, <xref ref-type="bibr" rid="B444">2014</xref>).</p>
<p>Generative models can also be trained via backpropagation. Recent advances have shown how to perform variational approximations to Bayesian inference inside backpropagation-based neural networks (Kingma and Welling, <xref ref-type="bibr" rid="B224">2013</xref>), and how to exploit this to create generative models (Goodfellow et al., <xref ref-type="bibr" rid="B148">2014a</xref>; Gregor et al., <xref ref-type="bibr" rid="B153">2015</xref>; Radford et al., <xref ref-type="bibr" rid="B356">2015</xref>; Eslami et al., <xref ref-type="bibr" rid="B112">2016</xref>). Through either explicitly statistical or gradient descent based learning, the brain can thus obtain a probabilistic model that simulates features of the world.</p>
</sec>
<sec>
<title>3.2.2. Cost functions that approximate properties of the world</title>
<p>A perceiving system should exploit statistical regularities in the world that are not present in an arbitrary dataset or input distribution. For example, objects are sparse, at least in certain representations: there are far fewer objects than there are potential places in the world, and of all possible objects there is only a small subset visible at any given time. As such, we know that the output of an object recognition system must have sparse activations. Building the assumption of sparseness into simulated systems replicates a number of representational properties of the early visual system (Olshausen and Field, <xref ref-type="bibr" rid="B329">1997</xref>; Rozell et al., <xref ref-type="bibr" rid="B374">2008</xref>), and indeed the original paper on sparse coding obtained sparsity by gradient descent optimization of a cost function (Olshausen and Field, <xref ref-type="bibr" rid="B328">1996</xref>). A range of unsupervised machine learning techniques, such as the sparse autoencoders (Le et al., <xref ref-type="bibr" rid="B254">2012</xref>) used to discover cats in YouTube videos, build sparseness into neural networks. Building in such spatio-temporal sparseness priors should serve as an &#x0201C;inductive bias&#x0201D; (Mitchell, <xref ref-type="bibr" rid="B310">1980</xref>) that can accelerate learning.</p>
<p>But we know much more about the regularities of objects. As young babies, we already know (Bremner et al., <xref ref-type="bibr" rid="B57">2015</xref>) that objects tend to persist over time. The emergence or disappearance of an object from a region of space is a rare event. Moreover, object locations and configurations tend to be coherent in time. We can formulate this prior knowledge as a cost function, for example by penalizing representations which are not temporally continuous. This idea of continuity is used in a great number of artificial neural networks and related models (Wiskott and Sejnowski, <xref ref-type="bibr" rid="B471">2002</xref>; F&#x000F6;ldi&#x000E1;k, <xref ref-type="bibr" rid="B124">2008</xref>; Mobahi et al., <xref ref-type="bibr" rid="B314">2009</xref>). Imposing continuity within certain models gives rise to aspects of the visual system including complex cells (K&#x000F6;rding et al., <xref ref-type="bibr" rid="B231">2004</xref>), specific properties of visual invariance (Isik et al., <xref ref-type="bibr" rid="B202">2012</xref>), and even other representational properties such as the existence of place cells (Wyss et al., <xref ref-type="bibr" rid="B474">2006</xref>; Franzius et al., <xref ref-type="bibr" rid="B130">2007</xref>). Unsupervised learning mechanisms that maximize temporal coherence or slowness are increasingly used in machine learning<xref ref-type="fn" rid="fn0035"><sup>35</sup></xref>.</p>
<p>We also know that objects tend to undergo predictable sequences of transformations, and it is possible to build this assumption into unsupervised neural learning systems (George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>). The minimization of prediction error explains a number of properties of the nervous system (Friston and Stephan, <xref ref-type="bibr" rid="B132">2007</xref>; Huang and Rao, <xref ref-type="bibr" rid="B200">2011</xref>), and biologically plausible theories are available for how cortex could learn using prediction errors by exploiting temporal differences (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B337">2014b</xref>) or top-down feedback (George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>). In one implementation, a system can simply predict the next input delivered to the system and can then use the difference between the actual next input and the predicted next input as a full vectorial error signal for supervised gradient descent. Thus, rather than optimization of prediction error being implicitly implemented by the network dynamics, the prediction error is used as an explicit cost function in the manner of supervised learning, leading to error derivatives which can be back-propagated. Then, no special learning rules beyond simple backpropagation are needed. This approach has recently been advanced within machine learning (Lotter et al., <xref ref-type="bibr" rid="B266">2015</xref>, <xref ref-type="bibr" rid="B267">2016</xref>). Recently, combining such prediction-based learning with a specific gating mechanism has been shown to lead to unsupervised learning of disentangled representations (Whitney et al., <xref ref-type="bibr" rid="B465">2016</xref>). Neural networks can also be designed to learn to invert spatial transformations (Jaderberg et al., <xref ref-type="bibr" rid="B207">2015</xref>). Statistically describing transformations or sequences is thus an unsupervised way of learning representations.</p>
<p>Furthermore, there are multiple modalities of input to the brain. Each sensory modality is primarily connected to one part of the brain<xref ref-type="fn" rid="fn0036"><sup>36</sup></xref>. But higher levels of cortex in each modality are heavily connected to the other modalities. This can enable forms of self-supervised learning: with a developing visual understanding of the world we can predict its sounds, and then test those predictions with the auditory input, and vice versa. The same is true about multiple parts of the same modality: if we understand the left half of the visual field, it tells us an awful lot about the right. Indeed, we can use observations of one part of a visual scene to predict the contents of other parts (Noroozi and Favaro, <xref ref-type="bibr" rid="B324">2016</xref>; van den Oord et al., <xref ref-type="bibr" rid="B446">2016</xref>), and optimize a cost function that reflects the discrepancy. Maximizing mutual information is a natural way of improving learning (Becker and Hinton, <xref ref-type="bibr" rid="B34">1992</xref>; Mohamed and Rezende, <xref ref-type="bibr" rid="B316">2015</xref>), and there are many other ways in which multiple modalities or processing streams could mutually train one another. This way, each modality effectively produces training signals for the others<xref ref-type="fn" rid="fn0037"><sup>37</sup></xref>. Evidence from psychophysics suggests that some kind of training via detection of sensory conflicts may be occurring in children (Nardini et al., <xref ref-type="bibr" rid="B320">2010</xref>).</p>
</sec>
</sec>
<sec>
<title>3.3. Cost functions for supervised learning</title>
<p>In what cases might the brain use supervised learning, given that it requires the system to &#x0201C;already know&#x0201D; the exact target pattern to train toward? One possibility is that the brain can store records of states that led to good outcomes. For example, if a baby reaches for a target and misses, and then tries again and successfully hits the target, then the difference in the neural representations of these two tries reflects the direction in which the system should change. The brain could potentially use a comparator circuit to directly compute this vectorial difference in the neural population codes and then apply this difference vector as an error signal.</p>
<p>Another possibility is that the brain uses supervised learning to implement a form of &#x0201C;chunking,&#x0201D; i.e., a consolidation of something the brain already knows how to do: routines that are initially learned as multi-step, deliberative procedures could be compiled down to more rapid and automatic functions by using supervised learning to train a network to mimic the overall input-output behavior of the original multi-step process. Such a process is assumed to occur in cognitive models like ACT-R (Servan-Schreiber and Anderson, <xref ref-type="bibr" rid="B392">1990</xref>), and methods for compressing the knowledge in neural networks into smaller networks are also being developed (Ba and Caruana, <xref ref-type="bibr" rid="B24">2014</xref>). Thus supervised learning can be used to train a network to do in &#x0201C;one step&#x0201D; what would otherwise require long-range routing and sequential recruitment of multiple systems.</p>
</sec>
<sec>
<title>3.4. Repurposing reinforcement learning for diverse internal cost functions</title>
<p>Certain generalized forms of reinforcement learning may be ubiquitous throughout the brain. Such reinforcement signals may be repurposed to optimize diverse internal cost functions. These internal cost functions could be specified at least in part by genetics.</p>
<p>Some brain systems such as in the striatum appear to learn via some form of temporal difference reinforcement learning (Tesauro, <xref ref-type="bibr" rid="B434">1995</xref>; Foster et al., <xref ref-type="bibr" rid="B125">2000</xref>). This is reinforcement learning based on a global value function (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B335">2014a</xref>) that predicts total future reward or utility for the agent. Reward-driven signaling is not restricted to the striatum, and is present even in primary visual cortex (Chubykin et al., <xref ref-type="bibr" rid="B77">2013</xref>; Stanisor et al., <xref ref-type="bibr" rid="B411">2013</xref>). Remarkably, the reward signaling in primary visual cortex is mediated in part by glial cells (Takata et al., <xref ref-type="bibr" rid="B425">2011</xref>), rather than neurons, and involves the neurotransmitter acetylcholine (Chubykin et al., <xref ref-type="bibr" rid="B77">2013</xref>; Hangya et al., <xref ref-type="bibr" rid="B163">2015</xref>). On the other hand, some studies have suggested that visual cortex learns the basics of invariant object recognition in the absence of reward (Li and Dicarlo, <xref ref-type="bibr" rid="B263">2012</xref>), perhaps using reinforcement only for more refined perceptual learning (Roelfsema et al., <xref ref-type="bibr" rid="B368">2010</xref>).</p>
<p>But beyond these well-known global reward signals, we argue that the basic mechanisms of reinforcement learning may be widely re-purposed to train local networks using a variety of internally generated error signals. These internally generated signals may allow a learning system to go beyond what can be learned via standard unsupervised methods, effectively guiding or steering the system to learn specific features or computations (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>).</p>
<sec>
<title>3.4.1. Cost functions for bootstrapping learning in the human environment</title>
<p>Special, internally-generated signals are needed specifically for learning problems where standard unsupervised methods&#x02014;based purely on matching the statistics of the world, or on optimizing simple mathematical objectives like temporal continuity or sparsity&#x02014;will fail to discover properties of the world which are statistically weak in an objective sense but nevertheless have special significance to the organism (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>). Indigo bunting birds, for example, learn a template for the constellations of the night sky long before ever leaving the nest to engage in navigation-dependent tasks (Emlen, <xref ref-type="bibr" rid="B109">1967</xref>). This memory template is directly used to determine the direction of flight during migratory periods, a process that is modulated hormonally so that winter and summer flights are reversed. Learning is therefore a multi-phase process in which navigational cues are memorized prior to the acquisition of motor control.</p>
<p>In humans, we suspect that similar multi-stage bootstrapping processes are arranged to occur. Humans have innate specializations for social learning. We need to be able to read one another&#x00027;s expressions as indicated with hands and faces. Hands are important because they allow us to learn about the set of actions that can be produced by agents (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>). Faces are important because they give us insight into what others are thinking. People have intentions and personalities that differ from one another, and their feelings are important. How could we hack together cost functions, built on simple genetically specifiable mechanisms, to make it easier for a learning system to discover such behaviorally relevant variables?</p>
<p>Some preliminary studies are beginning to suggest specific mechanisms and heuristics that humans may be using to bootstrap more sophisticated knowledge. In a groundbreaking study, Ullman et al. (<xref ref-type="bibr" rid="B443">2012</xref>) asked how could we explain hands, to a system that does not already know about them, in a cheap way, without the need for labeled training examples? Hands are common in our visual space and have special roles in the scene: they move objects, collect objects, and caress babies. Building these biases into an area specialized to detect hands could guide the right kind of learning, by providing a downstream learning system with many likely positive examples of hands on the basis of innately-stored, heuristic signatures about how hands tend to look or behave (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>). Indeed, an internally supervised learning algorithm containing specialized, hard-coded biases to detect hands, on the basis of their typical motion properties, can be used to bootstrap the training of an image recognition module that learns to recognize hands based on their appearance. Thus, a simple, hard-coded module bootstraps the training of a much more complex algorithm for visual recognition of hands.</p>
<p>Ullman et al. (<xref ref-type="bibr" rid="B443">2012</xref>) then further exploits a combination of hand and face detection to bootstrap a predictor for gaze direction, based on the heuristic that faces tend to be looking toward hands. Of course, given a hand detector, it also becomes much easier to train a system for reaching, crawling, and so forth. Efforts are underway in psychology to determine whether the heuristics discovered to be useful computationally are, in fact, being used by human children during learning (Yu and Smith, <xref ref-type="bibr" rid="B483">2013</xref>; Fausey et al., <xref ref-type="bibr" rid="B113">2016</xref>).</p>
<p>Ullman refers to such primitive, inbuilt detectors as innate &#x0201C;proto-concepts&#x0201D; (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>). Their broader claim is that such pre-specification of mutual supervision signals can make learning the relevant features of the world far easier, by giving an otherwise unsupervised learner the right kinds of hints or heuristic biases at the right times. Here we call these approximate, heuristic cost functions &#x0201C;bootstrap cost functions.&#x0201D; The purpose of the bootstrap cost functions is to reduce the amount of data required to learn a specific feature or task, but at the same time to avoid a need for fully unsupervised learning.</p>
<p>Could the neural circuitry for such a bootstrap hand-detector be pre-specified genetically? The precedent from other organisms is strong: for example, it is famously known that the frog retina contains circuitry sufficient to implement a kind of &#x0201C;bug detector&#x0201D; (Lettvin et al., <xref ref-type="bibr" rid="B255">1959</xref>). Ullman&#x00027;s hand detector, in fact, operates via a simple local optical flow calculation to detect &#x0201C;mover&#x0201D; events. This type of simple, local calculation could potentially be implemented in genetically-specified and/or spontaneously self-organized neural circuitry in the retina or early dorsal visual areas (B&#x000FC;lthoff et al., <xref ref-type="bibr" rid="B64">1989</xref>), perhaps similarly to the frog&#x00027;s &#x0201C;bug detector.&#x0201D;</p>
<p>How could we explain faces without any training data? Faces tend to have two dark dots in their upper half, a line in the lower half and tend to be symmetric about a vertical axis. Indeed, we know that babies are very much attracted to things with these generic features of upright faces starting from birth, and that they will acquire face-specific cortical areas<xref ref-type="fn" rid="fn0038"><sup>38</sup></xref> in their first few years of life if not earlier (McKone et al., <xref ref-type="bibr" rid="B295">2009</xref>). It is easy to define a local rule that produces a kind of crude face detector (e.g., detecting two dots on top of a horizontal line), and indeed some evidence suggests that the brain can rapidly detect faces without even a single feed-forward pass through the ventral visual stream (Crouzet and Thorpe, <xref ref-type="bibr" rid="B88">2011</xref>). The crude detection of human faces used together with statistical learning should be analogous to semi-supervised learning (Sukhbaatar et al., <xref ref-type="bibr" rid="B416">2014</xref>) and could allow identifying faces with high certainty.</p>
<p>Humans have areas devoted to emotional processing, and the brain seems to embody prior knowledge about the structure of emotional expressions and how they relate to causes in the world: emotions should have specific types of strong couplings to various other higher-level variables such as goal-satisfaction, should be expressed through the face, and so on (Phillips et al., <xref ref-type="bibr" rid="B348">2002</xref>; Skerry and Spelke, <xref ref-type="bibr" rid="B404">2014</xref>; Baillargeon et al., <xref ref-type="bibr" rid="B23">2016</xref>; Lyons and Cheries, <xref ref-type="bibr" rid="B270">2016</xref>). What about agency? It makes sense to describe, when dealing with high-level thinking, other beings as optimizers of their own goal functions. It appears that heuristically specified notions of goals and agency are infused into human psychological development from early infancy and that notions of agency are used to bootstrap heuristics for ethical evaluation (Hamlin et al., <xref ref-type="bibr" rid="B162">2007</xref>; Skerry and Spelke, <xref ref-type="bibr" rid="B404">2014</xref>). Algorithms for establishing more complex, innately-important social relationships such as joint attention are under study (Gao et al., <xref ref-type="bibr" rid="B135">2014</xref>), building upon more primitive proto-concepts like face detectors and Ullman&#x00027;s hand detectors (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>). The brain can thus use innate detectors to create cost functions and training procedures to train the next stages of learning. This prior knowledge, encoded into brain structure via evolution, could allow learning signals to come from the right places and to appear developmentally at the right times.</p>
<p>It is intuitive to ask whether this type of bootstrapping poses a kind of &#x0201C;chicken and egg&#x0201D; problem: if the brain already has an inbuilt heuristic hand detector, how can it be used to train a detector that performs any better than those heuristics? After all, isn&#x00027;t a trained system only as good as its training data? The work of Ullman et al. (<xref ref-type="bibr" rid="B443">2012</xref>) illustrates why this is not the case. First, the &#x0201C;innate detector&#x0201D; can be used to train a downstream detector that operates based on different cues: for example, based on the spatial and body context of the hand, rather than its motion. Second, once multiple such pathways of detection come into existence, they can be used to improve each other. In Ullman et al. (<xref ref-type="bibr" rid="B443">2012</xref>), appearance, body context, and mover motion are all used to bootstrap off of one another, creating a detector that is better than any of its training heuristics. In effect, the innate detectors are used not as supervision signals <italic>per se</italic>, but rather to guide or steer the learning process, enabling it to discover features that would otherwise be difficult. If such affordances can be found in other domains, it seems likely that the brain would make extensive use of them to ensure that developing animals learn the precise patterns of perception and behavior needed to ensure their later survival and reproduction.</p>
<p>Thus, generalizing previous ideas (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>; Poggio, <xref ref-type="bibr" rid="B353">2015</xref>), we suggest that the brain uses optimization with respect to internally generated heuristic<xref ref-type="fn" rid="fn0039"><sup>39</sup></xref> detection signals to bootstrap learning of biologically relevant features which would otherwise be missed by an unsupervised learner. In one possible implementation, such bootstrapping may occur via reinforcement learning, using the outputs of the innate detectors as local reinforcement signals, and perhaps using mechanisms similar to Stanisor et al. (<xref ref-type="bibr" rid="B411">2013</xref>), Rombouts et al. (<xref ref-type="bibr" rid="B371">2015</xref>), Brosch et al. (<xref ref-type="bibr" rid="B59">2015</xref>), and Roelfsema and van Ooyen (<xref ref-type="bibr" rid="B369">2005</xref>) to perform reinforcement learning through a multi-layer network. It is also possible that the brain could use such internally generated heuristic detectors in other ways, for example to bias the inputs delivered to an unsupervised learning network toward entities of interest to humans via an attentional process (Joscha Bach, personal communication), to bias hippocampal replay (Kumaran et al., <xref ref-type="bibr" rid="B240">2016</xref>) or other aspects of memory access, or to directly train simple classifiers (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>).</p>
</sec>
<sec>
<title>3.4.2. Cost functions for learning by imitation and through social feedback</title>
<p>It has been widely observed that the capacity for imitation and social learning may be a feature that is uniquely human, and that enables other human traits (Ramachandran, <xref ref-type="bibr" rid="B358">2000</xref>). Humans need to learn more from the environment by than trial and error can provide for, and more than genetically orchestrated internal bootstrapping signals can effectively guide. Hence, babies spend a long time watching adults, especially adults they are attached to Meltzoff (<xref ref-type="bibr" rid="B298">1999</xref>), and later use specific kinds of social cues from their parents to shape their development. Babies and children learn about cause and effect through models based on goals, outcomes and agents, not just pure statistical inference. For example, young children make inferences about causality selectively in situations where a human is trying to achieve an outcome (Meltzoff et al., <xref ref-type="bibr" rid="B299">2012</xref>, <xref ref-type="bibr" rid="B300">2013</xref>). Minsky (<xref ref-type="bibr" rid="B306">2006</xref>) discusses how we derive not just skills but also goals from our attachment figures, through socially induced emotions like pride and shame. To do all this requires a powerful infrastructure of mental abilities: we must attribute social feedback to particular aspects of our goals or actions, and hence we need to signal to each other positively and negatively, to draw attention to these aspects. Minsky speculates (Minsky, <xref ref-type="bibr" rid="B306">2006</xref>) that the development of such &#x0201C;learning by being told&#x0201D; led to language by selecting for the development of increasingly precise parsing of synatatic structures in relation to our representations of agents and action-plans.</p>
<p>How does this connect with cost functions? The idea of goals is central here, as we need to be able to identify the goals of others, update our own goals based on feedback, and measure the success of actions relative to goals. It has been proposed that human intrinsically use a model based on abstract goal and costs to underpin learning about the social world (Jara-Ettinger et al., <xref ref-type="bibr" rid="B210">2016</xref>). Perhaps we even learn about our &#x0201C;selves&#x0201D; by inferring a model of our own goals and cost functions. Relatedly, machine learning in some settings can infer their cost functions from samples of behavior (Ho and Ermon, <xref ref-type="bibr" rid="B194">2016</xref>).</p>
</sec>
<sec>
<title>3.4.3. Cost functions for story generation and understanding</title>
<p>It has been widely noticed in cognitive science and AI that the generation and understanding of stories are crucial to human cognition. Researchers such as Winston have framed story understanding as the key to human-like intelligence (Winston, <xref ref-type="bibr" rid="B470">2011</xref>). Stories consist of a linear sequence of episodes, in which one episode refers to another through cause and effect relationships, with these relationships often involving the implicit goals of agents. Many other cognitive faculties, such as conceptual grounding of language, could conceivably emerge from an underlying internal representation in terms of stories.</p>
<p>Perhaps the ultimate series of bootstrap cost functions would be those which would direct the brain to utilize its learning networks and specialized systems so as to construct representations that are specifically useful as components of stories, to spontaneously chain these representations together, and to update them through experience and communication. How could such cost functions arise? One possibility is that they are bootstrapped through imitation and communication, where a child learns to mimic the story-telling behavior of others. Another possibility is that useful representations and primitives for stories emerge spontaneously from mechanisms for learning state and action chunking in hierarchical reinforcement learning and planning. Yet another is that stories emerge from learned patterns of saliency-directed memory storage and recall (e.g., Xiong et al., <xref ref-type="bibr" rid="B477">2016</xref>). In addition, priors that direct the developing child&#x00027;s brain to learn about and attend to social agency seem to be important for stories.</p>
<p>In this section, we have seen how cost functions can be specified that could lead to the learning of increasingly sophisticated mental abilities in a biologically plausible manner. Importantly, however, cost functions and optimization are not the whole story. To achieve more complex forms of optimization, e.g., for learning to understand complex patterns of cause and effect over long timescales, to plan and reason prospectively, or to effectively coordinate many widely distributed brain resources, the brain seems to invoke specialized, pre-constructed data structures, algorithms and communication systems, which in turn facilitate specific kinds of optimization. Moreover, optimization occurs in a tightly orchestrated multi-stage process, and specialized, pre-structured brain systems need to be invoked to account for this meta-level of control over when, where and how each optimization problem is set up. We now turn to how these pre-specialized systems may orchestrate and facilitate optimization.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Optimization occurs in the context of specialized structures</title>
<p>Optimization of initially unstructured &#x0201C;blank slate&#x0201D; networks is not sufficient to generate complex cognition in the brain, we argue, even given a diversity of powerful genetically-specified cost functions and local learning rules, as we have posited above. Instead, in Hypothesis 3, we suggest that specialized, pre-structured architectures are needed for at least two purposes.</p>
<p>First, pre-structured architectures are needed to allow the brain to find efficient solutions to certain types of problems. When we write computer code, there are a broad range of algorithms and data structures employed for different purposes: we may use dynamic programming to solve planning problems, trees to efficiently implement nearest neighbor search, or stacks to implement recursion. Having the right kind of algorithm and data structure in place to solve a problem allows it to be solved efficiently, robustly and with a minimum amount of learning or optimization needed. This observation is concordant with the increasing use of pre-specialized architectures and specialized computational components in machine learning (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>; Weston et al., <xref ref-type="bibr" rid="B464">2014</xref>; Neelakantan et al., <xref ref-type="bibr" rid="B321">2015</xref>). In particular, to enable the learning of efficient computational solutions, the brain may need pre-specialized systems for planning and executing sequential multi-step processes, for accessing memories, and for forming and manipulating compositional and recursive structures<xref ref-type="fn" rid="fn0040"><sup>40</sup></xref>.</p>
<p>Second, the training of optimization modules may need to be coordinated in a complex and dynamic fashion, including delivering the right training signals and activating the right learning rules in the right places and at the right times. To allow this, the brain may need specialized systems for storing and routing data, and for flexibly routing training signals such as target patterns, training data, reinforcement signals, attention signals, and modulatory signals. These mechanisms may need to be at least partially in place in advance of learning.</p>
<p>Looking at the brain, we indeed seem to find highly conserved structures, e.g., cortex, where it is theorized that a similar type of learning and/or computation is happening in multiple places (Braitenberg and Schutz, <xref ref-type="bibr" rid="B54">1991</xref>; Douglas and Martin, <xref ref-type="bibr" rid="B101">2004</xref>). But we also see a large number of specialized structures, including thalamus, hippocampus, basal ganglia and cerebellum (Solari and Stoner, <xref ref-type="bibr" rid="B406">2011</xref>). These structures evolutionarily pre-date (Lee et al., <xref ref-type="bibr" rid="B248">2015</xref>) the cortex, and hence the cortex may have evolved to work in the context of such specialized mechanisms. For example, the cortex may have evolved as a trainable module for which the training is orchestrated by these older structures.</p>
<p>Even within the cortex itself, microcircuitry within different areas may be specialized: tinkered variations on a common ancestral microcircuit scaffold could potentially allow different cortical areas, such as sensory areas vs. prefrontal areas, to be configured to adopt a number of qualitatively distinct computational and learning configurations (Yuste et al., <xref ref-type="bibr" rid="B484">2005</xref>; Marcus et al., <xref ref-type="bibr" rid="B284">2014a</xref>,<xref ref-type="bibr" rid="B285">b</xref>), even while sharing a common gross physical layout and communication interface. Within cortex, over forty distinct cell types&#x02014;differing in such aspects as dendritic organization, distribution throughout the six cortical layers, connectivity pattern, gene expression, and electrophysiological properties&#x02014;have already been found (Markram et al., <xref ref-type="bibr" rid="B289">2015</xref>; Zeisel et al., <xref ref-type="bibr" rid="B486">2015</xref>). Central pattern generator circuits provide an example of the kinds of architectures that can be pre-wired into neural microcircuitry, and may have evolutionary relationships with cortical circuits (Yuste et al., <xref ref-type="bibr" rid="B484">2005</xref>). Thus, while the precise degree of architectural specificity of particular cortical regions is still under debate (Marcus et al., <xref ref-type="bibr" rid="B284">2014a</xref>,<xref ref-type="bibr" rid="B285">b</xref>), various mechanisms could offer pre-specified heterogeneity.</p>
<p>In this section, we explore the kinds of computational problems for which specialized structures may be useful, and attempt to map these to putative elements within the brain. Our preliminary sketch of a functional decomposition can be viewed as a summary of suggestions for specialized functions that have been made throughout the computational neuroscience literature, and is influenced strongly by the models of O&#x00027;Reilly, Eliasmith, Grossberg, Marcus, Hayworth and others (Marcus, <xref ref-type="bibr" rid="B282">2001</xref>; O&#x00027;Reilly, <xref ref-type="bibr" rid="B333">2006</xref>; Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>; Hayworth, <xref ref-type="bibr" rid="B178">2012</xref>; Grossberg, <xref ref-type="bibr" rid="B155">2013</xref>). The correspondence between these models and actual neural circuitry is, of course, still the subject of extensive debate.</p>
<p>Many of the computational and neural concepts sketched here are preliminary and will need to be made more rigorous through future study. Our knowledge of the functions of particular brain areas, and thus our proposed mappings of certain computations onto neuroanatomy, also remains tentative. Finally, it is still far from established which processes in the brain emerge from optimization of cost functions, which emerge from other forms of self-organization, which are pre-structured through genetics and development, and which rely on an interplay of all these mechanisms<xref ref-type="fn" rid="fn0041"><sup>41</sup></xref>. Our discussion here should therefore be viewed as a sketch of potential directions for further study.</p>
<sec>
<title>4.1. Structured forms of memory</title>
<p>One of the central elements of computation is memory. Importantly, multiple different kinds of memory are needed (Squire, <xref ref-type="bibr" rid="B408">2004</xref>). For example, we need memory that is stored for a long period of time and that can be retrieved in a number of ways, such as in situations similar to the time when the memory was first stored (content addressable memory). We also need memory that we can keep for a short period of time and that we can rapidly rewrite (working memory). Lastly, we need the kind of implicit memory that we cannot explicitly recall, similar to the kind of memory that is classically learned using gradient descent on errors, i.e., sculpted into the weight matrix of a neural network.</p>
<sec>
<title>4.1.1. Content addressable memories</title>
<p>Content addressable memories<xref ref-type="fn" rid="fn0042"><sup>42</sup></xref> are classic models in neuroscience (Hopfield, <xref ref-type="bibr" rid="B196">1982</xref>). Most simply, they allow us to recognize a situation similar to one that we have seen before, and to &#x0201C;fill in&#x0201D; stored patterns based on partial or noisy information, but they may also be put to use as sub-components of many other functions. Recent research has shown that including such memories allows deep networks to learn to solve problems that previously were out of reach, even of LSTM networks that already have a simpler form of local memory and are already capable of learning long-term dependencies (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>; Weston et al., <xref ref-type="bibr" rid="B464">2014</xref>). Hippocampal area CA3 may act as an auto-associative memory<xref ref-type="fn" rid="fn0043"><sup>43</sup></xref> capable of content-addressable pattern completion, with pattern separation occurring in the dentate gyrus (Rolls, <xref ref-type="bibr" rid="B370">2013</xref>). If no similar pattern is available, an unfamiliar input will be stored as a new memory (Kumaran et al., <xref ref-type="bibr" rid="B240">2016</xref>). Such systems could permit the retrieval of complete memories from partial cues, enabling networks to perform operations similar to database retrieval or to instantiate lookup tables of historical stimulus-response mappings, among numerous other possibilities.</p>
<p>Of course, memory systems may be organized&#x02014;through cost function optimization or other mechanisms&#x02014;into higher-order structures. Cost functions might be used to bias memory representations to adopt particular structures, e.g., to be organized into data structures like like Minskys frames and trans-frames (Minsky, <xref ref-type="bibr" rid="B306">2006</xref>).</p>
</sec>
<sec>
<title>4.1.2. Working memory buffers</title>
<p>Cognitive science has long characterized properties of the working memory. Its capacity is somewhat limited, with the old idea being that verbal working memory has a capacity of &#x0201C;seven plus or minus two&#x0201D; (Miller, <xref ref-type="bibr" rid="B301">1956</xref>), while visual working memory has a capacity of four (Luck and Vogel, <xref ref-type="bibr" rid="B268">1997</xref>) (or, other authors defend, one). There are many models of working memory (O&#x00027;Reilly and Frank, <xref ref-type="bibr" rid="B334">2006</xref>; Singh and Eliasmith, <xref ref-type="bibr" rid="B401">2006</xref>; Warden and Miller, <xref ref-type="bibr" rid="B455">2007</xref>; Wang, <xref ref-type="bibr" rid="B454">2012</xref>; Buschman and Miller, <xref ref-type="bibr" rid="B66">2014</xref>), some of which attribute it to persistent, self-reinforcing patterns of neural activation (Goldman et al., <xref ref-type="bibr" rid="B145">2003</xref>) in the recurrent networks of the prefrontal cortex. Prefrontal working memory appears to be made up of multiple functionally distinct subsystems (Markowitz et al., <xref ref-type="bibr" rid="B287">2015</xref>). Neural models of working memory can store not only scalar variables (Seung, <xref ref-type="bibr" rid="B393">1998</xref>), but also high-dimensional vectors (Eliasmith and Anderson, <xref ref-type="bibr" rid="B106">2004</xref>; Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>) or sequences of vectors (Choo and Eliasmith, <xref ref-type="bibr" rid="B76">2010</xref>). Working memory buffers seem crucial for human-like cognition, e.g., reasoning, as they allow short-term storage while also&#x02014;in conjunction with other mechanisms&#x02014;enabling generalization of operations across anything that can fill the buffer.</p>
</sec>
<sec>
<title>4.1.3. Storing state in association with saliency</title>
<p>Saliency, or interestingness, measures can be used to tag the importance of a memory (Gonzalez Andino and Grave de Peralta Menendez, <xref ref-type="bibr" rid="B146">2012</xref>). This can allow removal of the boring data from the training set, allowing a mechanism that is more like optimal experimentation. Moreover, saliency can guide memory replay or sampling from generative models, to generate more training data drawn from a distribution useful for learning (Ji and Wilson, <xref ref-type="bibr" rid="B214">2007</xref>; Mnih et al., <xref ref-type="bibr" rid="B313">2015</xref>). Conceivably, hippocampal replay could allow a batch-like training process, similar to how most machine learning systems are trained, rather than requiring all training to occur in an online fashion. Plasticity mechanisms in memory systems which are gated by saliency are starting to be uncovered in neuroscience (Dudman et al., <xref ref-type="bibr" rid="B103">2007</xref>). Importantly, the notions of &#x0201C;saliency&#x0201D; computed by the brain could be quite intricate and multi-faceted, potentially leading to complex schemes by which specific kinds of memories would be tagged for later context-dependent retrieval. As a hypothetical example, representations of both timing and importance associated with memories could perhaps allow retrieval only of important memories that happened within a certain window of time (MacDonald et al., <xref ref-type="bibr" rid="B275">2011</xref>; Kraus et al., <xref ref-type="bibr" rid="B233">2013</xref>; Rubin et al., <xref ref-type="bibr" rid="B375">2015</xref>). Storing and retrieving information selectively based on specific properties of the information itself, or of &#x0201C;tags&#x0201D; appended to that information, is a powerful computational primitive that could enable learning of more complex tasks. Relatedly, we know that certain pathways become associated with certain kinds of memories, e.g., specific pathways for fear-related memory in mice.</p>
</sec>
</sec>
<sec>
<title>4.2. Structured routing systems</title>
<p>To use its information flexibly, the brain needs structured systems for routing data. Such systems need to address multiple temporal and spatial scales, and multiple modalities of control. Thus, there are several different kinds of information routing systems in the brain which operate by different mechanisms and under different constraints.</p>
<sec>
<title>4.2.1. Attention</title>
<p>If we can focus on one thing at a time, we may be able to allocate more computational resources to processing it, make better use of scarce data to learn about it, and more easily store and retrieve it from memory<xref ref-type="fn" rid="fn0044"><sup>44</sup></xref>. Notably in this context, attention allows improvements in learning: if we can focus on just a single object, instead of an entire scene, we can learn about it more easily using limited data. Formal accounts in a Bayesian framework talk about attention reducing the sample complexity of learning (Chikkerur et al., <xref ref-type="bibr" rid="B75">2010</xref>). Likewise, in models, the processes of applying attention, and of effectively making use of incoming attentional signals to appropriately modulate local circuit activity, can themselves be learned by optimizing cost functions (Jaramillo and Pearlmutter, <xref ref-type="bibr" rid="B211">2004</xref>; Mnih et al., <xref ref-type="bibr" rid="B312">2014</xref>). The right kinds of attention make processing and learning more efficient, and also allow for a kind of programmatic control over multi-step perceptual tasks.</p>
<p>How does the brain determine where to allocate attention, and how is the attentional signal physically mediated? Answering this question is still an active area of neuroscience. Higher-level cortical areas may be specialized in allocating attention. The problem is made complex by the fact that there seem to be many different types of attention&#x02014;such as object-based, feature-based and spatial attention in vision&#x02014;that may be mediated by interactions between different brain areas. The frontal eye fields (area FEF), for example, are important in visual attention, specifically for controlling saccades of the eyes to attended locations. Area FEF contains &#x0201C;retinotopic&#x0201D; spatial maps whose activation determines the saccade targets in the visual field. Other prefrontal areas such as the dorsolateral prefrontal cortex and inferior frontal junction are also involved in maintaining representations that specify the targets of certain types of attention. Certain forms of attention may require a complex interaction between brain areas, e.g., to determine targets of attention based on higher-level properties that are represented across multiple areas, like the identity and spatial location of a specific face (Baldauf and Desimone, <xref ref-type="bibr" rid="B26">2014</xref>).</p>
<p>There are many proposed neural mechanisms of attention, including the idea that synchrony plays a role (Baldauf and Desimone, <xref ref-type="bibr" rid="B26">2014</xref>), perhaps by creating resonances that facilitate the transfer of information between synchronously oscillating neural populations in different areas<xref ref-type="fn" rid="fn0045"><sup>45</sup></xref>. Other proposed mechanisms include specific circuits for attention-dependent signal routing (Anderson and Van Essen, <xref ref-type="bibr" rid="B7">1987</xref>; Olshausen et al., <xref ref-type="bibr" rid="B327">1993</xref>). Various forms of attention also have specific neurophysiological signatures, such as enhancements in synchrony among neural spikes and with the ambient local field potential, changes in the sharpness of neural tuning curves, and other properties. These diverse effects and signatures of attention may be consequences of underlying pathways that wire up to particular elements of cortical microcircuits to mediate different attentional effects (Bobier et al., <xref ref-type="bibr" rid="B48">2014</xref>).</p>
</sec>
<sec>
<title>4.2.2. Buffers</title>
<p>One possibility is that the brain uses distinct groups of neurons, which we can call &#x0201C;buffers,&#x0201D; to store distinct variables, such as the subject or object in a sentence (Frankland and Greene, <xref ref-type="bibr" rid="B128">2015</xref>). Having memory buffers allows the abstraction of a variable.</p>
<p>Once we establish that the brain has a number of memory buffers, we need ways for those buffers to interact. We need to be able to take a buffer, do a computation on its contents and store the output into another buffer. But if the representations in each of two groups of neurons are learned, and hence are coded differently, how can the brain &#x0201C;copy and paste&#x0201D; information between these groups of neurons? Malsburg argued that such a system of separate buffers is impossible because the neural pattern for &#x0201C;chair&#x0201D; in buffer 1 has nothing in common with the neural pattern for &#x0201C;chair&#x0201D; in buffer 2&#x02014;any learning that occurs for the contents of buffer 1 would not automatically be transferable to buffer 2. Various mechanisms have been proposed to allow such transferability, which focus on ways in which all buffers could be trained jointly and then later separated so that they can work independently when they need to<xref ref-type="fn" rid="fn0046"><sup>46</sup></xref>.</p>
</sec>
<sec>
<title>4.2.3. Discrete gating of information flow between buffers</title>
<p>Dense connectivity is only achieved locally, but it would be desirable to have a way for any two cortical units to talk to one another, if needed, regardless of their distance from one another, and without introducing crosstalk<xref ref-type="fn" rid="fn0047"><sup>47</sup></xref>. It is therefore critical to be able to dynamically turn on and off the transfer of information between different source and destination regions, in much the manner of a switchboard. Together with attention, such dedicated routing systems can make sure that a brain area receives exactly the information it needs. Such a discrete routing system is, of course, central to cognitive architectures like ACT-R (Anderson, <xref ref-type="bibr" rid="B8">2007</xref>). The key feature of ACT-R is the ability to evaluate the IF clauses of tens of thousands of symbolic rules (called &#x0201C;productions&#x0201D;), in parallel, approximately every 50 ms. Each rule requires equality comparisons between the contents of many constant and variable memory buffers, and the execution of a rule leads to the conditional routing of information from one buffer to another.</p>
<p>What controls which long-range routing operations occur when, i.e., where is the switchboad and what controls it? Several models, including ACT-R, have attributed such parallel rule-based control of routing to the action selection circuitry (Gurney et al., <xref ref-type="bibr" rid="B160">2001</xref>; Terrence Stewart, <xref ref-type="bibr" rid="B432">2010</xref>) of the basal ganglia (BG) (O&#x00027;Reilly and Frank, <xref ref-type="bibr" rid="B334">2006</xref>; Stocco et al., <xref ref-type="bibr" rid="B413">2010</xref>), and its interaction with working memory buffers in the prefrontal cortex. In conventional models of thalamo-cortico-striatal loops, competing actions of the direct and indirect pathways through the basal ganglia can inhibit or disinhibit an area of motor cortex, thereby gating a motor action<xref ref-type="fn" rid="fn0048"><sup>48</sup></xref>. Models like (O&#x00027;Reilly and Frank, <xref ref-type="bibr" rid="B334">2006</xref>; Stocco et al., <xref ref-type="bibr" rid="B413">2010</xref>; Terrence Stewart, <xref ref-type="bibr" rid="B432">2010</xref>) propose further that the basal ganglia can gate not just the transfer of information from motor cortex to downstream actuators, but also the transfer of information between cortical areas. To do so, the basal ganglia would dis-inhibit a thalamic relay (Sherman, <xref ref-type="bibr" rid="B397">2005</xref>, <xref ref-type="bibr" rid="B398">2007</xref>) linking two cortical areas. Dopamine-related activity is thought to lead to temporal difference reinforcement learning of such gating policies in the basal ganglia (Frank and Badre, <xref ref-type="bibr" rid="B129">2012</xref>). Beyond the basal ganglia, there are also other, separate pathways involved in action selection, e.g., in the prefrontal cortex (Daw et al., <xref ref-type="bibr" rid="B92">2006</xref>). Thus, multiple systems including basal ganglia and cortex could control the gating of long-range information transfer between cortical areas, with the thalamus perhaps largely constituting the switchboard itself.</p>
<p>How is such routing put to use in a learning context? One possibility is that the basal ganglia acts to orchestrate the training of the cortex. The basal ganglia may exert tight control<xref ref-type="fn" rid="fn0049"><sup>49</sup></xref> over the cortex, helping to determine when and how it is trained. Indeed, because the basal ganglia pre-dates the cortex evolutionarily, it is possible that the cortex evolved as a flexible, trainable resource that could be harnessed by existing basal ganglia circuitry. All of the main regions and circuits of the basal ganglia are conserved from our common ancestor with the lamprey more than five hundred million years ago. The major part of the basal ganglia even seems to be conserved from our common ancestor with insects (Strausfeld and Hirth, <xref ref-type="bibr" rid="B415">2013</xref>). Thus, in addition to its real-time action selection and routing functions, the basal ganglia may sculpt how the cortex learns.</p>
</sec>
</sec>
<sec>
<title>4.3. Structured state representations to enable efficient algorithms</title>
<p>Certain algorithmic problems benefit greatly from particular types of representation and transformation, such as a grid-like representation of space. In some cases, rather than just waiting for them to emerge via gradient descent optimization of appropriate cost functions, the brain may be pre-structured to facilitate their creation.</p>
<sec>
<title>4.3.1. Continuous predictive control</title>
<p>We often have to plan and execute complicated sequences of actions on the fly, in response to a new situation. At the lowest level, that of motor control, our body and our immediate environment change all the time. As such, it is important for us to maintain knowledge about this environment in a continuous way. The deviations between our planned movements and those movements that we actually execute continuously provide information about the properties of the environment. Therefore, it seems important to have a specialized system, optimized for high-speed continuous processing, that takes all our motor errors and uses them to update a dynamical model of our body and our immediate environment that can predict the delayed sensory results of our motor actions (McKinstry et al., <xref ref-type="bibr" rid="B294">2006</xref>).</p>
<p>It appears that the cerebellum is such a structure, and lesions to it abolish our way of dealing successfully with a changing body. Incidentally, the cerebellum has more connections than the rest of the brain taken together, apparently in a largely feedforward architecture, and the tiny cerebellar granule cells, which may form a randomized high-dimensional input representation (Marr, <xref ref-type="bibr" rid="B290">1969</xref>; Jacobson and Friedrich, <xref ref-type="bibr" rid="B205">2013</xref>), outnumber all other neurons. The brain clearly needs a dedicated way of quickly and continuously correcting movements to minimize errors, without needing to rely on slow and complex association learning in the neocortex in order to do so.</p>
<p>Newer research shows that the cerebellum is involved in a broad range of cognitive problems (Moberget et al., <xref ref-type="bibr" rid="B315">2014</xref>) as well, potentially because they share computational problems with motor control. For example, when subjects estimate time intervals, which are naturally important for movement, it appears that the brain uses the cerebellum even if no movements are involved (Gooch et al., <xref ref-type="bibr" rid="B147">2010</xref>). Even individual cerebellar Purkinjie cells may learn to generate precise timings of their outputs (Johansson et al., <xref ref-type="bibr" rid="B215">2014</xref>). The brain also appears to use inverse models to rapidly predict motor activity that would give rise to a given sensory target (Hanuschkin et al., <xref ref-type="bibr" rid="B164">2013</xref>; Giret et al., <xref ref-type="bibr" rid="B143">2014</xref>). Such mechanisms could be put to use far beyond motor control, in bootstrapping the training of a larger architecture by exploiting continuously changing error signals to update a real-time model of the system state.</p>
</sec>
<sec>
<title>4.3.2. Hierarchical control</title>
<p>Importantly, many of the control problems we appear to be solving are hierarchical. We have a spinal cord, which deals with the fast signals coming from our muscles and proprioception. Within neuroscience, it is generally assumed that this system deals with fast feedback loops and that this behavior is learned to optimize its own cost function. The nature of cost functions in motor control is still under debate. In particular, the timescale over which cost functions operate remains unclear: motor optimization may occur via real-time responses to a cost function that is computed and optimized online, or via policy choices that change over time more slowly in response to the cost function (K&#x000F6;rding, <xref ref-type="bibr" rid="B228">2007</xref>). Nevertheless, the effect is that central processing in the brain has an effectively simplified physical system to control, e.g., one that is far more linear. So the spinal cord itself already suggests the existence of two levels of a hierarchy, each trained using different cost functions.</p>
<p>However, within the computational motor control literature (see e.g., DeWolf and Eliasmith, <xref ref-type="bibr" rid="B99">2011</xref>), this idea can be pushed far further, e.g., with a hierarchy including spinal cord, M1, PMd, frontal, prefrontal areas. A low level may deal with muscles, the next level may deal with getting our limbs to places or moving objects, a next layer may deal with solving simple local problems (e.g., navigating across a room) while the highest levels may deal with us planning our path through life. This factorization of the problem comes with multiple aspects: First, each level can be solved with its own cost functions, and second, every layer has a characteristic timescale. Some levels, e.g., the spinal cord, must run at a high speed. Other levels, e.g., high-level planning, only need to be touched much more rarely. Converting the computationally hard optimal control problem into a hierarchical approximation promises to make it dramatically easier.</p>
<p>Does the brain solve control problems hierarchically? There is evidence that the brain uses such a strategy (Botvinick et al., <xref ref-type="bibr" rid="B50">2009</xref>; Botvinick and Weinstein, <xref ref-type="bibr" rid="B51">2014</xref>), beside neural network demonstrations (Wayne and Abbott, <xref ref-type="bibr" rid="B458">2014</xref>). The brain may use specialized structures at each hierarchical level to ensure that each operates efficiently given the nature of its problem space and available training signals. At higher levels, these systems may use an abstract syntax for combining sequences of actions in pursuit of goals (Allen et al., <xref ref-type="bibr" rid="B6">2010</xref>). Subroutines in such processes could be derived by a process of chunking sequences of actions into single actions (Graybiel, <xref ref-type="bibr" rid="B152">1998</xref>; Botvinick and Weinstein, <xref ref-type="bibr" rid="B51">2014</xref>). Some brain areas like Broca&#x00027;s area, known for its involvement in language, also appear to be specifically involved in processing the hierarchical structure of behavior, as such, as opposed to its detailed temporal structure (Koechlin and Jubault, <xref ref-type="bibr" rid="B226">2006</xref>).</p>
<p>At the highest level of the decision making and control hierarchy, human reward systems reflect changing goals and subgoals, and we are only beginning to understand how goals are actually coded in the brain, how we switch between goals, and how the cost functions used in learning depend on goal state (Buschman and Miller, <xref ref-type="bibr" rid="B66">2014</xref>; O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B337">2014b</xref>; Pezzulo et al., <xref ref-type="bibr" rid="B346">2014</xref>). Goal hierarchies are beginning to be incorporated into deep learning (Kulkarni et al., <xref ref-type="bibr" rid="B238">2016</xref>).</p>
<p>Given this hierarchical structure, the optimization algorithms can be fine-tuned. For the low levels, there is sheer unlimited training data. For the high levels, a simulation of the world may be simple, with a tractable number of high-level actions to choose from. Finally, each area needs to give reinforcement to other areas, e.g., high levels need to punish lower levels for making planning complicated. Thus this type of architecture can simplify the learning of control problems.</p>
<p>Progress is being made in both neuroscience and machine learning on finding potential mechanisms for this type of hierarchical planning and goal-seeking. This is beginning to reveal mechanisms for chunking goals and actions and for searching and pruning decision trees (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B335">2014a</xref>; Huys et al., <xref ref-type="bibr" rid="B201">2015</xref>; Balaguer et al., <xref ref-type="bibr" rid="B25">2016</xref>; Krishnamurthy et al., <xref ref-type="bibr" rid="B236">2016</xref>; Tamar et al., <xref ref-type="bibr" rid="B426">2016</xref>). The study of model-based hierarchical reinforcement learning and prospective optimization (Sejnowski and Poizner, <xref ref-type="bibr" rid="B389">2014</xref>), which concerns the planning and evaluation of nested sequences of actions, implicates a network coupling the dorsolateral prefontral and orbitofrontal cortex, and the ventral and dorsolateral striatum (Botvinick et al., <xref ref-type="bibr" rid="B50">2009</xref>). Hierarchical RL relies on a hierarchical representation of state and action spaces, and it has been suggested that error-driven learning of an optimal such representation in the hippocampus<xref ref-type="fn" rid="fn0050"><sup>50</sup></xref> gives rise to place and grid cell properties (Stachenfeld, <xref ref-type="bibr" rid="B410">2014</xref>), with goal representations themselves emerging in the amygdala, prefrontal cortex and other areas (O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B335">2014a</xref>).</p>
<p>The question of how control problems can be successfully divided into component problems remains one of the central questions in neuroscience (Wolpert and Flanagan, <xref ref-type="bibr" rid="B472">2016</xref>) and machine learning (Kulkarni et al., <xref ref-type="bibr" rid="B238">2016</xref>), and the cost functions involved in learning to create such decompositions are still unknown. These considerations may begin to make plausible, however, how the brain could not only achieve its remarkable feats of motor learning&#x02014;such as generating complex &#x0201C;innate&#x0201D; motor programs, like walking in the newborn gazelle almost immediately after birth&#x02014;but also the kind of planning that allows a human to prepare a meal or travel from London to Chicago.</p>
</sec>
<sec>
<title>4.3.3. Spatial planning</title>
<p>Spatial planning requires solving shortest-path problems subject to constraints. If we want to get from one location to another, there are an arbitrarily large number of simple paths that could be taken. Most naive implementations of such shortest paths problems are grossly inefficient. It appears that, in animals, the hippocampus aids&#x02014;at least in part through &#x0201C;place cell&#x0201D; and &#x0201C;grid cell&#x0201D; systems&#x02014;in efficient learning about new environments and in targeted navigation in such environments (Brown et al., <xref ref-type="bibr" rid="B61">2016</xref>). Interestingly, once an environment becomes familiar, it appears that areas of the neocortex can take over the role of navigation (Hasselmo and Stern, <xref ref-type="bibr" rid="B171">2015</xref>).</p>
<p>In some simple models, targeted navigation in the hippocampus is achieved via the dynamics of &#x0201C;bump attractors&#x0201D; or propagating waves in a place cell network with Hebbian plasticity and adaptation (Hopfield, <xref ref-type="bibr" rid="B198">2009</xref>; Buzs&#x000E1;ki and Moser, <xref ref-type="bibr" rid="B69">2013</xref>; Ponulak and Hopfield, <xref ref-type="bibr" rid="B355">2013</xref>), which allows the network to effectively chart out a path in the space of place cell representations. Other navigation models make use of the grid cell system. The place cell network may<xref ref-type="fn" rid="fn0051"><sup>51</sup></xref> take input from a grid cell network that computes precise distances and directions, perhaps by integrating head direction and velocity signals&#x02014;grid cells fire when the animal is on any node of a regularly spaced hexagonal grid. Different parts of the entorhinal cortex contain grid cells with different grid spacings, and place cells may combine information from multiple such grids in order to build up responses to particular single positions. These systems are highly structured temporally, e.g., containing nested gamma and theta oscillation structures that are phased locked to sequences of place-cell responses, interfering oscillators frequency-shifted by the animal&#x00027;s motion velocity (Zilli and Hasselmo, <xref ref-type="bibr" rid="B488">2010</xref>), tuned cellular resonances (Giocomo et al., <xref ref-type="bibr" rid="B142">2007</xref>; Buzs&#x000E1;ki, <xref ref-type="bibr" rid="B68">2010</xref>), and other neural phenomena that lie far outside a conventional artificial neural network description. It seems that an intricate interplay of spatial and temporal network structures may be essential for encoding sequences of spatiotemporal events across multiple scales, and using them to drive multiple forms of learning, e.g., supporting forward and reverse sequence replay with various temporal compression factors (Buzs&#x000E1;ki, <xref ref-type="bibr" rid="B68">2010</xref>).</p>
<p>Higher-level cognitive tasks such as prospective planning appear to share computational sub-problems with path-finding (Hassabis and Maguire, <xref ref-type="bibr" rid="B167">2009</xref>)<xref ref-type="fn" rid="fn0052"><sup>52</sup></xref>. Interaction between hippocampus and prefrontal cortex could perhaps support a more abstract notion of &#x0201C;navigation&#x0201D; in a space of goals and sub-goals. Interestingly, there is preliminary evidence from fMRI that abstract concepts are also represented according to grid-cell-like hexagonal grid structures in humans (Constantinescu et al., <xref ref-type="bibr" rid="B84">2016</xref>), as well as preliminary evidence that social relationships may also be represented through a hippocampal map (Tavares et al., <xref ref-type="bibr" rid="B430">2015</xref>). Having specialized structures for path-finding could thus simplify a variety of computational problems at different levels of abstraction.</p>
</sec>
<sec>
<title>4.3.4. Variable binding</title>
<p>Language and reasoning appear to present a problem for neural networks (Minsky, <xref ref-type="bibr" rid="B307">1991</xref>; Marcus, <xref ref-type="bibr" rid="B282">2001</xref>; Hadley, <xref ref-type="bibr" rid="B161">2009</xref>): we seem to be able to apply common grammatical rules to sentences regardless of the content of those sentences, and regardless of whether we have ever seen even remotely similar sentences in the training data. While this is achieved automatically in a computer with fixed registers, location addressable memories, and hard-coded operations, how it could be achieved in a biological brain, or emerge from an optimization algorithm, has been under debate for decades.</p>
<p>As the putative key capability underlying such operations, variable binding has been defined as &#x0201C;the transitory or permanent tying together of two bits of information: a variable (such as an X or Y in algebra, or a placeholder like subject or verb in a sentence) and an arbitrary instantiation of that variable (say, a single number, symbol, vector, or word)&#x0201D; (Marcus et al., <xref ref-type="bibr" rid="B284">2014a</xref>,<xref ref-type="bibr" rid="B285">b</xref>). A number of potential biologically plausible binding mechanisms (Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>; Hayworth, <xref ref-type="bibr" rid="B178">2012</xref>; Kriete et al., <xref ref-type="bibr" rid="B235">2013</xref>; Goertzel, <xref ref-type="bibr" rid="B144">2014</xref>) are reviewed in Marcus et al. (<xref ref-type="bibr" rid="B284">2014a</xref>) and Marcus et al. (<xref ref-type="bibr" rid="B285">2014b</xref>). Some, such as vector symbolic architectures<xref ref-type="fn" rid="fn0053"><sup>53</sup></xref>, which were proposed in cognitive science (Plate, <xref ref-type="bibr" rid="B352">1995</xref>; Stewart and Eliasmith, <xref ref-type="bibr" rid="B412">2009</xref>; Eliasmith, <xref ref-type="bibr" rid="B105">2013</xref>), are also being considered in the context of efficiently-trainable artificial neural networks (Danihelka et al., <xref ref-type="bibr" rid="B91">2016</xref>)&#x02014;in effect, these systems learn how to use variable binding.</p>
<p>Variable binding could potentially emerge from simpler memory systems. For example, the Scrub-Jay can remember the place and time of last visit for hundreds of different locations, e.g., to determine whether high-quality food is currently buried at any given location (Clayton and Dickinson, <xref ref-type="bibr" rid="B80">1998</xref>). It is conceivable that such spatially-grounded memory systems enabled a more general binding mechanism to emerge during evolution, perhaps through integration with routing systems or other content-addressable or working memory systems.</p>
</sec>
<sec>
<title>4.3.5. Hierarchical syntax</title>
<p>Fixed, static hierarchies (e.g., the hierarchical organization of cortical areas Felleman and Van Essen, <xref ref-type="bibr" rid="B114">1991</xref>) only take us so far: to deal with long chains of arbitrary nested references, we need <italic>dynamic</italic> hierarchies that can implement recursion on the fly. Human language syntax has a hierarchical structure, which Berwick et al described as &#x0201C;composition of smaller forms like words and phrases into larger ones&#x0201D; (Berwick et al., <xref ref-type="bibr" rid="B43">2012</xref>; Miyagawa et al., <xref ref-type="bibr" rid="B311">2013</xref>). The extent of recursion in human language and thought may be captured by a class of automata known as higher-order pushdown automata, which can be implemented via finite state machines with access to nested stacks (Rodriguez and Granger, <xref ref-type="bibr" rid="B366">2016</xref>). Specific fronto-temporal networks may be involved in representing and generating such hierarchies (Dehaene et al., <xref ref-type="bibr" rid="B95">2015</xref>), e.g., with the hippocampal system playing a key role in implementing some analog of a pushdown stack (Rodriguez and Granger, <xref ref-type="bibr" rid="B366">2016</xref>)<xref ref-type="fn" rid="fn0054"><sup>54</sup></xref>.</p>
<p>Little is known about the underlying circuit mechanisms for such dynamic hierarchies, but it is clear that specific affordances for representing such hierarchies in an efficient way would be beneficial. This may be closely connected with the issue of variable binding, and it is possible that operations similar to pointers could be useful in this context, in both the brain and artificial neural networks (Kriete et al., <xref ref-type="bibr" rid="B235">2013</xref>; Kurach et al., <xref ref-type="bibr" rid="B242">2015</xref>). Augmenting neural networks with a differentiable analog of a push-down stack is another such affordance being pursued in machine learning (Joulin and Mikolov, <xref ref-type="bibr" rid="B217">2015</xref>).</p>
</sec>
<sec>
<title>4.3.6. Mental programs and imagination</title>
<p>Humans excel at stitching together sub-actions to form larger actions (Verwey, <xref ref-type="bibr" rid="B451">1996</xref>; Acuna et al., <xref ref-type="bibr" rid="B4">2014</xref>; Sejnowski and Poizner, <xref ref-type="bibr" rid="B389">2014</xref>). Structured, serial, hierarchical probabilistic programs have recently been shown to model aspects of human conceptual representation and compositional learning (Lake et al., <xref ref-type="bibr" rid="B243">2015</xref>). In particular, sequential programs were found to enable one-shot learning of new geometric/visual concepts (Lake et al., <xref ref-type="bibr" rid="B243">2015</xref>). Generative programs have also been proposed in the context of scene understanding (Battaglia et al., <xref ref-type="bibr" rid="B33">2013</xref>). The ability to deal with problems in terms of sub-problems is central both in human thought and in many successful algorithms.</p>
<p>One possibility is that the hippocampus supports the rapid construction and learning of sequential programs, e.g., in multi-step planning. An influential idea&#x02014;known as the &#x0201C;complementary learning systems hypothesis&#x0201D;&#x02014;is that the hippocampus plays a key role in certain processes where learning must occur quickly on the basis of single episodes, whereas the cortex learns more slowly by aggregating and integrating patterns across large amounts of data (Herd et al., <xref ref-type="bibr" rid="B181">2013</xref>; Leibo et al., <xref ref-type="bibr" rid="B252">2015a</xref>; Blundell et al., <xref ref-type="bibr" rid="B47">2016</xref>; Kumaran et al., <xref ref-type="bibr" rid="B240">2016</xref>). The hippocampus appears to explore, in simulation, possible future trajectories to a goal, even those involving previously unvisited locations (&#x000D3;lafsd&#x000F3;ttir et al., <xref ref-type="bibr" rid="B325">2015</xref>). Hippocampal-prefrontal interaction has been suggested to allow rapid, subconscious evaluation of potential action sequences during decision-making, with the hippocampus in effect simulating the expected outcomes of potential actions that are generated and evaluated in the prefrontal (Mushiake et al., <xref ref-type="bibr" rid="B319">2006</xref>; Wang et al., <xref ref-type="bibr" rid="B452">2015</xref>). The role of the hippocampus in imagination, concept generation (Kumaran et al., <xref ref-type="bibr" rid="B241">2009</xref>), scene construction (Hassabis and Maguire, <xref ref-type="bibr" rid="B168">2007</xref>), mental exploration and goal-directed path planning (Hopfield, <xref ref-type="bibr" rid="B198">2009</xref>; &#x000D3;lafsd&#x000F3;ttir et al., <xref ref-type="bibr" rid="B325">2015</xref>; Brown et al., <xref ref-type="bibr" rid="B61">2016</xref>) suggests that it could help to create generative models to underpin more complex inference such as program induction (Lake et al., <xref ref-type="bibr" rid="B243">2015</xref>) or common-sense world simulation (Battaglia et al., <xref ref-type="bibr" rid="B33">2013</xref>). For example, a sequential, programmatic process, mediated jointly by the basal ganglia, hippocampus and prefrontal cortex might allow one-shot learning of a new concept, as in the sequential computations underlying a process like Bayesian Program Learning (Lake et al., <xref ref-type="bibr" rid="B243">2015</xref>).</p>
<p>Another related possibility is that the cortex itself intrinsically supports the construction and learning of sequential programs (Bach and Herger, <xref ref-type="bibr" rid="B22">2015</xref>). Recurrent neural networks have been used for image generation through a sequential, attention-based process (Gregor et al., <xref ref-type="bibr" rid="B153">2015</xref>), although their correspondence with the brain is unclear<xref ref-type="fn" rid="fn0055"><sup>55</sup></xref>.</p>
</sec>
</sec>
<sec>
<title>4.4. Other specialized structures</title>
<p>Importantly, there are many other specialized structures known in neuroscience, which arguably receive less attention than they deserve, even for those interested in higher cognition. In the above, in addition to the hippocampus, basal ganglia and cortex, we emphasized the key roles of the thalamus in routing, of the cerebellum as a fast and rapidly trainable control and modeling system, of the amygdala and other areas as a potential source of utility functions, of the retina or early visual areas as a means to generate detectors for motion and other features to bootstrap more complex visual learning, and of the frontal eye fields and other areas as a possible source of attention control. We ignored other structures entirely, whose functions are only beginning to be uncovered, such as the claustrum (Crick and Koch, <xref ref-type="bibr" rid="B87">2005</xref>), which has been speculated to be important for rapidly binding together information from many modalities. Our overall understanding of the functional decomposition of brain circuitry still seems very preliminary.</p>
</sec>
<sec>
<title>4.5. Relationships with other cognitive frameworks involving specialized systems</title>
<p>A recent analysis (Lake et al., <xref ref-type="bibr" rid="B244">2016</xref>) suggested directions by which to modify and enhance existing neural-net-based machine learning toward more powerful and human-like cognitive capabilities, particularly by introducing new structures and systems which go beyond data-driven optimization. This analysis emphasized that systems should construct generative models of the world that incorporate compositionality (discrete construction from re-usable parts), inductive biases reflecting causality, intuitive physics and intuitive psychology, and the capacity for probabilistic inference over discrete structured models (e.g., structured as graphs, trees, or programs) (Tervo et al., <xref ref-type="bibr" rid="B433">2016</xref>) to harness abstractions and enable transfer learning.</p>
<p>We view these ideas as consistent with and complementary to the framework of cost functions, optimization and specialized systems discussed here. One might seek to understand how optimization and specialized systems could be used to implement some of the mechanisms proposed in Lake et al. (<xref ref-type="bibr" rid="B244">2016</xref>) inside neural networks. Lake et al. (<xref ref-type="bibr" rid="B244">2016</xref>) emphasize how incorporating additional structure into trainable neural networks can potentially give rise to systems that use compositional, causal and intuitive inductive biases and that &#x0201C;learn to learn&#x0201D; using structured models and shared data structures. For example, sub-dividing networks into units that can be modularly and dynamically combined, where representations can be copied and routed, may present a path toward improved compositionality and transfer learning (Andreas et al., <xref ref-type="bibr" rid="B9">2015</xref>). The control flow for recombining pre-existing modules and representations could be learned via reinforcement learning (Andreas et al., <xref ref-type="bibr" rid="B10">2016</xref>). How to implement the broad set of mechanisms discussed in Lake et al. (<xref ref-type="bibr" rid="B244">2016</xref>) is a key computational problem, and it remains open at which levels (e.g., cost functions and training procedures vs. specialized computational structures vs. underlying neural primitives) architectural innovations will need to be introduced to capture these phenomena.</p>
<p>Primitives that are more complex than those used in conventional neural networks&#x02014;for instance, primitives that act as state machines with complex message passing (Bach and Herger, <xref ref-type="bibr" rid="B22">2015</xref>) or networks that intrinsically implement Bayesian inference (George and Hawkins, <xref ref-type="bibr" rid="B137">2009</xref>)&#x02014;could potentially be useful, and it is plausible that some of these may be found in the brain. Recent findings on the power of generic optimization also do not rule out the idea that the brain may explicitly generate and use particular types of structured representations to constrain its inferences; indeed, the specialized brain systems discussed here might provide a means to enforce such constraints. It might be possible to further map the concepts of Lake et al. (<xref ref-type="bibr" rid="B244">2016</xref>) onto neuroscience via an infrastructure of interacting cost functions and specialized brain systems under rich genetic control, coupled to a powerful and generic neurally implemented capacity for optimization. For example, it was recently shown that complex probabilistic population coding and inference can arise automatically from backpropagation-based training of simple neural networks (Orhan and Ma, <xref ref-type="bibr" rid="B338">2016</xref>), without needing to be built in by hand. The nature of the underlying primitives in the brain, on top of which learning can operate, is a key question for neuroscience.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Machine learning inspired neuroscience</title>
<p>Hypotheses are primarily useful if they lead to concrete, experimentally testable predictions. As such, we now want to go through the hypotheses and see to which level they can be directly tested, as well as refined, through neuroscience.</p>
<sec>
<title>5.1. <italic>Hypothesis 1</italic>&#x02013; existence of cost functions</title>
<p>There are multiple general strategies for addressing whether and how the brain optimizes cost functions. A first strategy is based on observing the endpoint of learning. If the brain uses a cost function, and we can guess its identity, then the final state of the brain should be close to optimal for the cost function. We could thus compare (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B156">2015</xref>) receptive fields that are optimized in a simulation, according to a particular cost function, with the measured receptive fields. Various techniques exist to carry out such comparisons in fRMI studies, including population receptive field estimation (Dumoulin and Wandell, <xref ref-type="bibr" rid="B104">2008</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B156">2015</xref>) and representational dissimilarity matrices (Kriegeskorte et al., <xref ref-type="bibr" rid="B234">2008</xref>; Khaligh-Razavi and Kriegeskorte, <xref ref-type="bibr" rid="B223">2014</xref>). This strategy is only beginning to be used at the moment, perhaps because it has been difficult to measure the receptive fields or other representational properties across a large population of <italic>individual</italic> neurons (fMRI operates at a much coarser level), but this situation is beginning to improve technologically with the emergence of large-scale recording methods (Hasselmo, <xref ref-type="bibr" rid="B170">2015</xref>).</p>
<p>A second strategy could directly quantify how well a cost function describes learning. If the dynamics of learning minimize a cost function then the underlying vector field should have a strong gradient descent type component and a weak rotational component, i.e., weight changes will primarily move down the gradient rather than drifting in the nullspace. If we could somehow continuously monitor the synaptic strengths, while externally manipulating them, then we could, in principle, measure the vector field in the space of synaptic weights, and calculate its divergence as well as its rotation. For at least the subset of synapses that are being trained via some approximation to gradient descent, the divergence component should be strong relative to the rotational component. This strategy has not been developed yet due to experimental difficulties with monitoring large numbers of synaptic weights<xref ref-type="fn" rid="fn0056"><sup>56</sup></xref>.</p>
<p>A third strategy is based on perturbations: cost function based learning should undo the effects of perturbations which disrupt optimality, i.e., the system should return to local minima after a perturbation, and indeed perhaps to the same local minimum after a sufficiently small perturbation. If we change synaptic connections, e.g., in the context of a brain machine interface, we should be able to produce a reorganization that can be predicted based on a guess of the relevant cost function. This strategy is starting to be feasible in motor areas.</p>
<p>Lastly, if we knew structurally which cell types and connections mediated the delivery of error signals vs. input data or other types of connections, then we could stimulate specific connections so as to impose a user-defined cost function. In effect, we would use the brain&#x00027;s own networks as a trainable deep learning substrate, and then study how the network responds to training. Brain machine interfaces can be used to set up specific local learning problems, in which the brain is asked to create certain user-specified representations, and the dynamics of this process can be monitored (Sadtler et al., <xref ref-type="bibr" rid="B378">2014</xref>). Likewise, brain machine interfaces can be used to give the brain access to new datastreams, and to investigate how those datastreams are incorporated into task performance, and whether such incorporation is governed by optimality principles (Dadarlat et al., <xref ref-type="bibr" rid="B90">2015</xref>). In order to do this kind of experiment fully and optimally, we must first understand more about how the system is wired to deliver cost signals. Much of the structure that would be found in connectomic circuit maps, for example, would not just be relevant for short-timescale computing, but also for creating the infrastructure that supports cost functions and their optimization.</p>
<p>Many of the learning mechanisms that we have discussed in this paper make specific predictions about connectivity or dynamics. For example, the &#x0201C;feedback alignment&#x0201D; approach to biological backpropagation suggests that cortical feedback connections should, at some level of neuronal grouping, be largely sign-concordant with the corresponding feedforward connections, although not necessarily of concordant weight (Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>), and feedback alignment also makes predictions for synaptic normalization mechanisms (Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>). The Kickback model for biologically plausible backpropagation has a specific role for NMDA receptors (Balduzzi et al., <xref ref-type="bibr" rid="B29">2014</xref>). Some models that incorporate dendritic coincidence detection for learning temporal sequences predict that a given axon should make only a small number of synapses on a given dendritic segment (Hawkins and Ahmad, <xref ref-type="bibr" rid="B174">2016</xref>). Models that involve STDP learning will make predictions about the dynamics of changing firing rates (Hinton, <xref ref-type="bibr" rid="B184">2007</xref>, <xref ref-type="bibr" rid="B185">2016</xref>; Bengio et al., <xref ref-type="bibr" rid="B38">2015a</xref>,<xref ref-type="bibr" rid="B40">b</xref>; Bengio and Fischer, <xref ref-type="bibr" rid="B37">2015</xref>), as well as about the particular network structures, such as those based on autoencoders or recirculation, in which STDP can give rise to a form of backpropagation.</p>
<p>It is critical to establish the unit of optimization. We want to know the scale of the modules that are trainable by some approximation of gradient descent optimization. How large are the networks which share a given error signal or cost function? On what scales can appropriate training signals be delivered? It could be that the whole brain is optimized end-to-end, in principle. In this case we would expect to find connections that carry training signals from each layer to the preceding ones. On successively smaller scales, optimization could be within a brain area, a microcircuit<xref ref-type="fn" rid="fn0057"><sup>57</sup></xref>, or an individual neuron (Mel, <xref ref-type="bibr" rid="B297">1992</xref>; K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B229">2000</xref>, <xref ref-type="bibr" rid="B230">2001</xref>; Hawkins and Ahmad, <xref ref-type="bibr" rid="B174">2016</xref>). Importantly, optimization may co-exist across these scales. There may be some slow optimization end-to-end, with stronger optimization within a local area and very efficient algorithms within each cell. Careful experiments should be able to identify the scale of optimization, e.g., by quantifying the extent of learning induced by a local perturbation.</p>
<p>The tightness of the structure-function relationship is the hallmark of molecular and to some extent cellular biology, but in large connectionist learning systems, this relationship can become difficult to extract: the same initial network can be driven to compute many different functions by subjecting it to different training<xref ref-type="fn" rid="fn0058"><sup>58</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0059"><sup>59</sup></xref>. It can be hard to understand the way a neural network solves its problems.</p>
<p>How could one tell the difference, then, between a gradient-descent trained network vs. untrained or random networks vs. a network that has been trained against a different kind of task? One possibility would be to train artificial neural networks against various candidate cost functions, study the resulting neural tuning properties (Todorov, <xref ref-type="bibr" rid="B437">2002</xref>), and compare them with those found in the circuit of interest (Zipser and Andersen, <xref ref-type="bibr" rid="B489">1988</xref>). This has already been done to aid the interpretation of the neural dynamics underlying decision making in the PFC (Sussillo, <xref ref-type="bibr" rid="B418">2014</xref>), working memory in the posterior parietal cortex (Rajan et al., <xref ref-type="bibr" rid="B357">2016</xref>) and object or action representation in the visual system (Tacchetti et al., <xref ref-type="bibr" rid="B424">2016</xref>; Yamins and DiCarlo, <xref ref-type="bibr" rid="B479">2016a</xref>,<xref ref-type="bibr" rid="B480">b</xref>). Some have gone on to suggest a direct correspondence between cortical circuits and optimized, appropriately regularized (Sussillo et al., <xref ref-type="bibr" rid="B420">2015</xref>), recurrent neural networks (Liao and Poggio, <xref ref-type="bibr" rid="B260">2016</xref>). In any case, effective analytical methods to reverse engineer complex machine learning systems (Jonas and Kording, <xref ref-type="bibr" rid="B216">2016</xref>), and methods to reverse engineer biological brains, may have some commonalities.</p>
<p>Does this emphasis on function optimization and trainable substrates mean that we should give up on reverse engineering the brain based on detailed measurements and models of its specific connectivity and dynamics? On the contrary: we should use large-scale brain maps to try to better understand (a) how the brain implements optimization, (b) where the training signals come from and what cost functions they embody, and (c) what structures exist, at different levels of organization, to constrain this optimization to efficiently find solutions to specific kinds of problems. The answers may be influenced by diverse local properties of neurons and networks, such as homeostatic rules of neural structure, gene expression and function (Marder and Goaillard, <xref ref-type="bibr" rid="B286">2006</xref>), the diversity of synapse types, cell-type-specific connectivity (Jiang et al., <xref ref-type="bibr" rid="B213">2015</xref>), patterns of inter-laminar projection, distributions of inhibitory neuron types, dendritic targeting and local dendritic physiology and plasticity (Markram et al., <xref ref-type="bibr" rid="B289">2015</xref>; Bloss et al., <xref ref-type="bibr" rid="B46">2016</xref>; Morgan et al., <xref ref-type="bibr" rid="B318">2016</xref>; Sandler et al., <xref ref-type="bibr" rid="B380">2016</xref>) or local glial networks (Perea et al., <xref ref-type="bibr" rid="B344">2009</xref>). They may also be influenced by the integrated nature of higher-level brain systems, including mechanisms for developmental bootstrapping (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>), information routing (Gurney et al., <xref ref-type="bibr" rid="B160">2001</xref>; Stocco et al., <xref ref-type="bibr" rid="B413">2010</xref>), attention (Buschman and Miller, <xref ref-type="bibr" rid="B65">2010</xref>) and hierarchical decision making (Lee et al., <xref ref-type="bibr" rid="B248">2015</xref>). Mapping these systems in detail is of paramount importance to understanding how the brain works, down to the nanoscale dendritic organization of ion channels and up to the real-time global coordination of cortex, striatum and hippocampus, all of which are computationally relevant in the framework we have explicated here. We thus expect that large-scale, multi-resolution brain maps would be useful in testing these framework-level ideas, in inspiring their refinements, and in using them to guide more detailed analysis.</p>
</sec>
<sec>
<title>5.2. <italic>Hypothesis 2</italic>&#x02013; biological fine-structure of cost functions</title>
<p>Clearly, we can map differences in structure, dynamics and representation across brain areas. When we find such differences, the question remains as to whether we can interpret these as resulting from differences in the internally-generated cost functions, as opposed to differences in the input data, or from differences that reflect other constraints unrelated to cost functions. If we can directly measure aspects of the cost function in different areas, then we can also compare them across areas. For example, methods from inverse reinforcement learning<xref ref-type="fn" rid="fn0060"><sup>60</sup></xref> might allow backing out the cost function from observed plasticity (Ng and Russell, <xref ref-type="bibr" rid="B323">2000</xref>).</p>
<p>Moreover, as we begin to understand the &#x0201C;neural correlates&#x0201D; of particular cost functions&#x02014;perhaps encoded in particular synaptic or neuromodulatory learning rules, genetically-guided local wiring patterns, or patterns of interaction between brain areas&#x02014;we can also begin to understand when differences in observed neural circuit architecture reflect differences in cost functions.</p>
<p>We expect that, for each distinct learning rule or cost function, there may be specific molecularly identifiable types of cells and/or synapses. Moreover, for each specialized system there may be specific molecularly identifiable developmental programs that tune it or otherwise set its parameters. This would make sense if evolution has needed to tune the parameters of one cost function without impacting others.</p>
<p>How many different types of internal training signals does the brain generate? When thinking about error signals, we are not just talking about dopamine and serotonin, or other classical reward-related pathways. The error signals that may be used to train specific sub-networks in the brain, via some approximation of gradient descent or otherwise, are not necessarily equivalent to reward signals. It is important to distinguish between cost functions that may be used to drive optimization of specific sub-circuits in the brain, and what are referred to as &#x0201C;value functions&#x0201D; or &#x0201C;utility functions,&#x0201D; i.e., functions that predict the agent&#x00027;s aggregate future reward. In both cases, similar reinforcement learning mechanisms may be used, but the interpretation of the cost functions is different. We have not emphasized global utility functions for the animal here, since they are extensively studied elsewhere (e.g., O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B335">2014a</xref>; Bach, <xref ref-type="bibr" rid="B21">2015</xref>), and since we argue that, though important, they are only a part of the picture, i.e., that the brain is not solely an end-to-end reinforcement trained system.</p>
<p>Progress in brain mapping could soon allow us to classify the types of reward signals in the brain, follow the detailed anatomy and connectivity of reward pathways throughout the brain, and map in detail how reward pathways are integrated into striatal, cortical, hippocampal and cerebellar microcircuits. This program is beginning to be carried out in the fly brain, in which twenty specific types of dopamine neuron project to distinct anatomical compartments of the mushroom body to train distinct odor classifiers operating on a set of high-dimensional odor representations (Caron et al., <xref ref-type="bibr" rid="B73">2013</xref>; Aso et al., <xref ref-type="bibr" rid="B19">2014a</xref>,<xref ref-type="bibr" rid="B20">b</xref>; Cohn et al., <xref ref-type="bibr" rid="B82">2015</xref>). It is known that, even within the same system, such as the fly olfactory pathway, some neuronal wiring is highly specific and molecularly programmed (Hattori et al., <xref ref-type="bibr" rid="B173">2007</xref>; Hong and Luo, <xref ref-type="bibr" rid="B195">2014</xref>), while other wiring is effectively random (Caron et al., <xref ref-type="bibr" rid="B73">2013</xref>), and yet other wiring is learned (Aso et al., <xref ref-type="bibr" rid="B19">2014a</xref>). The interplay between such design principles could give rise to many forms of &#x0201C;division of labor&#x0201D; between genetics and learning. Likewise, it is believed that birdsong learning is driven by reinforcement learning using a specialized cost function that relies on comparison with a memorized version of a tutor&#x00027;s song (Fiete et al., <xref ref-type="bibr" rid="B118">2007</xref>), and also that it involves specialized structures for controlling song variability during learning (Aronov et al., <xref ref-type="bibr" rid="B15">2011</xref>). These detailed pathways underlying the construction of cost functions for vocal learning are beginning to be mapped (Mandelblat-Cerf et al., <xref ref-type="bibr" rid="B279">2014</xref>). Starting with simple systems, it should become possible to map the reward pathways and how they evolved and diversified, which would be a step on the way to understanding how the system learns.</p>
<p>These types of mapping efforts would be a first step toward the ability to create a concrete model of the brain&#x00027;s optimization architecture. Our discussion here has focused on trying to anticipate, based on known neuroscience knowledge and on approaches becoming successful in machine learning, the <italic>kinds</italic> of local cost functions that the brain may rely on, and how specialized brain systems may enable efficient solutions to optimization problems. However, this framework-level discussion is not a formal specification, either of the architecture, or of a notion of biologically applied cost function that could be directly measured based on neural data. In order to move toward a more formal specification of the kind of model we are proposing here, it would be useful to map the architecture of the brain&#x00027;s reward systems and to identify other biological pathways that may mediate the generation and delivery of error signals. Based on such maps, one could identify regions which are proposed to be subject to a single cost function. Otherwise, the problem of inference of the cost function, e.g., based on neural dynamics becomes ill-posed: one can define a local cost function for an <italic>arbitrary</italic> dynamics by integrating the trajectory of the system, but this approach in general lacks explanatory power and also, crucially, lacks any circuit-level relationship with the brain&#x00027;s actual neural mechanisms of optimization, i.e., such a defined cost function does not necessarily correspond to the cost functions that the biological machinery is actually organized to optimize. Notably, some of the relevant biological pathways mediating cost functions and error signals may involve key biomolecular or gene expression aspects, not just real-time patterns of neural activity.</p>
<p>Another related consideration, in trying to formalize this type approach and to infer cost functions from neural measurements, is that not all neurons in the circuit may be subject to optimization: after all, some neurons may be needed to generate the error signals themselves, or to mediate the optimization process for other neurons, or to perform other unrelated functions. Furthermore, within a given region, there may be multiple sub-circuits subject to different optimization pressures. It is the claim that the brain actually has structured biological machinery to generate, route and apply specific cost functions that gives substance to our proposal, over and above the trivial claim that many kinds of dynamics can be viewed as optimizations, but our knowledge of this machinery is still limited. This is not to mention the difficulties involved in inferring cost functions in the presence of noise or constraints on the dynamics. Thus, one cannot blindly collect the neurons in an arbitrary region, measure their dynamics, and hope to infer their cost function by solving an inverse problem&#x02014;instead, a rich interplay between structural mapping, dynamic mapping, hypothesis generation, modeling and perturbation is likely to be necessary in order to gain a detailed knowledge of which cost functions the brain uses and how it does so.</p>
</sec>
<sec>
<title>5.3. <italic>Hypothesis 3</italic>&#x02013; embedding within a pre-structured architecture</title>
<p>If different brain structures are performing distinct types of computations with a shared goal, then optimization of a joint cost function will take place with different dynamics in each area. If we focus on a higher level task, e.g., maximizing the probability of correctly detecting something, then we should find that basic feature detection circuits should learn when the features were insufficient for detection, that attentional routing structures should learn when a different allocation of attention would have improved detection and that memory structures should learn when items that matter for detection were not remembered. If we assume that multiple structures are participating in a joint computation, which optimizes an overall cost function (but see Hypothesis 2), then an understanding of the computational function of each area leads to a prediction of the measurable plasticity rules.</p>
</sec>
</sec>
<sec id="s6">
<title>6. Neuroscience inspired machine learning</title>
<p>Machine learning may be equally transformed by neuroscience. Within the brain, a myriad of subsystems and layers work together to produce an agent that exhibits general intelligence. The brain is able to show intelligent behavior across a broad range of problems using only relatively small amounts of data. As such, progress at understanding the brain promises to improve machine learning. In this section, we review our three hypotheses about the brain and discuss how their elaboration might contribute to more powerful machine learning systems.</p>
<sec>
<title>6.1. <italic>Hypothesis 1</italic>&#x02013; existence of cost functions</title>
<p>A good practitioner of machine learning should have a broad range of optimization methods at their disposal as different problems ask for different approaches. The brain, we have argued, is an implicit machine learning mechanism which has been evolved over millions of years. Consequently, we should expect the brain to be able to optimize cost functions efficiently, across many domains and kinds of data. Indeed, across different animal phyla, we even see <italic>convergent</italic> evolution of certain brain structures (Shimizu and Karten, <xref ref-type="bibr" rid="B399">2013</xref>; G&#x000FC;nt&#x000FC;rk&#x000FC;n and Bugnyar, <xref ref-type="bibr" rid="B159">2016</xref>), e.g., the bird brain has no cortex yet has developed homologous structures which&#x02014;as the linguistic feats of the African Gray Parrot demonstrate&#x02014;can give rise to quite complex intelligence. It seems reasonable to hope to learn how to do truly general-purpose optimization by looking at the brain.</p>
<p>Indeed, there are multiple kinds of optimization that we may expect to discover by looking at the brain. At the hardware level, the brain clearly manages to optimize functions efficiently despite having slow hardware subject to molecular fluctuations, suggesting directions for improving the hardware of machine learning to be more energy efficient. At the level of learning rules, the brain solves an optimization problem in a highly nonlinear, non-differentiable, temporally stochastic, spiking system with massive numbers of feedback connections, a problem that we arguably still do not know how to efficiently solve for neural networks. At the architectural level, the brain can optimize certain kinds of functions based on very few stimulus presentations, operates over diverse timescales, and clearly uses advanced forms of active learning to infer causal structure in the world.</p>
<p>While we have discussed a range of theories (O&#x00027;Reilly, <xref ref-type="bibr" rid="B332">1996</xref>; K&#x000F6;rding and K&#x000F6;nig, <xref ref-type="bibr" rid="B230">2001</xref>; Hinton, <xref ref-type="bibr" rid="B184">2007</xref>, <xref ref-type="bibr" rid="B185">2016</xref>; Roelfsema et al., <xref ref-type="bibr" rid="B368">2010</xref>; Balduzzi et al., <xref ref-type="bibr" rid="B29">2014</xref>; Lillicrap et al., <xref ref-type="bibr" rid="B262">2014</xref>; O&#x00027;Reilly et al., <xref ref-type="bibr" rid="B335">2014a</xref>; Bengio et al., <xref ref-type="bibr" rid="B38">2015a</xref>) for how the brain can carry out optimization, these theories are still preliminary. Thus, the first step is to understand whether the brain indeed performs multi-layer credit assignment in a manner that approximates full gradient descent, and if so, how it does this. Either way, we can expect that answer to impact machine learning. If the brain does <italic>not</italic> do some form of backpropagation, this suggests that machine learning may benefit from understanding the tricks that the brain uses to avoid having to do so. If, on the other hand, the brain does do backpropagation, then the underlying mechanisms clearly can support a very wide range of efficient optimization processes across many domains, including learning from rich temporal data-streams and via unsupervised mechanisms, and the architectures behind this will likely be of long-term value to machine learning<xref ref-type="fn" rid="fn0061"><sup>61</sup></xref>. Moreover, the search for biologically plausible forms of backpropagation has already led to interesting insights, such as the possibility of using random feedback weights (feedback alignment) in backpropagation (Lillicrap et al., <xref ref-type="bibr" rid="B262">2014</xref>), or the unexpected power of internal FORCE learning in chaotic, spontaneously active recurrent networks (Sussillo and Abbott, <xref ref-type="bibr" rid="B419">2009</xref>). This and other findings discussed here suggest that there are still fundamental things we don&#x00027;t understand about backpropagation&#x02014;which could potentially lead not only to more biologically plausible ways to train recurrent neural networks, but also to fundamentally simpler and more powerful ones.</p>
</sec>
<sec>
<title>6.2. <italic>Hypothesis 2</italic>&#x02013; biological fine-structure of cost functions</title>
<p>A good practitioner of machine learning has access to a broad range of learning techniques and thus implicitly is able to use many different cost functions. Some problems ask for clustering, others for extracting sparse variables, and yet others for prediction quality to be maximized. The brain also needs to be able to deal with many different kinds of datasets. As such, it makes sense for the brain to use a broad range of cost functions appropriate for the diverse set of tasks it has to solve to thrive in this world.</p>
<p>Many of the most notable successes of deep learning, from language modeling (Sutskever et al., <xref ref-type="bibr" rid="B422">2011</xref>), to vision (Krizhevsky et al., <xref ref-type="bibr" rid="B237">2012</xref>), to motor control (Levine et al., <xref ref-type="bibr" rid="B257">2015</xref>), have been driven by end-to-end optimization of single task objectives. We have highlighted cases where machine learning has opened the door to multiplicities of cost functions that shape network modules into specialized roles. We expect that machine learning will increasingly adopt these practices in the future.</p>
<p>In computer vision, we have begun to see researchers re-appropriate neural networks trained for one task (e.g., ImageNet classification) and then deploy them on new tasks other than the ones they were trained for or for which more limited training data is available (Oquab et al., <xref ref-type="bibr" rid="B331">2014</xref>; Yosinski et al., <xref ref-type="bibr" rid="B481">2014</xref>; Noroozi and Favaro, <xref ref-type="bibr" rid="B324">2016</xref>). We imagine this procedure will be generalized, whereby, in series and in parallel, diverse training problems, each with an associated cost function, are used to shape visual representations. For example, visual data streams can be segmented into elements like foreground vs. background, objects that can move of their own accord vs. those that cannot, all using diverse unsupervised criteria (Ullman et al., <xref ref-type="bibr" rid="B443">2012</xref>; Poggio, <xref ref-type="bibr" rid="B353">2015</xref>). Networks so trained can then be shared, augmented, and retrained on new tasks. They can be introduced as front-ends for systems that perform more complex objectives or even serve to produce cost functions for training other circuits (Watter et al., <xref ref-type="bibr" rid="B457">2015</xref>). As a simple example, a network that can discriminate between images of different kinds of architectural structures (pyramid, staircase, etc.) could act as a critic for a building-construction network.</p>
<p>Scientifically, determining the order in which cost functions are engaged in the biological brain will inform machine learning about how to construct systems with intricate and hierarchical behaviors via divide-and-conquer approaches to learning problems, active learning, and more.</p>
</sec>
<sec>
<title>6.3. <italic>Hypothesis 3</italic>&#x02013; embedding within a pre-structured architecture</title>
<p>A good practitioner of machine learning should have a broad range of algorithms at their disposal. Some problems are efficiently solved through dynamic programming, others through hashing, and yet others through multi-layer backpropagation. The brain needs to be able to solve a broad range of learning problems without the luxury of being reprogrammed. As such, it makes sense for the brain to have specialized structures that allow it to rapidly learn to approximate a broad range of algorithms.</p>
<p>The first neural networks were simple single-layer systems, either linear or with limited non-linearities (Rashevsky, <xref ref-type="bibr" rid="B360">1939</xref>). The explosion of neural network research in the 1980s (Rumelhart et al., <xref ref-type="bibr" rid="B377">1986</xref>) saw the advent of multilayer networks, followed by networks with layer-wise specializations as in convolutional networks (Fukushima, <xref ref-type="bibr" rid="B133">1980</xref>; LeCun and Bengio, <xref ref-type="bibr" rid="B246">1995</xref>). In the last two decades, architectures with specializations for holding variables stable in memory like the LSTM (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B192">1997</xref>), the control of content-addressable memory (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>; Weston et al., <xref ref-type="bibr" rid="B464">2014</xref>), and game playing by reinforcement learning (Mnih et al., <xref ref-type="bibr" rid="B313">2015</xref>) have been developed. These networks, though formerly exotic, are now becoming mainstream algorithms in the toolbox of any deep learning practitioner. There is no sign that progress in developing new varieties of structured architectures is halting, and the heterogeneity and modularity of the brain&#x00027;s circuitry suggests that diverse, specialized architectures are needed to solve the diverse challenges that confront a behaving animal.</p>
<p>The brain combines a jumble of specialized structures in a way that works. Solving this problem <italic>de novo</italic> in machine learning promises to be very difficult, making it attractive to be inspired by observations about how the brain does it. An understanding of the breadth of specialized structures, as well as the architecture that combines them, should be quite useful.</p>
</sec>
</sec>
<sec id="s7">
<title>7. Did evolution separate cost functions from optimization algorithms?</title>
<p>Deep learning methods have taken the field of machine learning by storm. Driving the success is the separation of the problem of learning into two pieces: (1) An algorithm, backpropagation, that allows efficient distributed optimization, and (2) Approaches to turn any given problem into an optimization problem, by designing a cost function and training procedure which will result in the desired computation. If we want to apply deep learning to a new domain, e.g., playing Jeopardy, we do not need to change the optimization algorithm&#x02014;we just need to cleverly set up the right cost function. A lot of work in deep learning, perhaps the majority, is now focused on setting up the right cost functions.</p>
<p>We hypothesize that the brain also acquired such a separation between optimization mechanisms and cost functions. If neural circuits, such as in cortex, implement a general-purpose optimization algorithm, then any improvement to that algorithm will improve function across the cortex. At the same time, different cortical areas solve different problems, so tinkering with each area&#x00027;s cost function is likely to improve its performance. As such, functionally and evolutionarily separating the problems of optimization and cost function generation could allow evolution to produce better computations, faster. For example, common unsupervised mechanisms could be combined with area-specific reinforcement-based or supervised mechanisms and error signals, much as recent advances in machine learning have found natural ways to combine supervised and unsupervised objectives in a single system (Rasmus and Berglund, <xref ref-type="bibr" rid="B361">2015</xref>).</p>
<p>This suggests interesting questions<xref ref-type="fn" rid="fn0062"><sup>62</sup></xref>: When did the division between cost functions and optimization algorithms occur? How is this separation implemented? How did innovations in cost functions and optimization algorithms evolve? And how do our own cost functions and learning algorithms differ from those of other animals?</p>
<p>There are many possibilities for how such a separation might be achieved in the brain. Perhaps the six-layered cortex represents a common optimization algorithm, which in different cortical areas is supplied with different cost functions. This claim is different from the claim that all cortical areas use a single unsupervised learning algorithm and achieve functional specificity by tuning the inputs to that algorithm. In that case, both the optimization mechanism and the implicit unsupervised cost function would be the same across areas (e.g., minimization of prediction error), with only the training data differing between areas, whereas in our suggestion, the optimization mechanism would be the same across areas but the cost function, <italic>as well as</italic> the training data, would differ. Thus the cost function itself would be like an ancillary input to a cortical area, in addition to its input and output data. Some cortical microcircuits could then, perhaps, compute the cost functions that are to be delivered to other cortical microcircuits. Another possibility is that, within the same circuitry, certain aspects of the wiring and learning rules specify an optimization mechanism and are relatively fixed across areas, while others specify the cost function and are more variable. This latter possibility would be similar to the notion of cortical microcircuits as molecularly and structurally configurable elements, akin to the cells in a field-programmable gate array (FPGA) (Marcus et al., <xref ref-type="bibr" rid="B284">2014a</xref>,<xref ref-type="bibr" rid="B285">b</xref>), rather than a homogenous substrate. The biological nature of such a separation, if any exists, remains an open question. For example, individual parts of a neuron may separately deal with optimization and with the specification of the cost function, or different parts of a microcircuit may specialize in this way, or there may be specialized types of cells, some of which deal with signal processing and others with cost functions.</p>
</sec>
<sec sec-type="conclusions" id="s8">
<title>8. Conclusions</title>
<p>Due to the complexity and variability of the brain, pure &#x0201C;bottom up&#x0201D; analysis of neural data faces potential challenges of interpretation (Robinson, <xref ref-type="bibr" rid="B365">1992</xref>; Jonas and Kording, <xref ref-type="bibr" rid="B216">2016</xref>). Theoretical frameworks can potentially be used to constrain the space of hypotheses being evaluated, allowing researchers to first address higher-level principles and structures in the system, and then &#x0201C;zoom in&#x0201D; to address the details. Proposed &#x0201C;top down&#x0201D; frameworks for understanding neural computation include entropy maximization, efficient encoding, faithful approximation of Bayesian inference, minimization of prediction error, attractor dynamics, modularity, the ability to subserve symbolic operations, and many others (Pinker, <xref ref-type="bibr" rid="B351">1999</xref>; Marcus, <xref ref-type="bibr" rid="B282">2001</xref>; Bialek, <xref ref-type="bibr" rid="B44">2002</xref>; Knill and Pouget, <xref ref-type="bibr" rid="B225">2004</xref>; Bialek et al., <xref ref-type="bibr" rid="B45">2006</xref>; Friston, <xref ref-type="bibr" rid="B131">2010</xref>). Interestingly, many of the &#x0201C;top down&#x0201D; frameworks boil down to assuming that the brain simply optimizes a single, given cost function for a single computational architecture. We generalize these proposals assuming both a heterogeneous combination of cost functions unfolding over development, and a diversity of specialized sub-systems.</p>
<p>Much of neuroscience has focused on the search for &#x0201C;the neural code,&#x0201D; i.e., it has asked which stimuli are good at driving activity in individual neurons, regions, or brain areas. But, if the brain is capable of generic optimization of cost functions, then we need to be aware that rather simple cost functions can give rise to complicated stimulus responses. This potentially leads to a different set of questions. Are differing cost functions indeed a useful way to think about the differing functions of brain areas? How does the optimization of cost functions in the brain actually occur, and how is this different from the implementations of gradient descent in artificial neural networks? What additional constraints are present in the circuitry that remain fixed while optimization occurs? How does optimization interact with a structured architecture, and is this architecture similar to what we have sketched? Which computations are wired into the architecture, which emerge through optimization, and which arise from a mixture of those two extremes? To what extent are cost functions explicitly computed in the brain, vs. implicit in its local learning rules? Did the brain evolve to separate the mechanisms involved in cost function generation from those involved in the optimization of cost functions, and if so how? What kinds of meta-level learning might the brain apply, to learn when and how to invoke different cost functions or specialized systems, among the diverse options available, to solve a given task? What crucial mechanisms are left out of this framework? A more in-depth dialog between neuroscience and machine learning could help elucidate some of these questions.</p>
<p>Much of machine learning has focused on finding ever faster ways of doing end-to-end gradient descent in neural networks. Neuroscience may inform machine learning at multiple levels. The optimization algorithms in the brain have undergone a couple of hundred million years of evolution. Moreover, the brain may have found ways of using heterogeneous cost functions that interact over development so as to simplify learning problems by guiding and shaping the outcomes of unsupervised learning. Lastly, the specialized structures evolved in the brain may inform us about ways of making learning efficient in a world that requires a broad range of computational problems to be solved over multiple timescales. Looking at the insights from neuroscience may help machine learning move toward general intelligence in a structured heterogeneous world with access to only small amounts of supervised data.</p>
<p>In some ways our proposal is opposite to many popular theories of neural computation. There is not one mechanism of optimization but (potentially) many, not one cost function but a host of them, not one kind of a representation but a representation of whatever is useful, and not one homogeneous structure but a large number of them. All these elements are held together by the optimization of internally generated cost functions, which allows these systems to make good use of one another. Rejecting simple unifying theories is in line with a broad range of previous approaches in AI. For example, Minsky and Papert&#x00027;s work on the Society of Mind (Minsky, <xref ref-type="bibr" rid="B305">1988</xref>)&#x02014;and more broadly on ideas of genetically staged and internally bootstrapped development in connectionist systems (Minsky, <xref ref-type="bibr" rid="B304">1977</xref>)&#x02014;emphasizes the need for a system of internal monitors and critics, specialized communication and storage mechanisms, and a hierarchical organization of simple control systems.</p>
<p>At the time these early works were written, it was not yet clear that gradient-based optimization could give rise to powerful feature representations and behavioral policies. One can view our proposal as a renewed argument against simple end-to-end training and in favor of a heterogeneous approach. In other words, this framework could be viewed as proposing a kind of &#x0201C;society&#x0201D; of cost functions and trainable networks, permitting internal bootstrapping processes reminiscent of the Society of Mind (Minsky, <xref ref-type="bibr" rid="B305">1988</xref>). In this view, intelligence is enabled by many computationally specialized structures, each trained with its own developmentally regulated cost function, where both the structures and the cost functions are themselves optimized by evolution like the hyperparameters in neural networks.</p>
</sec>
<sec id="s9">
<title>Author contribution</title>
<p>All authors contributed ideas and co-wrote the paper.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>We thank Ken Hayworth for key discussions that led to this paper. We thank Ed Boyden, Chris Eliasmith, Gary Marcus, Shimon Ullman, Tomaso Poggio, Josh Tenenbaum, Dario Amodei, Alex Williams, Erik Peterson, Tom Dean, David Sussillo, Matthew Botvinick, Joscha Bach, Mohammad Gheshlaghi Azar, Joshua Glaser, Marco Nardini, Ali Hummos, David Markowitz, David Rolnick, Sam Rodriques, Nick Barry, Matthew Larkum, Walter Senn, Eric Drexler, Vikash Mansinghka, Darcy Wayne, Lyra and Neo Marblestone, and all of the participants of a Kavli Salon on Cortical Computation (Feb/Oct 2015) for helpful comments. We thank Miles Brundage for an excellent Twitter feed of deep learning papers. We acknowledge the support of NIH grant R01MH103910.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Abbott</surname> <given-names>L.</given-names></name> <name><surname>DePasquale</surname> <given-names>B.</given-names></name> <name><surname>Memmesheimer</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <source>Building Functional Networks of Spiking Model Neurons</source>. Available online at: neurotheory.columbia.edu.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abbott</surname> <given-names>L. F.</given-names></name> <name><surname>Blum</surname> <given-names>K. I.</given-names></name></person-group> (<year>1996</year>). <article-title>Functional significance of long-term potentiation for sequence learning and prediction</article-title>. <source>Cereb. Cortex</source> <volume>6</volume>, <fpage>406</fpage>&#x02013;<lpage>416</lpage>. <pub-id pub-id-type="pmid">8670667</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ackley</surname> <given-names>D.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T.</given-names></name></person-group> (<year>1958</year>). <article-title>A learning algorithm for Boltzmann machines</article-title>. <source>Cogn. Sci.</source> <volume>9</volume>, <fpage>147</fpage>&#x02013;<lpage>169</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Acuna</surname> <given-names>D. E.</given-names></name> <name><surname>Wymbs</surname> <given-names>N. F.</given-names></name> <name><surname>Reynolds</surname> <given-names>C. A.</given-names></name> <name><surname>Picard</surname> <given-names>N.</given-names></name> <name><surname>Turner</surname> <given-names>R. S.</given-names></name> <name><surname>Strick</surname> <given-names>P. L.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Multifaceted aspects of chunking enable robust algorithms</article-title>. <source>J. Neurophysiol.</source> <volume>112</volume>, <fpage>1849</fpage>&#x02013;<lpage>1856</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00028.2014</pub-id><pub-id pub-id-type="pmid">25080566</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Alain</surname> <given-names>G.</given-names></name> <name><surname>Lamb</surname> <given-names>A.</given-names></name> <name><surname>Sankar</surname> <given-names>C.</given-names></name> <name><surname>Courville</surname> <given-names>A.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <source>Variance reduction in SGD by distributed importance sampling</source>. arXiv:1511.06481.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Allen</surname> <given-names>K.</given-names></name> <name><surname>Ibara</surname> <given-names>S.</given-names></name> <name><surname>Seymour</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Abstract structural representations of goal-directed behavior</article-title>. <source>Psychol. Sci.</source> <volume>21</volume>, <fpage>1518</fpage>&#x02013;<lpage>1524</lpage>. <pub-id pub-id-type="doi">10.1177/0956797610383434</pub-id><pub-id pub-id-type="pmid">20855906</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>C. H.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name></person-group> (<year>1987</year>). <article-title>Shifter circuits: a computational strategy for dynamic aspects of visual processing</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>84</volume>, <fpage>6297</fpage>&#x02013;<lpage>6301</lpage>. <pub-id pub-id-type="pmid">3114747</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>J. R.</given-names></name></person-group> (<year>2007</year>). <source>How Can the Human Mind Occur in the Physical Universe?</source> <publisher-loc>Oxford, UK</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B9">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Andreas</surname> <given-names>J.</given-names></name> <name><surname>Rohrbach</surname> <given-names>M.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Klein</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <source>Deep compositional question answering with neural module networks</source>. arXiv:1511.02799.</citation>
</ref>
<ref id="B10">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Andreas</surname> <given-names>J.</given-names></name> <name><surname>Rohrbach</surname> <given-names>M.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Klein</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <source>Learning to compose neural networks for question answering</source>. arXiv:1601.01705.</citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Angelucci</surname> <given-names>A.</given-names></name> <name><surname>Levitt</surname> <given-names>J. B.</given-names></name> <name><surname>Walton</surname> <given-names>E. J.</given-names></name> <name><surname>Hupe</surname> <given-names>J.-M.</given-names></name> <name><surname>Bullier</surname> <given-names>J.</given-names></name> <name><surname>Lund</surname> <given-names>J. S.</given-names></name></person-group> (<year>2002</year>). <article-title>Circuits for local and global signal integration in primary visual cortex</article-title>. <source>J. Neurosci.</source> <volume>22</volume>, <fpage>8633</fpage>&#x02013;<lpage>8646</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.jneurosci.org/content/22/19/8633.long">http://www.jneurosci.org/content/22/19/8633.long</ext-link></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anselmi</surname> <given-names>F.</given-names></name> <name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Rosasco</surname> <given-names>L.</given-names></name> <name><surname>Mutch</surname> <given-names>J.</given-names></name> <name><surname>Tacchetti</surname> <given-names>A.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>Unsupervised learning of invariant representations</article-title>. <source>Theor. Comput. Sci.</source> <volume>633</volume>, <fpage>112</fpage>&#x02013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.1016/j.tcs.2015.06.048</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Antic</surname> <given-names>S. D.</given-names></name> <name><surname>Zhou</surname> <given-names>W.-L.</given-names></name> <name><surname>Moore</surname> <given-names>A. R.</given-names></name> <name><surname>Short</surname> <given-names>S. M.</given-names></name> <name><surname>Ikonomu</surname> <given-names>K. D.</given-names></name></person-group> (<year>2010</year>). <article-title>The decade of the dendritic nmda spike</article-title>. <source>J. Neurosci. Res.</source> <volume>88</volume>, <fpage>2991</fpage>&#x02013;<lpage>3001</lpage>. <pub-id pub-id-type="doi">10.1002/jnr.22444</pub-id><pub-id pub-id-type="pmid">20544831</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arancio</surname> <given-names>O.</given-names></name> <name><surname>Kiebler</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>C. J.</given-names></name> <name><surname>Lev-Ram</surname> <given-names>V.</given-names></name> <name><surname>Tsien</surname> <given-names>R. Y.</given-names></name> <name><surname>Kandel</surname> <given-names>E. R.</given-names></name> <etal/></person-group>. (<year>1996</year>). <article-title>Nitric oxide acts directly in the presynaptic neuron to produce long-term potentiation in cultured hippocampal neurons</article-title>. <source>Cell</source> <volume>87</volume>, <fpage>1025</fpage>&#x02013;<lpage>1035</lpage>. <pub-id pub-id-type="pmid">8978607</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aronov</surname> <given-names>D.</given-names></name> <name><surname>Veit</surname> <given-names>L.</given-names></name> <name><surname>Goldberg</surname> <given-names>J. H.</given-names></name> <name><surname>Fee</surname> <given-names>M. S.</given-names></name></person-group> (<year>2011</year>). <article-title>Two distinct modes of forebrain circuit dynamics underlie temporal patterning in the vocalizations of young songbirds</article-title>. <source>J. Neurosci.</source> <volume>31</volume>, <fpage>16353</fpage>&#x02013;<lpage>16368</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3009-11.2011</pub-id><pub-id pub-id-type="pmid">22072687</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Arora</surname> <given-names>S.</given-names></name> <name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Ma</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <source>Why are deep nets reversible: a simple theory, with implications for training</source>. arXiv:1511.05653.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashby</surname> <given-names>F. G.</given-names></name> <name><surname>Ennis</surname> <given-names>J. M.</given-names></name> <name><surname>Spiering</surname> <given-names>B. J.</given-names></name></person-group> (<year>2007</year>). <article-title>A neurobiological theory of automaticity in perceptual categorization</article-title>. <source>Psychol. Rev.</source> <volume>114</volume>:<fpage>632</fpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.114.3.632</pub-id><pub-id pub-id-type="pmid">17638499</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashby</surname> <given-names>F. G.</given-names></name> <name><surname>Turner</surname> <given-names>B. O.</given-names></name> <name><surname>Horvitz</surname> <given-names>J. C.</given-names></name></person-group> (<year>2010</year>). <article-title>Cortical and basal ganglia contributions to habit learning and automaticity</article-title>. <source>Trends Cogn. Sci.</source> <volume>14</volume>, <fpage>208</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.02.001</pub-id><pub-id pub-id-type="pmid">20207189</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aso</surname> <given-names>Y.</given-names></name> <name><surname>Hattori</surname> <given-names>D.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Johnston</surname> <given-names>R. M.</given-names></name> <name><surname>Iyer</surname> <given-names>N. A.</given-names></name> <name><surname>Ngo</surname> <given-names>T.-T. B.</given-names></name> <etal/></person-group>. (<year>2014a</year>). <article-title>The neuronal architecture of the mushroom body provides a logic for associative learning</article-title>. <source>eLife</source> <volume>3</volume>:<fpage>e04577</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.04577</pub-id><pub-id pub-id-type="pmid">25535793</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aso</surname> <given-names>Y.</given-names></name> <name><surname>Sitaraman</surname> <given-names>D.</given-names></name> <name><surname>Ichinose</surname> <given-names>T.</given-names></name> <name><surname>Kaun</surname> <given-names>K. R.</given-names></name> <name><surname>Vogt</surname> <given-names>K.</given-names></name> <name><surname>Belliart-Gu&#x000E9;rin</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2014b</year>). <article-title>Mushroom body output neurons encode valence and guide memory-based action selection in Drosophila</article-title>. <source>eLife</source> <volume>3</volume>:<fpage>e04580</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.04580</pub-id><pub-id pub-id-type="pmid">25535794</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bach</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Modeling motivation in MicroPsi 2</article-title>, in <source>8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, Proceedings</source>, <volume>Vol. 9205</volume> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>3</fpage>&#x02013;<lpage>13</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bach</surname> <given-names>J.</given-names></name> <name><surname>Herger</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Request confirmation networks for neuro-symbolic script execution</article-title>, in <source>Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches at NIPS</source>, eds <person-group person-group-type="editor"><name><surname>Besold</surname> <given-names>T.</given-names></name> <name><surname>Avila Garcez</surname> <given-names>A.</given-names></name> <name><surname>Marcus</surname> <given-names>G.</given-names></name> <name><surname>Miikkulainen</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Montreal, QC</publisher-loc>).</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baillargeon</surname> <given-names>R.</given-names></name> <name><surname>Scott</surname> <given-names>R. M.</given-names></name> <name><surname>Bian</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Psychological reasoning in infancy</article-title>. <source>Annu. Rev. Psychol.</source> <volume>67</volume>, <fpage>159</fpage>&#x02013;<lpage>186</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-psych-010213-115033</pub-id><pub-id pub-id-type="pmid">26393869</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Ba</surname> <given-names>J.</given-names></name> <name><surname>Caruana</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Do deep nets really need to be deep?</article-title> <source>Adv. Neural Inform. Process.</source> <volume>27</volume>, <fpage>2654</fpage>&#x02013;<lpage>2662</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://papers.nips.cc/paper/5484-do-deep-nets-really-need-to-be-deep">https://papers.nips.cc/paper/5484-do-deep-nets-really-need-to-be-deep</ext-link></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balaguer</surname> <given-names>J.</given-names></name> <name><surname>Spiers</surname> <given-names>H.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Summerfield</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>Neural mechanisms of hierarchical planning in a virtual subway network</article-title>. <source>Neuron</source> <volume>90</volume>, <fpage>893</fpage>&#x02013;<lpage>903</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.03.037</pub-id><pub-id pub-id-type="pmid">27196978</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baldauf</surname> <given-names>D.</given-names></name> <name><surname>Desimone</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural mechanisms of object-based attention</article-title>. <source>Science</source> <volume>344</volume>, <fpage>424</fpage>&#x02013;<lpage>427</lpage>. <pub-id pub-id-type="doi">10.1126/science.1247003</pub-id><pub-id pub-id-type="pmid">24763592</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Baldi</surname> <given-names>P.</given-names></name> <name><surname>Sadowski</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <source>The Ebb and flow of deep learning: a theory of local learning</source>. arXiv:1506.06472.</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Balduzzi</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>Cortical prediction markets</article-title>, in <source>Proceedings of the 2014 International Conference on Autonomous AgentsMultiagent Systems (AAMAS)</source> (<publisher-loc>Paris</publisher-loc>).</citation>
</ref>
<ref id="B29">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Balduzzi</surname> <given-names>D.</given-names></name> <name><surname>Vanchinathan</surname> <given-names>H.</given-names></name> <name><surname>Buhmann</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <source>Kickback cuts Backprop&#x00027;s red-tape: biologically plausible credit assignment in neural networks</source>. arXiv:1411.6191.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bargmann</surname> <given-names>C. I.</given-names></name></person-group> (<year>2012</year>). <article-title>Beyond the connectome: how neuromodulators shape neural circuits</article-title>. <source>Bioessays</source> <volume>34</volume>, <fpage>458</fpage>&#x02013;<lpage>465</lpage>. <pub-id pub-id-type="doi">10.1002/bies.201100185</pub-id><pub-id pub-id-type="pmid">22396302</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bargmann</surname> <given-names>C. I.</given-names></name> <name><surname>Marder</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>From the connectome to brain function</article-title>. <source>Nat. Methods</source> <volume>10</volume>, <fpage>483</fpage>&#x02013;<lpage>490</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2451</pub-id><pub-id pub-id-type="pmid">23866325</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bastos</surname> <given-names>A. M.</given-names></name> <name><surname>Vezoli</surname> <given-names>J.</given-names></name> <name><surname>Bosman</surname> <given-names>C. A.</given-names></name> <name><surname>Schoffelen</surname> <given-names>J.-M.</given-names></name> <name><surname>Oostenveld</surname> <given-names>R.</given-names></name> <name><surname>Dowdall</surname> <given-names>J. R.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Visual areas exert feedforward and feedback influences through distinct frequency channels</article-title>. <source>Neuron</source> <volume>85</volume>, <fpage>390</fpage>&#x02013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2014.12.018</pub-id><pub-id pub-id-type="pmid">25556836</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Battaglia</surname> <given-names>P. W.</given-names></name> <name><surname>Hamrick</surname> <given-names>J. B.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2013</year>). <article-title>Simulation as an engine of physical scene understanding</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>110</volume>, <fpage>18327</fpage>&#x02013;<lpage>18332</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1306572110</pub-id><pub-id pub-id-type="pmid">24145417</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Becker</surname> <given-names>S.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>1992</year>). <article-title>Self-organizing neural network that discovers surfaces in random-dot stereograms</article-title>. <source>Nature</source> <volume>355</volume>, <fpage>161</fpage>&#x02013;<lpage>163</lpage>. <pub-id pub-id-type="pmid">1729650</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bekolay</surname> <given-names>T.</given-names></name> <name><surname>Kolbeck</surname> <given-names>C.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Simultaneous unsupervised and supervised learning of cognitive functions in biologically plausible spiking neural networks</article-title>, in <source>Proceedings of the 35th Annual Conference of the Cognitive Science Society</source> (<publisher-loc>Berlin</publisher-loc>), <fpage>169</fpage>&#x02013;<lpage>174</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <source>How auto-encoders could provide credit assignment in deep networks via target propagation</source>. arXiv:1407.7906.</citation>
</ref>
<ref id="B37">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Fischer</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <source>Early inference in energy-based models approximates back-propagation</source>. arXiv:1510.02777.</citation>
</ref>
<ref id="B38">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Lee</surname> <given-names>D.-H.</given-names></name> <name><surname>Bornschein</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>Z.</given-names></name></person-group> (<year>2015a</year>). <source>Towards biologically plausible deep learning</source>. arXiv:1502.04156.</citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Louradour</surname> <given-names>J.</given-names></name> <name><surname>Collobert</surname> <given-names>R.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>Curriculum learning</article-title>, in <source>Proceedings of the 26th Annual International Conference on Machine Learning</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>41</fpage>&#x02013;<lpage>48</lpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Mesnard</surname> <given-names>T.</given-names></name> <name><surname>Fischer</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name></person-group> (<year>2015b</year>). <source>STDP as presynaptic activity times rate of change of postsynaptic activity</source>. arXiv:1509.05936.</citation>
</ref>
<ref id="B41">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Scellier</surname> <given-names>B.</given-names></name> <name><surname>Bilaniuk</surname> <given-names>O.</given-names></name> <name><surname>Sacramento</surname> <given-names>J.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <source>Feedforward initialization for fast inference of deep generative networks is biologically plausible</source>. arXiv:1606.01651.</citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berezovskii</surname> <given-names>V. K.</given-names></name> <name><surname>Nassi</surname> <given-names>J. J.</given-names></name> <name><surname>Born</surname> <given-names>R. T.</given-names></name></person-group> (<year>2011</year>). <article-title>Segregation of feedforward and feedback projections in mouse visual cortex</article-title>. <source>J. Comp. Neurol.</source> <volume>519</volume>, <fpage>3672</fpage>&#x02013;<lpage>3683</lpage>. <pub-id pub-id-type="doi">10.1002/cne.22675</pub-id><pub-id pub-id-type="pmid">21618232</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berwick</surname> <given-names>R. C.</given-names></name> <name><surname>Beckers</surname> <given-names>G. J. L.</given-names></name> <name><surname>Okanoya</surname> <given-names>K.</given-names></name> <name><surname>Bolhuis</surname> <given-names>J. J.</given-names></name></person-group> (<year>2012</year>). <article-title>A bird&#x00027;s eye view of human language evolution</article-title>. <source>Front. Evol. Neurosci.</source> <volume>4</volume>:<issue>5</issue>. <pub-id pub-id-type="doi">10.3389/fnevo.2012.00005</pub-id><pub-id pub-id-type="pmid">22518103</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bialek</surname> <given-names>W.</given-names></name></person-group> (<year>2002</year>). <article-title>Thinking about the brain</article-title>, in <source>Physics of Bio-Molecules and Cells</source>, <volume>Vol. 75</volume>, eds <person-group person-group-type="editor"><name><surname>Flyvbjerg</surname> <given-names>F.</given-names></name> <name><surname>J&#x000FC;licher</surname> <given-names>F.</given-names></name> <name><surname>Ormos</surname> <given-names>P.</given-names></name> <name><surname>David</surname> <given-names>F.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>485</fpage>&#x02013;<lpage>578</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bialek</surname> <given-names>W.</given-names></name> <name><surname>De Ruyter Van Steveninck</surname> <given-names>R.</given-names></name> <name><surname>Tishby</surname> <given-names>N.</given-names></name></person-group> (<year>2006</year>). <article-title>Efficient representation as a design principle for neural coding and computation</article-title>, in <source>2006 IEEE International Symposium on Information Theory</source>, <volume>Vol. 75</volume> (<publisher-loc>Los Alamitos</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>659</fpage>&#x02013;<lpage>663</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bloss</surname> <given-names>E. B.</given-names></name> <name><surname>Cembrowski</surname> <given-names>M. S.</given-names></name> <name><surname>Karsh</surname> <given-names>B.</given-names></name> <name><surname>Colonell</surname> <given-names>J.</given-names></name> <name><surname>Fetter</surname> <given-names>R. D.</given-names></name> <name><surname>Spruston</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>Structured dendritic inhibition supports branch-selective integration in CA1 pyramidal cells</article-title>. <source>Neuron</source> <volume>89</volume>, <fpage>1016</fpage>&#x02013;<lpage>1030</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.01.029</pub-id><pub-id pub-id-type="pmid">26898780</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Blundell</surname> <given-names>C.</given-names></name> <name><surname>Uria</surname> <given-names>B.</given-names></name> <name><surname>Pritzel</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Ruderman</surname> <given-names>A.</given-names></name> <name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>Model-free episodic control</source>. arXiv:1606.04460.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bobier</surname> <given-names>B.</given-names></name> <name><surname>Stewart</surname> <given-names>T. C.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>A unifying mechanistic model of selective attention in spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003577</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003577</pub-id><pub-id pub-id-type="pmid">24921249</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bostrom</surname> <given-names>N.</given-names></name></person-group> (<year>1996</year>). <source>Cortical integration: possible solutions to the binding and linking problems in perception, reasoning and long term memory</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.nickbostrom.com/old/cortical.html">http://www.nickbostrom.com/old/cortical.html</ext-link></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Niv</surname> <given-names>Y.</given-names></name> <name><surname>Barto</surname> <given-names>A. C.</given-names></name></person-group> (<year>2009</year>). <article-title>Hierarchically organized behavior and its neural foundations: a reinforcement learning perspective</article-title>. <source>Cognition</source> <volume>113</volume>, <fpage>262</fpage>&#x02013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2008.08.011</pub-id><pub-id pub-id-type="pmid">18926527</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M.</given-names></name> <name><surname>Weinstein</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Model-based hierarchical reinforcement learning and human action control</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>369</volume>:<fpage>20130480</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2013.0480</pub-id><pub-id pub-id-type="pmid">25267822</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bouchard</surname> <given-names>G.</given-names></name> <name><surname>Trouillon</surname> <given-names>T.</given-names></name> <name><surname>Perez</surname> <given-names>J.</given-names></name> <name><surname>Gaidon</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <source>Accelerating stochastic gradient descent via online learning to sample</source>. arXiv:1506.09016.</citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bourdoukan</surname> <given-names>R.</given-names></name> <name><surname>Den&#x000E8;ve</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Enforcing balance allows local supervised learning in spiking recurrent networks</article-title>,&#x0201D; in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>982</fpage>&#x02013;<lpage>990</lpage>.</citation>
</ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Braitenberg</surname> <given-names>V.</given-names></name> <name><surname>Schutz</surname> <given-names>A.</given-names></name></person-group> (<year>1991</year>). <source>Anatomy of the Cortex: Studies of Brain Function</source>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brea</surname> <given-names>J.</given-names></name> <name><surname>Ga&#x000E1;l</surname> <given-names>A. T.</given-names></name> <name><surname>Urbanczik</surname> <given-names>R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Prospective coding by spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1005003</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1005003</pub-id><pub-id pub-id-type="pmid">27341100</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brea</surname> <given-names>J.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Does computational neuroscience need new synaptic learning paradigms?</article-title> <source>Curr. Opin. Behav. Sci.</source> <volume>11</volume>, <fpage>61</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1016/j.cobeha.2016.05.012</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bremner</surname> <given-names>J. G.</given-names></name> <name><surname>Slater</surname> <given-names>A. M.</given-names></name> <name><surname>Johnson</surname> <given-names>S. P.</given-names></name></person-group> (<year>2015</year>). <article-title>Perception of object persistence: the origins of object permanence in infancy</article-title>. <source>Child Dev. Perspect.</source> <volume>9</volume>, <fpage>7</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1111/cdep.12098</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Brito</surname> <given-names>C. S.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <source>Nonlinear hebbian learning as a unifying principle in receptive field formation</source>. arXiv:1601.00701.</citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brosch</surname> <given-names>T.</given-names></name> <name><surname>Neumann</surname> <given-names>H.</given-names></name> <name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name></person-group> (<year>2015</year>). <article-title>Reinforcement learning of linking and tracing contours in recurrent neural networks</article-title>. <source>PLoS Comput. Biol.</source> <volume>11</volume>:<fpage>e1004489</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004489</pub-id><pub-id pub-id-type="pmid">26496502</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brownstone</surname> <given-names>R. M.</given-names></name> <name><surname>Bui</surname> <given-names>T. V.</given-names></name> <name><surname>Stifani</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Spinal circuits for motor learning</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>33</volume>, <fpage>166</fpage>&#x02013;<lpage>173</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2015.04.007</pub-id><pub-id pub-id-type="pmid">25978563</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>T. I.</given-names></name> <name><surname>Carr</surname> <given-names>V. A.</given-names></name> <name><surname>LaRocque</surname> <given-names>K. F.</given-names></name> <name><surname>Favila</surname> <given-names>S. E.</given-names></name> <name><surname>Gordon</surname> <given-names>A. M.</given-names></name> <name><surname>Bowles</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Prospective representation of navigational goals in the human hippocampus</article-title>. <source>Science</source> <volume>352</volume>, <fpage>1323</fpage>&#x02013;<lpage>1326</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaf0784</pub-id><pub-id pub-id-type="pmid">27284194</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Bill</surname> <given-names>J.</given-names></name> <name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>Neural dynamics as sampling: a model for stochastic computation in recurrent networks of spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1002211</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002211</pub-id><pub-id pub-id-type="pmid">22096452</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buonomano</surname> <given-names>D. V.</given-names></name> <name><surname>Merzenich</surname> <given-names>M. M.</given-names></name></person-group> (<year>1995</year>). <article-title>Temporal information transformed into a spatial code by a neural network with realistic properties</article-title>. <source>Science</source> <volume>267</volume>, <fpage>1028</fpage>&#x02013;<lpage>1030</lpage>. <pub-id pub-id-type="pmid">7863330</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>B&#x000FC;lthoff</surname> <given-names>H.</given-names></name> <name><surname>Little</surname> <given-names>J.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>1989</year>). <article-title>A parallel algorithm for real-time computation of optical flow</article-title>. <source>Nature</source> <volume>337</volume>, <fpage>549</fpage>&#x02013;<lpage>553</lpage>. <pub-id pub-id-type="pmid">2915704</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buschman</surname> <given-names>T. J.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2010</year>). <article-title>Shifting the spotlight of attention: evidence for discrete computations in cognition</article-title>. <source>Front. Hum. Neurosci.</source> <volume>4</volume>:<issue>194</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2010.00194</pub-id><pub-id pub-id-type="pmid">21119775</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buschman</surname> <given-names>T. J.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2014</year>). <article-title>Goal-direction and top-down control</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>369</volume>:<fpage>20130471</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2013.0471</pub-id><pub-id pub-id-type="pmid">25267814</pub-id></citation>
</ref>
<ref id="B67">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Bush</surname> <given-names>K. A.</given-names></name></person-group> (<year>2007</year>). <source>An Echo State Model of Non-markovian Reinforcement Learning</source>. Doctoral Dissertation. Colorado State University.</citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buzs&#x000E1;ki</surname> <given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>Neural syntax: cell assemblies, synapsembles, and readers</article-title>. <source>Neuron</source> <volume>68</volume>, <fpage>362</fpage>&#x02013;<lpage>385</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2010.09.023</pub-id><pub-id pub-id-type="pmid">21040841</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buzs&#x000E1;ki</surname> <given-names>G.</given-names></name> <name><surname>Moser</surname> <given-names>E. I.</given-names></name></person-group> (<year>2013</year>). <article-title>Memory, navigation and theta rhythm in the hippocampal-entorhinal system</article-title>. <source>Nat. Neurosci.</source> <volume>16</volume>, <fpage>130</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3304</pub-id><pub-id pub-id-type="pmid">23354386</pub-id></citation>
</ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>D. J.</given-names></name> <name><surname>Aharoni</surname> <given-names>D.</given-names></name> <name><surname>Shuman</surname> <given-names>T.</given-names></name> <name><surname>Shobe</surname> <given-names>J.</given-names></name> <name><surname>Biane</surname> <given-names>J.</given-names></name> <name><surname>Song</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>A shared neural ensemble links distinct contextual memories encoded close in time</article-title>. <source>Nature</source> <volume>534</volume>, <fpage>115</fpage>&#x02013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.1038/nature17955</pub-id><pub-id pub-id-type="pmid">27251287</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Callaway</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Feedforward, feedback and inhibitory connections in primate visual cortex</article-title>. <source>Neural Netw.</source> <volume>17</volume>, <fpage>625</fpage>&#x02013;<lpage>632</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2004.04.004</pub-id><pub-id pub-id-type="pmid">15288888</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cappe</surname> <given-names>C.</given-names></name> <name><surname>Rouiller</surname> <given-names>E. M.</given-names></name> <name><surname>Barone</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <article-title>The neural bases of multisensory processes</article-title>, in <source>Cortical and Thalamic Pathways for Multisensory and Sensorimotor Interplay</source>, eds <person-group person-group-type="editor"><name><surname>Murray</surname> <given-names>M. M.</given-names></name> <name><surname>Wallace</surname> <given-names>M. T.</given-names></name></person-group> (<publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press/Taylor &#x00026; Francis</publisher-name>).</citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caron</surname> <given-names>S. J. C.</given-names></name> <name><surname>Ruta</surname> <given-names>V.</given-names></name> <name><surname>Abbott</surname> <given-names>L. F.</given-names></name> <name><surname>Axel</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>Random convergence of olfactory inputs in the Drosophila mushroom body</article-title>. <source>Nature</source> <volume>497</volume>, <fpage>113</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1038/nature12063</pub-id><pub-id pub-id-type="pmid">23615618</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.-C.</given-names></name> <name><surname>Schwing</surname> <given-names>A. G.</given-names></name> <name><surname>Yuille</surname> <given-names>A. L.</given-names></name> <name><surname>Urtasun</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <source>Learning deep structured models</source>. arXiv:1407.2538.</citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chikkerur</surname> <given-names>S.</given-names></name> <name><surname>Serre</surname> <given-names>T.</given-names></name> <name><surname>Tan</surname> <given-names>C.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2010</year>). <article-title>What and where: a Bayesian inference theory of attention</article-title>. <source>Vis. Res.</source> <volume>50</volume>, <fpage>2233</fpage>&#x02013;<lpage>2247</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2010.05.013</pub-id><pub-id pub-id-type="pmid">20493206</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Choo</surname> <given-names>F. X.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>A spiking neuron model of serial-order recall</article-title>, in <source>32nd Annual Conference of the Cognitive Science Society</source> (<publisher-loc>Portland</publisher-loc>).</citation>
</ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chubykin</surname> <given-names>A. A.</given-names></name> <name><surname>Roach</surname> <given-names>E. B.</given-names></name> <name><surname>Bear</surname> <given-names>M. F.</given-names></name> <name><surname>Shuler</surname> <given-names>M. G. H.</given-names></name></person-group> (<year>2013</year>). <article-title>A cholinergic mechanism for reward timing within primary visual cortex</article-title>. <source>Neuron</source> <volume>77</volume>, <fpage>723</fpage>&#x02013;<lpage>735</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.12.039</pub-id><pub-id pub-id-type="pmid">23439124</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>J.</given-names></name> <name><surname>Gulcehre</surname> <given-names>C.</given-names></name> <name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <source>Empirical evaluation of gated recurrent neural networks on sequence modeling</source>. arXiv:1412.3555.</citation>
</ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cichon</surname> <given-names>J.</given-names></name> <name><surname>Gan</surname> <given-names>W.-B.</given-names></name></person-group> (<year>2015</year>). <article-title>Branch-specific dendritic ca2<sup>&#x0002B;</sup> spikes cause persistent synaptic plasticity</article-title>. <source>Nature</source> <volume>520</volume>, <fpage>180</fpage>&#x02013;<lpage>185</lpage>. <pub-id pub-id-type="doi">10.1038/nature14251</pub-id><pub-id pub-id-type="pmid">25822789</pub-id></citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clayton</surname> <given-names>N. S.</given-names></name> <name><surname>Dickinson</surname> <given-names>A.</given-names></name></person-group> (<year>1998</year>). <article-title>Episodic-like memory during cache recovery by scrub jays</article-title>. <source>Nature</source> <volume>395</volume>, <fpage>272</fpage>&#x02013;<lpage>274</lpage>. <pub-id pub-id-type="pmid">9751053</pub-id></citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clopath</surname> <given-names>C.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2010</year>). <article-title>Voltage and spike timing interact in STDP&#x02013;a unified model</article-title>. <source>Front. Synaptic Neurosci.</source> <volume>2</volume>:<issue>25</issue>. <pub-id pub-id-type="doi">10.3389/fnsyn.2010.00025</pub-id><pub-id pub-id-type="pmid">21423511</pub-id></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohn</surname> <given-names>R.</given-names></name> <name><surname>Morantte</surname> <given-names>I.</given-names></name> <name><surname>Ruta</surname> <given-names>V.</given-names></name></person-group> (<year>2015</year>). <article-title>Coordinated and compartmentalized neuromodulation shapes sensory processing in drosophila</article-title>. <source>Cell</source> <volume>163</volume>, <fpage>1742</fpage>&#x02013;<lpage>1755</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.11.019</pub-id><pub-id pub-id-type="pmid">26687359</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colino</surname> <given-names>A.</given-names></name> <name><surname>Halliwell</surname> <given-names>J.</given-names></name></person-group> (<year>1987</year>). <article-title>Differential modulation of three separate K-conductances in hippocampal CA1 neurons by serotonin</article-title>. <source>Nature</source> <volume>328</volume>, <fpage>73</fpage>&#x02013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1038/328073a0</pub-id><pub-id pub-id-type="pmid">3600775</pub-id></citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Constantinescu</surname> <given-names>A. O.</given-names></name> <name><surname>O&#x00027;Reilly</surname> <given-names>J. X.</given-names></name> <name><surname>Behrens</surname> <given-names>T. E.</given-names></name></person-group> (<year>2016</year>). <article-title>Organizing conceptual knowledge in humans with a gridlike code</article-title>. <source>Science</source> <volume>352</volume>, <fpage>1464</fpage>&#x02013;<lpage>1468</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaf0941</pub-id><pub-id pub-id-type="pmid">27313047</pub-id></citation>
</ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. D.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural networks and neuroscience-inspired computer vision</article-title>. <source>Curr. Biol.</source> <volume>24</volume>, <fpage>R921</fpage>&#x02013;<lpage>R929</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2014.08.026</pub-id><pub-id pub-id-type="pmid">25247371</pub-id></citation>
</ref>
<ref id="B86">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crick</surname> <given-names>F.</given-names></name></person-group> (<year>1989</year>). <article-title>The recent excitement about neural networks</article-title>. <source>Nature</source> <volume>337</volume>, <fpage>129</fpage>&#x02013;<lpage>132</lpage>. <pub-id pub-id-type="pmid">2911347</pub-id></citation>
</ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crick</surname> <given-names>F. C.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title>What is the function of the claustrum?</article-title> <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>360</volume>, <fpage>1271</fpage>&#x02013;<lpage>1279</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2005.1661</pub-id><pub-id pub-id-type="pmid">16147522</pub-id></citation>
</ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crouzet</surname> <given-names>S. M.</given-names></name> <name><surname>Thorpe</surname> <given-names>S. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Low-level cues and ultra-fast face detection</article-title>. <source>Front. Psychol.</source> <volume>2</volume>:<issue>342</issue>. <pub-id pub-id-type="doi">10.3389/fpsyg.2011.00342</pub-id><pub-id pub-id-type="pmid">22125544</pub-id></citation>
</ref>
<ref id="B89">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Cui</surname> <given-names>Y.</given-names></name> <name><surname>Surpur</surname> <given-names>C.</given-names></name> <name><surname>Ahmad</surname> <given-names>S.</given-names></name> <name><surname>Hawkins</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <source>Continuous online sequence learning with an unsupervised neural network model</source>. arXiv:1512.05463.</citation>
</ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dadarlat</surname> <given-names>M. C.</given-names></name> <name><surname>O&#x00027;Doherty</surname> <given-names>J. E.</given-names></name> <name><surname>Sabes</surname> <given-names>P. N.</given-names></name></person-group> (<year>2015</year>). <article-title>A learning-based approach to artificial sensory feedback leads to optimal integration</article-title>. <source>Nat. Neurosci.</source> <volume>18</volume>, <fpage>138</fpage>&#x02013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3883</pub-id><pub-id pub-id-type="pmid">25420067</pub-id></citation>
</ref>
<ref id="B91">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Danihelka</surname> <given-names>I.</given-names></name> <name><surname>Wayne</surname> <given-names>G.</given-names></name> <name><surname>Uria</surname> <given-names>B.</given-names></name> <name><surname>Kalchbrenner</surname> <given-names>N.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <source>Associative long short-term memory</source>. arXiv:1602.03032.</citation>
</ref>
<ref id="B92">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>Niv</surname> <given-names>Y.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>Actions, policies, values and the basal ganglia</article-title>, in <source>Recent Breakthroughs in Basal Ganglia Research</source>, ed <person-group person-group-type="editor"><name><surname>Bezard</surname> <given-names>E.</given-names></name></person-group> (<publisher-loc>Hauppauge, NY</publisher-loc>: <publisher-name>Nova Science</publisher-name>), <fpage>91</fpage>&#x02013;<lpage>106</lpage>.</citation>
</ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <article-title>Twenty-five lessons from computational neuromodulation</article-title>. <source>Neuron</source> <volume>76</volume>, <fpage>240</fpage>&#x02013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.09.027</pub-id><pub-id pub-id-type="pmid">23040818</pub-id></citation>
</ref>
<ref id="B94">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <article-title>A computational model of the cerebral cortex</article-title>, in <source>Proceedings of the 20th National Conference on Artificial Intelligence</source> (<publisher-loc>Pittsburg, PA</publisher-loc>).</citation>
</ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Meyniel</surname> <given-names>F.</given-names></name> <name><surname>Wacongne</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Pallier</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>The neural representation of sequences: from transition probabilities to algebraic patterns and linguistic trees</article-title>. <source>Neuron</source> <volume>88</volume>, <fpage>2</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.09.019</pub-id><pub-id pub-id-type="pmid">26447569</pub-id></citation>
</ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dekker</surname> <given-names>T. M.</given-names></name> <name><surname>Nardini</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Risky visuomotor choices during rapid reaching in childhood</article-title>. <source>Dev. Sci</source>. <volume>19</volume>, <fpage>427</fpage>&#x02013;<lpage>439</lpage>. <pub-id pub-id-type="doi">10.1111/desc.12322</pub-id><pub-id pub-id-type="pmid">26190343</pub-id></citation>
</ref>
<ref id="B97">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Delalleau</surname> <given-names>O.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). <article-title>Shallow vs. deep sum-product networks</article-title>,&#x0201D; in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Grenada</publisher-loc>), <fpage>666</fpage>&#x02013;<lpage>674</lpage>.</citation>
</ref>
<ref id="B98">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>DePasquale</surname> <given-names>B.</given-names></name> <name><surname>Churchland</surname> <given-names>M.</given-names></name> <name><surname>Abbott</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <source>Using firing-rate dynamics to train recurrent networks of spiking model neurons</source>. arXiv:1601.07620.</citation>
</ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeWolf</surname> <given-names>T.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>The neural optimal control hierarchy for motor control</article-title>. <source>J. Neural Eng.</source> <volume>8</volume>:<fpage>065009</fpage>. <pub-id pub-id-type="doi">10.1088/1741-2560/8/6/065009</pub-id><pub-id pub-id-type="pmid">22056418</pub-id></citation>
</ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name> <name><surname>Zoccolan</surname> <given-names>D.</given-names></name> <name><surname>Rust</surname> <given-names>N. C.</given-names></name></person-group> (<year>2012</year>). <article-title>How does the brain solve visual object recognition?</article-title> <source>Neuron</source> <volume>73</volume>, <fpage>415</fpage>&#x02013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.01.010</pub-id><pub-id pub-id-type="pmid">22325196</pub-id></citation>
</ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Douglas</surname> <given-names>R. J.</given-names></name> <name><surname>Martin</surname> <given-names>K. A. C.</given-names></name></person-group> (<year>2004</year>). <article-title>Neuronal circuits of the neocortex</article-title>. <source>Annu. Rev. Neurosci.</source> <volume>27</volume>, <fpage>419</fpage>&#x02013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.neuro.27.070203.144152</pub-id><pub-id pub-id-type="pmid">15217339</pub-id></citation>
</ref>
<ref id="B102">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doya</surname> <given-names>K.</given-names></name></person-group> (<year>1999</year>). <article-title>What are the computations of the cerebellum, the basal ganglia and the cerebral cortex?</article-title> <source>Neural Netw.</source> <volume>12</volume>, <fpage>961</fpage>&#x02013;<lpage>974</lpage>. <pub-id pub-id-type="pmid">12662639</pub-id></citation>
</ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dudman</surname> <given-names>J. T.</given-names></name> <name><surname>Tsay</surname> <given-names>D.</given-names></name> <name><surname>Siegelbaum</surname> <given-names>S. A.</given-names></name></person-group> (<year>2007</year>). <article-title>A role for synaptic inputs at distal dendrites: instructive signals for hippocampal long-term plasticity</article-title>. <source>Neuron</source> <volume>56</volume>, <fpage>866</fpage>&#x02013;<lpage>879</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2007.10.020</pub-id><pub-id pub-id-type="pmid">18054862</pub-id></citation>
</ref>
<ref id="B104">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dumoulin</surname> <given-names>S. O.</given-names></name> <name><surname>Wandell</surname> <given-names>B. A.</given-names></name></person-group> (<year>2008</year>). <article-title>Population receptive field estimates in human visual cortex</article-title>. <source>Neuroimage</source> <volume>39</volume>, <fpage>647</fpage>&#x02013;<lpage>660</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.09.034</pub-id><pub-id pub-id-type="pmid">17977024</pub-id></citation>
</ref>
<ref id="B105">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <source>How to Build a Brain: A Neural Architecture for Biological Cognition</source>. <publisher-loc>Oxford, UK</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B106">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eliasmith</surname> <given-names>C.</given-names></name> <name><surname>Anderson</surname> <given-names>C. H.</given-names></name></person-group> (<year>2004</year>). <source>Neural Engineering: Computation, Representation, and Dynamics in Neurobiological Systems</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B107">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eliasmith</surname> <given-names>C.</given-names></name> <name><surname>Martens</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Normalization for probabilistic inference with neurons</article-title>. <source>Biol. Cybern.</source> <volume>104</volume>, <fpage>251</fpage>&#x02013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-011-0433-y</pub-id><pub-id pub-id-type="pmid">21573688</pub-id></citation>
</ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eliasmith</surname> <given-names>C.</given-names></name> <name><surname>Stewart</surname> <given-names>T. C.</given-names></name> <name><surname>Choo</surname> <given-names>X.</given-names></name> <name><surname>Bekolay</surname> <given-names>T.</given-names></name> <name><surname>DeWolf</surname> <given-names>T.</given-names></name> <name><surname>Tang</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>A large-scale model of the functioning brain</article-title>. <source>Science</source> <volume>338</volume>, <fpage>1202</fpage>&#x02013;<lpage>1205</lpage>. <pub-id pub-id-type="pmid">23197532</pub-id></citation>
</ref>
<ref id="B109">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Emlen</surname> <given-names>S. T.</given-names></name></person-group> (<year>1967</year>). <article-title>Migratory orientation in the indigo bunting, passerina cyanea: part i: evidence for use of celestial cues</article-title>. <source>Auk</source> <volume>84</volume>, <fpage>309</fpage>&#x02013;<lpage>342</lpage>.</citation>
</ref>
<ref id="B110">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Enel</surname> <given-names>P.</given-names></name> <name><surname>Procyk</surname> <given-names>E.</given-names></name> <name><surname>Quilodran</surname> <given-names>R.</given-names></name> <name><surname>Dominey</surname> <given-names>P. F.</given-names></name></person-group> (<year>2016</year>). <article-title>Reservoir computing properties of neural dynamics in prefrontal cortex</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1004967</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004967</pub-id><pub-id pub-id-type="pmid">27286251</pub-id></citation>
</ref>
<ref id="B111">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Erhan</surname> <given-names>D.</given-names></name> <name><surname>Manzagol</surname> <given-names>P.</given-names></name></person-group> (<year>2009</year>). <article-title>The difficulty of training deep architectures and the effect of unsupervised pre-training</article-title>. <source>Twelfth International Conference on Artificial Intelligence and Statistics (AISTATS), JMLR Workshop and Conference Procedings</source> (<publisher-loc>Clearwater Beach, FL</publisher-loc>), <fpage>153</fpage>&#x02013;<lpage>160</lpage>.</citation>
</ref>
<ref id="B112">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Eslami</surname> <given-names>S.</given-names></name> <name><surname>Heess</surname> <given-names>N.</given-names></name> <name><surname>Weber</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>Attend, infer, repeat: fast scene understanding with generative models</source>. arXiv:1603.08575.</citation>
</ref>
<ref id="B113">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fausey</surname> <given-names>C. M.</given-names></name> <name><surname>Jayaraman</surname> <given-names>S.</given-names></name> <name><surname>Smith</surname> <given-names>L. B.</given-names></name></person-group> (<year>2016</year>). <article-title>From faces to hands: changing visual input in the first two years</article-title>. <source>Cognition</source> <volume>152</volume>, <fpage>101</fpage>&#x02013;<lpage>107</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2016.03.005</pub-id><pub-id pub-id-type="pmid">27043744</pub-id></citation>
</ref>
<ref id="B114">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Felleman</surname> <given-names>D. J.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name></person-group> (<year>1991</year>). <article-title>Distributed hierarchical processing in the primate cerebral cortex</article-title>. <source>Cereb. cortex</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="pmid">1822724</pub-id></citation>
</ref>
<ref id="B115">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferster</surname> <given-names>D.</given-names></name> <name><surname>Miller</surname> <given-names>K. D.</given-names></name></person-group> (<year>2000</year>). <article-title>Neural mechanisms of orientation selectivity in the visual cortex</article-title>. <source>Annu. Rev. Neurosci.</source> <volume>23</volume>, <fpage>441</fpage>&#x02013;<lpage>471</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.neuro.23.1.441</pub-id><pub-id pub-id-type="pmid">10845071</pub-id></citation>
</ref>
<ref id="B116">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fetz</surname> <given-names>E. E.</given-names></name></person-group> (<year>1969</year>). <article-title>Operant conditioning of cortical unit activity</article-title>. <source>Science</source> <volume>163</volume>, <fpage>955</fpage>&#x02013;<lpage>958</lpage>. <pub-id pub-id-type="pmid">4974291</pub-id></citation>
</ref>
<ref id="B117">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fetz</surname> <given-names>E. E.</given-names></name></person-group> (<year>2007</year>). <article-title>Volitional control of neural activity: implications for brain&#x02013;computer interfaces</article-title>. <source>J. Physiol.</source> <volume>579</volume>, <fpage>571</fpage>&#x02013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1113/jphysiol.2006.127142</pub-id><pub-id pub-id-type="pmid">17234689</pub-id></citation>
</ref>
<ref id="B118">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiete</surname> <given-names>I. R.</given-names></name> <name><surname>Fee</surname> <given-names>M. S.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2007</year>). <article-title>Model of birdsong learning based on gradient estimation by dynamic perturbation of neural conductances</article-title>. <source>J. Neurophysiol.</source> <volume>98</volume>, <fpage>2038</fpage>&#x02013;<lpage>2057</lpage>. <pub-id pub-id-type="doi">10.1152/jn.01311.2006</pub-id><pub-id pub-id-type="pmid">17652414</pub-id></citation>
</ref>
<ref id="B119">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiete</surname> <given-names>I. R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>C. Z. H.</given-names></name> <name><surname>Hahnloser</surname> <given-names>R. H. R.</given-names></name></person-group> (<year>2010</year>). <article-title>Spike-time-dependent plasticity and heterosynaptic competition organize networks to produce long scale-free sequences of neural activity</article-title>. <source>Neuron</source> <volume>65</volume>, <fpage>563</fpage>&#x02013;<lpage>576</lpage>. <pub-id pub-id-type="pmid">20188660</pub-id></citation>
</ref>
<ref id="B120">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiete</surname> <given-names>I. R.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2006</year>). <article-title>Gradient learning in spiking neural networks by dynamic perturbation of conductances</article-title>. <source>Phys. Rev. Lett.</source> <volume>97</volume>:<fpage>048104</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.97.048104</pub-id><pub-id pub-id-type="pmid">16907616</pub-id></citation>
</ref>
<ref id="B121">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Finnerty</surname> <given-names>G.</given-names></name> <name><surname>Shadlen</surname> <given-names>M.</given-names></name> <name><surname>Jazayeri</surname> <given-names>M</given-names></name> <name><surname>Nobre</surname> <given-names>A. C</given-names></name> <name><surname>Buonomano</surname> <given-names>D. V.</given-names></name></person-group> (<year>2015</year>). <article-title>Time in Cortical Circuits</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>13912</fpage>&#x02013;<lpage>13916</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2654-15.2015</pub-id><pub-id pub-id-type="pmid">26468192</pub-id></citation>
</ref>
<ref id="B122">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Finn</surname> <given-names>C.</given-names></name> <name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <source>Guided cost learning: deep inverse optimal control via policy optimization</source>. arXiv:1603.00448.</citation>
</ref>
<ref id="B123">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fodor</surname> <given-names>J. D.</given-names></name> <name><surname>Crowther</surname> <given-names>C.</given-names></name></person-group> (<year>2002</year>). <article-title>Understanding stimulus poverty arguments</article-title>. <source>Ling. Rev.</source> <volume>18</volume>, <fpage>105</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1515/tlir.19.1-2.105</pub-id></citation>
</ref>
<ref id="B124">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>F&#x000F6;ldi&#x000E1;k</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>Learning invariance from transformation sequences</article-title>. <source>J. Neural Comput.</source> <volume>3</volume>, <fpage>194</fpage>&#x02013;<lpage>200</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1991.3.2.194</pub-id></citation>
</ref>
<ref id="B125">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Foster</surname> <given-names>D. J.</given-names></name> <name><surname>Morris</surname> <given-names>R. G. M.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>2000</year>). <article-title>Models of hippocampally dependent navigation, using the temporal difference learning rule</article-title>. <source>Hippocampus</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1002/(SICI)1098-1063(2000)10:1&#x0003C;1::AID-HIPO1&#x0003C;3.0.CO;2-1</pub-id><pub-id pub-id-type="pmid">10706212</pub-id></citation>
</ref>
<ref id="B126">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fournier</surname> <given-names>J.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>C. M.</given-names></name> <name><surname>Laurent</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Looking for the roots of cortical sensory computation in three-layered cortices</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>31</volume>, <fpage>119</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2014.09.006</pub-id><pub-id pub-id-type="pmid">25291080</pub-id></citation>
</ref>
<ref id="B127">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Franconeri</surname> <given-names>S. L.</given-names></name> <name><surname>Pylyshyn</surname> <given-names>Z. W.</given-names></name> <name><surname>Scholl</surname> <given-names>B. J.</given-names></name></person-group> (<year>2012</year>). <article-title>A simple proximity heuristic allows tracking of multiple objects through occlusion</article-title>. <source>Atten. Percept. Psychophys.</source> <volume>74</volume>, <fpage>691</fpage>&#x02013;<lpage>702</lpage>. <pub-id pub-id-type="doi">10.3758/s13414-011-0265-9</pub-id><pub-id pub-id-type="pmid">22271165</pub-id></citation>
</ref>
<ref id="B128">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frankland</surname> <given-names>S. M.</given-names></name> <name><surname>Greene</surname> <given-names>J. D.</given-names></name></person-group> (<year>2015</year>). <article-title>An architecture for encoding sentence meaning in left mid-superior temporal cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>11732</fpage>&#x02013;<lpage>11737</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1421236112</pub-id><pub-id pub-id-type="pmid">26305927</pub-id></citation>
</ref>
<ref id="B129">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>M. J.</given-names></name> <name><surname>Badre</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Mechanisms of hierarchical reinforcement learning in corticostriatal circuits 1: computational analysis</article-title>. <source>Cereb. Cortex</source> <volume>22</volume>, <fpage>509</fpage>&#x02013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhr114</pub-id><pub-id pub-id-type="pmid">21693490</pub-id></citation>
</ref>
<ref id="B130">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Franzius</surname> <given-names>M.</given-names></name> <name><surname>Sprekeler</surname> <given-names>H.</given-names></name> <name><surname>Wiskott</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). <article-title>Slowness and sparseness lead to place, head-direction, and spatial-view cells</article-title>. <source>PLoS Comput. Biol.</source> <volume>3</volume>:<fpage>e166</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.0030166</pub-id><pub-id pub-id-type="pmid">17784780</pub-id></citation>
</ref>
<ref id="B131">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>11</volume>, <fpage>127</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id><pub-id pub-id-type="pmid">20068583</pub-id></citation>
</ref>
<ref id="B132">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name></person-group> (<year>2007</year>). <article-title>Free-energy and the brain</article-title>. <source>Synthese</source> <volume>159</volume>, <fpage>417</fpage>&#x02013;<lpage>458</lpage>. <pub-id pub-id-type="doi">10.1007/s11229-007-9237-y</pub-id><pub-id pub-id-type="pmid">19325932</pub-id></citation>
</ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fukushima</surname> <given-names>K.</given-names></name></person-group> (<year>1980</year>). <article-title>Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position</article-title>. <source>Biol. Cybern.</source> <volume>36</volume>, <fpage>193</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1007/BF00344251</pub-id><pub-id pub-id-type="pmid">7370364</pub-id></citation>
</ref>
<ref id="B134">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galtier</surname> <given-names>M. N.</given-names></name> <name><surname>Wainrib</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>A biological gradient descent for prediction through a combination of stdp and homeostatic plasticity</article-title>. <source>Neural Comput.</source> <volume>25</volume>, <fpage>2815</fpage>&#x02013;<lpage>2832</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00512</pub-id><pub-id pub-id-type="pmid">24001342</pub-id></citation>
</ref>
<ref id="B135">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>T.</given-names></name> <name><surname>Harari</surname> <given-names>D.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J.</given-names></name> <name><surname>Ullman</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <source>When computer vision gazes at cognition</source>. arXiv:1412.2672.</citation>
</ref>
<ref id="B136">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gemp</surname> <given-names>I.</given-names></name> <name><surname>Mahadevan</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Modeling context in cognition using variational inequalities</article-title>, in <source>Modeling Changing Perspectives&#x02014;Reconceptualizing Sensorimotor Experiences: Papers from the 2014 AAAI Fall Symposium</source> (<publisher-loc>Arlington, TX</publisher-loc>).</citation>
</ref>
<ref id="B137">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>George</surname> <given-names>D.</given-names></name> <name><surname>Hawkins</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>Towards a mathematical theory of cortical micro-circuits</article-title>. <source>PLoS Comput. Biol.</source> <volume>5</volume>:<fpage>e1000532</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000532</pub-id><pub-id pub-id-type="pmid">19816557</pub-id></citation>
</ref>
<ref id="B138">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <name><surname>Beck</surname> <given-names>J. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Complex probabilistic inference: from cognition to neural computation</article-title>, in <source>Computational Models of Brain and Behavior</source>, ed <person-group person-group-type="editor"><name><surname>Moustafa</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>Wiley-Blackwell</publisher-name>).</citation>
</ref>
<ref id="B139">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <name><surname>Moore</surname> <given-names>C. D.</given-names></name> <name><surname>Todd</surname> <given-names>M. T.</given-names></name> <name><surname>Norman</surname> <given-names>K. A.</given-names></name> <name><surname>Sederberg</surname> <given-names>P. B.</given-names></name></person-group> (<year>2012</year>). <article-title>The successor representation and temporal context</article-title>. <source>Neural Comput.</source> <volume>24</volume>, <fpage>1553</fpage>&#x02013;<lpage>1568</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00282</pub-id><pub-id pub-id-type="pmid">22364500</pub-id></citation>
</ref>
<ref id="B140">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <name><surname>Moustafa</surname> <given-names>A. A.</given-names></name> <name><surname>Ludvig</surname> <given-names>E. A.</given-names></name></person-group> (<year>2014</year>). <article-title>Time representation in reinforcement learning models of the basal ganglia</article-title>. <source>Front. Comput. Neurosci.</source> <volume>7</volume>:<issue>194</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2013.00194</pub-id><pub-id pub-id-type="pmid">24409138</pub-id></citation>
</ref>
<ref id="B141">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ghahramani</surname> <given-names>I. M. Z.</given-names></name></person-group> (<year>2005</year>). <source>A Note On the Evidence and Bayesian Occam&#x00027;s Razor</source>. Gatsby Unit Technical Report GCNU-TR 2005&#x02013;003.</citation>
</ref>
<ref id="B142">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giocomo</surname> <given-names>L. M.</given-names></name> <name><surname>Zilli</surname> <given-names>E. A.</given-names></name> <name><surname>Frans&#x000E9;n</surname> <given-names>E.</given-names></name> <name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name></person-group> (<year>2007</year>). <article-title>Temporal frequency of subthreshold oscillations scales with entorhinal grid cell field spacing</article-title>. <source>Science</source> <volume>315</volume>, <fpage>1719</fpage>&#x02013;<lpage>1722</lpage>. <pub-id pub-id-type="doi">10.1126/science.1139207</pub-id><pub-id pub-id-type="pmid">17379810</pub-id></citation>
</ref>
<ref id="B143">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giret</surname> <given-names>N.</given-names></name> <name><surname>Kornfeld</surname> <given-names>J.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Hahnloser</surname> <given-names>R. H. R.</given-names></name></person-group> (<year>2014</year>). <article-title>Evidence for a causal inverse model in an avian cortico-basal ganglia circuit</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>6063</fpage>&#x02013;<lpage>6068</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1317087111</pub-id><pub-id pub-id-type="pmid">24711417</pub-id></citation>
</ref>
<ref id="B144">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Goertzel</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>How might the brain represent complex symbolic knowledge?</article-title> in <source>International Joint Conference on Neural Networks (IJCNN)</source> (<publisher-loc>Beijing</publisher-loc>).</citation>
</ref>
<ref id="B145">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldman</surname> <given-names>M. S.</given-names></name> <name><surname>Levine</surname> <given-names>J. H.</given-names></name> <name><surname>Major</surname> <given-names>G.</given-names></name> <name><surname>Tank</surname> <given-names>D. W.</given-names></name> <name><surname>Seung</surname> <given-names>H.</given-names></name></person-group> (<year>2003</year>). <article-title>Robust persistent neural activity in a model integrator with multiple hysteretic dendrites per neuron</article-title>. <source>Cereb. Cortex</source> <volume>13</volume>, <fpage>1185</fpage>&#x02013;<lpage>1195</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhg095</pub-id><pub-id pub-id-type="pmid">14576210</pub-id></citation>
</ref>
<ref id="B146">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gonzalez Andino</surname> <given-names>S. L. Grave de Peralta Menendez, R.</given-names></name></person-group> (<year>2012</year>). <article-title>Coding of saliency by ensemble bursting in the amygdala of primates</article-title>. <source>Front. Behav. Neurosci.</source> <volume>6</volume>:<issue>38</issue>. <pub-id pub-id-type="doi">10.3389/fnbeh.2012.00038</pub-id><pub-id pub-id-type="pmid">22848193</pub-id></citation>
</ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gooch</surname> <given-names>C. M.</given-names></name> <name><surname>Wiener</surname> <given-names>M.</given-names></name> <name><surname>Wencil</surname> <given-names>E. B.</given-names></name> <name><surname>Coslett</surname> <given-names>H. B.</given-names></name></person-group> (<year>2010</year>). <article-title>Interval timing disruptions in subjects with cerebellar lesions</article-title>. <source>Neuropsychologia</source> <volume>48</volume>, <fpage>1022</fpage>&#x02013;<lpage>1031</lpage>. <pub-id pub-id-type="pmid">19962999</pub-id></citation>
</ref>
<ref id="B148">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I. J.</given-names></name> <name><surname>Pouget-Abadie</surname> <given-names>J.</given-names></name> <name><surname>Mirza</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Warde-Farley</surname> <given-names>D.</given-names></name> <name><surname>Ozair</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014a</year>). <source>Generative adversarial networks</source>. arXiv:1406.2661.</citation>
</ref>
<ref id="B149">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I. J.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Saxe</surname> <given-names>A. M.</given-names></name></person-group> (<year>2014b</year>). <source>Qualitatively characterizing neural network optimization problems</source>. arXiv:1412.6544.</citation>
</ref>
<ref id="B150">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gopnik</surname> <given-names>A.</given-names></name> <name><surname>Meltzoff</surname> <given-names>A. N.</given-names></name> <name><surname>Kuhl</surname> <given-names>P. K.</given-names></name></person-group> (<year>2000</year>). <source>The Scientist in the Crib: What Early Learning Tells us About the Mind</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Harper Paperbacks</publisher-name>.</citation>
</ref>
<ref id="B151">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Wayne</surname> <given-names>G.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name></person-group> (<year>2014</year>). <source>Neural Turing Machines</source>. arXiv:1410.5401.</citation>
</ref>
<ref id="B152">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Graybiel</surname> <given-names>A. M.</given-names></name></person-group> (<year>1998</year>). <article-title>The basal ganglia and chunking of action repertoires</article-title>. <source>Neurobiol. Learn. Mem.</source> <volume>70</volume>, <fpage>119</fpage>&#x02013;<lpage>136</lpage>. <pub-id pub-id-type="pmid">9753592</pub-id></citation>
</ref>
<ref id="B153">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Gregor</surname> <given-names>K.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Rezende</surname> <given-names>D. J.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <source>DRAW: a recurrent neural network for image generation</source>. arXiv:1502.04623.</citation>
</ref>
<ref id="B154">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grillner</surname> <given-names>S.</given-names></name> <name><surname>Hellgren</surname> <given-names>J.</given-names></name> <name><surname>M&#x000E9;nard</surname> <given-names>A.</given-names></name> <name><surname>Saitoh</surname> <given-names>K.</given-names></name> <name><surname>Wikstr&#x000F6;m</surname> <given-names>M. A.</given-names></name></person-group> (<year>2005</year>). <article-title>Mechanisms for selection of basic motor programs&#x02013;roles for the striatum and pallidum</article-title>. <source>Trends Neurosci.</source> <volume>28</volume>, <fpage>364</fpage>&#x02013;<lpage>370</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2005.05.004</pub-id><pub-id pub-id-type="pmid">15935487</pub-id></citation>
</ref>
<ref id="B155">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grossberg</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Adaptive resonance theory: how a brain learns to consciously attend, learn, and recognize a changing world</article-title>. <source>Neural Netw.</source> <volume>37</volume>, <fpage>1</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2012.09.017</pub-id><pub-id pub-id-type="pmid">23149242</pub-id></citation>
</ref>
<ref id="B156">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>10005</fpage>&#x02013;<lpage>10014</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5023-14.2015</pub-id><pub-id pub-id-type="pmid">26157000</pub-id></citation>
</ref>
<ref id="B157">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Guez</surname> <given-names>A.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <source>Efficient bayes-adaptive reinforcement learning using sample-based search</source>. arXiv:1205.3109.</citation>
</ref>
<ref id="B158">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>G&#x000FC;l&#x000E7;ehre</surname> <given-names>&#x000C7;.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Knowledge matters: importance of prior information for optimization</article-title>. <source>J. Mach. Learn. Res.</source> <volume>17</volume>, <fpage>1</fpage>&#x02013;<lpage>32</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://jmlr.org/papers/v17/gulchere16a.html">http://jmlr.org/papers/v17/gulchere16a.html</ext-link></citation>
</ref>
<ref id="B159">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;nt&#x000FC;rk&#x000FC;n</surname> <given-names>O.</given-names></name> <name><surname>Bugnyar</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Cognition without cortex</article-title>. <source>Trends Cogn. Sci.</source> <volume>20</volume>, <fpage>291</fpage>&#x02013;<lpage>303</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2016.02.001</pub-id><pub-id pub-id-type="pmid">26944218</pub-id></citation>
</ref>
<ref id="B160">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gurney</surname> <given-names>K.</given-names></name> <name><surname>Prescott</surname> <given-names>T. J.</given-names></name> <name><surname>Redgrave</surname> <given-names>P.</given-names></name></person-group> (<year>2001</year>). <article-title>A computational model of action selection in the basal ganglia. I. A new functional anatomy</article-title>. <source>Biol. Cybern.</source> <volume>84</volume>, <fpage>401</fpage>&#x02013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1007/PL00007984</pub-id><pub-id pub-id-type="pmid">11417052</pub-id></citation>
</ref>
<ref id="B161">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hadley</surname> <given-names>R. F.</given-names></name></person-group> (<year>2009</year>). <article-title>The problem of rapid variable creation</article-title>. <source>Neural Comput.</source> <volume>21</volume>, <fpage>510</fpage>&#x02013;<lpage>532</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.07-07-572</pub-id><pub-id pub-id-type="pmid">19431268</pub-id></citation>
</ref>
<ref id="B162">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hamlin</surname> <given-names>J. K.</given-names></name> <name><surname>Wynn</surname> <given-names>K.</given-names></name> <name><surname>Bloom</surname> <given-names>P.</given-names></name></person-group> (<year>2007</year>). <article-title>Social evaluation by preverbal infants</article-title>. <source>Nature</source> <volume>450</volume>, <fpage>557</fpage>&#x02013;<lpage>559</lpage>. <pub-id pub-id-type="doi">10.1038/nature06288</pub-id><pub-id pub-id-type="pmid">18033298</pub-id></citation>
</ref>
<ref id="B163">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hangya</surname> <given-names>B.</given-names></name> <name><surname>Ranade</surname> <given-names>S.</given-names></name> <name><surname>Lorenc</surname> <given-names>M.</given-names></name> <name><surname>Kepecs</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Central cholinergic neurons are rapidly recruited by reinforcement feedback</article-title>. <source>Cell</source> <volume>162</volume>, <fpage>1155</fpage>&#x02013;<lpage>1168</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.07.057</pub-id><pub-id pub-id-type="pmid">26317475</pub-id></citation>
</ref>
<ref id="B164">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanuschkin</surname> <given-names>A.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Hahnloser</surname> <given-names>R. H. R.</given-names></name></person-group> (<year>2013</year>). <article-title>A hebbian learning rule gives rise to mirror neurons and links them to control theoretic inverse models</article-title>. <source>Front. Neural Circ.</source> <volume>7</volume>:<issue>106</issue>. <pub-id pub-id-type="doi">10.3389/fncir.2013.00106</pub-id><pub-id pub-id-type="pmid">23801941</pub-id></citation>
</ref>
<ref id="B165">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harris</surname> <given-names>C. M.</given-names></name> <name><surname>Wolpert</surname> <given-names>D. M.</given-names></name></person-group> (<year>1998</year>). <article-title>Signal-dependent noise determines motor planning</article-title>. <source>Nature</source> <volume>394</volume>, <fpage>780</fpage>&#x02013;<lpage>784</lpage>. <pub-id pub-id-type="pmid">9723616</pub-id></citation>
</ref>
<ref id="B166">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harris</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <article-title>Stability of the fittest: organizing learning through retroaxonal signals</article-title>. <source>Trends Neurosci.</source> <volume>31</volume>, <fpage>130</fpage>&#x02013;<lpage>136</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2007.12.002</pub-id><pub-id pub-id-type="pmid">18255165</pub-id></citation>
</ref>
<ref id="B167">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Maguire</surname> <given-names>E.</given-names></name></person-group> (<year>2009</year>). <article-title>The construction system of the brain</article-title>. <source>Philos. Trans. R. Soc. B.</source> <volume>364</volume>, <fpage>1263</fpage>&#x02013;<lpage>1271</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2008.0296</pub-id><pub-id pub-id-type="pmid">19528007</pub-id></citation>
</ref>
<ref id="B168">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Maguire</surname> <given-names>E. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Deconstructing episodic memory with construction</article-title>. <source>Trends Cogn. Sci.</source> <volume>11</volume>, <fpage>299</fpage>&#x02013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2007.05.001</pub-id><pub-id pub-id-type="pmid">17548229</pub-id></citation>
</ref>
<ref id="B169">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name></person-group> (<year>2006</year>). <article-title>The role of acetylcholine in learning and memory</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>16</volume>, <fpage>710</fpage>&#x02013;<lpage>715</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2006.09.002</pub-id><pub-id pub-id-type="pmid">17011181</pub-id></citation>
</ref>
<ref id="B170">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name></person-group> (<year>2015</year>). <article-title>If i had a million neurons: Potential tests of cortico-hippocampal theories</article-title>. <source>Progr. Brain Res.</source> <volume>219</volume>, <fpage>1</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/bs.pbr.2015.03.009</pub-id><pub-id pub-id-type="pmid">26072231</pub-id></citation>
</ref>
<ref id="B171">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name> <name><surname>Stern</surname> <given-names>C. E.</given-names></name></person-group> (<year>2015</year>). <article-title>Current questions on space and time encoding</article-title>. <source>Hippocampus</source> <volume>25</volume>, <fpage>744</fpage>&#x02013;<lpage>752</lpage>. <pub-id pub-id-type="doi">10.1002/hipo.22454</pub-id><pub-id pub-id-type="pmid">25786389</pub-id></citation>
</ref>
<ref id="B172">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name> <name><surname>Wyble</surname> <given-names>B. P.</given-names></name></person-group> (<year>1997</year>). <article-title>Free recall and recognition in a network model of the hippocampus: simulating effects of scopolamine on human memory function</article-title>. <source>Behav. Brain Res.</source> <volume>89</volume>, <fpage>1</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="pmid">9475612</pub-id></citation>
</ref>
<ref id="B173">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hattori</surname> <given-names>D.</given-names></name> <name><surname>Demir</surname> <given-names>E.</given-names></name> <name><surname>Kim</surname> <given-names>H. W.</given-names></name> <name><surname>Viragh</surname> <given-names>E.</given-names></name> <name><surname>Zipursky</surname> <given-names>S. L.</given-names></name> <name><surname>Dickson</surname> <given-names>B. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Dscam diversity is essential for neuronal wiring and self-recognition</article-title>. <source>Nature</source> <volume>449</volume>, <fpage>223</fpage>&#x02013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1038/nature06099</pub-id><pub-id pub-id-type="pmid">17851526</pub-id></citation>
</ref>
<ref id="B174">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hawkins</surname> <given-names>J.</given-names></name> <name><surname>Ahmad</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Why neurons have thousands of synapses, a theory of sequence memory in neocortex</article-title>. <source>Front. Neural Circ.</source> <volume>10</volume>:<issue>23</issue>. <pub-id pub-id-type="doi">10.3389/fncir.2016.00023</pub-id><pub-id pub-id-type="pmid">27065813</pub-id></citation>
</ref>
<ref id="B175">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hawkins</surname> <given-names>J.</given-names></name> <name><surname>Blakeslee</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <source>On Intelligence</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Henry Holt and Company</publisher-name>.</citation>
</ref>
<ref id="B176">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayashi-Takagi</surname> <given-names>A.</given-names></name> <name><surname>Yagishita</surname> <given-names>S.</given-names></name> <name><surname>Nakamura</surname> <given-names>M.</given-names></name> <name><surname>Shirai</surname> <given-names>F.</given-names></name> <name><surname>Wu</surname> <given-names>Y. I.</given-names></name> <name><surname>Loshbaugh</surname> <given-names>A. L.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Labelling and optical erasure of synaptic memory traces in the motor cortex</article-title>. <source>Nature</source> <volume>525</volume>, <fpage>333</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1038/nature15257</pub-id><pub-id pub-id-type="pmid">26352471</pub-id></citation>
</ref>
<ref id="B177">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haykin</surname> <given-names>S. S.</given-names></name></person-group> (<year>1994</year>). <source>Neural Networks: A Comprehensive Foundation</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Macmillan</publisher-name>.</citation>
</ref>
<ref id="B178">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayworth</surname> <given-names>K. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Dynamically partitionable autoassociative networks as a solution to the neural binding problem</article-title>. <source>Front. Comput. Neurosci.</source> <volume>6</volume>:<issue>73</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2012.00073</pub-id><pub-id pub-id-type="pmid">23060784</pub-id></citation>
</ref>
<ref id="B179">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayworth</surname> <given-names>K. J.</given-names></name> <name><surname>Lescroart</surname> <given-names>M. D.</given-names></name> <name><surname>Biederman</surname> <given-names>I.</given-names></name></person-group> (<year>2011</year>). <article-title>Neural encoding of relative position</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform.</source> <volume>37</volume>, <fpage>1032</fpage>&#x02013;<lpage>1050</lpage>. <pub-id pub-id-type="doi">10.1037/a0022338</pub-id><pub-id pub-id-type="pmid">21517211</pub-id></citation>
</ref>
<ref id="B180">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hennequin</surname> <given-names>G.</given-names></name> <name><surname>Vogels</surname> <given-names>T. P.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <article-title>Optimal control of transient dynamics in balanced networks supports generation of complex movements</article-title>. <source>Neuron</source> <volume>82</volume>, <fpage>1394</fpage>&#x02013;<lpage>1406</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2014.04.045</pub-id><pub-id pub-id-type="pmid">24945778</pub-id></citation>
</ref>
<ref id="B181">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herd</surname> <given-names>S. A.</given-names></name> <name><surname>Krueger</surname> <given-names>K. A.</given-names></name> <name><surname>Kriete</surname> <given-names>T. E.</given-names></name> <name><surname>Huang</surname> <given-names>T. R</given-names></name> <name><surname>Hazy</surname> <given-names>T. E</given-names></name> <name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>2013</year>). <article-title>Strategic cognitive sequencing: a computational cognitive neuroscience approach</article-title>. <source>Comput. Intell. Neurosci.</source> <volume>2013</volume>:<fpage>149329</fpage>. <pub-id pub-id-type="doi">10.1155/2013/149329</pub-id><pub-id pub-id-type="pmid">23935605</pub-id></citation>
</ref>
<ref id="B182">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Higgins</surname> <given-names>I.</given-names></name> <name><surname>Matthey</surname> <given-names>L.</given-names></name> <name><surname>Glorot</surname> <given-names>X.</given-names></name> <name><surname>Pal</surname> <given-names>A.</given-names></name> <name><surname>Uria</surname> <given-names>B.</given-names></name> <name><surname>Blundell</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>Early visual concept learning with unsupervised deep learning</source>. arXiv:1606.05579.</citation>
</ref>
<ref id="B183">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>1989</year>). <article-title>Connectionist learning procedures</article-title>. <source>Artif. Intell.</source> <volume>40</volume>, <fpage>185</fpage>&#x02013;<lpage>234</lpage>. <pub-id pub-id-type="doi">10.1016/0004-3702(89)90049-0</pub-id></citation>
</ref>
<ref id="B184">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2007</year>). <article-title>How to do backpropagation in a brain</article-title>,&#x0201D; in <source>Invited Talk at the NIPS&#x00027;2007 Deep Learning Workshop</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation>
</ref>
<ref id="B185">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Can the brain do back-propagation?</article-title>, in <source>Invited talk at Stanford University Colloquium on Computer Systems</source> (<publisher-loc>Stanford, CA</publisher-loc>).</citation>
</ref>
<ref id="B186">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Frey</surname> <given-names>B. J.</given-names></name> <name><surname>Neal</surname> <given-names>R. M.</given-names></name></person-group> (<year>1995</year>). <article-title>The &#x0201C;wake-sleep&#x0201D; algorithm for unsupervised neural networks</article-title>. <source>Science</source> <volume>268</volume>, <fpage>1158</fpage>&#x02013;<lpage>1161</lpage>. <pub-id pub-id-type="pmid">7761831</pub-id></citation>
</ref>
<ref id="B187">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Osindero</surname> <given-names>S.</given-names></name> <name><surname>Teh</surname> <given-names>Y.-W.</given-names></name></person-group> (<year>2006</year>). <article-title>A fast learning algorithm for deep belief nets</article-title>. <source>Neural Comput.</source> <volume>18</volume>, <fpage>1527</fpage>&#x02013;<lpage>1554</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2006.18.7.1527</pub-id><pub-id pub-id-type="pmid">16764513</pub-id></citation>
</ref>
<ref id="B188">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>McClelland</surname> <given-names>J.</given-names></name></person-group> (<year>1988</year>). <article-title>Learning representations by recirculation</article-title>. <source>Neural information processing</source>. <pub-id pub-id-type="pmid">11387044</pub-id></citation>
</ref>
<ref id="B189">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Transforming auto-encoders</article-title>, in <source>Artificial Neural Networks and Machine Leaning</source>, eds <person-group person-group-type="editor"><name><surname>Honkela</surname> <given-names>T.</given-names></name> <name><surname>Duch</surname> <given-names>W.</given-names></name> <name><surname>Girolami</surname> <given-names>M.</given-names></name> <name><surname>Kaski</surname> <given-names>S.</given-names></name></person-group> (<publisher-loc>Helsinki</publisher-loc>), <fpage>44</fpage>&#x02013;<lpage>51</lpage>.</citation>
</ref>
<ref id="B190">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hires</surname> <given-names>S. A.</given-names></name> <name><surname>Gutnisky</surname> <given-names>D. A.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name> <name><surname>O&#x00027;Connor</surname> <given-names>D. H.</given-names></name> <name><surname>Svoboda</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>Low-noise encoding of active touch by layer 4 in the somatosensory cortex</article-title>. <source>eLife</source> <fpage>4</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.06619</pub-id><pub-id pub-id-type="pmid">26245232</pub-id></citation>
</ref>
<ref id="B191">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Histed</surname> <given-names>M. H.</given-names></name> <name><surname>Ni</surname> <given-names>A. M.</given-names></name> <name><surname>Maunsell</surname> <given-names>J. H.</given-names></name></person-group> (<year>2013</year>). <article-title>Insights into cortical mechanisms of behavior from microstimulation experiments</article-title>. <source>Progr. Neurobiol.</source> <volume>103</volume>, <fpage>115</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1016/j.pneurobio.2012.01.006</pub-id><pub-id pub-id-type="pmid">22307059</pub-id></citation>
</ref>
<ref id="B192">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>. <pub-id pub-id-type="pmid">9832927</pub-id></citation>
</ref>
<ref id="B193">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoerzer</surname> <given-names>G. M.</given-names></name> <name><surname>Legenstein</surname> <given-names>R.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <article-title>Emergence of complex computational structures from chaotic neural networks through reward-modulated Hebbian learning</article-title>. <source>Cereb. Cortex</source> <volume>24</volume>, <fpage>677</fpage>&#x02013;<lpage>690</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhs348</pub-id><pub-id pub-id-type="pmid">23146969</pub-id></citation>
</ref>
<ref id="B194">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ho</surname> <given-names>J.</given-names></name> <name><surname>Ermon</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <source>Generative adversarial imitation learning</source>. arXiv:1606.03476.</citation>
</ref>
<ref id="B195">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hong</surname> <given-names>W.</given-names></name> <name><surname>Luo</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Genetic control of wiring specificity in the fly olfactory system</article-title>. <source>Genetics</source> <volume>196</volume>, <fpage>17</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.113.154336</pub-id><pub-id pub-id-type="pmid">24395823</pub-id></citation>
</ref>
<ref id="B196">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>1982</year>). <article-title>Neural networks and physical systems with emergent collective computational abilities</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>79</volume>, <fpage>2554</fpage>&#x02013;<lpage>2558</lpage>. <pub-id pub-id-type="pmid">6953413</pub-id></citation>
</ref>
<ref id="B197">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>1984</year>). <article-title>Neurons with graded response have collective computational properties like those of two-state neurons</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>81</volume>, <fpage>3088</fpage>&#x02013;<lpage>3092</lpage>. <pub-id pub-id-type="pmid">6587342</pub-id></citation>
</ref>
<ref id="B198">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Neurodynamics of mental exploration</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>107</volume>, <fpage>1648</fpage>&#x02013;<lpage>1653</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0913991107</pub-id><pub-id pub-id-type="pmid">20080534</pub-id></citation>
</ref>
<ref id="B199">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hosoya</surname> <given-names>T.</given-names></name> <name><surname>Baccus</surname> <given-names>S. A.</given-names></name> <name><surname>Meister</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>Dynamic predictive coding by the retina</article-title>. <source>Nature</source> <volume>436</volume>, <fpage>71</fpage>&#x02013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1038/nature03689</pub-id><pub-id pub-id-type="pmid">16001064</pub-id></citation>
</ref>
<ref id="B200">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Y.</given-names></name> <name><surname>Rao</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Predictive coding</article-title>. <source>Wiley Interdiscip. Rev. Cogn. Sci.</source> <volume>2</volume>, <fpage>580</fpage>&#x02013;<lpage>593</lpage>. <pub-id pub-id-type="doi">10.1002/wcs.142</pub-id></citation>
</ref>
<ref id="B201">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huys</surname> <given-names>Q. J.</given-names></name> <name><surname>Lally</surname> <given-names>N.</given-names></name> <name><surname>Faulkner</surname> <given-names>P.</given-names></name> <name><surname>Eshel</surname> <given-names>N.</given-names></name> <name><surname>Seifritz</surname> <given-names>E.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Interplay of approximate planning strategies</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>3098</fpage>&#x02013;<lpage>3103</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1414219112.</pub-id><pub-id pub-id-type="pmid">25675480</pub-id></citation>
</ref>
<ref id="B202">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Isik</surname> <given-names>L.</given-names></name> <name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>Learning and disrupting invariance in visual recognition with a temporal association rule</article-title>. <source>Front. Comput. Neurosci.</source> <volume>6</volume>:<issue>37</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2012.00037</pub-id><pub-id pub-id-type="pmid">22754523</pub-id></citation>
</ref>
<ref id="B203">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Izhikevich</surname> <given-names>E. M.</given-names></name></person-group> (<year>2006</year>). <article-title>Polychronization: computation with spikes</article-title>. <source>Neural Comput.</source> <volume>18</volume>, <fpage>245</fpage>&#x02013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1162/089976606775093882</pub-id><pub-id pub-id-type="pmid">16378515</pub-id></citation>
</ref>
<ref id="B204">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Izhikevich</surname> <given-names>E. M.</given-names></name></person-group> (<year>2007</year>). <article-title>Solving the distal reward problem through linkage of STDP and dopamine signaling</article-title>. <source>Cereb. Cortex</source> <volume>17</volume>, <fpage>2443</fpage>&#x02013;<lpage>2452</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhl152</pub-id><pub-id pub-id-type="pmid">17220510</pub-id></citation>
</ref>
<ref id="B205">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jacobson</surname> <given-names>G. A.</given-names></name> <name><surname>Friedrich</surname> <given-names>R. W.</given-names></name></person-group> (<year>2013</year>). <article-title>Neural circuits: random design of a higher-order olfactory projection</article-title>. <source>Curr. Biol.</source> <volume>23</volume>, <fpage>R448</fpage>&#x02013;<lpage>R451</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2013.04.016</pub-id><pub-id pub-id-type="pmid">23701688</pub-id></citation>
</ref>
<ref id="B206">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Jaderberg</surname> <given-names>M.</given-names></name> <name><surname>Czarnecki</surname> <given-names>W. M.</given-names></name> <name><surname>Osindero</surname> <given-names>S.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <source>Decoupled neural interfaces using synthetic gradients</source>. arXiv:1608.05343.</citation>
</ref>
<ref id="B207">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaderberg</surname> <given-names>M.</given-names></name> <name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>Spatial transformer networks</article-title>, in <source>Advances in Neural Information Processing Systems 28 (NIPS 2015)</source>. <fpage>arXiv</fpage>:1506.02025.</citation>
</ref>
<ref id="B209">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaeger</surname> <given-names>H.</given-names></name> <name><surname>Haas</surname> <given-names>H.</given-names></name></person-group> (<year>2004</year>). <article-title>Harnessing nonlinearity: predicting chaotic systems and saving energy in wireless communication</article-title>. <source>Science</source> <volume>304</volume>, <fpage>78</fpage>&#x02013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1126/science.1091277</pub-id><pub-id pub-id-type="pmid">15064413</pub-id></citation>
</ref>
<ref id="B210">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jara-Ettinger</surname> <given-names>J.</given-names></name> <name><surname>Gweon</surname> <given-names>H.</given-names></name> <name><surname>Schulz</surname> <given-names>L. E.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2016</year>). <article-title>The na&#x000EF;ve utility calculus: computational principles underlying commonsense psychology</article-title>. <source>Trends Cogn. Sci.</source> <volume>20</volume>, <fpage>589</fpage>&#x02013;<lpage>604</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2016.05.011</pub-id><pub-id pub-id-type="pmid">27388875</pub-id></citation>
</ref>
<ref id="B211">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaramillo</surname> <given-names>S.</given-names></name> <name><surname>Pearlmutter</surname> <given-names>B. A.</given-names></name></person-group> (<year>2004</year>). <article-title>A normative model of attention: receptive field modulation</article-title>. <source>Neurocomputing</source> <volume>58</volume>, <fpage>613</fpage>&#x02013;<lpage>618</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2004.01.103</pub-id></citation>
</ref>
<ref id="B212">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jhuang</surname> <given-names>H.</given-names></name> <name><surname>Serre</surname> <given-names>T.</given-names></name> <name><surname>Wolf</surname> <given-names>L.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2007</year>). <article-title>A biologically inspired system for action recognition</article-title>, in <source>IEEE 11th International Conference on Computer Vision, 2007</source> (<publisher-loc>Rio de Janeiro</publisher-loc>: <publisher-name>ICCV</publisher-name>), <volume>2007</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B213">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Shen</surname> <given-names>S.</given-names></name> <name><surname>Cadwell</surname> <given-names>C. R.</given-names></name> <name><surname>Berens</surname> <given-names>P.</given-names></name> <name><surname>Sinz</surname> <given-names>F.</given-names></name> <name><surname>Ecker</surname> <given-names>A. S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Principles of connectivity among morphologically defined cell types in adult neocortex</article-title>. <source>Science</source> <volume>350</volume>, <fpage>aac9462</fpage>. <pub-id pub-id-type="doi">10.1126/science.aac9462</pub-id><pub-id pub-id-type="pmid">26612957</pub-id></citation>
</ref>
<ref id="B214">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ji</surname> <given-names>D.</given-names></name> <name><surname>Wilson</surname> <given-names>M. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Coordinated memory replay in the visual cortex and hippocampus during sleep</article-title>. <source>Nat. Neurosci.</source> <volume>10</volume>, <fpage>100</fpage>&#x02013;<lpage>107</lpage>. <pub-id pub-id-type="doi">10.1038/nn1825</pub-id><pub-id pub-id-type="pmid">17173043</pub-id></citation>
</ref>
<ref id="B215">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johansson</surname> <given-names>F.</given-names></name> <name><surname>Jirenhed</surname> <given-names>D.-A.</given-names></name> <name><surname>Rasmussen</surname> <given-names>A.</given-names></name> <name><surname>Zucca</surname> <given-names>R.</given-names></name> <name><surname>Hesslow</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>Memory trace and timing mechanism localized to cerebellar Purkinje cells</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>14930</fpage>&#x02013;<lpage>14934</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1415371111</pub-id><pub-id pub-id-type="pmid">25267641</pub-id></citation>
</ref>
<ref id="B216">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jonas</surname> <given-names>E.</given-names></name> <name><surname>Kording</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Could a neuroscientist understand a microprocessor?</article-title> <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/055624</pub-id></citation>
</ref>
<ref id="B217">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Joulin</surname> <given-names>A.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <source>Inferring algorithmic patterns with stack-augmented recurrent nets</source>. arXiv:1503.01007.</citation>
</ref>
<ref id="B218">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kalisman</surname> <given-names>N.</given-names></name> <name><surname>Silberberg</surname> <given-names>G.</given-names></name> <name><surname>Markram</surname> <given-names>H.</given-names></name></person-group> (<year>2005</year>). <article-title>The neocortical microcircuit as a tabula rasa</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>102</volume>, <fpage>880</fpage>&#x02013;<lpage>885</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0407088102</pub-id><pub-id pub-id-type="pmid">15630093</pub-id></citation>
</ref>
<ref id="B219">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kanwisher</surname> <given-names>N.</given-names></name> <name><surname>McDermott</surname> <given-names>J.</given-names></name> <name><surname>Chun</surname> <given-names>M. M.</given-names></name></person-group> (<year>1997</year>). <article-title>The fusiform face area: a module in human extrastriate cortex specialized for face perception</article-title>. <source>J. Neurosci.</source> <volume>17</volume>, <fpage>4302</fpage>&#x02013;<lpage>4311</lpage>. <pub-id pub-id-type="pmid">9151747</pub-id></citation>
</ref>
<ref id="B220">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kappel</surname> <given-names>D.</given-names></name> <name><surname>Habenschuss</surname> <given-names>S.</given-names></name> <name><surname>Legenstein</surname> <given-names>R.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2015</year>). <article-title>Network plasticity as bayesian inference</article-title>. <source>PLoS Comput. Biol.</source> <volume>11</volume>:<fpage>e1004485</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004485</pub-id><pub-id pub-id-type="pmid">26545099</pub-id></citation>
</ref>
<ref id="B221">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kappel</surname> <given-names>D.</given-names></name> <name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <article-title>STDP installs in Winner-Take-All circuits an online approximation to hidden Markov model learning</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003511</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003511</pub-id><pub-id pub-id-type="pmid">24675787</pub-id></citation>
</ref>
<ref id="B222">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kempter</surname> <given-names>R.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>van Hemmen</surname> <given-names>J. L.</given-names></name></person-group> (<year>2001</year>). <article-title>Intrinsic stabilization of output rates by spike-based Hebbian learning</article-title>. <source>Neural Comput.</source> <volume>13</volume>, <fpage>2709</fpage>&#x02013;<lpage>2741</lpage>. <pub-id pub-id-type="doi">10.1162/089976601317098501</pub-id><pub-id pub-id-type="pmid">11705408</pub-id></citation>
</ref>
<ref id="B223">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khaligh-Razavi</surname> <given-names>S.-M.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2014</year>). <article-title>Deep supervised, but not unsupervised, models may explain it cortical representation</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003915</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003915</pub-id><pub-id pub-id-type="pmid">25375136</pub-id></citation>
</ref>
<ref id="B224">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <source>Auto-Encoding Variational Bayes</source>. arXiv:1312.6114.</citation>
</ref>
<ref id="B225">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Knill</surname> <given-names>D.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>The Bayesian brain: the role of uncertainty in neural coding and computation</article-title>. <source>Trends Neurosci</source>. <volume>27</volume>, <fpage>712</fpage>&#x02013;<lpage>719</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2004.10.007</pub-id><pub-id pub-id-type="pmid">15541511</pub-id></citation>
</ref>
<ref id="B226">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koechlin</surname> <given-names>E.</given-names></name> <name><surname>Jubault</surname> <given-names>T.</given-names></name></person-group> (<year>2006</year>). <article-title>Broca&#x00027;s area and the hierarchical organization of human behavior</article-title>. <source>Neuron</source> <volume>50</volume>, <fpage>963</fpage>&#x02013;<lpage>974</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2006.05.017</pub-id><pub-id pub-id-type="pmid">16772176</pub-id></citation>
</ref>
<ref id="B227">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Komer</surname> <given-names>B.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>A unified theoretical approach for biological cognition and learning</article-title>. <source>Curr. Opin. Behav. Sci.</source> <volume>11</volume>, <fpage>14</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1016/j.cobeha.2016.03.006</pub-id></citation>
</ref>
<ref id="B228">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;rding</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <article-title>Decision theory: what &#x0201C;should&#x0201D; the nervous system do?</article-title> <source>Science</source> <volume>318</volume>, <fpage>606</fpage>&#x02013;<lpage>610</lpage>. <pub-id pub-id-type="doi">10.1126/science.1142998</pub-id><pub-id pub-id-type="pmid">17962554</pub-id></citation>
</ref>
<ref id="B229">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;rding</surname> <given-names>K.</given-names></name> <name><surname>K&#x000F6;nig</surname> <given-names>P.</given-names></name></person-group> (<year>2000</year>). <article-title>A learning rule for dynamic recruitment and decorrelation</article-title>. <source>Neural Netw.</source> <volume>13</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/S0893-6080(99)00088-X</pub-id><pub-id pub-id-type="pmid">10935452</pub-id></citation>
</ref>
<ref id="B230">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;rding</surname> <given-names>K. P.</given-names></name> <name><surname>K&#x000F6;nig</surname> <given-names>P.</given-names></name></person-group> (<year>2001</year>). <article-title>Supervised and unsupervised learning with two sites of synaptic integration</article-title>. <source>J. Comput. Neurosci.</source> <volume>11</volume>, <fpage>207</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1023/A:1013776130161</pub-id><pub-id pub-id-type="pmid">11796938</pub-id></citation>
</ref>
<ref id="B231">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;rding</surname> <given-names>K. P.</given-names></name> <name><surname>Kayser</surname> <given-names>C.</given-names></name> <name><surname>Einh&#x000E4;user</surname> <given-names>W.</given-names></name> <name><surname>K&#x000F6;nig</surname> <given-names>P.</given-names></name></person-group> (<year>2004</year>). <article-title>How are complex cell properties adapted to the statistics of natural stimuli?</article-title> <source>J. Neurophysiol.</source> <volume>91</volume>, <fpage>206</fpage>&#x02013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00149.2003</pub-id><pub-id pub-id-type="pmid">12904330</pub-id></citation>
</ref>
<ref id="B232">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kouh</surname> <given-names>M.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2008</year>). <article-title>A canonical neural circuit for cortical nonlinear operations</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>1427</fpage>&#x02013;<lpage>1451</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.02-07-466</pub-id><pub-id pub-id-type="pmid">18254695</pub-id></citation>
</ref>
<ref id="B233">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kraus</surname> <given-names>B. J.</given-names></name> <name><surname>Robinson</surname> <given-names>R. J. II</given-names></name> <name><surname>White</surname> <given-names>J. A.</given-names></name> <name><surname>Eichenbaum</surname> <given-names>H.</given-names></name> <name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Hippocampal time cells: time versus path integration</article-title>. <source>Neuron</source> <volume>78</volume>, <fpage>1090</fpage>&#x02013;<lpage>1101</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2013.04.015</pub-id><pub-id pub-id-type="pmid">23707613</pub-id></citation>
</ref>
<ref id="B234">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Mur</surname> <given-names>M.</given-names></name> <name><surname>Bandettini</surname> <given-names>P. A.</given-names></name></person-group> (<year>2008</year>). <article-title>Representational similarity analysis-connecting the branches of systems neuroscience</article-title>. <source>Front. Syst. Neurosci.</source> <volume>2</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.3389/neuro.06.004.2008</pub-id><pub-id pub-id-type="pmid">19104670</pub-id></citation>
</ref>
<ref id="B235">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriete</surname> <given-names>T.</given-names></name> <name><surname>Noelle</surname> <given-names>D. C.</given-names></name> <name><surname>Cohen</surname> <given-names>J. D.</given-names></name> <name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>2013</year>). <article-title>Indirection and symbol-like processing in the prefrontal cortex and basal ganglia</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>110</volume>, <fpage>16390</fpage>&#x02013;<lpage>16935</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1303547110</pub-id><pub-id pub-id-type="pmid">24062434</pub-id></citation>
</ref>
<ref id="B236">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Krishnamurthy</surname> <given-names>R.</given-names></name> <name><surname>Lakshminarayanan</surname> <given-names>A. S.</given-names></name> <name><surname>Kumar</surname> <given-names>P.</given-names></name> <name><surname>Ravindran</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <source>Hierarchical reinforcement learning using spatio-temporal abstractions and deep neural networks</source>. arXiv:1605.05359.</citation>
</ref>
<ref id="B237">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Lake Tahoe, CL</publisher-loc>), <fpage>1097</fpage>&#x02013;<lpage>1105</lpage>.</citation>
</ref>
<ref id="B238">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kulkarni</surname> <given-names>T. D.</given-names></name> <name><surname>Narasimhan</surname> <given-names>K. R.</given-names></name> <name><surname>Saeedi</surname> <given-names>A.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2016</year>). <source>Hierarchical deep reinforcement learning: integrating temporal abstraction and intrinsic motivation</source>. arXiv:1604.06057.</citation>
</ref>
<ref id="B239">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kulkarni</surname> <given-names>T. D.</given-names></name> <name><surname>Whitney</surname> <given-names>W.</given-names></name> <name><surname>Kohli</surname> <given-names>P.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <source>Deep Convolutional Inverse Graphics Network</source>. arXiv:1503.03167.</citation>
</ref>
<ref id="B240">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumaran</surname> <given-names>D.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<year>2016</year>). <article-title>What learning systems do intelligent agents need? complementary learning systems theory updated</article-title>. <source>Trends Cogn. Sci.</source> <volume>20</volume>, <fpage>512</fpage>&#x02013;<lpage>534</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2016.05.004</pub-id><pub-id pub-id-type="pmid">27315762</pub-id></citation>
</ref>
<ref id="B241">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumaran</surname> <given-names>D.</given-names></name> <name><surname>Summerfield</surname> <given-names>J. J.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Maguire</surname> <given-names>E. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Tracking the emergence of conceptual knowledge during human decision making</article-title>. <source>Neuron</source> <volume>63</volume>, <fpage>889</fpage>&#x02013;<lpage>901</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2009.07.030</pub-id><pub-id pub-id-type="pmid">19778516</pub-id></citation>
</ref>
<ref id="B242">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kurach</surname> <given-names>K.</given-names></name> <name><surname>Andrychowicz</surname> <given-names>M.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2015</year>). <source>Neural Random-Access Machines</source>. arXiv:1511.06392.</citation>
</ref>
<ref id="B243">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lake</surname> <given-names>B. M.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <article-title>Human-level concept learning through probabilistic program induction</article-title>. <source>Science</source> <volume>350</volume>, <fpage>1332</fpage>&#x02013;<lpage>1338</lpage>. <pub-id pub-id-type="doi">10.1126/science.aab3050</pub-id><pub-id pub-id-type="pmid">26659050</pub-id></citation>
</ref>
<ref id="B244">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lake</surname> <given-names>B. M.</given-names></name> <name><surname>Ullman</surname> <given-names>T. D.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2016</year>). <source>Building machines that learn and think like people</source>. arXiv:1604.00289.</citation>
</ref>
<ref id="B245">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larkum</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex</article-title>. <source>Trends Neurosci.</source> <volume>36</volume>, <fpage>141</fpage>&#x02013;<lpage>151</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2012.11.006</pub-id><pub-id pub-id-type="pmid">23273272</pub-id></citation>
</ref>
<ref id="B246">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>1995</year>). <article-title>Convolutional networks for images, speech, and time series</article-title>, in <source>The Handbook of Brain Theory and Neural Networks</source>, ed <person-group person-group-type="editor"><name><surname>Arbib</surname> <given-names>M. A.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>3361</fpage>.</citation>
</ref>
<ref id="B247">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation>
</ref>
<ref id="B248">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>A. M.</given-names></name> <name><surname>Tai</surname> <given-names>L.-H.</given-names></name> <name><surname>Zador</surname> <given-names>A.</given-names></name> <name><surname>Wilbrecht</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Between the primate and &#x00027;reptilian&#x00027; brain: rodent models demonstrate the role of corticostriatal circuits in decision making</article-title>. <source>Neuroscience</source> <volume>296</volume>, <fpage>66</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroscience.2014.12.042</pub-id><pub-id pub-id-type="pmid">25575943</pub-id></citation>
</ref>
<ref id="B249">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>T.</given-names></name> <name><surname>Yuille</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Efficient coding of visual scenes by grouping and segmentation: theoretical predictions and biological evidence</article-title>. <source>Department of Statistics, UCLA.</source></citation>
</ref>
<ref id="B250">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>T. S.</given-names></name> <name><surname>Mumford</surname> <given-names>D.</given-names></name></person-group> (<year>2003</year>). <article-title>Hierarchical Bayesian inference in the visual cortex</article-title>. <source>J. Opt. Soc. Am. A Opt. Image Sci. Vis.</source> <volume>20</volume>, <fpage>1434</fpage>&#x02013;<lpage>1448</lpage>. <pub-id pub-id-type="doi">10.1364/JOSAA.20.001434</pub-id><pub-id pub-id-type="pmid">12868647</pub-id></citation>
</ref>
<ref id="B251">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Legenstein</surname> <given-names>R.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>Branch-specific plasticity enables self-organization of nonlinear computation in single neurons</article-title>. <source>J. Neurosci.</source> <volume>31</volume>, <fpage>10787</fpage>&#x02013;<lpage>10802</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5684-10.2011</pub-id><pub-id pub-id-type="pmid">21795531</pub-id></citation>
</ref>
<ref id="B252">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Cornebise</surname> <given-names>J.</given-names></name> <name><surname>G&#x000F3;mez</surname> <given-names>S.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name></person-group> (<year>2015a</year>). <source>Approximate hubel-wiesel modules and the data structures of neural computation</source>. arXiv:1512.08457v1.</citation>
</ref>
<ref id="B253">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Anselmi</surname> <given-names>F.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2015b</year>). <article-title>The invariance hypothesis implies domain-specific regions in visual cortex</article-title>. <source>PLoS Comput. Biol.</source> <volume>11</volume>:<fpage>e1004390</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004390</pub-id><pub-id pub-id-type="pmid">26496457</pub-id></citation>
</ref>
<ref id="B254">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>Q.</given-names></name> <name><surname>Ranzato</surname> <given-names>M.</given-names></name> <name><surname>Monga</surname> <given-names>R.</given-names></name> <name><surname>Devin</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Building high-level features using large scale unsupervised learning</article-title>, in <source>International Conference in Machine Learning</source> (<publisher-loc>Edinburg</publisher-loc>).</citation>
</ref>
<ref id="B255">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lettvin</surname> <given-names>J.</given-names></name> <name><surname>Maturana</surname> <given-names>H.</given-names></name> <name><surname>McCulloch</surname> <given-names>W.</given-names></name> <name><surname>Pitts</surname> <given-names>W.</given-names></name></person-group> (<year>1959</year>). <article-title>What the frog&#x00027;s eye tells the frog&#x00027;s brain</article-title>. <source>Proc. IRE</source> <volume>47</volume>, <fpage>1940</fpage>&#x02013;<lpage>1951</lpage>.</citation>
</ref>
<ref id="B256">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Letzkus</surname> <given-names>J. J.</given-names></name> <name><surname>Kampa</surname> <given-names>B. M.</given-names></name> <name><surname>Stuart</surname> <given-names>G. J.</given-names></name></person-group> (<year>2006</year>). <article-title>Learning rules for spike timing-dependent plasticity depend on dendritic synapse location</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>10420</fpage>&#x02013;<lpage>10429</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2650-06.2006</pub-id><pub-id pub-id-type="pmid">17035526</pub-id></citation>
</ref>
<ref id="B257">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Finn</surname> <given-names>C.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <source>End-to-end training of deep visuomotor policies</source>. arXiv:1504.00702.</citation>
</ref>
<ref id="B258">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lewicki</surname> <given-names>M. S.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Learning overcomplete representations</article-title>. <source>Neural Comput.</source> <volume>12</volume>, <fpage>337</fpage>&#x02013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300015826</pub-id><pub-id pub-id-type="pmid">10636946</pub-id></citation>
</ref>
<ref id="B259">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lewis</surname> <given-names>S. N.</given-names></name> <name><surname>Harris</surname> <given-names>K. D.</given-names></name></person-group> (<year>2014</year>). <source>The Neural Marketplace: I. General Formalism and Linear Theory</source>. Technical Report. bioRxiv:013185.</citation>
</ref>
<ref id="B260">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>Bridging the gaps between residual learning, recurrent neural networks and visual cortex</source>. arXiv:1604.03640.</citation>
</ref>
<ref id="B261">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>Q.</given-names></name> <name><surname>Leibo</surname> <given-names>J. Z.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <source>How important is weight symmetry in backpropagation?</source> arXiv:1510.05067.</citation>
</ref>
<ref id="B262">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lillicrap</surname> <given-names>T. P.</given-names></name> <name><surname>Cownden</surname> <given-names>D.</given-names></name> <name><surname>Tweed</surname> <given-names>D. B.</given-names></name> <name><surname>Akerman</surname> <given-names>C. J.</given-names></name></person-group> (<year>2014</year>). <source>Random feedback weights support learning in deep neural networks</source>. arXiv:1411.0247.</citation>
</ref>
<ref id="B263">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>N.</given-names></name> <name><surname>Dicarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Neuronal learning of invariant object representation in the ventral visual stream is not dependent on reward</article-title>. <source>J. Neurosci.</source> <volume>32</volume>, <fpage>6611</fpage>&#x02013;<lpage>6620</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3786-11.2012</pub-id><pub-id pub-id-type="pmid">22573683</pub-id></citation>
</ref>
<ref id="B264">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J. K.</given-names></name> <name><surname>Buonomano</surname> <given-names>D. V.</given-names></name></person-group> (<year>2009</year>). <article-title>Embedding multiple trajectories in simulated recurrent neural networks in a self-organizing manner</article-title>. <source>J. Neurosci.</source> <volume>29</volume>, <fpage>13172</fpage>&#x02013;<lpage>13181</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2358-09.2009</pub-id><pub-id pub-id-type="pmid">19846705</pub-id></citation>
</ref>
<ref id="B265">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Livni</surname> <given-names>R.</given-names></name> <name><surname>Shalev-Shwartz</surname> <given-names>S.</given-names></name> <name><surname>Shamir</surname> <given-names>O.</given-names></name></person-group> (<year>2013</year>). <source>An algorithm for training polynomial networks</source>. arXiv:1304.7045.</citation>
</ref>
<ref id="B266">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lotter</surname> <given-names>W.</given-names></name> <name><surname>Kreiman</surname> <given-names>G.</given-names></name> <name><surname>Cox</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <source>Unsupervised learning of visual structure using predictive generative networks</source>. arXiv:1511.06380.</citation>
</ref>
<ref id="B267">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Lotter</surname> <given-names>W.</given-names></name> <name><surname>Kreiman</surname> <given-names>G.</given-names></name> <name><surname>Cox</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <source>Deep predictive coding networks for video prediction and unsupervised learning</source>. arXiv:1605.08104.</citation>
</ref>
<ref id="B268">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luck</surname> <given-names>S. J.</given-names></name> <name><surname>Vogel</surname> <given-names>E. K.</given-names></name></person-group> (<year>1997</year>). <article-title>The capacity of visual working memory for features and conjunctions</article-title>. <source>Nature</source> <volume>390</volume>, <fpage>279</fpage>&#x02013;<lpage>281</lpage>. <pub-id pub-id-type="pmid">9384378</pub-id></citation>
</ref>
<ref id="B269">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Boix</surname> <given-names>X.</given-names></name> <name><surname>Roig</surname> <given-names>G.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name></person-group> (<year>2015</year>). <source>Foveation-based mechanisms alleviate adversarial examples</source>. arXiv:1511.06292.</citation>
</ref>
<ref id="B270">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lyons</surname> <given-names>A. B.</given-names></name> <name><surname>Cheries</surname> <given-names>E. W.</given-names></name></person-group> (<year>2016</year>). <article-title>Inferring social disposition by sound and surface appearance in infancy</article-title>. <source>J. Cogn. Dev.</source> <pub-id pub-id-type="doi">10.1080/15248372.2016.1200048</pub-id></citation>
</ref>
<ref id="B271">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>W. J.</given-names></name> <name><surname>Beck</surname> <given-names>J. M.</given-names></name> <name><surname>Latham</surname> <given-names>P. E.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Bayesian inference with probabilistic population codes</article-title>. <source>Nat. Neurosci.</source> <volume>9</volume>, <fpage>1432</fpage>&#x02013;<lpage>1438</lpage>. <pub-id pub-id-type="doi">10.1038/nn1790</pub-id><pub-id pub-id-type="pmid">17057707</pub-id></citation>
</ref>
<ref id="B272">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Searching for principles of brain computation</article-title>. <source>Curr. Opin. Behav. Sci.</source> <volume>11</volume>, <fpage>81</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1016/j.cobeha.2016.06.003.</pub-id></citation>
</ref>
<ref id="B273">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maass</surname> <given-names>W.</given-names></name> <name><surname>Joshi</surname> <given-names>P.</given-names></name> <name><surname>Sontag</surname> <given-names>E. D.</given-names></name></person-group> (<year>2007</year>). <article-title>Computational aspects of feedback in neural circuits</article-title>. <source>PLoS Comput. Biol.</source> <volume>3</volume>:<fpage>e165</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.0020165</pub-id><pub-id pub-id-type="pmid">17238280</pub-id></citation>
</ref>
<ref id="B274">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maass</surname> <given-names>W.</given-names></name> <name><surname>Natschl&#x000E4;ger</surname> <given-names>T.</given-names></name> <name><surname>Markram</surname> <given-names>H.</given-names></name></person-group> (<year>2002</year>). <article-title>Real-time computing without stable states: a new framework for neural computation based on perturbations</article-title>. <source>Neural Comput.</source> <volume>14</volume>, <fpage>2531</fpage>&#x02013;<lpage>2560</lpage>. <pub-id pub-id-type="doi">10.1162/089976602760407955</pub-id><pub-id pub-id-type="pmid">12433288</pub-id></citation>
</ref>
<ref id="B275">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>MacDonald</surname> <given-names>C. J.</given-names></name> <name><surname>Lepage</surname> <given-names>K. Q.</given-names></name> <name><surname>Eden</surname> <given-names>U. T.</given-names></name> <name><surname>Eichenbaum</surname> <given-names>H.</given-names></name></person-group> (<year>2011</year>). <article-title>Hippocampal time cells bridge the gap in memory for discontiguous events</article-title>. <source>Neuron</source> <volume>71</volume>, <fpage>737</fpage>&#x02013;<lpage>749</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2011.07.012</pub-id><pub-id pub-id-type="pmid">21867888</pub-id></citation>
</ref>
<ref id="B276">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Maclaurin</surname> <given-names>D.</given-names></name> <name><surname>Duvenaud</surname> <given-names>D.</given-names></name> <name><surname>Adams</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <source>Gradient-based hyperparameter optimization through reversible learning</source>. arXiv:1502.03492.</citation>
</ref>
<ref id="B277">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Makin</surname> <given-names>J. G.</given-names></name> <name><surname>Dichter</surname> <given-names>B. K.</given-names></name> <name><surname>Sabes</surname> <given-names>P. N.</given-names></name></person-group> (<year>2016</year>). <source>Recurrent exponential-family harmoniums without backprop-through-time</source>. arXiv:1605.05799.</citation>
</ref>
<ref id="B278">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Makin</surname> <given-names>J. G.</given-names></name> <name><surname>Fellows</surname> <given-names>M. R.</given-names></name> <name><surname>Sabes</surname> <given-names>P. N.</given-names></name></person-group> (<year>2013</year>). <article-title>Learning multisensory integration and coordinate transformation via density estimation</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003035</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003035</pub-id><pub-id pub-id-type="pmid">23637588</pub-id></citation>
</ref>
<ref id="B279">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mandelblat-Cerf</surname> <given-names>Y.</given-names></name> <name><surname>Las</surname> <given-names>L.</given-names></name> <name><surname>Denisenko</surname> <given-names>N.</given-names></name> <name><surname>Fee</surname> <given-names>M. S.</given-names></name></person-group> (<year>2014</year>). <article-title>A role for descending auditory cortical projections in songbird vocal learning</article-title>. <source>eLife</source> <volume>3</volume>:<fpage>e02152</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.02152</pub-id><pub-id pub-id-type="pmid">24935934</pub-id></citation>
</ref>
<ref id="B280">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mansinghka</surname> <given-names>V.</given-names></name> <name><surname>Jonas</surname> <given-names>E.</given-names></name></person-group> (<year>2014</year>). <source>Building fast bayesian computing machines out of intentionally stochastic, digital parts</source>. arXiv:1402.4914.</citation>
</ref>
<ref id="B281">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marblestone</surname> <given-names>A. H.</given-names></name> <name><surname>Boyden</surname> <given-names>E. S.</given-names></name></person-group> (<year>2014</year>). <article-title>Designing tools for assumption-proof brain mapping</article-title>. <source>Neuron</source> <volume>83</volume>, <fpage>1239</fpage>&#x02013;<lpage>1241</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2014.09.004</pub-id><pub-id pub-id-type="pmid">25233303</pub-id></citation>
</ref>
<ref id="B282">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Marcus</surname> <given-names>G.</given-names></name></person-group> (<year>2001</year>). <source>The Algebraic Mind: Integrating Connectionism and Cognitive Science</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B283">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Marcus</surname> <given-names>G.</given-names></name></person-group> (<year>2004</year>). <source>The Birth of the Mind: How a Tiny Number of Genes Creates the Complexities of Human Thought</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Basic Books</publisher-name>.</citation>
</ref>
<ref id="B284">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Marcus</surname> <given-names>G.</given-names></name> <name><surname>Marblestone</surname> <given-names>A.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014a</year>). <source>Frequently asked question for: the atoms of neural computation</source>. arXiv:1410.8826.</citation>
</ref>
<ref id="B285">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcus</surname> <given-names>G.</given-names></name> <name><surname>Marblestone</surname> <given-names>A.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014b</year>). <article-title>The atoms of neural computation</article-title>. <source>Science</source> <volume>346</volume>, <fpage>551</fpage>&#x02013;<lpage>552</lpage>. <pub-id pub-id-type="doi">10.1126/science.1261661</pub-id><pub-id pub-id-type="pmid">25359953</pub-id></citation>
</ref>
<ref id="B286">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marder</surname> <given-names>E.</given-names></name> <name><surname>Goaillard</surname> <given-names>J.-M.</given-names></name></person-group> (<year>2006</year>). <article-title>Variability, compensation and homeostasis in neuron and network function</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>7</volume>, <fpage>563</fpage>&#x02013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1949</pub-id><pub-id pub-id-type="pmid">16791145</pub-id></citation>
</ref>
<ref id="B287">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markowitz</surname> <given-names>D. A.</given-names></name> <name><surname>Curtis</surname> <given-names>C. E.</given-names></name> <name><surname>Pesaran</surname> <given-names>B.</given-names></name></person-group> (<year>2015</year>). <article-title>Multiple component networks support working memory in prefrontal cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>11084</fpage>&#x02013;<lpage>11089</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1504172112</pub-id><pub-id pub-id-type="pmid">26371320</pub-id></citation>
</ref>
<ref id="B288">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markram</surname> <given-names>H.</given-names></name> <name><surname>L&#x000FC;bke</surname> <given-names>J.</given-names></name> <name><surname>Frotscher</surname> <given-names>M.</given-names></name> <name><surname>Sakmann</surname> <given-names>B.</given-names></name></person-group> (<year>1997</year>). <article-title>Regulation of synaptic efficacy by coincidence of postsynaptic APs and EPSPs</article-title>. <source>Science</source> <volume>275</volume>, <fpage>213</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="pmid">8985014</pub-id></citation>
</ref>
<ref id="B289">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markram</surname> <given-names>H.</given-names></name> <name><surname>Muller</surname> <given-names>E.</given-names></name> <name><surname>Ramaswamy</surname> <given-names>S.</given-names></name> <name><surname>Reimann</surname> <given-names>M. W.</given-names></name> <name><surname>Abdellah</surname> <given-names>M.</given-names></name> <name><surname>Sanchez</surname> <given-names>C. A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Reconstruction and simulation of neocortical microcircuitry</article-title>. <source>Cell</source> <volume>163</volume>, <fpage>456</fpage>&#x02013;<lpage>492</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2015.09.029</pub-id><pub-id pub-id-type="pmid">26451489</pub-id></citation>
</ref>
<ref id="B290">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name></person-group> (<year>1969</year>). <article-title>A theory of cerebellar cortex</article-title>. <source>J. Physiol.</source> <volume>202</volume>, <fpage>437</fpage>&#x02013;<lpage>470</lpage>. <pub-id pub-id-type="pmid">5784296</pub-id></citation>
</ref>
<ref id="B291">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Martens</surname> <given-names>J.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2011</year>). <article-title>Learning recurrent neural networks with hessian-free optimization</article-title>,&#x0201D; in <source>Proceedings of the 28th International Conference on Machine Learning</source> (<publisher-loc>Bellevue, WA</publisher-loc>).</citation>
</ref>
<ref id="B292">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCandliss</surname> <given-names>B. D.</given-names></name> <name><surname>Cohen</surname> <given-names>L.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name></person-group> (<year>2003</year>). <article-title>The visual word form area: expertise for reading in the fusiform gyrus</article-title>. <source>Trends Cogn. Sci.</source> <volume>7</volume>, <fpage>293</fpage>&#x02013;<lpage>299</lpage>. <pub-id pub-id-type="doi">10.1016/S1364-6613(03)00134-7</pub-id><pub-id pub-id-type="pmid">12860187</pub-id></citation>
</ref>
<ref id="B293">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCulloch</surname> <given-names>W. S.</given-names></name> <name><surname>Pitts</surname> <given-names>W.</given-names></name></person-group> (<year>1943</year>). <article-title>A logical calculus of the ideas immanent in nervous activity</article-title>. <source>Bull. Math. Biophys.</source> <volume>5</volume>, <fpage>115</fpage>&#x02013;<lpage>133</lpage>. <pub-id pub-id-type="doi">10.1007/BF02478259</pub-id></citation>
</ref>
<ref id="B294">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKinstry</surname> <given-names>J. L.</given-names></name> <name><surname>Edelman</surname> <given-names>G. M.</given-names></name> <name><surname>Krichmar</surname> <given-names>J. L.</given-names></name></person-group> (<year>2006</year>). <article-title>A cerebellar model for predictive motor control tested in a brain-based device</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>103</volume>, <fpage>3387</fpage>&#x02013;<lpage>3392</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0511281103</pub-id><pub-id pub-id-type="pmid">16488974</pub-id></citation>
</ref>
<ref id="B295">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McKone</surname> <given-names>E.</given-names></name> <name><surname>Crookes</surname> <given-names>K.</given-names></name> <name><surname>Kanwisher</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <article-title>The cognitive and neural development of face recognition in humans</article-title>, in <source>The Cognitive Neurosciences, 4th Edn</source>, eds <person-group person-group-type="editor"><name><surname>Michael</surname> <given-names>S. G.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <volume>4</volume>, <fpage>467</fpage>&#x02013;<lpage>482</lpage>.</citation>
</ref>
<ref id="B296">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McLeod</surname> <given-names>P.</given-names></name> <name><surname>Dienes</surname> <given-names>Z.</given-names></name></person-group> (<year>1996</year>). <article-title>Do fielders know where to go to catch the ball or only how to get there?</article-title> <source>J. Exp. Psychol. Hum. Percept. Perform.</source> <volume>22</volume>, <fpage>531</fpage>. <pub-id pub-id-type="pmid">26626544</pub-id></citation>
</ref>
<ref id="B297">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mel</surname> <given-names>B.</given-names></name></person-group> (<year>1992</year>). <article-title>The clusteron: toward a simple abstraction for a complex neuron</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>4</volume>, <fpage>35</fpage>&#x02013;<lpage>42</lpage>.</citation>
</ref>
<ref id="B298">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meltzoff</surname> <given-names>A. N.</given-names></name></person-group> (<year>1999</year>). <article-title>Born to learn: what infants learn from watching us</article-title>. <source>Role Early Exp. Infant Dev.</source> <fpage>145</fpage>&#x02013;<lpage>164</lpage>.</citation>
</ref>
<ref id="B299">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meltzoff</surname> <given-names>A. N.</given-names></name> <name><surname>Waismeyer</surname> <given-names>A.</given-names></name> <name><surname>Gopnik</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Learning about causes from people: observational causal learning in 24-month-old infants</article-title>. <source>Dev. Psychol.</source> <volume>48</volume>, <fpage>1215</fpage>&#x02013;<lpage>1258</lpage>. <pub-id pub-id-type="doi">10.1037/a0027440</pub-id><pub-id pub-id-type="pmid">22369335</pub-id></citation>
</ref>
<ref id="B300">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Meltzoff</surname> <given-names>A. N.</given-names></name> <name><surname>Williamson</surname> <given-names>R. A.</given-names></name> <name><surname>Marshall</surname> <given-names>P. J.</given-names></name></person-group> (<year>2013</year>). <article-title>11 developmental perspectives on action science: lessons from infant imitation and cognitive neuroscience</article-title>, in <source>Action Science: Foundations of an Emerging Discipline</source>, eds <person-group person-group-type="editor"><name><surname>Prinz</surname> <given-names>W.</given-names></name> <name><surname>Beisert</surname> <given-names>M.</given-names></name> <name><surname>Herwig</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>281</fpage>&#x02013;<lpage>306</lpage>.</citation>
</ref>
<ref id="B301">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>G. A.</given-names></name></person-group> (<year>1956</year>). <article-title>The magical number seven, plus or minus two: some limits on our capacity for processing information</article-title>. <source>Psychol. Rev.</source> <volume>63</volume>, <fpage>81</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="pmid">13310704</pub-id></citation>
</ref>
<ref id="B302">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>K. D.</given-names></name> <name><surname>Keller</surname> <given-names>J. B.</given-names></name> <name><surname>Stryker</surname> <given-names>M. P.</given-names></name></person-group> (<year>1989</year>). <article-title>Ocular dominance column development: analysis and simulation</article-title>. <source>Science</source>, <volume>245</volume>, <fpage>605</fpage>&#x02013;<lpage>615</lpage>. <pub-id pub-id-type="pmid">2762813</pub-id></citation>
</ref>
<ref id="B303">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>K. D.</given-names></name> <name><surname>MacKay</surname> <given-names>D. J. C.</given-names></name></person-group> (<year>1994</year>). <article-title>The role of constraints in hebbian learning</article-title>. <source>Neural Comput.</source> <volume>6</volume>, <fpage>100</fpage>&#x02013;<lpage>126</lpage>.</citation>
</ref>
<ref id="B304">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M.</given-names></name></person-group> (<year>1977</year>). <article-title>Plain talk about neurodevelopmental epistemology</article-title>, in <source>IJCAI&#x00027;77 Proceedings of the 5th International Joint Conference on Artificial Intelligence</source> (<publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>), <fpage>1083</fpage>&#x02013;<lpage>1092</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://hdl.handle.net/1721.1/5763">http://hdl.handle.net/1721.1/5763</ext-link></citation>
</ref>
<ref id="B305">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M.</given-names></name></person-group> (<year>1988</year>). <source>Society of Mind</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Simon and Schuster</publisher-name>.</citation>
</ref>
<ref id="B306">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <source>The Emotion Machine</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Pantheon</publisher-name>.</citation>
</ref>
<ref id="B307">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M. L.</given-names></name></person-group> (<year>1991</year>). <article-title>Logical versus analogical or symbolic versus connectionist or neat versus scruffy</article-title>. <source>AI magazine</source> <volume>12</volume>, <fpage>34</fpage>&#x02013;<lpage>51</lpage>.</citation>
</ref>
<ref id="B308">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Minsky</surname> <given-names>M. L.</given-names></name> <name><surname>Papert</surname> <given-names>S.</given-names></name></person-group> (<year>1972</year>). <source>Perceptrons: An Introduction to Computational Geometry</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B309">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mishra</surname> <given-names>R. K.</given-names></name> <name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Guzman</surname> <given-names>S. J.</given-names></name> <name><surname>Jonas</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Symmetric spike timing-dependent plasticity at ca3-ca3 synapses optimizes storage and recall in autoassociative networks</article-title>. <source>Nat. Commun.</source> <volume>7</volume>:<fpage>1552</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms11552</pub-id><pub-id pub-id-type="pmid">27174042</pub-id></citation>
</ref>
<ref id="B310">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Mitchell</surname> <given-names>T. M.</given-names></name></person-group> (<year>1980</year>). <article-title>The need for biases in learning generalizations</article-title>, in <source>Readings in Machine Learning</source>, eds <person-group person-group-type="editor"><name><surname>Shavlik</surname> <given-names>J. W.</given-names></name> <name><surname>Dietterich</surname> <given-names>T. G.</given-names></name></person-group> (Morgan Kauffman), <fpage>184</fpage>&#x02013;<lpage>191</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.cs.nott.ac.uk/&#x0007E;bsl/G52HPA/articles/Mitchell:80a.pdf">http://www.cs.nott.ac.uk/&#x0007E;bsl/G52HPA/articles/Mitchell:80a.pdf</ext-link></citation>
</ref>
<ref id="B311">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Miyagawa</surname> <given-names>S.</given-names></name> <name><surname>Berwick</surname> <given-names>R. C.</given-names></name> <name><surname>Okanoya</surname> <given-names>K.</given-names></name></person-group> (<year>2013</year>). <article-title>The emergence of hierarchical structure in human language</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<issue>71</issue>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00071</pub-id></citation>
</ref>
<ref id="B312">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Heess</surname> <given-names>N.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <article-title>Recurrent models of visual attention</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>2204</fpage>&#x02013;<lpage>2212</lpage>.</citation>
</ref>
<ref id="B313">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source> <volume>518</volume>, <fpage>529</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id></citation>
</ref>
<ref id="B314">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mobahi</surname> <given-names>H.</given-names></name> <name><surname>Collobert</surname> <given-names>R.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>Deep learning from temporal coherence in video</article-title>,&#x0201D; in <source>Proceedings of the 26th Annual International Conference on Machine Learning - ICML &#x00027;09</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM Press</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B315">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moberget</surname> <given-names>T.</given-names></name> <name><surname>Gullesen</surname> <given-names>E. H.</given-names></name> <name><surname>Andersson</surname> <given-names>S.</given-names></name> <name><surname>Ivry</surname> <given-names>R. B.</given-names></name> <name><surname>Endestad</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Generalized role for the cerebellum in encoding internal models: evidence from semantic processing</article-title>. <source>J. Neurosci.</source> <volume>34</volume>, <fpage>2871</fpage>&#x02013;<lpage>2878</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2264-13.2014</pub-id><pub-id pub-id-type="pmid">24553928</pub-id></citation>
</ref>
<ref id="B316">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mohamed</surname> <given-names>S.</given-names></name> <name><surname>Rezende</surname> <given-names>D. J.</given-names></name></person-group> (<year>2015</year>). <source>Variational information maximisation for intrinsically motivated reinforcement learning</source>. arXiv:1509.08731</citation>
</ref>
<ref id="B317">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mordatch</surname> <given-names>I.</given-names></name> <name><surname>Todorov</surname> <given-names>E.</given-names></name> <name><surname>Popovi&#x00107;</surname> <given-names>Z.</given-names></name></person-group> (<year>2012</year>). <article-title>Discovery of complex behaviors through contact-invariant optimization</article-title>. <source>ACM Trans. Graph.</source> <volume>31</volume>:<fpage>43</fpage>. <pub-id pub-id-type="doi">10.1145/2185520.2185539</pub-id></citation>
</ref>
<ref id="B318">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morgan</surname> <given-names>J. L.</given-names></name> <name><surname>Berger</surname> <given-names>D. R.</given-names></name> <name><surname>Wetzel</surname> <given-names>A. W.</given-names></name> <name><surname>Lichtman</surname> <given-names>J. W.</given-names></name></person-group> (<year>2016</year>). <article-title>The fuzzy logic of network connectivity in mouse visual thalamus</article-title>. <source>Cell</source> <volume>165</volume>, <fpage>192</fpage>&#x02013;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2016.02.033</pub-id><pub-id pub-id-type="pmid">27015312</pub-id></citation>
</ref>
<ref id="B319">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mushiake</surname> <given-names>H.</given-names></name> <name><surname>Saito</surname> <given-names>N.</given-names></name> <name><surname>Sakamoto</surname> <given-names>K.</given-names></name> <name><surname>Itoyama</surname> <given-names>Y.</given-names></name> <name><surname>Tanji</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Activity in the lateral prefrontal cortex reflects multiple steps of future events in action plans</article-title>. <source>Neuron</source> <volume>50</volume>, <fpage>631</fpage>&#x02013;<lpage>641</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2006.03.045</pub-id><pub-id pub-id-type="pmid">16701212</pub-id></citation>
</ref>
<ref id="B320">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nardini</surname> <given-names>M.</given-names></name> <name><surname>Bedford</surname> <given-names>R.</given-names></name> <name><surname>Mareschal</surname> <given-names>D.</given-names></name></person-group> (<year>2010</year>). <article-title>Fusion of visual cues is not mandatory in children</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>107</volume>, <fpage>17041</fpage>&#x02013;<lpage>17046</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1001699107</pub-id><pub-id pub-id-type="pmid">20837526</pub-id></citation>
</ref>
<ref id="B321">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Neelakantan</surname> <given-names>A.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2015</year>). <source>Neural programmer: inducing latent programs with gradient descent</source>. arXiv:1511.04834.</citation>
</ref>
<ref id="B322">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name> <name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Bayesian computation emerges in generic cortical microcircuits through spike-timing-dependent plasticity</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003037</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003037</pub-id><pub-id pub-id-type="pmid">23633941</pub-id></citation>
</ref>
<ref id="B323">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>A.</given-names></name> <name><surname>Russell</surname> <given-names>S.</given-names></name></person-group> (<year>2000</year>). <article-title>Algorithms for inverse reinforcement learning</article-title>,&#x0201D; in <source>ICML &#x00027;00 Proceedings of the Seventeenth International Conference on Machine Learning</source> (<publisher-loc>San Francisco, CA</publisher-loc>).</citation>
</ref>
<ref id="B324">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Noroozi</surname> <given-names>M.</given-names></name> <name><surname>Favaro</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <source>Unsupervised learning of visual representations by solving jigsaw puzzles</source>. arXiv:1603.09246.</citation>
</ref>
<ref id="B325">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x000D3;lafsd&#x000F3;ttir</surname> <given-names>H. F.</given-names></name> <name><surname>Barry</surname> <given-names>C.</given-names></name> <name><surname>Saleem</surname> <given-names>A. B.</given-names></name> <name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Spiers</surname> <given-names>H. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Hippocampal place cells construct reward related sequences through unexplored space</article-title>. <source>eLife</source> <volume>4</volume>:<fpage>e06063</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.06063</pub-id><pub-id pub-id-type="pmid">26112828</pub-id></citation>
</ref>
<ref id="B326">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Ollivier</surname> <given-names>Y.</given-names></name> <name><surname>Charpiat</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <source>Training recurrent networks online without backtracking</source>. arXiv:1507.07680.</citation>
</ref>
<ref id="B327">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B. A.</given-names></name> <name><surname>Anderson</surname> <given-names>C. H.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name></person-group> (<year>1993</year>). <article-title>A neurobiological model of visual attention and invariant pattern recognition based on dynamic routing of information</article-title>. <source>J. Neurosci.</source> <volume>13</volume>, <fpage>4700</fpage>&#x02013;<lpage>4719</lpage>. <pub-id pub-id-type="pmid">8229193</pub-id></citation>
</ref>
<ref id="B328">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B. A.</given-names></name> <name><surname>Field</surname> <given-names>D. J.</given-names></name></person-group> (<year>1996</year>). <article-title>Emergence of simple-cell receptive field properties by learning a sparse code for natural images</article-title>. <source>Nature</source> <volume>381</volume>, <fpage>607</fpage>&#x02013;<lpage>609</lpage>. <pub-id pub-id-type="pmid">8637596</pub-id></citation>
</ref>
<ref id="B329">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B. A.</given-names></name> <name><surname>Field</surname> <given-names>D. J.</given-names></name></person-group> (<year>1997</year>). <article-title>Sparse coding with an overcomplete basis set: a strategy employed by V1?</article-title> <source>Vis. Res.</source> <volume>37</volume>, <fpage>3311</fpage>&#x02013;<lpage>3325</lpage>. <pub-id pub-id-type="pmid">9425546</pub-id></citation>
</ref>
<ref id="B330">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B. A.</given-names></name> <name><surname>Field</surname> <given-names>D. J.</given-names></name></person-group> (<year>2004</year>). <article-title>What is the other 85% of v1 doing</article-title>. <source>Prob. Syst. Neurosci.</source> <volume>4</volume>, <fpage>182</fpage>&#x02013;<lpage>211</lpage>. <pub-id pub-id-type="doi">10.1093/acprof:oso/9780195148220.003.0010</pub-id></citation>
</ref>
<ref id="B331">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oquab</surname> <given-names>M.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Laptev</surname> <given-names>I.</given-names></name> <name><surname>Sivic</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Learning and transferring mid-level image representations using convolutional neural networks</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Columbus, OH</publisher-loc>), <fpage>1717</fpage>&#x02013;<lpage>1724</lpage>.</citation>
</ref>
<ref id="B332">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>1996</year>). <article-title>Biologically plausible error-driven learning using local activation differences: the generalized recirculation algorithm</article-title>. <source>Neural Comput.</source> <volume>8</volume>, <fpage>895</fpage>&#x02013;<lpage>938</lpage>.</citation>
</ref>
<ref id="B333">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>2006</year>). <article-title>Biologically based computational models of high-level cognition</article-title>. <source>Science</source> <volume>314</volume>, <fpage>91</fpage>&#x02013;<lpage>94</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127242</pub-id><pub-id pub-id-type="pmid">17023651</pub-id></citation>
</ref>
<ref id="B334">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Frank</surname> <given-names>M. J.</given-names></name></person-group> (<year>2006</year>). <article-title>Making working memory work: a computational model of learning in the prefrontal cortex and basal ganglia</article-title>. <source>Neural Comput.</source> <volume>18</volume>, <fpage>283</fpage>&#x02013;<lpage>328</lpage>. <pub-id pub-id-type="doi">10.1162/089976606775093909</pub-id><pub-id pub-id-type="pmid">16378516</pub-id></citation>
</ref>
<ref id="B335">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Hazy</surname> <given-names>T. E.</given-names></name> <name><surname>Mollick</surname> <given-names>J.</given-names></name> <name><surname>Mackie</surname> <given-names>P.</given-names></name> <name><surname>Herd</surname> <given-names>S.</given-names></name></person-group> (<year>2014a</year>). <source>Goal-driven cognition in the brain: a computational framework</source>. arXiv:1404.7591.</citation>
</ref>
<ref id="B336">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Munakata</surname> <given-names>Y.</given-names></name> <name><surname>Frank</surname> <given-names>M. J.</given-names></name> <name><surname>Hazy</surname> <given-names>T. E.</given-names></name></person-group> (<year>2012</year>). <source>Computational Cognitive Neuroscience, 1st Edn.</source> Wiki Book. Available online at: <ext-link ext-link-type="uri" xlink:href="http://ccnbook.colorado.edu">http://ccnbook.colorado.edu</ext-link></citation>
</ref>
<ref id="B337">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Wyatte</surname> <given-names>D.</given-names></name> <name><surname>Rohrlich</surname> <given-names>J.</given-names></name></person-group> (<year>2014b</year>). <source>Learning through time in the thalamocortical loops</source>. arXiv:1407.3432, 37.</citation>
</ref>
<ref id="B338">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Orhan</surname> <given-names>A. E.</given-names></name> <name><surname>Ma</surname> <given-names>W. J.</given-names></name></person-group> (<year>2016</year>). <source>The inevitability of probability: probabilistic inference in generic neural networks trained with non-probabilistic feedback</source> arXiv:1601.03060.</citation>
</ref>
<ref id="B339">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Palmer</surname> <given-names>L. M.</given-names></name> <name><surname>Shai</surname> <given-names>A. S.</given-names></name> <name><surname>Reeve</surname> <given-names>J. E.</given-names></name> <name><surname>Anderson</surname> <given-names>H. L.</given-names></name> <name><surname>Paulsen</surname> <given-names>O.</given-names></name> <name><surname>Larkum</surname> <given-names>M. E.</given-names></name></person-group> (<year>2014</year>). <article-title>NMDA spikes enhance action potential generation during sensory input</article-title>. <source>Nat. Neurosci.</source> <volume>17</volume>, <fpage>383</fpage>&#x02013;<lpage>390</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3646</pub-id><pub-id pub-id-type="pmid">24487231</pub-id></citation>
</ref>
<ref id="B340">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parisien</surname> <given-names>C.</given-names></name> <name><surname>Anderson</surname> <given-names>C. H.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>Solving the problem of negative synaptic weights in cortical models</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>1473</fpage>&#x02013;<lpage>1494</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.07-06-295</pub-id><pub-id pub-id-type="pmid">18254696</pub-id></citation>
</ref>
<ref id="B341">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pasupathy</surname> <given-names>A.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2005</year>). <article-title>Different time courses of learning-related activity in the prefrontal cortex and striatum</article-title>. <source>Nature</source> <volume>433</volume>, <fpage>873</fpage>&#x02013;<lpage>876</lpage>. <pub-id pub-id-type="doi">10.1038/nature03287</pub-id><pub-id pub-id-type="pmid">15729344</pub-id></citation>
</ref>
<ref id="B342">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Patel</surname> <given-names>A.</given-names></name> <name><surname>Nguyen</surname> <given-names>T.</given-names></name> <name><surname>Baraniuk</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <source>A probabilistic theory of deep learning</source>. arXiv:1504.00641.</citation>
</ref>
<ref id="B343">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pehlevan</surname> <given-names>C.</given-names></name> <name><surname>Chklovskii</surname> <given-names>D. B.</given-names></name></person-group> (<year>2015</year>). <article-title>Optimization theory of hebbian/anti-hebbian networks for pca and whitening</article-title>,&#x0201D; in <source>53rd Annual Allerton Conference on Communication, Control, and Computing</source> (<publisher-loc>Monticello, IL</publisher-loc>), <fpage>1458</fpage>&#x02013;<lpage>1465</lpage>.</citation>
</ref>
<ref id="B344">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perea</surname> <given-names>G.</given-names></name> <name><surname>Navarrete</surname> <given-names>M.</given-names></name> <name><surname>Araque</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Tripartite synapses: astrocytes process and control synaptic information</article-title>. <source>Trends Neurosci.</source> <volume>32</volume>, <fpage>421</fpage>&#x02013;<lpage>431</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2009.05.001</pub-id><pub-id pub-id-type="pmid">19615761</pub-id></citation>
</ref>
<ref id="B345">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Petrov</surname> <given-names>A. A.</given-names></name> <name><surname>Jilk</surname> <given-names>D. J.</given-names></name> <name><surname>O&#x00027;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>2010</year>). <article-title>The Leabra architecture: specialization without modularity</article-title>. <source>Behav. Brain Sci.</source> <volume>33</volume>, <fpage>286</fpage>&#x02013;<lpage>287</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X10001160</pub-id></citation>
</ref>
<ref id="B346">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pezzulo</surname> <given-names>G.</given-names></name> <name><surname>Verschure</surname> <given-names>P. F. M. J.</given-names></name> <name><surname>Balkenius</surname> <given-names>C.</given-names></name> <name><surname>Pennartz</surname> <given-names>C. M. A.</given-names></name></person-group> (<year>2014</year>). <article-title>The principles of goal-directed decision-making: from neural mechanisms to computation and robotics</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>369</volume>, <fpage>20130470</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2013.0470</pub-id><pub-id pub-id-type="pmid">25267813</pub-id></citation>
</ref>
<ref id="B347">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pfister</surname> <given-names>J.-P.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2006</year>). <article-title>Triplets of spikes in a model of spike timing-dependent plasticity</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>9673</fpage>&#x02013;<lpage>9682</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1425-06.2006</pub-id><pub-id pub-id-type="pmid">16988038</pub-id></citation>
</ref>
<ref id="B348">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Phillips</surname> <given-names>A. T.</given-names></name> <name><surname>Wellman</surname> <given-names>H. M.</given-names></name> <name><surname>Spelke</surname> <given-names>E. S.</given-names></name></person-group> (<year>2002</year>). <article-title>Infants&#x00027; ability to connect gaze and emotional expression to intentional action</article-title>. <source>Cognition</source> <volume>85</volume>, <fpage>53</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1016/S0010-0277(02)00073-2</pub-id><pub-id pub-id-type="pmid">12086713</pub-id></citation>
</ref>
<ref id="B349">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Piekniewski</surname> <given-names>F.</given-names></name> <name><surname>Laurent</surname> <given-names>P.</given-names></name> <name><surname>Petre</surname> <given-names>C.</given-names></name> <name><surname>Richert</surname> <given-names>M.</given-names></name> <name><surname>Fisher</surname> <given-names>D.</given-names></name> <name><surname>Hylton</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>Unsupervised learning from continuous video in a scalable predictive recurrent network</source>. arXiv:1607.06854.</citation>
</ref>
<ref id="B350">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pineda</surname> <given-names>F. J.</given-names></name></person-group> (<year>1987</year>). <article-title>Generalization of back-propagation to recurrent neural networks</article-title>. <source>Phys. Rev. Lett.</source> <volume>59</volume>:<fpage>2229</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.59.2229</pub-id><pub-id pub-id-type="pmid">10035458</pub-id></citation>
</ref>
<ref id="B351">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pinker</surname> <given-names>S.</given-names></name></person-group> (<year>1999</year>). <article-title>How the mind works</article-title>. <source>Ann. N.Y. Acad. Sci</source>. <volume>882</volume>, <fpage>119</fpage>&#x02013;<lpage>127</lpage>. <pub-id pub-id-type="pmid">10415890</pub-id></citation>
</ref>
<ref id="B352">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plate</surname> <given-names>T. A.</given-names></name></person-group> (<year>1995</year>). <article-title>Holographic reduced representations</article-title>. <source>IEEE Trans. Neural Netw.</source> <volume>6</volume>, <fpage>623</fpage>&#x02013;<lpage>641</lpage>. <pub-id pub-id-type="pmid">18263348</pub-id></citation>
</ref>
<ref id="B353">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <source>What if&#x02026;</source>, <publisher-name>MIT Center for Brains Minds and Machines Memo</publisher-name>.</citation>
</ref>
<ref id="B354">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poggio</surname> <given-names>T.</given-names></name> <name><surname>Bizzi</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Generalization in vision and motor control</article-title>. <source>Nature</source> <volume>431</volume>, <fpage>768</fpage>&#x02013;<lpage>774</lpage>. <pub-id pub-id-type="doi">10.1038/nature03014</pub-id><pub-id pub-id-type="pmid">15483597</pub-id></citation>
</ref>
<ref id="B355">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ponulak</surname> <given-names>F.</given-names></name> <name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Rapid, parallel path planning by propagating wavefronts of spiking neural activity</article-title>. <source>Front. Comput. Neurosci.</source> <volume>7</volume>:<issue>98</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2013.00098</pub-id><pub-id pub-id-type="pmid">23882213</pub-id></citation>
</ref>
<ref id="B356">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Metz</surname> <given-names>L.</given-names></name> <name><surname>Chintala</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <source>Unsupervised representation learning with deep convolutional generative adversarial networks</source>. arXiv:1511.06434.</citation>
</ref>
<ref id="B357">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajan</surname> <given-names>K.</given-names></name> <name><surname>Harvey</surname> <given-names>C. D.</given-names></name> <name><surname>Tank</surname> <given-names>D. W.</given-names></name></person-group> (<year>2016</year>). <article-title>Recurrent network models of sequence generation and memory</article-title>. <source>Neuron</source> <volume>90</volume>, <fpage>128</fpage>&#x02013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.02.009</pub-id><pub-id pub-id-type="pmid">26971945</pub-id></citation>
</ref>
<ref id="B358">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Ramachandran</surname> <given-names>V. S.</given-names></name></person-group> (<year>2000</year>). <source>Mirror Neurons and Imitation Learning as the Driving Force Behind &#x0201C;the Great Leap Forward&#x0201D; in Human Evolution</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.edge.org/conversation/mirror-neurons-and-imitation-learning-as-the-driving-force-behind-the-great-leap-forward-in-human-evolution">https://www.edge.org/conversation/mirror-neurons-and-imitation-learning-as-the-driving-force-behind-the-great-leap-forward-in-human-evolution</ext-link></citation>
</ref>
<ref id="B359">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>R. P.</given-names></name></person-group> (<year>2004</year>). <article-title>Bayesian computation in recurrent neural circuits</article-title>. <source>Neural Comput.</source> <volume>16</volume>, <fpage>1</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1162/08997660460733976</pub-id><pub-id pub-id-type="pmid">15006021</pub-id></citation>
</ref>
<ref id="B360">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rashevsky</surname> <given-names>N.</given-names></name></person-group> (<year>1939</year>). <article-title>Mathematical biophysics: physico-mathematical foundations of biology</article-title>. <source>Bull. Amer. Math. Soc.</source> <volume>45</volume>, <fpage>223</fpage>&#x02013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1090/S0002-9904-1939-06963-2</pub-id></citation>
</ref>
<ref id="B361">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Rasmus</surname> <given-names>A.</given-names></name> <name><surname>Berglund</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <source>Semi-supervised learning with ladder networks</source>. arXiv:1507.02672.</citation>
</ref>
<ref id="B362">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reynolds</surname> <given-names>J. H.</given-names></name> <name><surname>Desimone</surname> <given-names>R.</given-names></name></person-group> (<year>1999</year>). <article-title>The role of neural mechanisms of attention in solving the binding problem</article-title>. <source>Neuron</source> <volume>24</volume>, <fpage>19</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="pmid">10677024</pub-id></citation>
</ref>
<ref id="B363">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Rezende</surname> <given-names>D. J.</given-names></name> <name><surname>Mohamed</surname> <given-names>S.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name> <name><surname>Gregor</surname> <given-names>K.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <source>One-shot generalization in deep generative models</source>. arXiv:1603.05106.</citation>
</ref>
<ref id="B364">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rigotti</surname> <given-names>M.</given-names></name> <name><surname>Barak</surname> <given-names>O.</given-names></name> <name><surname>Warden</surname> <given-names>M. R.</given-names></name> <name><surname>Wang</surname> <given-names>X.-J.</given-names></name> <name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>The importance of mixed selectivity in complex cognitive tasks</article-title>. <source>Nature</source> <volume>497</volume>, <fpage>585</fpage>&#x02013;<lpage>590</lpage>. <pub-id pub-id-type="doi">10.1038/nature12160</pub-id><pub-id pub-id-type="pmid">23685452</pub-id></citation>
</ref>
<ref id="B365">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robinson</surname> <given-names>D.</given-names></name></person-group> (<year>1992</year>). <article-title>Implications of neural networks for how we think about brain function</article-title>. <source>Behav. Brain Sci</source>. <volume>15</volume>, <fpage>644</fpage>&#x02013;<lpage>655</lpage>.</citation>
</ref>
<ref id="B366">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>A.</given-names></name> <name><surname>Granger</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>The grammar of mammalian brain capacity</article-title>. <source>Theor. Comput. Sci.</source> <volume>633</volume>, <fpage>100</fpage>&#x02013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1016/j.tcs.2016.03.021</pub-id></citation>
</ref>
<ref id="B367">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>A.</given-names></name> <name><surname>Whitson</surname> <given-names>J.</given-names></name> <name><surname>Granger</surname> <given-names>R.</given-names></name></person-group> (<year>2004</year>). <article-title>Derivation and analysis of basic computational operations of thalamocortical circuits</article-title>. <source>J. Cogn. Neurosci.</source> <volume>16</volume>, <fpage>856</fpage>&#x02013;<lpage>877</lpage>. <pub-id pub-id-type="doi">10.1162/089892904970690</pub-id><pub-id pub-id-type="pmid">15200713</pub-id></citation>
</ref>
<ref id="B368">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name> <name><surname>van Ooyen</surname> <given-names>A.</given-names></name> <name><surname>Watanabe</surname> <given-names>T.</given-names></name></person-group> (<year>2010</year>). <article-title>Perceptual learning rules based on reinforcers and attention</article-title>. <source>Trends Cogn. Sci</source>. <volume>14</volume>, <fpage>64</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2009.11.005</pub-id><pub-id pub-id-type="pmid">20060771</pub-id></citation>
</ref>
<ref id="B369">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name> <name><surname>van Ooyen</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Attention-gated reinforcement learning of internal representations for classification</article-title>. <source>Neural Comput.</source> <volume>17</volume>, <fpage>2176</fpage>&#x02013;<lpage>2214</lpage>. <pub-id pub-id-type="doi">10.1162/0899766054615699</pub-id><pub-id pub-id-type="pmid">16105222</pub-id></citation>
</ref>
<ref id="B370">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rolls</surname> <given-names>E. T.</given-names></name></person-group> (<year>2013</year>). <article-title>The mechanisms for pattern completion and pattern separation in the hippocampus</article-title>. <source>Front. Syst. Neurosci.</source> <volume>7</volume>:<issue>74</issue>. <pub-id pub-id-type="doi">10.3389/fnsys.2013.00074</pub-id><pub-id pub-id-type="pmid">24198767</pub-id></citation>
</ref>
<ref id="B371">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rombouts</surname> <given-names>J. O.</given-names></name> <name><surname>Bohte</surname> <given-names>S. M.</given-names></name> <name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name></person-group> (<year>2015</year>). <article-title>How attention can create synaptic tags for the learning of working memories in sequential tasks</article-title>. <source>PLoS Comput. Biol.</source> <volume>11</volume>:<fpage>e1004060</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004060</pub-id><pub-id pub-id-type="pmid">25742003</pub-id></citation>
</ref>
<ref id="B372">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero</surname> <given-names>A.</given-names></name> <name><surname>Ballas</surname> <given-names>N.</given-names></name> <name><surname>Kahou</surname> <given-names>S. E.</given-names></name> <name><surname>Chassang</surname> <given-names>A.</given-names></name> <name><surname>Gatta</surname> <given-names>C.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Fitnets: hints for thin deep nets</article-title>. <source>arXiv</source> arXiv:1412.6550.</citation>
</ref>
<ref id="B373">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roudi</surname> <given-names>Y.</given-names></name> <name><surname>Taylor</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning with hidden variables</article-title>. <source>Curr. Opin. Neurobiol</source>. <volume>35</volume>, <fpage>110</fpage>&#x02013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2015.07.006</pub-id><pub-id pub-id-type="pmid">26298193</pub-id></citation>
</ref>
<ref id="B374">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rozell</surname> <given-names>C. J.</given-names></name> <name><surname>Johnson</surname> <given-names>D. H.</given-names></name> <name><surname>Baraniuk</surname> <given-names>R. G.</given-names></name> <name><surname>Olshausen</surname> <given-names>B. A.</given-names></name></person-group> (<year>2008</year>). <article-title>Sparse coding via thresholding and local competition in neural circuits</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>2526</fpage>&#x02013;<lpage>2563</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.03-07-486</pub-id><pub-id pub-id-type="pmid">18439138</pub-id></citation>
</ref>
<ref id="B375">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rubin</surname> <given-names>A.</given-names></name> <name><surname>Geva</surname> <given-names>N.</given-names></name> <name><surname>Sheintuch</surname> <given-names>L.</given-names></name> <name><surname>Ziv</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Hippocampal ensemble dynamics timestamp events in long-term memory</article-title>. <source>eLife</source> <fpage>4:e12247</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.12247</pub-id><pub-id pub-id-type="pmid">26682652</pub-id></citation>
</ref>
<ref id="B376">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Williams</surname> <given-names>R. J.</given-names></name></person-group> (<year>1986</year>). <article-title>Learning representations by back-propagating errors</article-title>. <source>Nature</source> <volume>323</volume>, <fpage>533</fpage>&#x02013;<lpage>536</lpage>. <pub-id pub-id-type="doi">10.1038/323533a0</pub-id></citation>
</ref>
<ref id="B377">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>Zipser</surname> <given-names>D.</given-names></name></person-group> (<year>1986</year>). <article-title>Feature discovery by competitive learning</article-title>, in <source>Parallel Distributed Processing</source>, <volume>Vol. 1</volume>, eds <person-group person-group-type="editor"><name><surname>Rumel hart</surname> <given-names>D. E</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>151</fpage>&#x02013;<lpage>163</lpage>.</citation>
</ref>
<ref id="B378">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sadtler</surname> <given-names>P. T.</given-names></name> <name><surname>Quick</surname> <given-names>K. M.</given-names></name> <name><surname>Golub</surname> <given-names>M. D.</given-names></name> <name><surname>Chase</surname> <given-names>S. M.</given-names></name> <name><surname>Ryu</surname> <given-names>S. I.</given-names></name> <name><surname>Tyler-Kabara</surname> <given-names>E. C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Neural constraints on learning</article-title>. <source>Nature</source> <volume>512</volume>, <fpage>423</fpage>&#x02013;<lpage>426</lpage>. <pub-id pub-id-type="doi">10.1038/nature13665</pub-id><pub-id pub-id-type="pmid">25164754</pub-id></citation>
</ref>
<ref id="B379">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahani</surname> <given-names>M.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>2003</year>). <article-title>Doubly distributional population codes: simultaneous representation of uncertainty and multiplicity</article-title>. <source>Neural Comput.</source> <volume>15</volume>, <fpage>2255</fpage>&#x02013;<lpage>2279</lpage>. <pub-id pub-id-type="doi">10.1162/089976603322362356</pub-id><pub-id pub-id-type="pmid">14511521</pub-id></citation>
</ref>
<ref id="B380">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sandler</surname> <given-names>M.</given-names></name> <name><surname>Shulman</surname> <given-names>Y.</given-names></name> <name><surname>Schiller</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>A novel form of local plasticity in tuft dendrites of neocortical somatosensory layer 5 pyramidal neurons</article-title>. <source>Neuron</source> <volume>90</volume>, <fpage>1028</fpage>&#x02013;<lpage>1042</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2016.04.032</pub-id><pub-id pub-id-type="pmid">27210551</pub-id></citation>
</ref>
<ref id="B381">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Santoro</surname> <given-names>A.</given-names></name> <name><surname>Bartunov</surname> <given-names>S.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name> <name><surname>Lillicrap</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>One-shot learning with memory-augmented neural networks</source>. <volume>13</volume>. <fpage>arXiv:1605.06065</fpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1605.06065">https://arxiv.org/abs/1605.06065</ext-link></citation>
</ref>
<ref id="B382">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Saxe</surname> <given-names>A. M.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <source>Exact solutions to the nonlinear dynamics of learning in deep linear neural networks</source>. arXiv:1312.6120.</citation>
</ref>
<ref id="B383">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Scellier</surname> <given-names>B.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <source>Towards a biologically plausible backprop</source>. arXiv:1602.05179.</citation>
</ref>
<ref id="B384">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schiess</surname> <given-names>M.</given-names></name> <name><surname>Urbanczik</surname> <given-names>R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Somato-dendritic synaptic plasticity and error-backpropagation in active dendrites</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1004638</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004638</pub-id><pub-id pub-id-type="pmid">26841235</pub-id></citation>
</ref>
<ref id="B385">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Formal theory of creativity, fun, and intrinsic motivation (19902010)</article-title>. <source>Auton. Ment. Dev. IEEE</source>. <volume>2</volume>, <fpage>230</fpage>&#x02013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1109/TAMD.2010.2056368</pub-id></citation>
</ref>
<ref id="B386">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning in neural networks: an overview</article-title>. <source>Neural Netw.</source> <volume>61</volume>, <fpage>85</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2014.09.003</pub-id><pub-id pub-id-type="pmid">25462637</pub-id></citation>
</ref>
<ref id="B387">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scholl</surname> <given-names>B. J.</given-names></name></person-group> (<year>2004</year>). <article-title>Can infants&#x00027; object concepts be trained?</article-title> <source>Trends Cogn. Sci.</source> <volume>8</volume>, <fpage>49</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2003.12.006</pub-id><pub-id pub-id-type="pmid">15588805</pub-id></citation>
</ref>
<ref id="B388">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwabe</surname> <given-names>L.</given-names></name> <name><surname>Obermayer</surname> <given-names>K.</given-names></name> <name><surname>Angelucci</surname> <given-names>A.</given-names></name> <name><surname>Bressloff</surname> <given-names>P. C.</given-names></name></person-group> (<year>2006</year>). <article-title>The role of feedback in shaping the extra-classical receptive field of cortical neurons: a recurrent network model</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>9117</fpage>&#x02013;<lpage>9129</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1253-06.2006</pub-id><pub-id pub-id-type="pmid">16957068</pub-id></citation>
</ref>
<ref id="B389">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sejnowski</surname> <given-names>T.</given-names></name> <name><surname>Poizner</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Prospective optimization</article-title>. <source>Proc. IEEE</source>. <volume>102</volume>, <fpage>799</fpage>&#x02013;<lpage>811</lpage>. <pub-id pub-id-type="doi">10.1109/jproc.2014.2314297</pub-id><pub-id pub-id-type="pmid">25328167</pub-id></citation>
</ref>
<ref id="B390">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2013</year>). <article-title>Pedestrian detection with unsupervised multi-stage feature learning</article-title>,&#x0201D; in <source>Proceedings of Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Portland, OR</publisher-loc>).</citation>
</ref>
<ref id="B391">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Serre</surname> <given-names>T.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2007</year>). <article-title>A feedforward architecture accounts for rapid categorization</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>104</volume>, <fpage>6424</fpage>&#x02013;<lpage>6429</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0700622104</pub-id><pub-id pub-id-type="pmid">17404214</pub-id></citation>
</ref>
<ref id="B392">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Servan-Schreiber</surname> <given-names>E.</given-names></name> <name><surname>Anderson</surname> <given-names>J.</given-names></name></person-group> (<year>1990</year>). <article-title>Chunking as a mechanism of implicit learning</article-title>. <source>J. Exp. Psychol.</source> <volume>16</volume>, <fpage>592</fpage>&#x02013;<lpage>608</lpage>.</citation>
</ref>
<ref id="B393">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>1998</year>). <article-title>Continuous attractors and oculomotor control</article-title>. <source>Neural Netw.</source> <volume>11</volume>, <fpage>1253</fpage>&#x02013;<lpage>1258</lpage>. <pub-id pub-id-type="doi">10.1016/S0893-6080(98)00064-1</pub-id><pub-id pub-id-type="pmid">12662748</pub-id></citation>
</ref>
<ref id="B394">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2003</year>). <article-title>Learning in spiking neural networks by reinforcement of stochastic synaptic transmission</article-title>. <source>Neuron</source> <volume>40</volume>, <fpage>1063</fpage>&#x02013;<lpage>1073</lpage>. <pub-id pub-id-type="doi">10.1016/S0896-6273(03)00761-X</pub-id><pub-id pub-id-type="pmid">14687542</pub-id></citation>
</ref>
<ref id="B395">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shai</surname> <given-names>A. S.</given-names></name> <name><surname>Anastassiou</surname> <given-names>C. A.</given-names></name> <name><surname>Larkum</surname> <given-names>M. E.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Physiology of layer 5 pyramidal neurons in mouse primary visual cortex: coincidence detection through bursting</article-title>. <source>PLoS Comput. Biol.</source> <volume>11</volume>:<fpage>e1004090</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004090</pub-id><pub-id pub-id-type="pmid">25768881</pub-id></citation>
</ref>
<ref id="B396">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shepherd</surname> <given-names>G. M.</given-names></name></person-group> (<year>2014</year>). <article-title>The microcircuit concept applied to cortical evolution: from three-layer to six-layer cortex</article-title>. <source>Front. Neuroanat.</source> <volume>5</volume>:<issue>30</issue>. <pub-id pub-id-type="doi">10.3389/fnana.2011.00030</pub-id><pub-id pub-id-type="pmid">21647397</pub-id></citation>
</ref>
<ref id="B397">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sherman</surname> <given-names>S. M.</given-names></name></person-group> (<year>2005</year>). <article-title>Thalamic relays and cortical functioning</article-title>. <source>Prog. Brain Res.</source> <volume>149</volume>, <fpage>107</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1016/S0079-6123(05)49009-3</pub-id><pub-id pub-id-type="pmid">16226580</pub-id></citation>
</ref>
<ref id="B398">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sherman</surname> <given-names>S. M.</given-names></name></person-group> (<year>2007</year>). <article-title>The thalamus is more than just a relay</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>17</volume>, <fpage>417</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2007.07.003</pub-id><pub-id pub-id-type="pmid">17707635</pub-id></citation>
</ref>
<ref id="B399">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shimizu</surname> <given-names>T.</given-names></name> <name><surname>Karten</surname> <given-names>H. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Multiple origins of neocortex: contributions of the dorsal</article-title>. <source>Neocortex</source> <volume>200</volume>:<fpage>75</fpage>. <pub-id pub-id-type="doi">10.1007/978-1-4899-0652-6_8</pub-id></citation>
</ref>
<ref id="B400">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siegel</surname> <given-names>M.</given-names></name> <name><surname>Warden</surname> <given-names>M. R.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2009</year>). <article-title>Phase-dependent neuronal coding of objects in short-term memory</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>106</volume>, <fpage>21341</fpage>&#x02013;<lpage>21346</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0908193106</pub-id><pub-id pub-id-type="pmid">19926847</pub-id></citation>
</ref>
<ref id="B401">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>R.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2006</year>). <article-title>Higher-dimensional neurons explain the tuning and dynamics of working memory cells</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>3667</fpage>&#x02013;<lpage>3678</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4864-05.2006</pub-id><pub-id pub-id-type="pmid">16597721</pub-id></citation>
</ref>
<ref id="B402">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sj&#x000F6;str&#x000F6;m</surname> <given-names>J.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name></person-group> (<year>2010</year>). <article-title>Spike-timing dependent plasticity</article-title>. <source>Scholarpedia</source> <volume>5</volume>:<fpage>1362</fpage>. <pub-id pub-id-type="doi">10.4249/scholarpedia.1362</pub-id><pub-id pub-id-type="pmid">26374672</pub-id></citation>
</ref>
<ref id="B403">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sj&#x000F6;str&#x000F6;m</surname> <given-names>P. J.</given-names></name> <name><surname>H&#x000E4;usser</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>A cooperative switch determines the sign of synaptic plasticity in distal dendrites of neocortical pyramidal neurons</article-title>. <source>Neuron</source> <volume>51</volume>, <fpage>227</fpage>&#x02013;<lpage>238</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2006.06.017</pub-id><pub-id pub-id-type="pmid">16846857</pub-id></citation>
</ref>
<ref id="B404">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skerry</surname> <given-names>A. E.</given-names></name> <name><surname>Spelke</surname> <given-names>E. S.</given-names></name></person-group> (<year>2014</year>). <article-title>Preverbal infants identify emotional reactions that are incongruent with goal outcomes</article-title>. <source>Cognition</source> <volume>130</volume>, <fpage>204</fpage>&#x02013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2013.11.002</pub-id><pub-id pub-id-type="pmid">24321623</pub-id></citation>
</ref>
<ref id="B405">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Softky</surname> <given-names>W.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>1993</year>). <article-title>The highly irregular firing of cortical cells is inconsistent with temporal integration of random EPSPs</article-title>. <source>J. Neurosci</source>. <volume>13</volume>, <fpage>334</fpage>&#x02013;<lpage>350</lpage>. <pub-id pub-id-type="pmid">8423479</pub-id></citation>
</ref>
<ref id="B406">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Solari</surname> <given-names>S. V. H.</given-names></name> <name><surname>Stoner</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Cognitive consilience: primate non-primary neuroanatomical circuits underlying cognition</article-title>. <source>Front. Neuroanat.</source> <volume>5</volume>:<issue>65</issue>. <pub-id pub-id-type="doi">10.3389/fnana.2011.00065</pub-id><pub-id pub-id-type="pmid">22194717</pub-id></citation>
</ref>
<ref id="B407">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sountsov</surname> <given-names>P.</given-names></name> <name><surname>Miller</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Spiking neuron network Helmholtz machine</article-title>. <source>Front. Comput. Neurosci.</source> <volume>9</volume>:<issue>46</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2015.00046</pub-id><pub-id pub-id-type="pmid">25954191</pub-id></citation>
</ref>
<ref id="B408">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Squire</surname> <given-names>L. R.</given-names></name></person-group> (<year>2004</year>). <article-title>Memory systems of the brain: a brief history and current perspective</article-title>. <source>Neurobiol. Learn. Memory</source> <volume>82</volume>, <fpage>171</fpage>&#x02013;<lpage>177</lpage>. <pub-id pub-id-type="doi">10.1016/j.nlm.2004.06.005</pub-id><pub-id pub-id-type="pmid">15464402</pub-id></citation>
</ref>
<ref id="B409">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Srivastava</surname> <given-names>N.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>. <source>J. Mach. Learn. Res.</source> <volume>15</volume>, <fpage>1929</fpage>&#x02013;<lpage>1958</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.jmlr.org/papers/v15/srivastava14a.html">http://www.jmlr.org/papers/v15/srivastava14a.html</ext-link></citation>
</ref>
<ref id="B410">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stachenfeld</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <article-title>Design principles of the hippocampal cognitive map</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Montreal, QC</publisher-loc>).</citation>
</ref>
<ref id="B411">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stanisor</surname> <given-names>L.</given-names></name> <name><surname>van der Togt</surname> <given-names>C.</given-names></name> <name><surname>Pennartz</surname> <given-names>C. M. A.</given-names></name> <name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name></person-group> (<year>2013</year>). <article-title>A unified selection signal for attention and reward in primary visual cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>110</volume>, <fpage>9136</fpage>&#x02013;<lpage>9141</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1300117110</pub-id><pub-id pub-id-type="pmid">23676276</pub-id></citation>
</ref>
<ref id="B412">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stewart</surname> <given-names>T.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>Compositionality and biologically plausible models</article-title>, in <source>Oxford Handbook of Compositionality</source>, eds <person-group person-group-type="editor"><name><surname>Hinzen</surname> <given-names>W.</given-names></name> <name><surname>Machery</surname> <given-names>E.</given-names></name> <name><surname>Werning</surname> <given-names>M.</given-names></name></person-group> (<publisher-name>Oxford University Press</publisher-name>). <pub-id pub-id-type="doi">10.1093/oxfordhb/9780199541072.013.0029</pub-id></citation>
</ref>
<ref id="B413">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stocco</surname> <given-names>A.</given-names></name> <name><surname>Lebiere</surname> <given-names>C.</given-names></name> <name><surname>Anderson</surname> <given-names>J. R.</given-names></name></person-group> (<year>2010</year>). <article-title>Conditional routing of information to the cortex: a model of the basal ganglia&#x00027;s role in cognitive coordination</article-title>. <source>Psychol. Rev.</source> <volume>117</volume>, <fpage>541</fpage>&#x02013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1037/a0019077</pub-id><pub-id pub-id-type="pmid">20438237</pub-id></citation>
</ref>
<ref id="B414">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stork</surname> <given-names>D. G.</given-names></name></person-group> (<year>1989</year>). <article-title>Is backpropagation biologically plausible?</article-title>, in <source>International Joint Conference on Neural Networks</source>, <volume>Vol. 2</volume> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>241</fpage>&#x02013;<lpage>246</lpage>.</citation>
</ref>
<ref id="B415">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Strausfeld</surname> <given-names>N. J.</given-names></name> <name><surname>Hirth</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>Deep homology of arthropod central complex and vertebrate basal ganglia</article-title>. <source>Science</source> (<publisher-loc>New York, N.Y.</publisher-loc>) <volume>340</volume>, <fpage>157</fpage>&#x02013;<lpage>161</lpage>. <pub-id pub-id-type="doi">10.1126/science.1231828</pub-id></citation>
</ref>
<ref id="B416">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sukhbaatar</surname> <given-names>S.</given-names></name> <name><surname>Bruna</surname> <given-names>J.</given-names></name> <name><surname>Paluri</surname> <given-names>M.</given-names></name> <name><surname>Bourdev</surname> <given-names>L.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Training convolutional networks with noisy labels</article-title>. <source>arXiv preprint arXiv:1406.2080</source>.</citation>
</ref>
<ref id="B417">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Gomez</surname> <given-names>F.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Planning to be surprised: optimal bayesian exploration in dynamic environments</article-title>, in <source>Artificial General Intelligence</source> (<publisher-loc>Mountain View, CA</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>41</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-22887-2_5</pub-id></citation>
</ref>
<ref id="B418">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussillo</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural circuits as computational dynamical systems</article-title>. <source>Curr. Opin. Neurobiol</source>. <volume>25</volume>, <fpage>156</fpage>&#x02013;<lpage>163</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2014.01.008</pub-id><pub-id pub-id-type="pmid">24509098</pub-id></citation>
</ref>
<ref id="B419">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussillo</surname> <given-names>D.</given-names></name> <name><surname>Abbott</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>Generating coherent patterns of activity from chaotic neural networks</article-title>. <source>Neuron</source> <volume>63</volume>, <fpage>544</fpage>&#x02013;<lpage>557</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2009.07.018</pub-id><pub-id pub-id-type="pmid">19709635</pub-id></citation>
</ref>
<ref id="B420">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sussillo</surname> <given-names>D.</given-names></name> <name><surname>Churchland</surname> <given-names>M. M.</given-names></name> <name><surname>Kaufman</surname> <given-names>M. T.</given-names></name> <name><surname>Shenoy</surname> <given-names>K. V.</given-names></name></person-group> (<year>2015</year>). <article-title>A neural network that finds a naturalistic solution for the production of muscle activity</article-title>. <source>Nat. Neurosci.</source> <volume>18</volume>, <fpage>1025</fpage>&#x02013;<lpage>1033</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4042</pub-id><pub-id pub-id-type="pmid">26075643</pub-id></citation>
</ref>
<ref id="B421">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Martens</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>On the importance of initialization and momentum in deep learning</article-title>,&#x0201D; in <source>Proceedings of the 30 th International Conference on Machine Learning</source> (<publisher-loc>Atlanta</publisher-loc>: <publisher-name>JMLR:W&#x00026;CP</publisher-name>).</citation>
</ref>
<ref id="B422">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Martens</surname> <given-names>J.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2011</year>). <article-title>Generating text with recurrent neural networks</article-title>, in <source>Proceedings of the 28th International Conference on Machine Learning (ICML-11)</source> (<publisher-loc>Bellevue</publisher-loc>), <fpage>1017</fpage>&#x02013;<lpage>1024</lpage>.</citation>
</ref>
<ref id="B423">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Barto</surname> <given-names>A. G.</given-names></name></person-group> (<year>1998</year>). <source>Reinforcement Learning: An Introduction</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B424">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tacchetti</surname> <given-names>A.</given-names></name> <name><surname>Isik</surname> <given-names>L.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Spatio-temporal convolutional neural networks explain human neural representations of action recognition</article-title>. <source>arXiv preprint arXiv:1606.04698</source>.</citation>
</ref>
<ref id="B425">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Takata</surname> <given-names>N.</given-names></name> <name><surname>Mishima</surname> <given-names>T.</given-names></name> <name><surname>Hisatsune</surname> <given-names>C.</given-names></name> <name><surname>Nagai</surname> <given-names>T.</given-names></name> <name><surname>Ebisui</surname> <given-names>E.</given-names></name> <name><surname>Mikoshiba</surname> <given-names>K.</given-names></name> <name><surname>Hirase</surname> <given-names>H.</given-names></name></person-group> (<year>2011</year>). <article-title>Astrocyte calcium signaling transforms cholinergic modulation to cortical plasticity <italic>in vivo</italic></article-title>. <source>J. Neurosci.</source> <volume>31</volume>, <fpage>18155</fpage>&#x02013;<lpage>18165</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5289-11.2011</pub-id><pub-id pub-id-type="pmid">22159127</pub-id></citation>
</ref>
<ref id="B426">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tamar</surname> <given-names>A.</given-names></name> <name><surname>Levine</surname> <given-names>S.</given-names></name> <name><surname>Abbeel</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Value iteration networks</article-title>. <source>arXiv preprint arXiv:1602.02867</source>.</citation>
</ref>
<ref id="B427">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>Y.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <source>Deep mixtures of factor analysers</source>. arXiv:1206.4635</citation>
</ref>
<ref id="B428">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>Y.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>Tensor analyzers</article-title>, in <source>Proceedings of the 30th International Conference on Machine Learning (ICML-13)</source> (<publisher-loc>Atlanta, GA</publisher-loc>).</citation>
</ref>
<ref id="B429">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tapson</surname> <given-names>J.</given-names></name> <name><surname>van Schaik</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Learning the pseudoinverse solution to network weights</article-title>. <source>Neural Netw.</source> <volume>45</volume>, <fpage>94</fpage>&#x02013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2013.02.008</pub-id><pub-id pub-id-type="pmid">23541926</pub-id></citation>
</ref>
<ref id="B430">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tavares</surname> <given-names>R. M.</given-names></name> <name><surname>Mendelsohn</surname> <given-names>A.</given-names></name> <name><surname>Grossman</surname> <given-names>Y.</given-names></name> <name><surname>Williams</surname> <given-names>C. H.</given-names></name> <name><surname>Shapiro</surname> <given-names>M.</given-names></name> <name><surname>Trope</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>A map for social navigation in the human brain</article-title>. <source>Neuron</source> <volume>87</volume>, <fpage>231</fpage>&#x02013;<lpage>243</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.06.011</pub-id><pub-id pub-id-type="pmid">26139376</pub-id></citation>
</ref>
<ref id="B431">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taylor</surname> <given-names>S. V.</given-names></name> <name><surname>Faisal</surname> <given-names>A. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Does the cost function of human motor control depend on the internal metabolic state?</article-title> <source>BMC Neurosci.</source> <volume>12</volume>(<supplement>Suppl. 1</supplement>):<fpage>P99</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2202-12-S1-P99</pub-id></citation>
</ref>
<ref id="B432">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Terrence Stewart</surname> <given-names>C. E.</given-names></name> <name><surname>Choo</surname> <given-names>X.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Symbolic reasoning in spiking neurons: a model of the cortex/basal ganglia/thalamus loop</article-title>, in <source>32nd Annual Meeting of the Cognitive Science Society</source> (<publisher-loc>Portland, OR</publisher-loc>).</citation>
</ref>
<ref id="B433">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tervo</surname> <given-names>D. G. R.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Gershman</surname> <given-names>S. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Toward the neural implementation of structure learning</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>37</volume>, <fpage>99</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2016.01.014</pub-id><pub-id pub-id-type="pmid">26874471</pub-id></citation>
</ref>
<ref id="B434">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tesauro</surname> <given-names>G.</given-names></name></person-group> (<year>1995</year>). <article-title>Temporal difference learning and TD-Gammon</article-title>. <source>Commun. ACM</source>. <volume>38</volume>, <fpage>58</fpage>&#x02013;<lpage>68</lpage>.</citation>
</ref>
<ref id="B435">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Thalmeier</surname> <given-names>D.</given-names></name> <name><surname>Uhlmann</surname> <given-names>M.</given-names></name> <name><surname>Kappen</surname> <given-names>H. J.</given-names></name> <name><surname>Memmesheimer</surname> <given-names>R.-M.</given-names></name></person-group> (<year>2015</year>). <source>Learning universal computations with spikes</source>. arXiv:1505.07866.</citation>
</ref>
<ref id="B436">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tinbergen</surname> <given-names>N.</given-names></name></person-group> (<year>1965</year>). <article-title>Behavior and natural selection</article-title>, in <source>Ideas in Modern Biology: proceedings of the 16th International Zoological Congress</source>, ed <person-group person-group-type="editor"><name><surname>Moore</surname> <given-names>J. A.</given-names></name></person-group> (<publisher-loc>Washington, DC</publisher-loc>), <fpage>521</fpage>&#x02013;<lpage>542</lpage>.</citation>
</ref>
<ref id="B437">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name></person-group> (<year>2002</year>). <article-title>Cosine tuning minimizes motor errors</article-title>. <source>Neural Comput.</source> <volume>14</volume>, <fpage>1233</fpage>&#x02013;<lpage>1260</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2016.01.014</pub-id><pub-id pub-id-type="pmid">12020444</pub-id></citation>
</ref>
<ref id="B438">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name></person-group> (<year>2009</year>). <article-title>Efficient computation of optimal actions</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>106</volume>, <fpage>11478</fpage>&#x02013;<lpage>11483</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0710743106</pub-id><pub-id pub-id-type="pmid">19574462</pub-id></citation>
</ref>
<ref id="B439">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Todorov</surname> <given-names>E.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2002</year>). <article-title>Optimal feedback control as a theory of motor coordination</article-title>. <source>Nat. Neurosci.</source> <volume>5</volume>, <fpage>1226</fpage>&#x02013;<lpage>1235</lpage>. <pub-id pub-id-type="doi">10.1038/nn963</pub-id><pub-id pub-id-type="pmid">12404008</pub-id></citation>
</ref>
<ref id="B440">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tripp</surname> <given-names>B.</given-names></name> <name><surname>Eliasmith</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>Function approximation in inhibitory networks</article-title>. <source>Neural Netw.</source> <volume>77</volume>, <fpage>95</fpage>&#x02013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2016.01.010</pub-id><pub-id pub-id-type="pmid">26963256</pub-id></citation>
</ref>
<ref id="B441">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turner</surname> <given-names>R. S.</given-names></name> <name><surname>Desmurget</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Basal ganglia contributions to motor control: a vigorous tutor</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>20</volume>, <fpage>704</fpage>&#x02013;<lpage>716</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2010.08.022</pub-id><pub-id pub-id-type="pmid">20850966</pub-id></citation>
</ref>
<ref id="B442">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turrigiano</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>Homeostatic synaptic plasticity: local and global mechanisms for stabilizing neuronal function</article-title>. <source>Cold Spring Harb. Perspect. Biol.</source> <volume>4</volume>:<fpage>a005736</fpage>. <pub-id pub-id-type="doi">10.1101/cshperspect.a005736</pub-id><pub-id pub-id-type="pmid">22086977</pub-id></citation>
</ref>
<ref id="B443">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ullman</surname> <given-names>S.</given-names></name> <name><surname>Harari</surname> <given-names>D.</given-names></name> <name><surname>Dorfman</surname> <given-names>N.</given-names></name></person-group> (<year>2012</year>). <article-title>From simple innate biases to complex visual concepts</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>109</volume>, <fpage>18215</fpage>&#x02013;<lpage>18220</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1207690109</pub-id><pub-id pub-id-type="pmid">23012418</pub-id></citation>
</ref>
<ref id="B444">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Urbanczik</surname> <given-names>R.</given-names></name> <name><surname>Senn</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <article-title>Learning by the dendritic prediction of somatic spiking</article-title>. <source>Neuron</source> <volume>81</volume>, <fpage>521</fpage>&#x02013;<lpage>528</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2013.11.030</pub-id><pub-id pub-id-type="pmid">24507189</pub-id></citation>
</ref>
<ref id="B445">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Valpola</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <source>From neural PCA to deep unsupervised learning</source>. arXiv:1411.7783.</citation>
</ref>
<ref id="B446">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>van den Oord</surname> <given-names>A.</given-names></name> <name><surname>Kalchbrenner</surname> <given-names>N.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <source>Pixel recurrent neural networks</source>. arXiv:1601.06759.</citation>
</ref>
<ref id="B447">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Heijningen</surname> <given-names>C. A.</given-names></name> <name><surname>De Visser</surname> <given-names>J.</given-names></name> <name><surname>Zuidema</surname> <given-names>W.</given-names></name> <name><surname>Ten Cate</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>Simple rules can explain discrimination of putative recursive syntactic structures by a songbird species</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>106</volume>, <fpage>20538</fpage>&#x02013;<lpage>20543</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0908113106</pub-id><pub-id pub-id-type="pmid">19918074</pub-id></citation>
</ref>
<ref id="B448">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Kerkoerle</surname> <given-names>T.</given-names></name> <name><surname>Self</surname> <given-names>M. W.</given-names></name> <name><surname>Dagnino</surname> <given-names>B.</given-names></name> <name><surname>Gariel-Mathis</surname> <given-names>M.-A.</given-names></name> <name><surname>Poort</surname> <given-names>J.</given-names></name> <name><surname>Van Der Togt</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Alpha and gamma oscillations characterize feedback and feedforward processing in monkey visual cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>14332</fpage>&#x02013;<lpage>14341</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1402773111</pub-id><pub-id pub-id-type="pmid">25205811</pub-id></citation>
</ref>
<ref id="B449">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Veit</surname> <given-names>A.</given-names></name> <name><surname>Wilber</surname> <given-names>M.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <source>Residual networks are exponential ensembles of relatively shallow networks</source>. arXiv:1605.06431.</citation>
</ref>
<ref id="B450">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verney</surname> <given-names>C.</given-names></name> <name><surname>Baulac</surname> <given-names>M.</given-names></name> <name><surname>Berger</surname> <given-names>B.</given-names></name> <name><surname>Alvarez</surname> <given-names>C.</given-names></name> <name><surname>Vigny</surname> <given-names>A.</given-names></name> <name><surname>Helle</surname> <given-names>K.</given-names></name></person-group> (<year>1985</year>). <article-title>Morphological evidence for a dopaminergic terminal field in the hippocampal formation of young and adult rat</article-title>. <source>Neuroscience</source> <volume>14</volume>, <fpage>1039</fpage>&#x02013;<lpage>1052</lpage>. <pub-id pub-id-type="doi">10.1016/0306-4522(85)90275-1</pub-id><pub-id pub-id-type="pmid">2860616</pub-id></citation>
</ref>
<ref id="B451">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verwey</surname> <given-names>W. B.</given-names></name></person-group> (<year>1996</year>). <article-title>Buffer loading and chunking in sequential keypressing</article-title>. <source>J. Exp. Psychol.</source> <volume>22</volume>:<fpage>544</fpage>.</citation>
</ref>
<ref id="B452">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J. X.</given-names></name> <name><surname>Cohen</surname> <given-names>N. J.</given-names></name> <name><surname>Voss</surname> <given-names>J. L.</given-names></name></person-group> (<year>2015</year>). <article-title>Covert rapid action-memory simulation (CRAMS): a hypothesis of hippocampal-prefrontal interactions for adaptive behavior</article-title>. <source>Neurobiol. Learn. Memory</source> <volume>117</volume>, <fpage>22</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1016/j.nlm.2014.04.003</pub-id><pub-id pub-id-type="pmid">24752152</pub-id></citation>
</ref>
<ref id="B453">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yuille</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <source>Semantic part segmentation using compositional model combining shape and appearance</source>. arXiv:1412.6124.</citation>
</ref>
<ref id="B454">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.-J.</given-names></name></person-group> (<year>2012</year>). <article-title>The prefrontal cortex as a quintessential &#x0201C;cognitive-type&#x0201D; neural circuit</article-title>, in <source>Principles of Frontal Lobe Function</source>, Edited by <person-group person-group-type="editor"><name><surname>Stuss</surname> <given-names>D. T.</given-names></name> <name><surname>Knight</surname> <given-names>R. T.</given-names></name></person-group> (<publisher-name>Oxford University Press</publisher-name>), <fpage>226</fpage>&#x02013;<lpage>248</lpage>.</citation>
</ref>
<ref id="B455">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Warden</surname> <given-names>M. R.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2007</year>). <article-title>The representation of multiple objects in prefrontal neuronal delay activity</article-title>. <source>Cereb. Cortex</source> (<publisher-loc>New York, N.Y.</publisher-loc>: 1991) <volume>17</volume>(<supplement>Suppl. 1</supplement>):<fpage>i41</fpage>&#x02013;<lpage>i50</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhm070</pub-id></citation>
</ref>
<ref id="B456">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Warden</surname> <given-names>M. R.</given-names></name> <name><surname>Miller</surname> <given-names>E. K.</given-names></name></person-group> (<year>2010</year>). <article-title>Task-dependent changes in short-term memory in the prefrontal cortex</article-title>. <source>J. Neurosci.</source> <volume>30</volume>, <fpage>15801</fpage>&#x02013;<lpage>15810</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1569-10.2010</pub-id><pub-id pub-id-type="pmid">21106819</pub-id></citation>
</ref>
<ref id="B457">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Watter</surname> <given-names>M.</given-names></name> <name><surname>Springenberg</surname> <given-names>J.</given-names></name> <name><surname>Boedecker</surname> <given-names>J.</given-names></name> <name><surname>Riedmiller</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Embed to control: a locally linear latent dynamics model for control from raw images</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>2728</fpage>&#x02013;<lpage>2736</lpage>.</citation>
</ref>
<ref id="B458">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wayne</surname> <given-names>G.</given-names></name> <name><surname>Abbott</surname> <given-names>L. F.</given-names></name></person-group> (<year>2014</year>). <article-title>Hierarchical control using networks trained with higher-level forward models</article-title>. <source>Neural Comput.</source> <volume>26</volume>, <fpage>2163</fpage>&#x02013;<lpage>2193</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00639</pub-id><pub-id pub-id-type="pmid">25058706</pub-id></citation>
</ref>
<ref id="B459">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Werbos</surname> <given-names>P.</given-names></name></person-group> (<year>1974</year>). <source>Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences</source>. Doctoral Dissertation, Harvard University, Harvard.</citation>
</ref>
<ref id="B460">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Werbos</surname> <given-names>P.</given-names></name></person-group> (<year>1982</year>). <article-title>Applications of advances in nonlinear sensitivity analysis</article-title>. <source>Syst. Model. Optim</source>. <volume>38</volume>, <fpage>762</fpage>&#x02013;<lpage>770</lpage>. <pub-id pub-id-type="doi">10.1007/bfb0006203</pub-id></citation>
</ref>
<ref id="B461">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Werbos</surname> <given-names>P.</given-names></name></person-group> (<year>1990</year>). <article-title>Backpropagation through time: what it does and how to do it</article-title>. <source>Proc. IEEE</source>. <volume>78</volume>, <fpage>1550</fpage>&#x02013;<lpage>1560</lpage>. <pub-id pub-id-type="doi">10.1109/5.58337</pub-id></citation>
</ref>
<ref id="B462">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Werbos</surname> <given-names>P. J.</given-names></name> <name><surname>Si</surname> <given-names>J.</given-names></name></person-group> (eds.). (<year>2004</year>). <source>Handbook of Learning and Approximate Dynamic Programming</source>, <volume>Vol. 2</volume>. <publisher-loc>Playa del Carmen</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons.</publisher-name></citation>
</ref>
<ref id="B463">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Werfel</surname> <given-names>J.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2005</year>). <article-title>Learning curves for stochastic gradient descent in linear feedforward networks</article-title>. <source>Neural Comput.</source> <volume>17</volume>, <fpage>2699</fpage>&#x02013;<lpage>2718</lpage>. <pub-id pub-id-type="doi">10.1162/089976605774320539</pub-id><pub-id pub-id-type="pmid">16212768</pub-id></citation>
</ref>
<ref id="B464">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Weston</surname> <given-names>J.</given-names></name> <name><surname>Chopra</surname> <given-names>S.</given-names></name> <name><surname>Bordes</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <source>Memory networks</source>. arXiv:1410.3916.</citation>
</ref>
<ref id="B465">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Whitney</surname> <given-names>W. F.</given-names></name> <name><surname>Chang</surname> <given-names>M.</given-names></name> <name><surname>Kulkarni</surname> <given-names>T.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2016</year>). <source>Understanding visual concepts with continuation learning</source>. arXiv:1602.06822.</citation>
</ref>
<ref id="B466">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>R. J.</given-names></name></person-group> (<year>1992</year>). <article-title>Simple statistical gradient-following algorithms for connectionist reinforcement learning</article-title>. <source>Mach. Learn.</source> <volume>8</volume>, <fpage>229</fpage>&#x02013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1007/BF00992696</pub-id></citation>
</ref>
<ref id="B467">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>R. J.</given-names></name> <name><surname>Baird</surname> <given-names>L. C.</given-names></name></person-group> (<year>1993</year>). <source>Tight Performance Bounds on Greedy Policies based on Imperfect Value Functions</source>. Technical Report, Citeseer.</citation>
</ref>
<ref id="B468">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>S. R.</given-names></name> <name><surname>Stuart</surname> <given-names>G. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Backpropagation of physiological spike trains in neocortical pyramidal neurons: implications for temporal coding in dendrites</article-title>. <source>J. Neurosci.</source> <volume>20</volume>, <fpage>8238</fpage>&#x02013;<lpage>8246</lpage>. <pub-id pub-id-type="pmid">11069929</pub-id></citation>
</ref>
<ref id="B469">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilson</surname> <given-names>R. I.</given-names></name> <name><surname>Nicoll</surname> <given-names>R. A.</given-names></name></person-group> (<year>2001</year>). <article-title>Endogenous cannabinoids mediate retrograde signalling at hippocampal synapses</article-title>. <source>Nature</source> <volume>410</volume>, <fpage>588</fpage>&#x02013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.1038/35069076</pub-id><pub-id pub-id-type="pmid">11279497</pub-id></citation>
</ref>
<ref id="B470">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Winston</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>The strong story hypothesis and the directed perception hypothesis</article-title>, in <source>AAAI Fall Symposium Series</source> (Association for the Advancement of Artificial Intelligence).</citation>
</ref>
<ref id="B471">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wiskott</surname> <given-names>L.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2002</year>). <article-title>Slow feature analysis: unsupervised learning of invariances</article-title>. <source>Neural Comput.</source> <volume>14</volume>, <fpage>715</fpage>&#x02013;<lpage>770</lpage>. <pub-id pub-id-type="doi">10.1162/089976602317318938</pub-id><pub-id pub-id-type="pmid">11936959</pub-id></citation>
</ref>
<ref id="B472">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolpert</surname> <given-names>D. M.</given-names></name> <name><surname>Flanagan</surname> <given-names>J. R.</given-names></name></person-group> (<year>2016</year>). <article-title>Computations underlying sensorimotor learning</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>37</volume>, <fpage>7</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2015.12.003</pub-id><pub-id pub-id-type="pmid">26719992</pub-id></citation>
</ref>
<ref id="B473">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Womelsdorf</surname> <given-names>T.</given-names></name> <name><surname>Valiante</surname> <given-names>T. A.</given-names></name> <name><surname>Sahin</surname> <given-names>N. T.</given-names></name> <name><surname>Miller</surname> <given-names>K. J.</given-names></name> <name><surname>Tiesinga</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>Dynamic circuit motifs underlying rhythmic gain control, gating and integration</article-title>. <source>Nat. Neurosci.</source> <volume>17</volume>, <fpage>1031</fpage>&#x02013;<lpage>1039</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3764</pub-id><pub-id pub-id-type="pmid">25065440</pub-id></citation>
</ref>
<ref id="B474">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wyss</surname> <given-names>R.</given-names></name> <name><surname>K&#x000F6;nig</surname> <given-names>P.</given-names></name> <name><surname>Verschure</surname> <given-names>P. F. M. J.</given-names></name></person-group> (<year>2006</year>). <article-title>A model of the ventral visual system based on temporal stability and local memory</article-title>. <source>PLoS Biol.</source> <volume>4</volume>:<fpage>e120</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pbio.0040120</pub-id><pub-id pub-id-type="pmid">16605306</pub-id></citation>
</ref>
<ref id="B475">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Seung</surname> <given-names>H.</given-names></name></person-group> (<year>2000</year>). <article-title>Spike-based learning rules and stabilization of persistent neural activity</article-title>, in <source>Advances in Neural Information Processing System</source> (<publisher-loc>Denver</publisher-loc>).</citation>
</ref>
<ref id="B476">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2003</year>). <article-title>Equivalence of backpropagation and contrastive Hebbian learning in a layered network</article-title>. <source>Neural Comput.</source> <volume>15</volume>, <fpage>441</fpage>&#x02013;<lpage>454</lpage>. <pub-id pub-id-type="doi">10.1162/089976603762552988</pub-id><pub-id pub-id-type="pmid">12590814</pub-id></citation>
</ref>
<ref id="B477">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiong</surname> <given-names>C.</given-names></name> <name><surname>Merity</surname> <given-names>S.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>Dynamic memory networks for visual and textual question answering</article-title>. <source>arXiv preprint arXiv:1603.01417</source>.</citation>
</ref>
<ref id="B478">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>S.-Y.</given-names></name> <name><surname>Dan</surname> <given-names>Y.</given-names></name> <name><surname>Poo</surname> <given-names>M.-M.</given-names></name></person-group> (<year>2014</year>). <article-title>Representation of interval timing by temporally scalable firing patterns in rat prefrontal cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>480</fpage>&#x02013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1321314111</pub-id><pub-id pub-id-type="pmid">24367075</pub-id></citation>
</ref>
<ref id="B479">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016a</year>). <article-title>Using goal-driven deep learning models to understand sensory cortex</article-title>. <source>Nat. Neurosci</source>. <volume>19</volume>, <fpage>356</fpage>&#x02013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4244</pub-id><pub-id pub-id-type="pmid">26906502</pub-id></citation>
</ref>
<ref id="B480">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016b</year>). <article-title>Eight open questions in the computational modeling of higher sensory cortex</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>37</volume>, <fpage>114</fpage>&#x02013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2016.02.001</pub-id><pub-id pub-id-type="pmid">26921828</pub-id></citation>
</ref>
<ref id="B481">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yosinski</surname> <given-names>J.</given-names></name> <name><surname>Clune</surname> <given-names>J.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Lipson</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>How transferable are features in deep neural networks?</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>3320</fpage>&#x02013;<lpage>3328</lpage>.</citation>
</ref>
<ref id="B482">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yttri</surname> <given-names>E. A.</given-names></name> <name><surname>Dudman</surname> <given-names>J. T.</given-names></name></person-group> (<year>2016</year>). <article-title>Opponent and bidirectional control of movement velocity in the basal ganglia</article-title>. <source>Nature</source> <volume>533</volume>, <fpage>402</fpage>&#x02013;<lpage>406</lpage>. <pub-id pub-id-type="doi">10.1038/nature17639</pub-id><pub-id pub-id-type="pmid">27135927</pub-id></citation>
</ref>
<ref id="B483">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>C.</given-names></name> <name><surname>Smith</surname> <given-names>L. B.</given-names></name></person-group> (<year>2013</year>). <article-title>Joint attention without gaze following: human infants and their parents coordinate visual attention to objects through eye-hand coordination</article-title>. <source>PLoS ONE</source> <volume>8</volume>:<fpage>e79659</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0079659</pub-id><pub-id pub-id-type="pmid">24236151</pub-id></citation>
</ref>
<ref id="B484">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuste</surname> <given-names>R.</given-names></name> <name><surname>MacLean</surname> <given-names>J. N.</given-names></name> <name><surname>Smith</surname> <given-names>J.</given-names></name> <name><surname>Lansner</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>The cortex as a central pattern generator</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>6</volume>, <fpage>477</fpage>&#x02013;<lpage>483</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1686</pub-id><pub-id pub-id-type="pmid">15928717</pub-id></citation>
</ref>
<ref id="B485">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zaremba</surname> <given-names>W.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2015</year>). <article-title>Reinforcement learning neural turing machines</article-title>. <source>arXiv preprint arXiv:1505.00521</source>.</citation>
</ref>
<ref id="B486">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeisel</surname> <given-names>A.</given-names></name> <name><surname>Manchado</surname> <given-names>A. B. M.</given-names></name> <name><surname>Codeluppi</surname> <given-names>S.</given-names></name> <name><surname>L&#x000F6;nnerberg</surname> <given-names>P.</given-names></name> <name><surname>La Manno</surname> <given-names>G.</given-names></name> <name><surname>Jur&#x000E9;us</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Cell types in the mouse cortex and hippocampus revealed by single-cell RNA-seq</article-title>. <source>Science</source> <volume>347</volume>, <fpage>1138</fpage>&#x02013;<lpage>1142</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaa1934</pub-id><pub-id pub-id-type="pmid">25700174</pub-id></citation>
</ref>
<ref id="B487">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zemel</surname> <given-names>R. S.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name></person-group> (<year>1997</year>). <article-title>Combining probabilistic population codes</article-title>, in <source>International Joint Conference on Artificial Intelligence</source> (<publisher-loc>Nagoya</publisher-loc>), <fpage>1114</fpage>&#x02013;<lpage>1119</lpage>.</citation>
</ref>
<ref id="B488">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zilli</surname> <given-names>E. A.</given-names></name> <name><surname>Hasselmo</surname> <given-names>M. E.</given-names></name></person-group> (<year>2010</year>). <article-title>Coupled noisy spiking neurons as velocity-controlled oscillators in a model of grid cell spatial firing</article-title>. <source>J. Neurosci.</source> <volume>30</volume>, <fpage>13850</fpage>&#x02013;<lpage>13860</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0547-10.2010</pub-id><pub-id pub-id-type="pmid">20943925</pub-id></citation>
</ref>
<ref id="B489">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zipser</surname> <given-names>D.</given-names></name> <name><surname>Andersen</surname> <given-names>R.</given-names></name></person-group> (<year>1988</year>). <article-title>A back-propagation programmed network that simulates response properties of a subset of posterior parietal neurons</article-title>. <source>Nature</source> <volume>331</volume>, <fpage>679</fpage>&#x02013;<lpage>684</lpage>. <pub-id pub-id-type="doi">10.1038/331679a0</pub-id><pub-id pub-id-type="pmid">3344044</pub-id></citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>Hyper-parameter optimization shows that complicated schedules of training, which differ across parts of the network, lead to optimal performance (Maclaurin et al., <xref ref-type="bibr" rid="B276">2015</xref>).</p></fn>
<fn id="fn0002"><p><sup>2</sup>In adversarial networks, a generator network is trained to fool a discriminator network into being unable to distinguish generated samples from real data samples, while the discriminator network is trained to prevent the generator network from fooling it in this way.</p></fn>
<fn id="fn0003"><p><sup>3</sup>Psychologists have been quantifying the subtleties of many such developmental stagings, e.g., of our perceptual and motor performance, e.g., Nardini et al. (<xref ref-type="bibr" rid="B320">2010</xref>), Dekker and Nardini (<xref ref-type="bibr" rid="B96">2015</xref>), and McKone et al. (<xref ref-type="bibr" rid="B295">2009</xref>).</p></fn>
<fn id="fn0004"><p><sup>4</sup>Our point in this section will not be that all learning in the brain can be captured by cost function optimization, but rather, somewhat more narrowly, our claim is that the algorithms for optimization like backpropagation in deep learning may have correspondences in biological brains. We feel that it is an important task for neuroscience to determine whether and how brains implement these algorithms. The brain may also disclose dynamics that are unlike these algorithms, so we are not disclaiming the possibility of broader theories. In machine learning, many useful algorithms are not explicitly formulated as cost function optimization; for example, many algorithms are based on linear algebra procedures like singular value decomposition, rather than explicit optimization. Such methods can be made nonlinear by using nonlinear kernels&#x02014;relatedly, some brain circuits run specialized computations using fixed nonlinear basis functions (e.g., in cerebellum). Moreover, while an implicit cost function can be attributed to account for many dynamical processes, as well as many popular learning algorithms, our claim is not merely that the brain uses other learning procedures that lead to solutions which implicitly minimize a cost function, but rather that it actually finds its solutions by performing a powerful form of optimization as such.</p></fn>
<fn id="fn0005"><p><sup>5</sup>Of course, some circuits may also be heavily genetically pre-specified to minimize the burden on learning. For instance, particular cell adhesion molecules (Hattori et al., <xref ref-type="bibr" rid="B173">2007</xref>) expressed on particular parts of particular neurons defined by a genetic cell type (Zeisel et al., <xref ref-type="bibr" rid="B486">2015</xref>), and the detailed shapes and placements of neuronal arbors, may constrain connectivity in some cases, though in other cases local connectivity is thought to be only weakly constrained (Kalisman et al., <xref ref-type="bibr" rid="B218">2005</xref>). Genetics is sufficient to specify complex circuits involving hundreds of neurons, such as central pattern generators (Yuste et al., <xref ref-type="bibr" rid="B484">2005</xref>) which create complex self-stabilizing oscillations, or the entire nervous systems of small worms. Genetically guided wiring should not be thought of as fixed &#x0201C;hard-wiring&#x0201D; but rather as a programmatic construction process that can also accept external inputs and interact with learning mechanisms (Marcus, <xref ref-type="bibr" rid="B283">2004</xref>).</p></fn>
<fn id="fn0006"><p><sup>6</sup>Hebbian plasticity even has a well-understood biological basis in the form of the NMDA receptors, which are activated by the simultaneous occurrence of chemical transmitter delivered from the pre-synaptic neuron, and voltage depolarization of the post-synaptic neuron.</p></fn>
<fn id="fn0007"><p><sup>7</sup>The variance can be mitigated by averaging out many perturbations before making a change to the baseline value of the weights, but this would take significant time for a network of non-trivial size as the variance of weight perturbation&#x00027;s estimates scales in proportion to the number of synapses in the network.</p></fn>
<fn id="fn0008"><p><sup>8</sup>If the error derivatives of the cost function with respect to the last layer of unit activities are unknown, then they can be replaced with node-perturbation-like correlations, as is common in reinforcement learning.</p></fn>
<fn id="fn0009"><p><sup>9</sup>Interestingly, STDP is not a unitary phenomenon, but rather a diverse collection of different rules with different timescales and temporal asymmetries (Sj&#x000F6;str&#x000F6;m and Gerstner, <xref ref-type="bibr" rid="B402">2010</xref>; Mishra et al., <xref ref-type="bibr" rid="B309">2016</xref>). Effects include STDP with the inverse temporal asymmetry, symmetric STDP and STDP with different temporal window sizes. STDP is also frequency dependent, which can be explained by rules that depend on triplets rather than pairs of spikes (Pfister and Gerstner, <xref ref-type="bibr" rid="B347">2006</xref>). In some cortical neurons, STDP even switches its sign as the synapse moves away from the neuron&#x00027;s soma into the dendritic tree (Letzkus et al., <xref ref-type="bibr" rid="B256">2006</xref>). While STDP is often included explicitly in models, biophysical derivations of STDP from various underlying phenomena are also being attempted, some of which involve the post-synaptic voltage (Clopath and Gerstner, <xref ref-type="bibr" rid="B81">2010</xref>) or a local dendritic voltage (Urbanczik and Senn, <xref ref-type="bibr" rid="B444">2014</xref>). Meanwhile, other theories suggest that STDP may enable the use of precise timing codes based on temporal coincidence of inputs, the generation and unsupervised learning of temporal sequences (Abbott and Blum, <xref ref-type="bibr" rid="B2">1996</xref>; Fiete et al., <xref ref-type="bibr" rid="B119">2010</xref>), enhancements to distal reward processing in reinforcement learning (Izhikevich, <xref ref-type="bibr" rid="B204">2007</xref>), stabilization of neural responses (Kempter et al., <xref ref-type="bibr" rid="B222">2001</xref>), or many other higher-level properties (Nessler et al., <xref ref-type="bibr" rid="B322">2013</xref>; Kappel et al., <xref ref-type="bibr" rid="B221">2014</xref>).</p></fn>
<fn id="fn0010"><p><sup>10</sup>Hinton has suggested (Hinton, <xref ref-type="bibr" rid="B184">2007</xref>, <xref ref-type="bibr" rid="B185">2016</xref>) that this could take place in the context of autoencoders and recirculation (Hinton and McClelland, <xref ref-type="bibr" rid="B188">1988</xref>). Bengio and colleagues have proposed (Bengio, <xref ref-type="bibr" rid="B36">2014</xref>; Bengio and Fischer, <xref ref-type="bibr" rid="B37">2015</xref>; Scellier and Bengio, <xref ref-type="bibr" rid="B383">2016</xref>) another context in which the connection between STDP and plasticity rules that depend on the temporal derivative of the post-synaptic firing rate can be exploited for biologically plausible multilayer credit assignment. This setting relies on clamping of outputs and stochastic relaxation in energy-based models (Ackley et al., <xref ref-type="bibr" rid="B3">1958</xref>), which leads to a continuous network dynamics (Hopfield, <xref ref-type="bibr" rid="B197">1984</xref>) in which hidden units are perturbed toward target values (Bengio and Fischer, <xref ref-type="bibr" rid="B37">2015</xref>), loosely similar to that which occurs in XCAL. This dynamics then allows the STDP-based rule to correspond to gradient descent on the energy function with respect to the weights (Scellier and Bengio, <xref ref-type="bibr" rid="B383">2016</xref>). This scheme requires symmetric weights, but in an autoencoder context, Bengio notes that these can arise spontaneously (Arora et al., <xref ref-type="bibr" rid="B16">2015</xref>).</p></fn>
<fn id="fn0011"><p><sup>11</sup>Even BPTT has arguably not been completely successful in recurrent networks. The problems of vanishing and exploding gradients led to long short term memory networks with gated memory units. An alternative is to use optimization methods that go beyond first order derivatives (Martens and Sutskever, <xref ref-type="bibr" rid="B291">2011</xref>). This suggests the need for specialized systems and structures in the brain to mitigate problems of temporal credit assignment.</p></fn>
<fn id="fn0012"><p><sup>12</sup>Interestingly, the hippocampus seems to &#x0201C;time stamp&#x0201D; memories by encoding them into ensembles with cellular compositions and activity patterns that change gradually as a function of time on the scale of days (Rubin et al., <xref ref-type="bibr" rid="B375">2015</xref>; Cai et al., <xref ref-type="bibr" rid="B70">2016</xref>), and may use &#x0201C;time cells&#x0201D; to mark temporal positions within episodes on a timescale of seconds (Kraus et al., <xref ref-type="bibr" rid="B233">2013</xref>).</p></fn>
<fn id="fn0013"><p><sup>13</sup>Control theory concepts also appear to be useful for simplifying optimization problems in certain other settings (Todorov, <xref ref-type="bibr" rid="B438">2009</xref>; Hennequin et al., <xref ref-type="bibr" rid="B180">2014</xref>).</p></fn>
<fn id="fn0014"><p><sup>14</sup>In one intriguing study of interval timing, single neurons exhibited response patterns over time which were scaled to the interval duration, and cooling the brain to slow down neural dynamics led to longer intervals being computed by the brain (Xu et al., <xref ref-type="bibr" rid="B478">2014</xref>).</p></fn>
<fn id="fn0015"><p><sup>15</sup>Analogs of weight perturbation and node perturbation are known for spiking networks (Seung, <xref ref-type="bibr" rid="B394">2003</xref>; Fiete and Seung, <xref ref-type="bibr" rid="B120">2006</xref>). Seung (<xref ref-type="bibr" rid="B394">2003</xref>) also discusses implications of gradient based learning algorithms for neuroscience, echoing some of our considerations here.</p></fn>
<fn id="fn0016"><p><sup>16</sup>A related, but more general, question is how to learn over many layers of non-differentiable structures. One option is to perform updates via finite-sized rather than infinitesimal steps, e.g., via target-propagation (Bengio, <xref ref-type="bibr" rid="B36">2014</xref>).</p></fn>
<fn id="fn0017"><p><sup>17</sup>Eliasmith and others have shown (Eliasmith and Anderson, <xref ref-type="bibr" rid="B106">2004</xref>; Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>; Eliasmith, <xref ref-type="bibr" rid="B105">2013</xref>) that complex functions and control systems can be compiled onto such networks, using nonlinear encoding and linear decoding of high-dimensional vectors.</p></fn>
<fn id="fn0018"><p><sup>18</sup>Dendritic computation may also have other functions, e.g., competitive interactions between dendrites in a single neuron could also allow neurons to contribute to multiple different ensembles (Legenstein and Maass, <xref ref-type="bibr" rid="B251">2011</xref>).</p></fn>
<fn id="fn0019"><p><sup>19</sup>Localized activity in dendrites drives localized plasticity, with inhibitory interneurons, and interactions between inputs at different parts of the dendritic tree, controlling the local sign and spatial distribution of this plasticity (Sj&#x000F6;str&#x000F6;m and H&#x000E4;usser, <xref ref-type="bibr" rid="B403">2006</xref>; Cichon and Gan, <xref ref-type="bibr" rid="B79">2015</xref>).</p></fn>
<fn id="fn0020"><p><sup>20</sup>In the model of K&#x000F6;rding and K&#x000F6;nig (<xref ref-type="bibr" rid="B230">2001</xref>), single spikes are used to transmit activations and burst spikes are used to transmit error information. In other models, including the dendritic voltage in a plasticity rule leads to error-driven and predictive learning that can approximate backpropagation <italic>inside</italic> a single complex neuron (in effect backpropagating from the net somatic output, through nonlinearities at the dendritic branch points, all the way back to the individual input synaptic weights) and that generalize to a reinforcement learning context (Urbanczik and Senn, <xref ref-type="bibr" rid="B444">2014</xref>; Schiess et al., <xref ref-type="bibr" rid="B384">2016</xref>). Single neurons with active dendrites and many synapses may also embody learning rules of greater complexity, such as the storage and recall of temporal patterns (Hawkins and Ahmad, <xref ref-type="bibr" rid="B174">2016</xref>).</p></fn>
<fn id="fn0021"><p><sup>21</sup>Interestingly, some connectomic studies are finding more obvious connectivity structure at the level of dendritic organization than at the cellular level (Morgan et al., <xref ref-type="bibr" rid="B318">2016</xref>).</p></fn>
<fn id="fn0022"><p><sup>22</sup>An interesting recent study explored this idea in the context of a model of modular cortical-column-like units (Piekniewski et al., <xref ref-type="bibr" rid="B349">2016</xref>). Local units are multi-layer perceptrons trained to minimize a prediction error by gradient descent. Within each unit, predictive autoencoders form a data compression in their middle layers, which is then fed up to higher levels as well as laterally. This system is suggestive of the power of using modular units of intermediate complexity, each of which minimizes a prediction error locally, e.g., in a local few-layer network. The system currently uses a fixed format for transmission of vectors from one unit to another, but ideally the inter-module connections should also be trained by gradient descent as well or by reinforcement learning rather than being fixed. The cortical-column-like modules could also be made more complex and could be organized into higher-order structures like Minsky&#x00027;s semantic networks, frames and K-lines (Minsky, <xref ref-type="bibr" rid="B305">1988</xref>) rather than in simple hierarchies, or such an architecture could self-organize via reinforcement learning or other mechanisms for defining inter-column connections. Such a system also needs connections with specific kinds of memory and long-range information routing systems.</p></fn>
<fn id="fn0023"><p><sup>23</sup>This idea has been used by Hawkins and colleagues to suggest mechanisms for continuous online sequence learning (Cui et al., <xref ref-type="bibr" rid="B89">2015</xref>; Hawkins and Ahmad, <xref ref-type="bibr" rid="B174">2016</xref>) and by Larkum and colleagues for comparison of top-down and bottom-up signals (Larkum, <xref ref-type="bibr" rid="B245">2013</xref>). The Larkum model focuses on the layer 5 (L5) pyramidal neuron type. The cell body of this neuron lies in L5 but extends its &#x0201C;apical&#x0201D; dendritic tree all the way up to a tuft at the top of the cortex in layer 1 (L1), which is a primary target of feedback projections. In the model, interactions between local spiking in these different dendritic zones, which are targeted by different kinds of projections, are crucial to the learning function. The model of Hawkins (Cui et al., <xref ref-type="bibr" rid="B89">2015</xref>; Hawkins and Ahmad, <xref ref-type="bibr" rid="B174">2016</xref>) also focused on the unique dendritic structure of the L5 pyramidal neuron, and distinguishes internal states of the neuron, which impact its responsiveness to other inputs, from activation states, which directly translate into spike rates. Three integration zones in each neuron, and dendritic NMDA spikes (Palmer et al., <xref ref-type="bibr" rid="B339">2014</xref>) acting as local coincidence detectors (Shai et al., <xref ref-type="bibr" rid="B395">2015</xref>), allow temporal patterns of dendritic input to impact the cell&#x00027;s internal state. Intra-column inhibition is also used in this model. Other cortical models pay less attention to the details of dendritic computation, but still provide detailed interpretations of the inter-laminar projection patterns of the neocortex. For example, in O&#x00027;Reilly et al. (<xref ref-type="bibr" rid="B337">2014b</xref>), an architecture is presented for continuous learning based on prediction of the next input. Time is discretized into 100 ms bins via an alpha oscillation, and the deep vs. shallow layers maintain different information during these time bins, with deep layers maintaining a record of the previous time step, and shallow layers representing the current state. The stored information in the deep layers leads to a prediction of the current state, which is then compared with the actual current state. Periodic bursting locked to the oscillation provides a kind of clock that causes the current state to be shifted into the deep layers for maintenance during the subsequent time step, and recurrent loops with the thalamus allow this representation to remain stable for sufficiently long to be used to generate the prediction. Other theories utilize the biophysics of dendritic computation and spike timing dependent plasticity to explain how neurons could learn to make predictions (Brea et al., <xref ref-type="bibr" rid="B55">2016</xref>) on a timescale of seconds using neurons with intrinsic plasticity time constants of a few tens of milliseconds.</p></fn>
<fn id="fn0024"><p><sup>24</sup>I-theory can perhaps be viewed as a generalized alternative paradigm to the online optimization of cost functions via multi-layer gradient descent, as used in deep learning. It exploits similar network architectures as conventional deep learning, e.g., hierarchical convolutional networks for the case of feedforward vision, but rather than backpropagating errors, it uses local circuits and learning rules to store templates against which new inputs are compared. This relies on a theory of generalization in learning based on combinations of tuned units (Poggio and Bizzi, <xref ref-type="bibr" rid="B354">2004</xref>), which has been applied to both vision and motor control. Neurons with the required Gaussian-like tunings to stored templates could be obtained through canonical, local, normalization-based circuits (Kouh and Poggio, <xref ref-type="bibr" rid="B232">2008</xref>), which can also be tweaked to implement other aspects of a vision architecture like softmax operations and pooling.</p></fn>
<fn id="fn0025"><p><sup>25</sup>One alternative picture that contrasts with straightforward cost function optimization emphasizes the types of computation that appear most naturally suited to heterogeneous, stochastic, noisy, continually changing neural circuitry (Maass, <xref ref-type="bibr" rid="B272">2016</xref>). On this view, network plasticity is viewed as a sampling-based approximation to Bayesian inference (Kappel et al., <xref ref-type="bibr" rid="B220">2015</xref>) where transiently changing synapses sample from a posterior distribution of network configurations, rather than as gradient descent on a cost function. This view emphasizes Monte-Carlo sampling procedures, rather than cost function optimization.</p></fn>
<fn id="fn0026"><p><sup>26</sup>Sampling based inference procedures are used widely in Bayesian statistics, and efforts have been made to connect these procedures with circuit-based models of computations (Mansinghka and Jonas, <xref ref-type="bibr" rid="B280">2014</xref>). It currently appears difficult, however, to reconcile generic Marcov Chain Monte Carlo (MCMC) dynamics, which mix slowly, with the fast time scales of human psychophysics. But Bayesian methods are powerful and come with a methodology for model comparison (Ghahramani, <xref ref-type="bibr" rid="B141">2005</xref>). In machine learning, variational Bayesian methods have recently become popular precisely because they are capable of fast though approximate posterior inference (inferring causes from observables), but seem to be powerful enough to create strong models. For example, stochastic gradient descent optimization is beginning to be used for variational Bayesian inference (Kingma and Welling, <xref ref-type="bibr" rid="B224">2013</xref>). Restricted Boltzmann Machines (RBMs) also achieve fast inference in shallow architectures&#x02014;with only a small number of iterations of mixing required&#x02014;but they do not mix quickly when stacked into deep hierarchies as deep Boltzmann machines. The greedy, layer-wise pre-training of a deep belief network (Hinton et al., <xref ref-type="bibr" rid="B187">2006</xref>) provides a heuristic way to stack the RBMs by auto-encoding, but these have achieved less competitive results than current variational Bayesian models. The problem of fast inference in MCMC models is the subject of current research, including at the interface with biologically plausible models (Bengio et al., <xref ref-type="bibr" rid="B41">2016</xref>). When these models are made to perform fast inference, they actually become somewhat similar to variational Bayesian methods, since they rely on feedforward approximate inference, at least to initialize the system.</p></fn>
<fn id="fn0027"><p><sup>27</sup>Heuristics are widely used to simplify motor planning and control, e.g., McLeod and Dienes (<xref ref-type="bibr" rid="B296">1996</xref>).</p></fn>
<fn id="fn0028"><p><sup>28</sup>Many other reinforcement learning algorithms, including REINFORCE (Williams, <xref ref-type="bibr" rid="B466">1992</xref>), can be implemented as fully online algorithms using &#x0201C;eligibility traces,&#x0201D; which accumulate the sensitivity of action distributions to parameters in a temporally local manner (Sutton and Barto, <xref ref-type="bibr" rid="B423">1998</xref>).</p></fn>
<fn id="fn0029"><p><sup>29</sup>Zaremba and Sutskever (<xref ref-type="bibr" rid="B485">2015</xref>) also bridges reinforcement learning and backpropagation learning in the same system, in the context of a neural network controlling discrete interfaces, and illustrates some of the challenges of this approach: compared to an end-to-end backpropagation-trained Neural Turing Machine (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>), reinforcement based training allows training of only relatively simple algorithmic tasks. Special measures need to be taken to make reinforcement efficient, including limiting the number of possible actions, subtracting a baseline reward, and training the network using a curriculum schedule.</p></fn>
<fn id="fn0030"><p><sup>30</sup>This is distinct from a game-theoretic scenario in which multiple actors can achieve an equilibrium, e.g., Gemp and Mahadevan (<xref ref-type="bibr" rid="B136">2015</xref>).</p></fn>
<fn id="fn0031"><p><sup>31</sup>Single neurons act as comparators in the motor system, e.g., Brownstone et al. (<xref ref-type="bibr" rid="B60">2015</xref>), and networks in the retina adapt so as to report local differences in space or time rather than absolute values, a form of predictive coding (Hosoya et al., <xref ref-type="bibr" rid="B199">2005</xref>).</p></fn>
<fn id="fn0032"><p><sup>32</sup>Beginning with Hopfield&#x00027;s definition of an energy function for inference in certain classes of symmetric network (Hopfield, <xref ref-type="bibr" rid="B196">1982</xref>), researchers have discovered networks with inherent dynamics that implicitly optimizes certain objectives even while the connection weights are fixed, such as statistical reconstruction of the input via stochastic relaxation in Boltzmann machines (Ackley et al., <xref ref-type="bibr" rid="B3">1958</xref>). Fast approximations of some of these inference procedures are perhaps biologically plausible and could rely on dendritic computation (Bengio et al., <xref ref-type="bibr" rid="B41">2016</xref>). Iterative local Hebbian-like <italic>learning</italic> rules are often used to train the <italic>weights</italic> of such networks, without explicitly propagating error derivatives in the manner of backpropagation. In an appropriate network context, many other combinations of network dynamics and plasticity rules can give rise to inference and learning procedures that implicitly descend cost functions in activity space and/or weight space.</p></fn>
<fn id="fn0033"><p><sup>33</sup>Dreams arguably illustrate that the brain uses generative models which also involve selective recall and recombination of episodic memories.</p></fn>
<fn id="fn0034"><p><sup>34</sup>Much is known about the architecture of cortical feedback vs. feedforward connections. For example, canonically, feedforward connections project from superficial cortical layers to layer 4 of the recipient layer, while feedback connections terminate outside layer 4 and often originate in deeper layers. These types of relationships can be used anatomically to define the hierarchical organization of visual areas, as in Felleman and Van Essen (<xref ref-type="bibr" rid="B114">1991</xref>), although the original studies were performed in primates and the precise generalization to rodent cortex is not fully clear (Berezovskii et al., <xref ref-type="bibr" rid="B42">2011</xref>), and there may be various alternate or overlapping anatomical pathways (Callaway, <xref ref-type="bibr" rid="B71">2004</xref>), e.g., with some pathways involved in specific functions like gain control, others routed through specific gating mechanisms, and so forth. Advances in connectomics should allow this architecture to be studied more directly. The study of receptive field properties in the visual cortical hierarchy has led to many insights into this hierarchical system. For example, while each neuron in V1 has a classical local receptive field, neural responses at a given location in V1 also depend on visual locations far from the classical receptive field, e.g., through various forms of surround suppression. These studies have allowed an understanding of the spatial scales over which feedback connections operate in the early visual system (Angelucci et al., <xref ref-type="bibr" rid="B11">2002</xref>). In particular, feedback connections are invoked to account for longer-range receptive field interactions, whereas horizontal connections are invoked to account for shorter-range receptive field interactions (Schwabe et al., <xref ref-type="bibr" rid="B388">2006</xref>). Feedforward and feedback pathways are also distinguished dynamically, e.g., by propagating different oscillatory frequencies (Van Kerkoerle et al., <xref ref-type="bibr" rid="B448">2014</xref>; Bastos et al., <xref ref-type="bibr" rid="B32">2015</xref>), and moleculary, e.g., with NMDA receptors playing an important role in feedback processing.</p></fn>
<fn id="fn0035"><p><sup>35</sup>Temporal continuity is exploited in Poggio (<xref ref-type="bibr" rid="B353">2015</xref>), which analyzes many properties of deep convolutional networks with respect to their biological plausibility, including their apparent need for large amounts of supervised training data, and concludes that the environment may in fact provide a sufficient number of &#x0201C;implicitly,&#x0201D; though not explicitly, labeled examples to train a deep convolutional network for object recognition. Implicit labeling of object identity, in this case, arises from temporal continuity: successive frames of a video are likely to have the same objects in similar places and orientations. This allows the brain to derive an invariant signature of object identity which is independent of transformations like translations and rotations, but which does not yet associate the object with a specific name or label. Once such an invariant signature is established, however, it becomes basically trivial to associate the signature with a label for classification (Anselmi et al., <xref ref-type="bibr" rid="B12">2015</xref>). Poggio (<xref ref-type="bibr" rid="B353">2015</xref>) also suggests specific means, in the context of I-theory (Anselmi et al., <xref ref-type="bibr" rid="B12">2015</xref>), by which this training could occur via the storage of image templates using Hebbian mechanisms among simple and complex cells in the visual cortex. Thus, in this model, the brain has used its implicit knowledge of the temporal continuity of object motion to provide a kind of minimal labeling that is sufficient to bootstrap object recognition. Although not formulated as a cost function, this shows how usefully the assumption of temporal continuity could be exploited by the brain.</p></fn>
<fn id="fn0036"><p><sup>36</sup>Although, some multi-sensory integration appears to occur even in the early sensory cortices (Cappe et al., <xref ref-type="bibr" rid="B72">2012</xref>).</p></fn>
<fn id="fn0037"><p><sup>37</sup>Other brain-inspired unsupervised objectives are being developed for unsupervised visual learning. One recent paper (Higgins et al., <xref ref-type="bibr" rid="B182">2016</xref>) uses an objective function that seeks representations of statistically independent factors in images, by introducing a regularization term that pushes the distribution of latent factors learned in a generative model to be close to a unit Gaussian. This is based on a theory that the ventral visual stream is optimized to disentangle factors of variation in images.</p></fn>
<fn id="fn0038"><p><sup>38</sup>In the visual system, it is still unknown why a clustered spatial pattern of representational categories arises, e.g., a physically localized &#x0201C;area&#x0201D; that seems to correspond to representations of faces (Kanwisher et al., <xref ref-type="bibr" rid="B219">1997</xref>), another area for representations of visual word forms (McCandliss et al., <xref ref-type="bibr" rid="B292">2003</xref>), and so on. It is also unknown why this spatial pattern seems to be largely reproducible across individuals. Some theories are based on bottom-up correlation-based clustering or neuronal competition mechanisms, which generate category-selective regions as a byproduct. Other theories suggest a computational reason for this organization, in the context of I-theory (Anselmi et al., <xref ref-type="bibr" rid="B12">2015</xref>), involving the limited ability to generalize transformation-invariances learned for one class of objects to other classes (Leibo et al., <xref ref-type="bibr" rid="B253">2015b</xref>). Areas for abstract culture-dependent concepts, like the visual word form area, suggest that the decomposition cannot be &#x0201C;purely genetic.&#x0201D; But it is conceivable that these areas could at least in part reflect different local cost functions.</p></fn>
<fn id="fn0039"><p><sup>39</sup>Psychologists have postulated other innate heuristics, e.g., in the context of object tracking (Franconeri et al., <xref ref-type="bibr" rid="B127">2012</xref>). That infant object concepts are trainable but only along certain dimensions (Scholl, <xref ref-type="bibr" rid="B387">2004</xref>) also suggests the notion of a heuristically &#x0201C;guided&#x0201D; or &#x0201C;bootstrapped&#x0201D; learning process in this context.</p></fn>
<fn id="fn0040"><p><sup>40</sup>Of course, specialized architecture also enters the picture at the level of the pre-structuring of trainable/optimizable modules themselves. Just as in deep learning, convolutional networks, LSTMs, residual networks and other specific architectures are used to make learning efficient and fast, even though more generic architectures like multilayer perceptrons or generally RNNs are universal function approximators.</p></fn>
<fn id="fn0041"><p><sup>41</sup>It is interesting to consider how standard neural network models of vision would fit into this categorization. Consider convolutional neural networks, for example, with the convolutional filters optimized via supervised backpropagation. This is by no means a completely unstructured prior to backpropagation-based training. Indeed, these networks typically contain max-pooling and normalization layers with fixed computations that are not altered during learning, as well as fixed architectural features such as number and arrangement of layers, size and stride of the sliding window, and so forth. Likewise &#x0201C;hierarchical max-pooling&#x0201D; (HMAX) models (Serre et al., <xref ref-type="bibr" rid="B391">2007</xref>) of the ventral stream are so-named because of these fixed architectural aspects. Thus, in a hypothetical biological implementation of such systems, these aspects would be pre-structured by genetics even if the convolutional weights would be trained via some kind of gradient descent optimization. There are some plausible neural circuits that would implement these standardized normalization and max pooling operations (Kouh and Poggio, <xref ref-type="bibr" rid="B232">2008</xref>). Moreover, in a biological implementation, the machinery necessary to carry out the optimization itself would need to be embodied by appropriate, genetically structured circuitry.</p></fn>
<fn id="fn0042"><p><sup>42</sup>Attractor models of memory in neuroscience tend to have the property that only one memory can be accessed at a time (although a brain can have many such memories that can be accessed in parallel). Recent machine learning systems, however, have constructed differentiable addressable memory (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>) and gating (Whitney et al., <xref ref-type="bibr" rid="B465">2016</xref>) systems by allowing weighted superpositions of memory registers or gates to be queried&#x02014;it is unclear whether the brain uses such mechanisms.</p></fn>
<fn id="fn0043"><p><sup>43</sup>Computational analogies have also been drawn between associative memory storage and object recognition (Leibo et al., <xref ref-type="bibr" rid="B252">2015a</xref>), suggesting the possibility of closely related computations occurring in parts of neocortex and hippocampus. Indeed, the hippocampus and olfactory cortex (a more ancient and simpler structure than the neocortex Shepherd, <xref ref-type="bibr" rid="B396">2014</xref>; Fournier et al., <xref ref-type="bibr" rid="B126">2015</xref>) are few-layer structures described in comparative anatomy as &#x0201C;allocortex,&#x0201D; as opposed to the six-layered &#x0201C;neocortex,&#x0201D; and both types of cortex have some anatomical similarities (particularly for CA1 and subiculum, though less so for CA3 and dentate gyrus) such as the presence of pyramidal neurons. It has been suggested that the hippocampus can be thought of as the top of the cortical hierarchy (Hawkins and Blakeslee, <xref ref-type="bibr" rid="B175">2007</xref>), responsible for handling and remembering information that could not be fully explained by lower levels of the hierarchy. These computational connections are still tentative.</p></fn>
<fn id="fn0044"><p><sup>44</sup>Attention also arguably solves certain types of perceptual binding problem (Reynolds and Desimone, <xref ref-type="bibr" rid="B362">1999</xref>).</p></fn>
<fn id="fn0045"><p><sup>45</sup>The precise roles of synchrony in information routing and other processes, and when it should be viewed as a causal factor vs. as an epiphenomenon of other mechanisms, is still being worked out. In some theories, oscillations occur as consequences of certain recurrent processing loops, e.g., thalamo-cortico-striatal loops (Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>). In other models, so-called &#x0201C;dynamic circuit motifs,&#x0201D; involving specific combinations of cellular and synaptic sub-types, both generate synchronies (e.g., in part via intrinsically rhythmic pacemaker neurons) and exploit them for specific computational roles, particularly in the rapid dynamic formation of communication networks (Womelsdorf et al., <xref ref-type="bibr" rid="B473">2014</xref>).</p></fn>
<fn id="fn0046"><p><sup>46</sup>One idea for achieving such transferability is that of a partitionable (Hayworth, <xref ref-type="bibr" rid="B178">2012</xref>) or annexable (Bostrom, <xref ref-type="bibr" rid="B49">1996</xref>) network. These models posit that a large associative memory network links all the different buffers. This large associative memory network has a number of stable attractor states. These are called &#x0201C;global&#x0201D; attractor states since they link across all the buffers. Forcing a given buffer into an activity pattern resembling that of its corresponding &#x0201C;piece&#x0201D; of an attractor state will cause the entire global network to enter that global attractor state. During training, all of the connections between buffers are turned on, so that their learned contents, though not identical, are kept in correspondence by being part of the same attractor. Later, the connections between specific buffers can be turned off to allow them to store different information. Copy and paste is then implemented by turning on the connections between a source buffer and a destination buffer (Hayworth, <xref ref-type="bibr" rid="B178">2012</xref>). Copying between a source and destination buffer can also be implemented, i.e., learned, in a deep learning system using methods similar to the addressing mechanisms of the Neural Turing Machine (Graves et al., <xref ref-type="bibr" rid="B151">2014</xref>).</p></fn>
<fn id="fn0047"><p><sup>47</sup>Micro-stimulation experiments, in which an animal learns to behaviorally report stimulation of electrode channels located in diverse cortical regions, suggest that many areas can be routed or otherwise linked to behavioral &#x0201C;outputs&#x0201D; (Histed et al., <xref ref-type="bibr" rid="B191">2013</xref>), although the mechanisms behind this&#x02014;e.g., whether this stimulation gives rise to a high-level percept that the animal then uses to make a decision&#x02014;are unclear. Likewise, it is possible to reinforcement-train an animal to control the activity of individual neurons (Fetz, <xref ref-type="bibr" rid="B116">1969</xref>, <xref ref-type="bibr" rid="B117">2007</xref>).</p></fn>
<fn id="fn0048"><p><sup>48</sup>Conventionally, models of the basal ganglia involve all or none gating of an action, but recent evidence suggests that the basal ganglia may also have continuous, analog outputs (Yttri and Dudman, <xref ref-type="bibr" rid="B482">2016</xref>).</p></fn>
<fn id="fn0049"><p><sup>49</sup>It has been suggested that the basic role of the BG is to provide tonic inhibition to other circuits (Grillner et al., <xref ref-type="bibr" rid="B154">2005</xref>). Release of this inhibition can then activate a &#x0201C;discrete&#x0201D; action, such as a motor command. A core function of the BG is thus to choose, based on patterns detected in its input, which of a finite set of actions to initiate via such release of inhibition. In many models of the basal ganglia&#x000E2;&#x00102;&#x00179;s role in cognitive control, the targets of inhibition are thalamic relays (Sherman, <xref ref-type="bibr" rid="B397">2005</xref>), which are set in a default &#x0201C;off&#x0201D; state by tonic inhibition from the basal ganglia. Upon disinhibition of a relay, information is transferred from one cortical location to another&#x02014;a form of conditional &#x0201C;gating&#x0201D; of information transfer. For example, the BG might be able to selectively &#x0201C;clamp&#x0201D; particular groups of cortical neurons in a fixed state, while leaving others free to learn and adapt. It could thereby enforce complex training routines, perhaps similar to those used to force the emergence of disentangled representations in (Kulkarni et al., <xref ref-type="bibr" rid="B239">2015</xref>). The idea that the basal ganglia can train the cortex is not new, and already appears to have considerable experimental and anatomical support (Pasupathy and Miller, <xref ref-type="bibr" rid="B341">2005</xref>; Ashby et al., <xref ref-type="bibr" rid="B17">2007</xref>, <xref ref-type="bibr" rid="B18">2010</xref>; Turner and Desmurget, <xref ref-type="bibr" rid="B441">2010</xref>).</p></fn>
<fn id="fn0050"><p><sup>50</sup>Like many brain areas, the hippocampus is richly innervated by a variety of reward-related and other neuromodulatory systems (Verney et al., <xref ref-type="bibr" rid="B450">1985</xref>; Colino and Halliwell, <xref ref-type="bibr" rid="B83">1987</xref>; Hasselmo and Wyble, <xref ref-type="bibr" rid="B172">1997</xref>).</p></fn>
<fn id="fn0051"><p><sup>51</sup>It remains unclear whether place cells take input from the grid cell system or vice versa (Hasselmo, <xref ref-type="bibr" rid="B170">2015</xref>).</p></fn>
<fn id="fn0052"><p><sup>52</sup>Other spatial problems such as mental rotation may require learning architectures specialized for geometric coordinate transformations (Hinton et al., <xref ref-type="bibr" rid="B189">2011</xref>; Jaderberg et al., <xref ref-type="bibr" rid="B207">2015</xref>) or binding mechanisms that support structural, compositional, parametric descriptions of a scene (Hayworth et al., <xref ref-type="bibr" rid="B179">2011</xref>).</p></fn>
<fn id="fn0053"><p><sup>53</sup>There is some direct fMRI evidence for anatomically separate registers representing the contents of different sentence roles in the human brain (Frankland and Greene, <xref ref-type="bibr" rid="B128">2015</xref>), which is suggestive of a possible anatomical binding mechanism, but also consistent with other mechanisms like vector symbolic architectures. More generally, the substrates of symbolic processing in the brain may bear an intimate connection with the representation of objects in working memory in the prefrontal cortex, and specifically with the question of how the PFC represents multiple objects in working memory simultaneously. This question is undergoing extensive study in primates (Warden and Miller, <xref ref-type="bibr" rid="B455">2007</xref>, <xref ref-type="bibr" rid="B456">2010</xref>; Siegel et al., <xref ref-type="bibr" rid="B400">2009</xref>; Rigotti et al., <xref ref-type="bibr" rid="B364">2013</xref>).</p></fn>
<fn id="fn0054"><p><sup>54</sup>There is controversy around claims that recursive syntax is also present in songbirds (Van Heijningen et al., <xref ref-type="bibr" rid="B447">2009</xref>).</p></fn>
<fn id="fn0055"><p><sup>55</sup>The above mechanisms are spontaneous and subconscious. In conscious thought, too, the brain can clearly visit the multiple layers of a program one after the other. We make high-level plans that we fill with lower-level plans. Humans also have memory for their own thought processes. We have some ability to put &#x0201C;on hold&#x0201D; our current state of mind, start a new train of thought, and then come back to our original thought. We also are able to ask, introspectively, whether we have had a given thought before. The neural basis of these processes is unclear, although one may speculate that the hippocampus is involved.</p></fn>
<fn id="fn0056"><p><sup>56</sup>Fluorescent techniques like (Hayashi-Takagi et al., <xref ref-type="bibr" rid="B176">2015</xref>) might be helpful.</p></fn>
<fn id="fn0057"><p><sup>57</sup>The use of structured microcircuits rather than individual neurons as the units of learning can ease the burden on the learning rules possessed by individual neurons, as exemplified by a study implementing Helmholtz machine learning in a network of spiking neurons using conventional plasticity rules (Roudi and Taylor, <xref ref-type="bibr" rid="B373">2015</xref>; Sountsov and Miller, <xref ref-type="bibr" rid="B407">2015</xref>). As a simpler example, the classical problem of how neurons with only one output axon could communicate both activation and error derivatives for backpropagation ceases to be a problem if the unit of optimization is not a single neuron. Similar considerations hold for the issue of weight symmetry, or approximate sign-concordance in the case of feedback alignment (Liao et al., <xref ref-type="bibr" rid="B261">2015</xref>).</p></fn>
<fn id="fn0058"><p><sup>58</sup>Within this framework, networks that adhere to the basic statistics of neural connectivity, electrophysiology and morphology, such as the initial cortical column models from the Blue Brain Project (Markram et al., <xref ref-type="bibr" rid="B289">2015</xref>), would recapitulate some properties of the cortex, but&#x02014;just like untrained neural networks&#x02014;would not spontaneously generate complex functional computation without being subjected to a multi-stage training process, naturalistic sensory data, signals arising from other brain areas and action-driven reinforcement signals.</p></fn>
<fn id="fn0059"><p><sup>59</sup>Not only in applied machine learning, but also in today&#x00027;s most advanced neuro-cognitive models such as SPAUN (Eliasmith et al., <xref ref-type="bibr" rid="B108">2012</xref>; Eliasmith, <xref ref-type="bibr" rid="B105">2013</xref>), the detailed local circuit connectivity is obtained through an optimization process of some kind to achieve a particular functionality. In the case of modern machine learning, training is often done via end-to-end backpropagation through an architecture that is only structured at the level of higher-level &#x0201C;blocks&#x0201D; of units, whereas in SPAUN each block is optimized (Eliasmith and Anderson, <xref ref-type="bibr" rid="B106">2004</xref>) separately according to a procedure that allows the blocks to subsequently be stitched together in a coherent way. Technically, the Neural Engineering Framework (Eliasmith and Anderson, <xref ref-type="bibr" rid="B106">2004</xref>) used in SPAUN uses singular value decomposition, rather than gradient descent, to compute the connections weights as optimal linear decoders. This is possible because of a nonlinear mapping into a high-dimensional space, in which approximating any desired function can be done via a hyperplane regression (Tapson and van Schaik, <xref ref-type="bibr" rid="B429">2013</xref>).</p></fn>
<fn id="fn0060"><p><sup>60</sup>There is a rich tradition of trying to estimate the cost function used by human beings (Ng and Russell, <xref ref-type="bibr" rid="B323">2000</xref>; Finn et al., <xref ref-type="bibr" rid="B122">2016</xref>; Ho and Ermon, <xref ref-type="bibr" rid="B194">2016</xref>). The idea is that we observe (by stipulation) behavior that is optimal for the human&#x00027;s cost function. We can then search for the cost function that makes the observed behavior most probable and simultaneously makes the behaviors that could have been observed, but were not, least probable. Extensions of such approaches could perhaps be used to ask which cost functions the brain is optimizing.</p></fn>
<fn id="fn0061"><p><sup>61</sup>Successes of deep learning are already being used, speculatively, to rationalize features of the brain. It has been suggested that large networks, with many more neurons available than are strictly needed for the target computation, make learning easier (Goodfellow et al., <xref ref-type="bibr" rid="B149">2014b</xref>). In concordance with this, visual cortex appears to be a 100-fold over-complete representation of the retinal output (Lewicki and Sejnowski, <xref ref-type="bibr" rid="B258">2000</xref>). Likewise, it has been suggested that biological neurons stabilized (Turrigiano, <xref ref-type="bibr" rid="B442">2012</xref>) to operate far below their saturating firing rates mirror the successful use of rectified linear units in facilitating the training of artificial neural networks (Roudi and Taylor, <xref ref-type="bibr" rid="B373">2015</xref>). Hinton and others have also suggested a biological motivation (Roudi and Taylor, <xref ref-type="bibr" rid="B373">2015</xref>) for &#x0201C;dropout&#x0201D; regularization (Srivastava et al., <xref ref-type="bibr" rid="B409">2014</xref>), in which a fraction of hidden units is stochastically set to zero during each round of training: such a procedure may correspond to the noisiness of neural spike trains, although other theories interpret spikes as sampling in probabilistic inference (Buesing et al., <xref ref-type="bibr" rid="B62">2011</xref>), or in many other ways. Randomness of spiking has some support in neuroscience (Softky and Koch, <xref ref-type="bibr" rid="B405">1993</xref>), although recent experiments suggest that spike trains in certain areas may be less noisy than previously thought (Hires et al., <xref ref-type="bibr" rid="B190">2015</xref>). The key role of proper initialization in enabling effective gradient descent is an important recent finding (Saxe et al., <xref ref-type="bibr" rid="B382">2013</xref>; Sutskever and Martens, <xref ref-type="bibr" rid="B421">2013</xref>) which may also be reflected by biological mechanisms of neural homeostasis or self-organization that would enforce appropriate initial conditions for learning. Retinal fixation has been tentatively connected with robustness of convolutional networks to adversarial perturbations in images (Luo et al., <xref ref-type="bibr" rid="B269">2015</xref>). But making these speculative claims of biological relevance more rigorous will require researchers to first evaluate <italic>whether</italic> biological neural circuits are performing multi-layer optimization of cost functions in the first place.</p></fn>
<fn id="fn0062"><p><sup>62</sup>It would be interesting to study these questions in specific brain systems. The primary visual cortex, for example, is still only understood very incompletely (Olshausen and Field, <xref ref-type="bibr" rid="B330">2004</xref>). It serves as a key input modality to both the ventral and dorsal visual pathways, one of which seems to specialize in object identity and the other in motion and manipulation. Higher-level areas like STP draw on both streams to perform tasks like complex action recognition. In some models (e.g., Jhuang et al., <xref ref-type="bibr" rid="B212">2007</xref>), both ventral and dorsal streams are structured hierarchically, but the ventral stream primarily makes use of the spatial filtering properties of V1, whereas the dorsal stream primarily makes use of its spatio-<italic>temporal</italic> filtering properties, e.g., temporal frequency filtering by the space-time receptive fields of V1 neurons. Given this, we can ask interesting questions about V1. Within a framework of multilayer optimization, do both dorsal and ventral pathways impose cost functions that help to shape V1&#x00027;s response properties? Or is V1 largely pre-structured by genetics and local self-organization, with different optimization principles in the ventral and dorsal streams only having effects at higher levels of the hierarchy? Or, more likely, is there some interplay between pre-structuring of the V1 circuitry and optimization according to multiple cost functions? Relatedly, what establishes the differing roles of the downstream ventral vs. dorsal cortical areas, and can their differences be attributed to differing cost functions? This relates to ongoing questions about the basic nature of cortical circuitry. For example, DiCarlo et al. (<xref ref-type="bibr" rid="B100">2012</xref>) suggests that visual cortical regions containing on the order of 10000 neurons are locally optimized to perform disentangling of the manifolds corresponding to their local views of the transformations of an object, allowing these manifolds to be linearly separated by readout areas. Yet, DiCarlo et al. (<xref ref-type="bibr" rid="B100">2012</xref>) also emphasizes the possibility that certain computations such as normalization are pre-initialized in the circuitry prior to learning-based optimization.</p></fn>
</fn-group>
</back>
</article>
