<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2016.00073</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Perspective</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Probabilistic Models and Generative Neural Networks: Towards an Unified Framework for Modeling Normal and Impaired Neurocognitive Functions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Testolin</surname> <given-names>Alberto</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/75580/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zorzi</surname> <given-names>Marco</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/9606/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of General Psychology and Center for Cognitive Neuroscience, University of Padova</institution> <country>Padua, Italy</country></aff>
<aff id="aff2"><sup>2</sup><institution>IRCCS San Camillo Neurorehabilitation Hospital</institution> <country>Venice-Lido, Italy</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Eirini Mavritsaki, Birmingham City University, UK</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Marcel Van Gerven, Radboud University Nijmegen, Netherlands; Julien Mayor, The University of Nottingham Malaysia Campus, Malaysia</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Alberto Testolin <email>alberto.testolin&#x00040;unipd.it</email> Marco Zorzi <email>marco.zorzi&#x00040;unipd.it</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>07</month>
<year>2016</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>10</volume>
<elocation-id>73</elocation-id>
<history>
<date date-type="received">
<day>27</day>
<month>05</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>06</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2016 Testolin and Zorzi.</copyright-statement>
<copyright-year>2016</copyright-year>
<copyright-holder>Testolin and Zorzi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution and reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract><p>Connectionist models can be characterized within the more general framework of probabilistic graphical models, which allow to efficiently describe complex statistical distributions involving a large number of interacting variables. This integration allows building more realistic computational models of cognitive functions, which more faithfully reflect the underlying neural mechanisms at the same time providing a useful bridge to higher-level descriptions in terms of Bayesian computations. Here we discuss a powerful class of graphical models that can be implemented as stochastic, generative neural networks. These models overcome many limitations associated with classic connectionist models, for example by exploiting unsupervised learning in hierarchical architectures <italic>(deep networks)</italic> and by taking into account top-down, predictive processing supported by feedback loops. We review some recent cognitive models based on generative networks, and we point out promising research directions to investigate neuropsychological disorders within this approach. Though further efforts are required in order to fill the gap between structured Bayesian models and more realistic, biophysical models of neuronal dynamics, we argue that generative neural networks have the potential to bridge these levels of analysis, thereby improving our understanding of the neural bases of cognition and of pathologies caused by brain damage.</p></abstract>
<kwd-group>
<kwd>connectionist modeling</kwd>
<kwd>unsupervised learning</kwd>
<kwd>deep neural networks</kwd>
<kwd>probabilistic generative models</kwd>
<kwd>computational neuropsychology</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="91"/>
<page-count count="9"/>
<word-count count="6658"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction" id="s1">
<title>Introduction</title>
<p>Despite the enormous progress in the prevention and treatment of neuropsychological disorders, traumatic brain injury and stroke are still among the major causes of adult disability and death (Mathers et al., <xref ref-type="bibr" rid="B56">2008</xref>; Feigin et al., <xref ref-type="bibr" rid="B24">2014</xref>). This social impact highlights the importance of neuropsychological research and the recent thrust in supporting empirical investigations with modern computational tools (Gerstner et al., <xref ref-type="bibr" rid="B29">2012</xref>). In particular, network-based models of brain function conceive cognitive processes as complex phenomena emerging from the simultaneous interaction of many constituent components, and are therefore particularly suited to study the effects of brain damage from a computational perspective (O&#x02019;Reilly and Munakata, <xref ref-type="bibr" rid="B65">2000</xref>).</p>
<p>One of the most successful attempts to ground neuropsychology within a computational framework has been achieved by parallel distributed processing (PDP) models (Rumelhart and McClelland, <xref ref-type="bibr" rid="B77">1986</xref>), which describe cognition as the evolution over time of a system of interconnected units that self-organize according to physical principles. Within this framework, the pattern seen in overt behavior (macroscopic dynamics of the system) reflects the operations of subcognitive processes (microscopic dynamics of the system), such as the propagation of activation and inhibition among simple processing units. A distinguishing feature of PDP models is their ability to adapt to the environment, which allows to simulate behavioral patterns associated with a broad range of cognitive functions and to study how learning mechanisms support cognitive development and knowledge acquisition (e.g., Elman et al., <xref ref-type="bibr" rid="B21">1996</xref>). Crucially, the tight link between structure and function in PDP models allows to investigate how changes in the underlying processing mechanisms are reflected by changes in overt behavior, thereby providing a principled way to simulate neuropsychological disorders following brain damage (e.g., Hinton and Shallice, <xref ref-type="bibr" rid="B39">1991</xref>; Plaut and Shallice, <xref ref-type="bibr" rid="B69">1993</xref>; McClelland et al., <xref ref-type="bibr" rid="B59">1995</xref>).</p>
<p>However, despite the broad range of cognitive functions (and cognitive disorders) investigated through this approach, many PDP models suffer from serious limitations. In particular, connectionist models are often trained in a supervised fashion using error backpropagation, but the assumption that learning is largely discriminative and that an external teaching signal is available at each learning event is implausible from a cognitive perspective (see Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>, for discussion). Moreover, besides the need for labeled patterns, classic PDP models usually entail an over-simplistic, &#x0201C;shallow&#x0201D; processing architecture, involving only one layer of hidden units and strictly feed-forward connectivity. This is in sharp contrast with well-known properties of cortical circuits, which exhibit a hierarchical organization (Felleman and Van Essen, <xref ref-type="bibr" rid="B25">1991</xref>) where information processing relies on both feed-forward and feedback mechanisms (Sillito et al., <xref ref-type="bibr" rid="B78">2006</xref>; Gilbert and Sigman, <xref ref-type="bibr" rid="B31">2007</xref>). Finally, these processing constraints (together with limitations in computational power) have prevented to extend &#x0201C;toy models&#x0201D; into large-scale simulations of neural networks composed by thousands of neurons and millions of connection weights that can be trained using realistic input patterns.</p>
<p>The aim of this article is to describe a new generation of PDP models that address these limitations. In particular, we discuss how they have been exploited for modeling a wide range of neurocognitive functions, and we highlight their potential for simulating neuropsychological deficits.</p>
</sec>
<sec id="s2">
<title>A New Generation of Parallel Distributed Processing Models</title>
<p>Probabilistic graphical models provide a general approach to model the stochastic behavior of a large number of interacting variables, whose relations are efficiently represented using graphical structures (Koller and Friedman, <xref ref-type="bibr" rid="B49">2009</xref>). Notably, many PDP models can be characterized within this probabilistic framework (Jordan and Sejnowski, <xref ref-type="bibr" rid="B45">2001</xref>). In particular, a powerful class of stochastic, recurrent neural networks can be characterized as fully-connected graphical models, where the undirected nature of the edges implies bidirectional flow of information between the nodes (Ackley et al., <xref ref-type="bibr" rid="B1">1985</xref>). This probabilistic interpretation of neural networks provides a useful bridge to more abstract computational descriptions of cognitive processes (Griffiths et al., <xref ref-type="bibr" rid="B34">2008</xref>), suggesting how high-level Bayesian computations might be implemented in neural circuits. Indeed, the problem of finding the best possible interpretation of an ambiguous stimulus can be formalized as an unconscious, statistical inference process. A possible role for recurrent feed-forward/feedback loops in the cerebral cortex might therefore be to integrate top-down, contextual priors with bottom-up, sensory observations, so as to implement concurrent probabilistic inference along the whole cortical hierarchy (Lee and Mumford, <xref ref-type="bibr" rid="B53">2003</xref>; McClelland, <xref ref-type="bibr" rid="B58">2013</xref>).</p>
<sec id="s2-1">
<title>Unsupervised Learning in Generative Neural Networks</title>
<p>Learning in probabilistic graphical models can be framed within two different settings. In <italic>discriminative</italic> learning, the goal is to model only conditional distributions over a set of target variables, whose values are specified by associating an explicit label to each observed pattern. In <italic>generative</italic> learning, instead, the aim is to model the joint distribution of all the variables in the model, thus including also the observed variables. Notably, generative models can be efficiently implemented as stochastic neural networks that learn to reconstruct the sensory input (maximum-likelihood learning) through feedback connections and Hebbian-like learning mechanisms (Hinton, <xref ref-type="bibr" rid="B36">2002</xref>). From a cognitive modeling perspective, these models are appealing because they can build high-level, distributed representations of the data by extracting statistical regularities in a completely unsupervised way (Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>). Moreover, feedback connections have a primary role in generative networks because they carry top-down expectations of the model, which are updated during learning in order to better reflect the observed sensory data (Hinton et al., <xref ref-type="bibr" rid="B40">1995</xref>).</p>
<p>Simple generative networks can be used as building blocks for more complex architectures, such as those used in <italic>deep learning</italic> systems, where the hidden variables of the generative model are hierarchically organized (Hinton and Salakhutdinov, <xref ref-type="bibr" rid="B38">2006</xref>). Hierarchical generative models efficiently structure the representation space by promoting features reuse: simple features extracted at lower levels can be successively combined to create more complex features, which eventually unveil the main causal factors underlying the data distribution (Hinton, <xref ref-type="bibr" rid="B37">2007</xref>). Moreover, these high-level, abstract representations of the sensory data can also easily support supervised read-outs (Testolin et al., <xref ref-type="bibr" rid="B87">2013</xref>; Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>; Figure <xref ref-type="fig" rid="F1">1A</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>(A)</bold> Graphical representation of a hierarchical generative model implemented as a deep neural network. Undirected edges entail bidirectional (recurrent) connections, which are encoded by different weight matrices at each processing layer (<italic>V</italic> represents the set of visible units, while <italic>H<sub>n</sub></italic> represents the set of hidden units at layer <italic>n</italic>). Dotted arrows with blue captions on the side of the hierarchy provide a Bayesian interpretation of bottom-up and top-down processing in terms of conditional probabilities. Multiple classification tasks (directed arrows on top) can be performed by applying supervised read-out modules (e.g., linear classifiers) to the top-level, abstract representations of the model. <bold>(B)</bold> Graphical representation of a sequential generative model implemented as a temporal, recurrent restricted Boltzmann machine (Sutskever et al., <xref ref-type="bibr" rid="B84">2008</xref>; Testolin et al., <xref ref-type="bibr" rid="B88">2016</xref>). At each timestep, directed connections are used to propagate temporal context over time through a hidden-to-hidden weight matrix. Blue captions provide a Bayesian interpretation of temporal prediction in terms of conditional probabilities: to differ from static, hierarchical models, here the activation probability <italic>H<sub>n</sub></italic> of hidden units is conditioned on both the previous hidden state <italic>H<sub>n&#x02212;1</sub></italic> and the current observed evidence <italic>V<sub>n</sub></italic>.</p></caption>
<graphic xlink:href="fncom-10-00073-g0001.tif"/>
</fig>
<p>Generative networks have also been extended to the temporal domain (e.g., Sutskever et al., <xref ref-type="bibr" rid="B84">2008</xref>), where input patterns appear in a precise, sequential order. In this case, statistical inference is performed by considering, besides the current observed evidence, also the history provided by the temporal context, which is propagated through delayed connections (Figure <xref ref-type="fig" rid="F1">1B</xref>). Extracting temporal dependencies is a formidable challenge for the brain (Dehaene et al., <xref ref-type="bibr" rid="B18">2015</xref>), but it leads to more powerful internal models of the environment that can be used to actively predict the sensory stream (Friston, <xref ref-type="bibr" rid="B27">2010</xref>; Clark, <xref ref-type="bibr" rid="B11">2013</xref>). The ability to anticipate external events is also crucial for attentional mechanisms, which efficiently select sensory information according to top-down expectations and current goals (Corbetta and Shulman, <xref ref-type="bibr" rid="B12">2002</xref>). In this respect, generative models allow to conceive attention as an intrinsic property of bidirectional processing networks (Casarotti et al., <xref ref-type="bibr" rid="B10">2012</xref>) and to use information theoretic measures to operationalize properties like novelty/surprise in terms of discrepancy between model&#x02019;s expectation and observed sensory evidence (Itti and Baldi, <xref ref-type="bibr" rid="B42">2009</xref>).</p>
<p>Finally, deep learning systems coupled with reinforcement learning algorithms have recently obtained state-of-the-art performance in extremely challenging cognitive tasks, for example by learning to play videogames at human-level (Mnih et al., <xref ref-type="bibr" rid="B61">2015</xref>) or by defeating professional players on difficult board games (Silver et al., <xref ref-type="bibr" rid="B79">2016</xref>). This powerful learning modality takes into account the effects of actions on the environment without requiring an explicit supervision signal, and therefore would constitute a cognitively (Botvinick et al., <xref ref-type="bibr" rid="B3">2009</xref>) and biologically (Gl&#x000E4;scher et al., <xref ref-type="bibr" rid="B32">2010</xref>) plausible way to couple unsupervised deep learning with goal-directed behavior.</p>
</sec>
<sec id="s2-2">
<title>Recent Neurocognitive Models</title>
<p>In the domain of numerical cognition, unsupervised deep learning has been successfully used to show how visual numerosity could emerge as a statistical property of images containing a variable number of items (Stoianov and Zorzi, <xref ref-type="bibr" rid="B81">2012</xref>; Figure <xref ref-type="fig" rid="F2">2A</xref>). Numerosity detectors developed by the network had response profiles resembling those of monkey parietal neurons (Roitman et al., <xref ref-type="bibr" rid="B75">2007</xref>), and supported numerosity estimation with the same behavioral signature shown by humans and animals. A subsequent study simulated typical and atypical developmental trajectories through incremental learning and manipulation of the computational resources (i.e., number of hidden units) of the generative model (Stoianov and Zorzi, <xref ref-type="bibr" rid="B82">2013</xref>), in line with the reduced gray matter density in the intraparietal sulcus observed in dyscalculic subjects (Rotzer et al., <xref ref-type="bibr" rid="B76">2008</xref>). Generative networks have also been used to model learning of arithmetic facts as joint distributions of operands and results, and to simulate acquired acalculia (Stoianov et al., <xref ref-type="bibr" rid="B83">2004</xref>; Zorzi et al., <xref ref-type="bibr" rid="B91">2005</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>(A)</bold> Graphical representation of the numerosity perception model of Stoianov and Zorzi (<xref ref-type="bibr" rid="B81">2012</xref>). A hierarchical generative model was first trained on a large set of realistic images containing visual sets with a varying number of objects. A linear read-out layer was then trained on the top-level internal representations on a numerosity comparison task. <bold>(B)</bold> Graphical representation of the letter perception model of Testolin et al. (<xref ref-type="bibr" rid="B86">under review</xref>). The bottom layer of the network receives the sensory signal encoded as gray-level activations of image pixels. Low-level processing occurring in the retina and thalamus is simulated using a biologically inspired whitening algorithm that captures local spatial correlations in the image and serves as a contrast-normalization step. Following generative learning on a set of patches of natural images, neurons in the first hidden layer (V1) encoded simple visual features which constitute a basic dictionary describing the statistical distribution of pixel intensities observed in natural environments. Specific learning about letters was then introduced in the model by training a second hidden layer with images containing a variety of uppercase letters. Neurons in the second hidden layer (V2/V4) learned to combine V1 features to represent letter fragments and in some cases, whole letter shapes. A linear read-out layer (OTS) was then trained on the top-level internal representations in order to decode letter classes. <bold>(C)</bold> Different types of high-level features (receptive fields) emerging from unsupervised deep learning. On the left side, a prototypical face (Le et al., <xref ref-type="bibr" rid="B51">2012</xref>), a prototypical handwritten digit (Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>) and a prototypical printed letter (Testolin et al., under <xref ref-type="bibr" rid="B86">review</xref>). In the middle panel, population activity of number-sensitive hidden neurons (mean activation value) as a function of number of objects in the display (Stoianov and Zorzi, <xref ref-type="bibr" rid="B81">2012</xref>). In the right panel, a prototypical hidden neuron with a retinotopic receptive field exhibiting gain modulation (De Filippo De Grazia et al., <xref ref-type="bibr" rid="B14">2012</xref>).</p></caption>
<graphic xlink:href="fncom-10-00073-g0002.tif"/>
</fig>
<p>Another major cognitive domain that has been modeled within this framework is that of visual object recognition, where the hierarchical representations emerging in deep networks show remarkable similarities with those recorded in the ventral visual pathway of the human brain (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B35">2015</xref>). Unsupervised deep learning has also been recently applied to model human-like letter perception (Testolin et al., under <xref ref-type="bibr" rid="B86">review</xref>), where visual primitives extracted from natural scenes are later recycled for learning letters (Figure <xref ref-type="fig" rid="F2">2B</xref>) thereby supporting the hypothesis that the shape of visual symbols has been culturally selected to match the statistical structure found in our visual environment (Dehaene and Cohen, <xref ref-type="bibr" rid="B17">2007</xref>). Perception of single letters can also be extended to model visual word recognition (Di Bono and Zorzi, <xref ref-type="bibr" rid="B20">2013</xref>; Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>), and a temporal version of the model has been used to learn the statistical structure of letter sequences and to simulate spontaneous generation of words and pseudowords (Testolin et al., <xref ref-type="bibr" rid="B88">2016</xref>). These generative networks can be used as building blocks to develop more realistic models of visual word recognition, paving the way for full-blown simulations of orthographic learning in both normal and atypical development, as well as of the impairments caused by brain damage, such as pure alexia (Plaut and Behrmann, <xref ref-type="bibr" rid="B68">2011</xref>).</p>
<p>Generative neural networks have also been used to study space coding for sensorimotor transformations and multisensory integration (De Filippo De Grazia et al., <xref ref-type="bibr" rid="B14">2012</xref>). The authors found that receptive fields reflecting those observed in the monkey posterior parietal cortex can emerge through unsupervised learning (Figure <xref ref-type="fig" rid="F2">2C</xref>), suggesting that gain modulation is an efficient coding strategy to integrate visual and postural information toward the generation of motor commands even though learning does not involve any explicit coordinate transformation. Notably, models of sensorimotor transformations building upon stipulated gain modulation have been used to account for visuospatial attention (Casarotti et al., <xref ref-type="bibr" rid="B10">2012</xref>) and neuropsychological deficits like hemineglect (Pouget and Driver, <xref ref-type="bibr" rid="B70">2000</xref>). Therefore, a promising venue for research will be to investigate these phenomena within the emergentist framework of deep generative networks.</p>
</sec>
<sec id="s2-3">
<title>Implications for Neuropsychology</title>
<p>From a neuropsychological modeling perspective, we discuss below a series of methodological advantages that this new generation of PDP models offers over more traditional connectionist models.</p>
<sec id="s2-3-1">
<title>Localized Damage Within a Hierarchical Architecture</title>
<p>The structured architecture of deep learning models allows to more carefully simulate cognitive deficits caused by localized brain damage, which may affect a specific representation level. Indeed, deep networks exploit multiple levels of representation, where low-level features are gradually combined in order to produce more abstract representations of the sensory data. For example, in the domain of visual object recognition, unsupervised deep learning can lead to the emergence of extremely high-level visual features (Figure <xref ref-type="fig" rid="F2">2C</xref>), such as those representing prototypical faces (Le et al., <xref ref-type="bibr" rid="B51">2012</xref>). By applying selective lesions to these models, we could assess the effect of damage to specific cortical regions, ranging from early visual processing to higher-level extrastriate areas, up to more anterior, associative areas. This would allow to simulate various forms of visual agnosia (Farah, <xref ref-type="bibr" rid="B23">2004</xref>) and investigate the emergence of category-specific deficits (Humphreys and Forde, <xref ref-type="bibr" rid="B41">2001</xref>). Most notably, the realistic scale of these models allows to evaluate the effect of damage using the same type of stimuli employed in patients&#x02019; testing (e.g., standardized pictures of Snodgrass and Vanderwart, <xref ref-type="bibr" rid="B80">1980</xref>).</p>
</sec>
<sec id="s2-3-2">
<title>Multiple Connection Pathways and Multimodal Learning</title>
<p>Deep learning architectures can also be used to simulate selective damage to specific connection pathways. For example, Cappelletti et al. (<xref ref-type="bibr" rid="B9">2014</xref>) simulated the declined performance of elderly population in numerosity comparison using the model of Stoianov and Zorzi (<xref ref-type="bibr" rid="B81">2012</xref>). Stochastic decay was applied to synaptic strengths to investigate two different types of impairment: a global degradation involving all network synapses, and a more selective degradation involving only the inhibitory synapses of a specific processing layer. The specific impairment of inhibition caused a large decrease of performance on stimuli in which irrelevant, continuous visual features competed with numerosity, mirroring the empirical data; conversely, the decline in performance following global impairment was identical across conditions. In line with an inhibition deficit hypothesis, the authors concluded that reduced inhibition of irrelevant information is critical to explain the specific pattern of impaired performance observed in aging. Selective damaging of connection pathways is also interesting in the context of multimodal deep learning (Ngiam et al., <xref ref-type="bibr" rid="B64">2011</xref>). For example, learning a shared representation for arithmetic facts presented in both semantic and symbolic formats produces two different subnetworks that can be selectively damaged to simulate different patterns of acquired acalculia (Stoianov et al., <xref ref-type="bibr" rid="B83">2004</xref>).</p>
</sec>
<sec id="s2-3-3">
<title>Balance Between Bottom-Up and Top-Down Processing</title>
<p>The prominent role of feedback connections in generative networks also allows to simulate unbalancing between top-down and bottom-up integration mechanisms, which are thought to underlie positive symptoms commonly observed in psychiatric disorders (Manford and Andermann, <xref ref-type="bibr" rid="B55">1998</xref>). Hierarchical generative models have been used to simulate visual hallucinations in the Charles Bonnet syndrome (Reichert et al., <xref ref-type="bibr" rid="B74">2013</xref>), suggesting that impaired homeostatic regulation of feed-forward and feedback neuronal activity might be responsible for a wide range of symptoms observed in patients.</p>
</sec>
<sec id="s2-3-4">
<title>Noise Might not Always be Detrimental</title>
<p>Another major difference with respect to traditional connectionist models relates to the role of noise in simulating brain damage. Injection of noise in the activation of hidden units has been often used as a way to simulate brain damage by disrupting internal representations (e.g., Joanisse and Seidenberg, <xref ref-type="bibr" rid="B44">1999</xref>). In stochastic models, instead, adding noise allows for a more efficient exploration of the network state space and helps settling into more stable attractors (Kirkpatrick et al., <xref ref-type="bibr" rid="B48">1983</xref>). This is compatible with the hypothesis that neuronal noise has a key computational role in the brain, for example by keeping it in a &#x0201C;metastable&#x0201D; state that facilitates flexible settling into the most appropriate configuration (Kelso, <xref ref-type="bibr" rid="B46">2012</xref>). Notably, this might also explain how structured fluctuations of brain activity, such as those observed during resting state, could emerge from noise-driven explorations of oscillatory states (Deco et al., <xref ref-type="bibr" rid="B15">2013</xref>).</p>
</sec>
<sec id="s2-3-5">
<title>From Toy Models to Realistic, Large-Scale Simulations</title>
<p>Finally, the appeal of generative neural networks has long been hindered by their high computational complexity. This has been radically changed by recent advances in parallel computing architectures, which allow to efficiently simulate large-scale neural networks composed by thousands of neurons (Raina et al., <xref ref-type="bibr" rid="B72">2009</xref>; Testolin et al., <xref ref-type="bibr" rid="B87">2013</xref>) that can be trained and tested using the same type of stimuli adopted in empirical research (Stoianov and Zorzi, <xref ref-type="bibr" rid="B81">2012</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B35">2015</xref>). This increased realism will have important benefits for neuropsychological modeling, which traditionally relied on small-scale, &#x0201C;toy-models&#x0201D; that cannot reproduce realistic experimental settings.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<title>Perspectives and Future Challenges</title>
<p>An important challenge will be to more closely link generative networks with structured Bayesian models (Ghahramani, <xref ref-type="bibr" rid="B30">2015</xref>), which can successfully simulate a wide variety of high-level cognitive functions ranging from one-shot learning (Lake et al., <xref ref-type="bibr" rid="B50">2015</xref>) to inferring causal relations, categories and hidden properties of objects, and meanings of words (see Tenenbaum et al., <xref ref-type="bibr" rid="B85">2011</xref>, for discussion).</p>
<p>At the opposite end, bridging generative networks to more realistic neuronal models that incorporate biophysical details is another major challenge. The popularity of <italic>supervised</italic> deep learning both in academic and industry research (LeCun et al., <xref ref-type="bibr" rid="B52">2015</xref>) has offset research on generative models, which nevertheless entail a more psychologically-plausible learning regimen as well as more biologically-plausible processing mechanisms (Zorzi et al., <xref ref-type="bibr" rid="B92">2013</xref>; Cox and Dean, <xref ref-type="bibr" rid="B13">2014</xref>). We believe, however, that generative networks will have an increasingly central role in neurocognitive modeling because they can simulate both evoked (feed-forward) and intrinsic (feedback) brain activity, where top-down mechanisms generate and maintain active representations that are modulated, rather than determined, by sensory information (Fiser et al., <xref ref-type="bibr" rid="B26">2010</xref>). In this respect, although the classical approach in cognitive neuroscience has been to study neuronal responses to stimuli during task performance, the importance of intrinsic activity in shaping brain dynamics is now widely recognized (Raichle, <xref ref-type="bibr" rid="B71">2015</xref>). Accordingly, spontaneous activity might not reflect trivial noisy fluctuations, because it is organized into clear spatiotemporal profiles that might reflect the functional architecture of the brain (Greicius et al., <xref ref-type="bibr" rid="B33">2003</xref>; Buckner et al., <xref ref-type="bibr" rid="B6">2008</xref>). The fact that intrinsic activity persists during sleep suggests its potential role in development and plasticity (Raichle, <xref ref-type="bibr" rid="B71">2015</xref>), which is in line with previous attempts to characterize learning in generative networks as being driven by &#x0201C;wake&#x0201D; and &#x0201C;sleep&#x0201D; phases (Hinton et al., <xref ref-type="bibr" rid="B40">1995</xref>). Nevertheless, resting activity is likely supported by dynamics emerging from synchronous oscillations of different brain areas over multiple frequency bands (Engel et al., <xref ref-type="bibr" rid="B22">2001</xref>; Varela et al., <xref ref-type="bibr" rid="B89">2001</xref>), but PDP models usually adopt processing units that are characterized by a single, real value representing the average activity of a neural ensemble. This implies that potentially important phase relations between spikes are completely lost. A possible way to address this limitation could be to integrate generative networks with spiking models, which can also perform near-optimal Bayesian inference (Rao, <xref ref-type="bibr" rid="B73">2004</xref>; Ma et al., <xref ref-type="bibr" rid="B54">2006</xref>; Deneve, <xref ref-type="bibr" rid="B19">2008</xref>) or implement efficient belief propagation schemes in generic graphical models (Pecevski et al., <xref ref-type="bibr" rid="B67">2011</xref>). Alternatively, networks of spiking neurons can perform probabilistic inference, thereby emulating Boltzmann machines, using an efficient but biologically realistic sampling scheme that explains many functional aspects of low-level brain dynamics, such as refractory mechanisms and finite durations of postsynaptic potentials (Buesing et al., <xref ref-type="bibr" rid="B7">2011</xref>). Moreover, related models have shown how maximum-likelihood learning might occur in this type of networks by exploiting spike-timing dependent plasticity, which could be facilitated by other physiological mechanisms such as background oscillations and synchronous activity (Nessler et al., <xref ref-type="bibr" rid="B62">2013</xref>). Notably, there have been other attempts to integrate models of spiking neurons with coarser mean-field models and neural masses, with the aim of providing multi-scale dynamical models of large-scale brain networks (Deco et al., <xref ref-type="bibr" rid="B16">2008</xref>; Mavritsaki et al., <xref ref-type="bibr" rid="B57">2011</xref>). Although these models are less easily interpretable in terms of high-level Bayesian learning and computation, they provide a more direct link to the vast amount of empirical data provided by modern neuroscience methods (e.g., Jirsa et al., <xref ref-type="bibr" rid="B43">2010</xref>).</p>
<p>Finally, a largely unexplored research frontier would be to study PDP models using the powerful analytical techniques developed by network science (Albert and Barabasi, <xref ref-type="bibr" rid="B2">2002</xref>; Newman, <xref ref-type="bibr" rid="B63">2010</xref>), which are rapidly becoming a standard tool in neuroscience research (e.g., Bullmore and Sporns, <xref ref-type="bibr" rid="B8">2009</xref>; Bressler and Menon, <xref ref-type="bibr" rid="B5">2010</xref>; Medaglia et al., <xref ref-type="bibr" rid="B60">2015</xref>). This would allow to more precisely characterize the relationship between structure and function in complex, self-organizing networks: indeed, in PDP models the initial processing architecture is fairly generic (e.g., for the restricted Boltzmann machine, a fully-connected bipartite graph with uniform random connections), and complex structural patterns gradually emerge as a product of learning. To the best of our knowledge, it is still unknown whether the emergent structure exhibits organizational principles that match those observed in brain networks, such as small-worldness and partial segregation into motifs (Park and Friston, <xref ref-type="bibr" rid="B66">2013</xref>). Notably, it has also been shown that a resilience index of complex networks can in fact be measured using a universal resilience function, thereby unveiling the network characteristics that can enhance or diminish its robustness to damage and external perturbations (Gao et al., <xref ref-type="bibr" rid="B28">2016</xref>). This surprising discovery could have a profound impact on neuropsychology, because it might allow to better understand how to improve fault-tolerance in neuronal networks, and how to more effectively recover network functions after damage.</p>
<p>In conclusion, we believe that stochastic, generative neural networks provide a unique interface between high-level descriptions of cognitive functions in terms of structured Bayesian computations and low-level, mechanistic explanations based on dynamical systems theory and simulations of networks whose connectivity and processing mechanisms can be constrained by neurobiological evidence. Such an integrated framework would allow building computational models spanning many levels of detail, capable of predicting salient aspects of behavior at varying levels of resolution at the same time guaranteeing interpretability according to different levels of abstractions (Gerstner et al., <xref ref-type="bibr" rid="B29">2012</xref>). If this ambitious enterprise will succeed (see <xref ref-type="boxed-text" rid="BX1">Box 1</xref> for a list of outstanding research questions) we would have the most valuable tools to understand how neuronal processes support complex behavior and cognition, how brain damage impairs performance, and how to devise intervention strategies to improve recovery of function.</p>
<boxed-text id="BX1" position="float">
<label>Box 1</label>
<title>Outstanding Questions</title>
<list list-type="bullet">
<list-item><p>Current deep learning research is mostly focused on <italic>supervised</italic> learning and feed-forward convolutional networks trained with error backpropagation (LeCun et al., <xref ref-type="bibr" rid="B52">2015</xref>), which have also been used to model cortical processing (e.g., Khaligh-Razavi and Kriegeskorte, <xref ref-type="bibr" rid="B47">2014</xref>). How well do generative/recurrent vs. discriminative/feed-forward models compare with respect to simulating neurophysiological data and the effect of network damage?</p></list-item>
<list-item><p>Feature detectors emerging in deep networks can be extremely complex and specialized. How does this relate to the theoretical debate on localist vs. distributed representations (e.g., Bowers, <xref ref-type="bibr" rid="B4">2009</xref>)? Is it possible to learn a form of explicit, localistic coding that retains the advantages provided by distributed representations? What is the theoretical implication for computational modeling in neuropsychology?</p></list-item>
<list-item><p>Is it possible to simulate the emergence of brain-like structural properties, such as small-worldness and rich-club organization, by starting from a general deep learning architecture? Do we need to include additional constraints (e.g., topological, metabolic)? How do learning regularizers (e.g., sparsity, weight decay, drop-out) compare with respect to organizational principles of biological neuronal networks?</p></list-item>
<list-item><p>Can we improve lesioning studies in PDP models by taking into account structural and functional properties of the network? Could deep learning systems exhibit the same universal resilience patterns observed in other types of complex networks (Gao et al., <xref ref-type="bibr" rid="B28">2016</xref>)?</p></list-item>
</list>
</boxed-text>
</sec>
<sec id="s4">
<title>Author Contributions</title>
<p>AT and MZ equally contributed to the conception and writing of the manuscript. AT and MZ are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.</p>
</sec>
<sec id="s5">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>This research was supported by grants from the European Research Council (No. 210922) and by the University of Padova (Strategic Grant NEURAT) to MZ.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ackley</surname> <given-names>D.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>1985</year>). <article-title>A learning algorithm for Boltzmann machines</article-title>. <source>Cogn. Sci.</source> <volume>9</volume>, <fpage>147</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1207/s15516709cog0901_7</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Albert</surname> <given-names>R.</given-names></name> <name><surname>Barabasi</surname> <given-names>A. L.</given-names></name></person-group> (<year>2002</year>). <article-title>Statistical mechanics of complex networks</article-title>. <source>Rev. Mod. Phys.</source> <volume>74</volume>, <fpage>47</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1103/RevModPhys.74.47</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Botvinick</surname> <given-names>M. M.</given-names></name> <name><surname>Niv</surname> <given-names>Y.</given-names></name> <name><surname>Barto</surname> <given-names>A. C.</given-names></name></person-group> (<year>2009</year>). <article-title>Hierarchically organized behavior and its neural foundations: a reinforcement learning perspective</article-title>. <source>Cognition</source> <volume>113</volume>, <fpage>262</fpage>&#x02013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2008.08.011</pub-id><pub-id pub-id-type="pmid">18926527</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bowers</surname> <given-names>J. S.</given-names></name></person-group> (<year>2009</year>). <article-title>On the biological plausibility of grandmother cells: implications for neural network theories in psychology and neuroscience</article-title>. <source>Psychol. Rev.</source> <volume>116</volume>, <fpage>220</fpage>&#x02013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.1037/a0014462</pub-id><pub-id pub-id-type="pmid">19159155</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bressler</surname> <given-names>S. L.</given-names></name> <name><surname>Menon</surname> <given-names>V.</given-names></name></person-group> (<year>2010</year>). <article-title>Large-scale brain networks in cognition: emerging methods and principles</article-title>. <source>Trends Cogn. Sci.</source> <volume>14</volume>, <fpage>277</fpage>&#x02013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.04.004</pub-id><pub-id pub-id-type="pmid">20493761</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buckner</surname> <given-names>R. L.</given-names></name> <name><surname>Andrews-Hanna</surname> <given-names>J. R.</given-names></name> <name><surname>Schacter</surname> <given-names>D. L.</given-names></name></person-group> (<year>2008</year>). <article-title>The brain&#x02019;s default network: anatomy, function and relevance to disease</article-title>. <source>Ann. N Y Acad. Sci.</source> <volume>1124</volume>, <fpage>1</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1196/annals.1440.011</pub-id><pub-id pub-id-type="pmid">18400922</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Bill</surname> <given-names>J.</given-names></name> <name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>Neural dynamics as sampling: a model for stochastic computation in recurrent networks of spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1002211</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002211</pub-id><pub-id pub-id-type="pmid">22096452</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bullmore</surname> <given-names>E.</given-names></name> <name><surname>Sporns</surname> <given-names>O.</given-names></name></person-group> (<year>2009</year>). <article-title>Complex brain networks: graph theoretical analysis of structural and functional systems</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>10</volume>, <fpage>186</fpage>&#x02013;<lpage>198</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2575</pub-id><pub-id pub-id-type="pmid">19190637</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cappelletti</surname> <given-names>M.</given-names></name> <name><surname>Didino</surname> <given-names>D.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Number skills are maintained in healthy ageing</article-title>. <source>Cogn. Psychol.</source> <volume>69</volume>, <fpage>25</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogpsych.2013.11.004</pub-id><pub-id pub-id-type="pmid">24423632</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casarotti</surname> <given-names>M.</given-names></name> <name><surname>Lisi</surname> <given-names>M.</given-names></name> <name><surname>Umilt&#x000E0;</surname> <given-names>C.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Paying attention through eye movements: a computational investigation of the premotor theory of spatial attention</article-title>. <source>J. Cogn. Neurosci.</source> <volume>24</volume>, <fpage>1519</fpage>&#x02013;<lpage>1531</lpage>. <pub-id pub-id-type="doi">10.1162/jocn_a_00231</pub-id><pub-id pub-id-type="pmid">22452561</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Whatever next? Predictive brains, situated agents and the future of cognitive science</article-title>. <source>Behav. Brain Sci.</source> <volume>36</volume>, <fpage>181</fpage>&#x02013;<lpage>204</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X12000477</pub-id><pub-id pub-id-type="pmid">23663408</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corbetta</surname> <given-names>M.</given-names></name> <name><surname>Shulman</surname> <given-names>G. L.</given-names></name></person-group> (<year>2002</year>). <article-title>Control of goal-directed and stimulus-driven attention in the brain</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>3</volume>, <fpage>201</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1038/nrn755</pub-id><pub-id pub-id-type="pmid">11994752</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. D.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural networks and neuroscience-inspired computer vision</article-title>. <source>Curr. Biol.</source> <volume>24</volume>, <fpage>R921</fpage>&#x02013;<lpage>R929</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2014.08.026</pub-id><pub-id pub-id-type="pmid">25247371</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>De Filippo De Grazia</surname> <given-names>M.</given-names></name> <name><surname>Cutini</surname> <given-names>S.</given-names></name> <name><surname>Lisi</surname> <given-names>M.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Space coding for sensorimotor transformations can emerge through unsupervised learning</article-title>. <source>Cogn. Process.</source> <volume>13</volume>, <fpage>S141</fpage>&#x02013;<lpage>S146</lpage>. <pub-id pub-id-type="doi">10.1007/s10339-012-0478-4</pub-id><pub-id pub-id-type="pmid">22802037</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>Jirsa</surname> <given-names>V. K.</given-names></name> <name><surname>McIntosh</surname> <given-names>A. R.</given-names></name></person-group> (<year>2013</year>). <article-title>Resting brains never rest: computational insights into potential cognitive architectures</article-title>. <source>Trends Neurosci.</source> <volume>36</volume>, <fpage>268</fpage>&#x02013;<lpage>274</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2013.03.001</pub-id><pub-id pub-id-type="pmid">23561718</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>Jirsa</surname> <given-names>V. K.</given-names></name> <name><surname>Robinson</surname> <given-names>P. A.</given-names></name> <name><surname>Breakspear</surname> <given-names>M.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2008</year>). <article-title>The dynamic brain: from spiking neurons to neural masses and cortical fields</article-title>. <source>PLoS Comput. Biol.</source> <volume>4</volume>:<fpage>e1000092</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000092</pub-id><pub-id pub-id-type="pmid">18769680</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Cohen</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). <article-title>Cultural recycling of cortical maps</article-title>. <source>Neuron</source> <volume>56</volume>, <fpage>384</fpage>&#x02013;<lpage>398</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2007.10.004</pub-id><pub-id pub-id-type="pmid">17964253</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Meyniel</surname> <given-names>F.</given-names></name> <name><surname>Wacongne</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Pallier</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>The neural representation of sequences: from transition probabilities to algebraic patterns and linguistic trees</article-title>. <source>Neuron</source> <volume>88</volume>, <fpage>2</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.09.019</pub-id><pub-id pub-id-type="pmid">26447569</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deneve</surname> <given-names>S.</given-names></name></person-group> (<year>2008</year>). <article-title>Bayesian spiking neurons I: inference</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>91</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.20.1.91</pub-id><pub-id pub-id-type="pmid">18045002</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Di Bono</surname> <given-names>M. G.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Deep generative learning of location-invariant visual word recognition</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>635</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00635</pub-id><pub-id pub-id-type="pmid">24065939</pub-id></citation></ref>
<ref id="B21"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name> <name><surname>Bates</surname> <given-names>E.</given-names></name> <name><surname>Johnson</surname> <given-names>M.</given-names></name> <name><surname>Karmiloff-smith</surname> <given-names>A.</given-names></name> <name><surname>Parisi</surname> <given-names>D.</given-names></name> <name><surname>Plunkett</surname> <given-names>K.</given-names></name></person-group> (<year>1996</year>). <source>Rethinking Innateness: A Connectionist Perspective on Development.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Engel</surname> <given-names>A. K.</given-names></name> <name><surname>Fries</surname> <given-names>P.</given-names></name> <name><surname>Singer</surname> <given-names>W.</given-names></name></person-group> (<year>2001</year>). <article-title>Dynamic predictions: oscillations and synchrony in top-down processing</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>2</volume>, <fpage>704</fpage>&#x02013;<lpage>716</lpage>. <pub-id pub-id-type="doi">10.1038/35094565</pub-id><pub-id pub-id-type="pmid">11584308</pub-id></citation></ref>
<ref id="B23"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Farah</surname> <given-names>M. J.</given-names></name></person-group> (<year>2004</year>). <source>Visual Agnosia.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feigin</surname> <given-names>V. L.</given-names></name> <name><surname>Forouzanfar</surname> <given-names>M. H.</given-names></name> <name><surname>Krishnamurthi</surname> <given-names>R.</given-names></name> <name><surname>Mensah</surname> <given-names>G. A.</given-names></name> <name><surname>Connor</surname> <given-names>M.</given-names></name> <name><surname>Bennett</surname> <given-names>D. A.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Global and regional burden of stroke during 1990&#x02013;2010: findings from the Global Burden of Disease Study 2010</article-title>. <source>Lancet</source> <volume>383</volume>, <fpage>245</fpage>&#x02013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1016/s0140-6736(13)61953-4</pub-id><pub-id pub-id-type="pmid">24449944</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Felleman</surname> <given-names>D. J.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name></person-group> (<year>1991</year>). <article-title>Distributed hierarchical processing in the primate cerebral cortex</article-title>. <source>Cereb. Cortex</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/1.1.1</pub-id><pub-id pub-id-type="pmid">1822724</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiser</surname> <given-names>J.</given-names></name> <name><surname>Berkes</surname> <given-names>P.</given-names></name> <name><surname>Orb&#x000E1;n</surname> <given-names>G.</given-names></name> <name><surname>Lengyel</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Statistically optimal perception and learning: from behavior to neural representations</article-title>. <source>Trends Cogn. Sci.</source> <volume>14</volume>, <fpage>119</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.01.003</pub-id><pub-id pub-id-type="pmid">20153683</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2010</year>). <article-title>The free-energy principle: a unified brain theory?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>11</volume>, <fpage>127</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2787</pub-id><pub-id pub-id-type="pmid">20068583</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Barzel</surname> <given-names>B.</given-names></name> <name><surname>Barab&#x000E1;si</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2016</year>). <article-title>Universal resilience patterns in complex networks</article-title>. <source>Nature</source> <volume>530</volume>, <fpage>307</fpage>&#x02013;<lpage>312</lpage>. <pub-id pub-id-type="doi">10.1038/nature16948</pub-id><pub-id pub-id-type="pmid">26887493</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Sprekeler</surname> <given-names>H.</given-names></name> <name><surname>Deco</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>Theory and simulation in neuroscience</article-title>. <source>Science</source> <volume>338</volume>, <fpage>60</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1126/science.1227356</pub-id><pub-id pub-id-type="pmid">23042882</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghahramani</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>Probabilistic machine learning and artificial intelligence</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>452</fpage>&#x02013;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.1038/nature14541</pub-id><pub-id pub-id-type="pmid">26017444</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilbert</surname> <given-names>C. D.</given-names></name> <name><surname>Sigman</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>Brain states: top-down influences in sensory processing</article-title>. <source>Neuron</source> <volume>54</volume>, <fpage>677</fpage>&#x02013;<lpage>696</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2007.05.019</pub-id><pub-id pub-id-type="pmid">17553419</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gl&#x000E4;scher</surname> <given-names>J.</given-names></name> <name><surname>Daw</surname> <given-names>N. D.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>O&#x02019;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2010</year>). <article-title>States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning</article-title>. <source>Neuron</source> <volume>66</volume>, <fpage>585</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2010.04.016</pub-id><pub-id pub-id-type="pmid">20510862</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greicius</surname> <given-names>M. D.</given-names></name> <name><surname>Krasnow</surname> <given-names>B.</given-names></name> <name><surname>Reiss</surname> <given-names>A. L.</given-names></name> <name><surname>Menon</surname> <given-names>V.</given-names></name></person-group> (<year>2003</year>). <article-title>Functional connectivity in the resting brain: a network analysis of the default mode hypothesis</article-title>. <source>Proc. Natl. Acad. Sci. U S A</source> <volume>100</volume>, <fpage>253</fpage>&#x02013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0135058100</pub-id><pub-id pub-id-type="pmid">12506194</pub-id></citation></ref>
<ref id="B34"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Griffiths</surname> <given-names>T. L.</given-names></name> <name><surname>Kemp</surname> <given-names>C.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2008</year>). &#x0201C;<article-title>Bayesian models of cognition</article-title>,&#x0201D; in <source>Cambridge Handbook of Computational Cognitive Modeling</source>, ed. <person-group person-group-type="editor"><name><surname>Sun</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>), <fpage>59</fpage>&#x02013;<lpage>100</lpage>.</citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>10005</fpage>&#x02013;<lpage>10014</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5023-14.2015</pub-id><pub-id pub-id-type="pmid">26157000</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2002</year>). <article-title>Training products of experts by minimizing contrastive divergence</article-title>. <source>Neural Comput.</source> <volume>14</volume>, <fpage>1771</fpage>&#x02013;<lpage>1800</lpage>. <pub-id pub-id-type="doi">10.1162/089976602760128018</pub-id><pub-id pub-id-type="pmid">12180402</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2007</year>). <article-title>Learning multiple layers of representation</article-title>. <source>Trends Cogn. Sci.</source> <volume>11</volume>, <fpage>428</fpage>&#x02013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2007.09.004</pub-id><pub-id pub-id-type="pmid">17921042</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Frey</surname> <given-names>B.</given-names></name> <name><surname>Neal</surname> <given-names>R. M.</given-names></name></person-group> (<year>1995</year>). <article-title>The &#x0201C;wake-sleep&#x0201D; algorithm for unsupervised neural networks</article-title>. <source>Science</source> <volume>268</volume>, <fpage>1158</fpage>&#x02013;<lpage>1161</lpage>. <pub-id pub-id-type="doi">10.1126/science.7761831</pub-id><pub-id pub-id-type="pmid">7761831</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2006</year>). <article-title>Reducing the dimensionality of data with neural networks</article-title>. <source>Science</source> <volume>313</volume>, <fpage>504</fpage>&#x02013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id><pub-id pub-id-type="pmid">16873662</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Shallice</surname> <given-names>T.</given-names></name></person-group> (<year>1991</year>). <article-title>Lesioning an attractor network: investigations of acquired dyslexia</article-title>. <source>Psychol. Rev.</source> <volume>98</volume>, <fpage>74</fpage>&#x02013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295x.98.1.74</pub-id><pub-id pub-id-type="pmid">2006233</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Humphreys</surname> <given-names>G. W.</given-names></name> <name><surname>Forde</surname> <given-names>E. M.</given-names></name></person-group> (<year>2001</year>). <article-title>Hierarchies, similarity and interactivity in object recognition: &#x0201C;category-specific&#x0201D; neuropsychological deficits</article-title>. <source>Behav. Brain Sci.</source> <volume>24</volume>, <fpage>453</fpage>&#x02013;<lpage>476</lpage>; discussion <fpage>476</fpage>&#x02013;<lpage>509</lpage>. <pub-id pub-id-type="pmid">11682799</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Itti</surname> <given-names>L.</given-names></name> <name><surname>Baldi</surname> <given-names>P.</given-names></name></person-group> (<year>2009</year>). <article-title>Bayesian surprise attracts human attention</article-title>. <source>Vision Res.</source> <volume>49</volume>, <fpage>1295</fpage>&#x02013;<lpage>1306</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2008.09.007</pub-id><pub-id pub-id-type="pmid">18834898</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jirsa</surname> <given-names>V. K.</given-names></name> <name><surname>Sporns</surname> <given-names>O.</given-names></name> <name><surname>Breakspear</surname> <given-names>M.</given-names></name> <name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>McIntosh</surname> <given-names>A. R.</given-names></name></person-group> (<year>2010</year>). <article-title>Towards the virtual brain: network modeling of the intact and the damaged brain</article-title>. <source>Arch. Ital. Biol.</source> <volume>148</volume>, <fpage>189</fpage>&#x02013;<lpage>205</lpage>. <pub-id pub-id-type="pmid">21175008</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joanisse</surname> <given-names>M.</given-names></name> <name><surname>Seidenberg</surname> <given-names>M. S.</given-names></name></person-group> (<year>1999</year>). <article-title>Impairments in verb morphology after brain injury: a connectionist model</article-title>. <source>Proc. Natl. Acad. Sci. U S A</source> <volume>96</volume>, <fpage>7592</fpage>&#x02013;<lpage>7597</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.96.13.7592</pub-id><pub-id pub-id-type="pmid">10377460</pub-id></citation></ref>
<ref id="B45"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Jordan</surname> <given-names>M. I.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T. J.</given-names></name></person-group> (<year>2001</year>). <source>Graphical Models: Foundations of Neural Computation.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kelso</surname> <given-names>J. A. S.</given-names></name></person-group> (<year>2012</year>). <article-title>Multistability and metastability: understanding dynamic coordination in the brain</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>367</volume>, <fpage>906</fpage>&#x02013;<lpage>918</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2011.0351</pub-id><pub-id pub-id-type="pmid">22371613</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khaligh-Razavi</surname> <given-names>S. M.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2014</year>). <article-title>Deep supervised, but not unsupervised, models may explain IT cortical representation</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003915</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003915</pub-id><pub-id pub-id-type="pmid">25375136</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kirkpatrick</surname> <given-names>S.</given-names></name> <name><surname>Gelatt</surname> <given-names>C.</given-names> <suffix>Jr.</suffix></name> <name><surname>Vecchi</surname> <given-names>M.</given-names></name></person-group> (<year>1983</year>). <article-title>Optimization by simmulated annealing</article-title>. <source>Science</source> <volume>220</volume>, <fpage>671</fpage>&#x02013;<lpage>680</lpage>. <pub-id pub-id-type="doi">10.1126/science.220.4598.671</pub-id><pub-id pub-id-type="pmid">17813860</pub-id></citation></ref>
<ref id="B49"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Koller</surname> <given-names>D.</given-names></name> <name><surname>Friedman</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <source>Probabilistic Graphical Models: Principles and Techniques.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lake</surname> <given-names>B. M.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <article-title>Humal-level concept learning through probabilistic program induction</article-title>. <source>Science</source> <volume>350</volume>, <fpage>1332</fpage>&#x02013;<lpage>1338</lpage>. <pub-id pub-id-type="doi">10.1126/science.aab3050</pub-id><pub-id pub-id-type="pmid">26659050</pub-id></citation></ref>
<ref id="B51"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Ranzato</surname> <given-names>M. A.</given-names></name> <name><surname>Monga</surname> <given-names>R.</given-names></name> <name><surname>Devin</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G. S.</given-names></name> <etal/></person-group>. (<year>2012</year>). &#x0201C;<article-title>Building high-level features using large scale unsupervised learning</article-title>,&#x0201D; in <source>Proceedings of the 29th International Conference on Machine Learning</source> (<publisher-loc>Edinburgh, Scotland, UK</publisher-loc>).</citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>T. S.</given-names></name> <name><surname>Mumford</surname> <given-names>D.</given-names></name></person-group> (<year>2003</year>). <article-title>Hierarchical bayesian inference in the visual cortex</article-title>. <source>J. Opt. Soc. Am. A Opt. Image Sci. Vis.</source> <volume>20</volume>, <fpage>1434</fpage>&#x02013;<lpage>1448</lpage>. <pub-id pub-id-type="doi">10.1364/josaa.20.001434</pub-id><pub-id pub-id-type="pmid">12868647</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>W. J.</given-names></name> <name><surname>Beck</surname> <given-names>J. M.</given-names></name> <name><surname>Latham</surname> <given-names>P. E.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Bayesian inference with probabilistic population codes</article-title>. <source>Nat. Neurosci.</source> <volume>9</volume>, <fpage>1432</fpage>&#x02013;<lpage>1438</lpage>. <pub-id pub-id-type="doi">10.1038/nn1790</pub-id><pub-id pub-id-type="pmid">17057707</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Manford</surname> <given-names>M.</given-names></name> <name><surname>Andermann</surname> <given-names>F.</given-names></name></person-group> (<year>1998</year>). <article-title>Complex visual hallucinations. Clinical and neurobiological insights</article-title>. <source>Brain</source> <volume>121</volume>, <fpage>1819</fpage>&#x02013;<lpage>1840</lpage>. <pub-id pub-id-type="doi">10.1093/brain/121.10.1819</pub-id><pub-id pub-id-type="pmid">9798740</pub-id></citation></ref>
<ref id="B56"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mathers</surname> <given-names>C.</given-names></name> <name><surname>Fat</surname> <given-names>D. M.</given-names></name> <name><surname>Boerma</surname> <given-names>T. J.</given-names></name></person-group> (<year>2008</year>). <source>The Global Burden of Disease: 2004 Update.</source> <publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>.</citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mavritsaki</surname> <given-names>E.</given-names></name> <name><surname>Heinke</surname> <given-names>D.</given-names></name> <name><surname>Allen</surname> <given-names>H.</given-names></name> <name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>Humphreys</surname> <given-names>G. W.</given-names></name></person-group> (<year>2011</year>). <article-title>Bridging the gap between physiology and behavior: evidence from the sSoTS model of human visual attention</article-title>. <source>Psychol. Rev.</source> <volume>118</volume>, <fpage>3</fpage>&#x02013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1037/a0021868</pub-id><pub-id pub-id-type="pmid">21244184</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<year>2013</year>). <article-title>Integrating probabilistic models of perception and interactive neural networks: a historical and tutorial review</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>503</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00503</pub-id><pub-id pub-id-type="pmid">23970868</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McClelland</surname> <given-names>J. L.</given-names></name> <name><surname>McNaughton</surname> <given-names>B.</given-names></name> <name><surname>O&#x02019;Reilly</surname> <given-names>R. C.</given-names></name></person-group> (<year>1995</year>). <article-title>Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory</article-title>. <source>Psychol. Rev.</source> <volume>102</volume>, <fpage>419</fpage>&#x02013;<lpage>457</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295x.102.3.419</pub-id><pub-id pub-id-type="pmid">7624455</pub-id></citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Medaglia</surname> <given-names>J. D.</given-names></name> <name><surname>Lynall</surname> <given-names>M.-E.</given-names></name> <name><surname>Bassett</surname> <given-names>D. S.</given-names></name></person-group> (<year>2015</year>). <article-title>Cognitive network neuroscience</article-title>. <source>J. Cogn. Neurosci.</source> <volume>27</volume>, <fpage>1471</fpage>&#x02013;<lpage>1491</lpage>. <pub-id pub-id-type="doi">10.1162/jocn_a_00810</pub-id><pub-id pub-id-type="pmid">25803596</pub-id></citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mnih</surname> <given-names>V.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Bellemare</surname> <given-names>M. G.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human-level control through deep reinforcement learning</article-title>. <source>Nature</source> <volume>518</volume>, <fpage>529</fpage>&#x02013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1038/nature14236</pub-id><pub-id pub-id-type="pmid">25719670</pub-id></citation></ref>
<ref id="B62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>M.</given-names></name> <name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Bayesian computation emerges in generic cortical microcircuits through spike-timing-dependent plasticity</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003037</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003037</pub-id><pub-id pub-id-type="pmid">23633941</pub-id></citation></ref>
<ref id="B63"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Newman</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <source>Networks: An Introduction.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B64"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ngiam</surname> <given-names>J.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name> <name><surname>Nam</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2011</year>). &#x0201C;<article-title>Multimodal deep learning</article-title>,&#x0201D; in <source>Proceedings of the 28th International Conference on Machine Learning (ICML-11)</source>, <publisher-loc>Bellevue, WA</publisher-loc>, <fpage>689</fpage>&#x02013;<lpage>696</lpage>.</citation></ref>
<ref id="B65"><citation citation-type="book"><person-group person-group-type="author"><name><surname>O&#x02019;Reilly</surname> <given-names>R. C.</given-names></name> <name><surname>Munakata</surname> <given-names>Y.</given-names></name></person-group> (<year>2000</year>). <source>Computational Exploration in Cognitive Neuroscience.</source> <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>H.-J.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Structural and functional brain networks: from connections to cognition</article-title>. <source>Science</source> <volume>342</volume>:<fpage>1238411</fpage>. <pub-id pub-id-type="doi">10.1126/science.1238411</pub-id><pub-id pub-id-type="pmid">24179229</pub-id></citation></ref>
<ref id="B67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pecevski</surname> <given-names>D.</given-names></name> <name><surname>Buesing</surname> <given-names>L.</given-names></name> <name><surname>Maass</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>Probabilistic inference in general graphical models through sampling in stochastic networks of spiking neurons</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1002294</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002294</pub-id><pub-id pub-id-type="pmid">22219717</pub-id></citation></ref>
<ref id="B68"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plaut</surname> <given-names>D. C.</given-names></name> <name><surname>Behrmann</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Complementary neural representations for faces and words: a computational exploration</article-title>. <source>Cogn. Neuropsychol.</source> <volume>28</volume>, <fpage>251</fpage>&#x02013;<lpage>275</lpage>. <pub-id pub-id-type="doi">10.1080/02643294.2011.609812</pub-id><pub-id pub-id-type="pmid">22185237</pub-id></citation></ref>
<ref id="B69"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plaut</surname> <given-names>D. C.</given-names></name> <name><surname>Shallice</surname> <given-names>T.</given-names></name></person-group> (<year>1993</year>). <article-title>Deep dyslexia: a case study of connectionist neuropsychology</article-title>. <source>Cogn. Neuropsychol.</source> <volume>10</volume>, <fpage>377</fpage>&#x02013;<lpage>500</lpage>. <pub-id pub-id-type="doi">10.1080/02643299308253469</pub-id></citation></ref>
<ref id="B70"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pouget</surname> <given-names>A.</given-names></name> <name><surname>Driver</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <article-title>Relating unilateral neglect to the neural coding of space</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>10</volume>, <fpage>242</fpage>&#x02013;<lpage>249</lpage>. <pub-id pub-id-type="doi">10.1016/s0959-4388(00)00077-5</pub-id><pub-id pub-id-type="pmid">10753799</pub-id></citation></ref>
<ref id="B71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Raichle</surname> <given-names>M. E.</given-names></name></person-group> (<year>2015</year>). <article-title>The restless brain: how intrinsic activity organizes brain function</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>370</volume>:<fpage>20140172</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2014.0172</pub-id><pub-id pub-id-type="pmid">25823869</pub-id></citation></ref>
<ref id="B72"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Raina</surname> <given-names>R.</given-names></name> <name><surname>Madhavan</surname> <given-names>A.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2009</year>). &#x0201C;<article-title>Large-scale deep unsupervised learning using graphics processors</article-title>,&#x0201D; in <source>International Conference on Machine Learning</source>, (<conf-loc>New York, NY</conf-loc>: <conf-name>ACM Press</conf-name>), <fpage>873</fpage>&#x02013;<lpage>880</lpage>.</citation></ref>
<ref id="B73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>R. P. N.</given-names></name></person-group> (<year>2004</year>). <article-title>Bayesian computation in recurrent neural circuits</article-title>. <source>Neural Comput.</source> <volume>16</volume>, <fpage>1</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1162/08997660460733976</pub-id><pub-id pub-id-type="pmid">15006021</pub-id></citation></ref>
<ref id="B74"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reichert</surname> <given-names>D. P.</given-names></name> <name><surname>Series</surname> <given-names>P.</given-names></name> <name><surname>Storkey</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Charles bonnet syndrome: evidence for a generative model in the cortex?</article-title> <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1003134</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003134</pub-id><pub-id pub-id-type="pmid">23874177</pub-id></citation></ref>
<ref id="B75"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roitman</surname> <given-names>J. D.</given-names></name> <name><surname>Brannon</surname> <given-names>E. M.</given-names></name> <name><surname>Platt</surname> <given-names>M. L.</given-names></name></person-group> (<year>2007</year>). <article-title>Monotonic coding of numerosity in macaque lateral intraparietal area</article-title>. <source>PLoS Biol.</source> <volume>5</volume>:<fpage>e208</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pbio.0050208</pub-id><pub-id pub-id-type="pmid">17676978</pub-id></citation></ref>
<ref id="B76"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rotzer</surname> <given-names>S.</given-names></name> <name><surname>Kucian</surname> <given-names>K.</given-names></name> <name><surname>Martin</surname> <given-names>E.</given-names></name> <name><surname>von Aster</surname> <given-names>M.</given-names></name> <name><surname>Klaver</surname> <given-names>P.</given-names></name> <name><surname>Loenneker</surname> <given-names>T.</given-names></name></person-group> (<year>2008</year>). <article-title>Optimized voxel-based morphometry in children with developmental dyscalculia</article-title>. <source>Neuroimage</source> <volume>39</volume>, <fpage>417</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.08.045</pub-id><pub-id pub-id-type="pmid">17928237</pub-id></citation></ref>
<ref id="B77"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>McClelland</surname> <given-names>J. L.</given-names></name></person-group> (<year>1986</year>). <source>Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1: Foundations</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B78"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sillito</surname> <given-names>A. M.</given-names></name> <name><surname>Cudeiro</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>H. E.</given-names></name></person-group> (<year>2006</year>). <article-title>Always returning: feedback and sensory processing in visual cortex and thalamus</article-title>. <source>Trends Neurosci.</source> <volume>29</volume>, <fpage>307</fpage>&#x02013;<lpage>316</lpage>. <pub-id pub-id-type="doi">10.1016/j.tins.2006.05.001</pub-id><pub-id pub-id-type="pmid">16713635</pub-id></citation></ref>
<ref id="B79"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Silver</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>A.</given-names></name> <name><surname>Maddison</surname> <given-names>C. J.</given-names></name> <name><surname>Guez</surname> <given-names>A.</given-names></name> <name><surname>Sifre</surname> <given-names>L.</given-names></name> <name><surname>van den Driessche</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Mastering the game of Go with deep neural networks and tree search</article-title>. <source>Nature</source> <volume>529</volume>, <fpage>484</fpage>&#x02013;<lpage>489</lpage>. <pub-id pub-id-type="doi">10.1038/nature16961</pub-id><pub-id pub-id-type="pmid">26819042</pub-id></citation></ref>
<ref id="B80"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Snodgrass</surname> <given-names>J. G.</given-names></name> <name><surname>Vanderwart</surname> <given-names>M.</given-names></name></person-group> (<year>1980</year>). <article-title>A standardized set of 260 pictures: norms for name agreement, image agreement, familiarity and visual complexity</article-title>. <source>J. Exp. Psychol. Hum. Learn.</source> <volume>6</volume>, <fpage>174</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.6.2.174</pub-id><pub-id pub-id-type="pmid">7373248</pub-id></citation></ref>
<ref id="B81"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Emergence of a &#x0201C;visual number sense&#x0201D; in hierarchical generative models</article-title>. <source>Nat. Neurosci.</source> <volume>15</volume>, <fpage>194</fpage>&#x02013;<lpage>196</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2996</pub-id><pub-id pub-id-type="pmid">22231428</pub-id></citation></ref>
<ref id="B82"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). &#x0201C;<article-title>Developmental trajectories of numerosity perception</article-title>,&#x0201D; in <source>Poster Presented at the Workshop: Interactions Between Space, Time and Number: 20 Years of Research</source> (<publisher-loc>Paris</publisher-loc>).</citation></ref>
<ref id="B83"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name> <name><surname>Umilt&#x000E0;</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <article-title>The role of semantic and symbolic representations in arithmetic processing: insights from simulated dyscalculia in a connectionist model</article-title>. <source>Cortex</source> <volume>40</volume>, <fpage>194</fpage>&#x02013;<lpage>196</lpage>. <pub-id pub-id-type="doi">10.1016/s0010-9452(08)70948-1</pub-id><pub-id pub-id-type="pmid">15174483</pub-id></citation></ref>
<ref id="B84"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Taylor</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>The recurrent temporal restricted Boltzmann machine</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>20</volume>, <fpage>1601</fpage>&#x02013;<lpage>1608</lpage>.</citation></ref>
<ref id="B85"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Kemp</surname> <given-names>C.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. L.</given-names></name> <name><surname>Goodman</surname> <given-names>N. D.</given-names></name></person-group> (<year>2011</year>). <article-title>How to grow a mind: statistics, structure and abstraction</article-title>. <source>Science</source> <volume>331</volume>, <fpage>1279</fpage>&#x02013;<lpage>1285</lpage>. <pub-id pub-id-type="doi">10.1126/science.1192788</pub-id><pub-id pub-id-type="pmid">21393536</pub-id></citation></ref>
<ref id="B87"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Testolin</surname> <given-names>A.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>De Filippo De Grazia</surname> <given-names>M.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Deep unsupervised learning on a desktop PC: a primer for cognitive scientists</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>251</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00251</pub-id><pub-id pub-id-type="pmid">23653617</pub-id></citation></ref>
<ref id="B88"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Testolin</surname> <given-names>A.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Sperduti</surname> <given-names>A.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Learning orthographic structure with sequential generative neural networks</article-title>. <source>Cogn. Sci.</source> <volume>40</volume>, <fpage>579</fpage>&#x02013;<lpage>606</lpage>. <pub-id pub-id-type="doi">10.1111/cogs.12258</pub-id><pub-id pub-id-type="pmid">26073971</pub-id></citation></ref>
<ref id="B86"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Testolin</surname> <given-names>A.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Zorzi</surname> <given-names>M.</given-names></name></person-group> (<year>under review</year>). <article-title>Human-like letter perception emerges from unsupervised deep learning and recycling of natural image statistics</article-title>.</citation></ref>
<ref id="B89"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varela</surname> <given-names>F.</given-names></name> <name><surname>Lachaux</surname> <given-names>J. P.</given-names></name> <name><surname>Rodriguez</surname> <given-names>E.</given-names></name> <name><surname>Martinerie</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>The brainweb: phase synchronization and large-scale integration</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>2</volume>, <fpage>229</fpage>&#x02013;<lpage>239</lpage>. <pub-id pub-id-type="doi">10.1038/35067550</pub-id><pub-id pub-id-type="pmid">11283746</pub-id></citation></ref>
<ref id="B91"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Zorzi</surname> <given-names>M.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name> <name><surname>Umilt&#x000E0;</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). &#x0201C;<article-title>Computational modeling of numerical cognition</article-title>,&#x0201D; in <source>Handbook of Mathematical Cognition</source>, ed. <person-group person-group-type="editor"><name><surname>Campbell</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Psychology Press</publisher-name>), <fpage>67</fpage>&#x02013;<lpage>84</lpage>.</citation></ref>
<ref id="B92"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zorzi</surname> <given-names>M.</given-names></name> <name><surname>Testolin</surname> <given-names>A.</given-names></name> <name><surname>Stoianov</surname> <given-names>I.</given-names></name></person-group> (<year>2013</year>). <article-title>Modeling language and cognition with deep unsupervised learning: a tutorial overview</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>515</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00515</pub-id><pub-id pub-id-type="pmid">23970869</pub-id></citation></ref>
</ref-list>
</back>
</article>