<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2016.00144</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>On the Maximum Storage Capacity of the Hopfield Model</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Folli</surname> <given-names>Viola</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/377835/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Leonetti</surname> <given-names>Marco</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/377838/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ruocco</surname> <given-names>Giancarlo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/377836/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Center for Life Nanoscience, Istituto Italiano di Tecnologia</institution> <country>Rome, Italy</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Physics, Sapienza University of Rome</institution> <country>Rome, Italy</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Marcel Van Gerven, Radboud University Nijmegen, Netherlands</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Simon R. Schultz, Imperial College London, UK; Fleur Zeldenrust, Donders Institute for Brain, Cognition and Behaviour, Netherlands</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Viola Folli <email>viola.folli&#x00040;iit.it</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2016</year>
</pub-date>
<volume>10</volume>
<elocation-id>144</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>09</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>12</month>
<year>2016</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Folli, Leonetti and Ruocco.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Folli, Leonetti and Ruocco</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Recurrent neural networks (RNN) have traditionally been of great interest for their capacity to store memories. In past years, several works have been devoted to determine the maximum storage capacity of RNN, especially for the case of the Hopfield network, the most popular kind of RNN. Analyzing the thermodynamic limit of the statistical properties of the Hamiltonian corresponding to the Hopfield neural network, it has been shown in the literature that the retrieval errors diverge when the number of stored memory patterns (<italic>P</italic>) exceeds a fraction (&#x02248; 14%) of the network size <italic>N</italic>. In this paper, we study the storage performance of a generalized Hopfield model, where the diagonal elements of the connection matrix are allowed to be different from zero. We investigate this model at finite <italic>N</italic>. We give an analytical expression for the number of retrieval errors and show that, by increasing the number of stored patterns over a certain threshold, the errors start to decrease and reach values below unit for <italic>P</italic> &#x0226B; <italic>N</italic>. We demonstrate that the strongest trade-off between efficiency and effectiveness relies on the number of patterns (<italic>P</italic>) that are stored in the network by appropriately fixing the connection weights. When <italic>P</italic>&#x0226B;<italic>N</italic> and the diagonal elements of the adjacency matrix are not forced to be zero, the optimal storage capacity is obtained with a number of stored memories much larger than previously reported. This theory paves the way to the design of RNN with high storage capacity and able to retrieve the desired pattern without distortions.</p></abstract>
<kwd-group>
<kwd>maximum storage memory</kwd>
<kwd>feed-forward structure</kwd>
<kwd>random recurrent network</kwd>
<kwd>Hopfield model</kwd>
<kwd>retrieval error</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="0"/>
<equation-count count="18"/>
<ref-count count="24"/>
<page-count count="9"/>
<word-count count="6695"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>A vast amount of literature deals with neural networks, both as model for brain functioning (Amit, <xref ref-type="bibr" rid="B2">1989</xref>), and as smart artificial systems for practical applications in computation and information handling (Haykin, <xref ref-type="bibr" rid="B13">1999</xref>).</p>
<p>Among the different possible applications of artificial neural networks, those referred to as &#x0201C;associative memory&#x0201D; are particularly important (Rojas, <xref ref-type="bibr" rid="B20">1996</xref>), i.e., circuits with the capability to store and retrieve specific information patterns. According to Amit et al. (<xref ref-type="bibr" rid="B3">1985a</xref>,<xref ref-type="bibr" rid="B4">b</xref>) there is a natural limit for the usage of an <italic>N</italic> nodes neural network built according to the Hebbian principle (Hebb, <xref ref-type="bibr" rid="B14">1949</xref>) as associative memory. The association is embedded within the connection matrix which has a dyadic form: the weight connecting neuron <italic>i</italic> to neuron <italic>j</italic> is the product of the respective signals. The limit of storage is linear with <italic>N</italic>: an attempt to store a number <italic>P</italic> of memory elements larger than &#x003B1;<sub><italic>c</italic></sub><italic>N</italic>, with &#x003B1;<sub><italic>c</italic></sub> &#x02248; 0.14, results in a &#x0201C;divergent&#x0201D; (order <italic>P</italic>) number of retrieval errors. In order to be effective (low retrieval error probability) a neural network working as associative memory cannot be efficient (i.e., it can store only a small number of memory elements). This is particularly frustrating in practical applications, as it strongly limits the use of artificial neural networks for information storage, especially since it is well known that the number of fixed points in randomly connected (symmetric) neural networks shows an exponential relation with <italic>N</italic> (Tanaka and Edwards, <xref ref-type="bibr" rid="B23">1980</xref>; Sompolinsky et al., <xref ref-type="bibr" rid="B22">1988</xref>; Wainrib and Touboul, <xref ref-type="bibr" rid="B24">2013</xref>).</p>
<p>Contemporaneous to Amit et al., Abu-Mostafa, and St. Jaques (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">1985</xref>) claimed that the number of fixed points that can be used for memory storage in a Hopfield model with a generic coupling matrix is limited to <italic>N</italic> (i.e., <italic>P</italic>&#x0003C;<italic>N</italic>). Soon after, Mc Eliece et al. (<xref ref-type="bibr" rid="B18">1987</xref>), considering only the Hebbian dyadic form for the coupling matrix, found a more severe limitation: the maximum <italic>P</italic> scales as <italic>N</italic>/log(<italic>N</italic>). In a more recent study, Sollacher et al. (<xref ref-type="bibr" rid="B21">2009</xref>) designed a network of specific topology, reaching &#x003B1;<sub><italic>c</italic></sub>-values larger than 0.14, but still maintaining the limit of a linear <italic>N</italic> dependence of the maximum storage capacity. The storage problem remains an open research question (Brunel, <xref ref-type="bibr" rid="B6">2016</xref>).</p>
<p>In this letter we show that the existence of a critical <italic>P</italic>/<italic>N</italic>- value in the Hebbian scheme for the coupling matrix is only part of the story. As demonstrated in Amit et al. (<xref ref-type="bibr" rid="B3">1985a</xref>,<xref ref-type="bibr" rid="B4">b</xref>), the limit <italic>P</italic>&#x0003C;&#x003B1;<sub><italic>c</italic></sub><italic>N</italic> holds in the region where <italic>P</italic>&#x0003C;<italic>N</italic>. In all previous studies, the diagonal elements are removed from the dyadic form of the coupling matrix. Here we show the existence of a not yet explored region in the parameter space, with <italic>P</italic>&#x0226B;<italic>N</italic>, where the number of retrieval errors decreases with increasing <italic>P</italic> and reaches values lower than one. This region can be found by not removing the diagonal elements. Strictly speaking the present model is not a &#x0201C;Hopfield model,&#x0201D; as in the latter case the diagonal elements are forced to vanish and&#x02014;as we will see- bring significant differences in the network behavior. In order to avoid confusion, let us call the present model as &#x0201C;Hopfield model with autapses&#x0201D; or &#x0201C;Generalized Hopfield model.&#x0201D; This strategy allows the design of effective and efficient associative memories based on artificial neural networks. In the following we will derive analytically the probability of retrieval errors, validate these results by their comparison with a numerical simulation and study the efficiency of the system as a function of <italic>P</italic> and <italic>N</italic>.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2. Methods</title>
<sec>
<title>2.1. Network model</title>
<p>In an artificial neural network working as associative memory, one deals with a network of <italic>N</italic> neurons of which each one has state <italic>s</italic><sub><italic>i</italic></sub> (<italic>i</italic> &#x0003D; 1&#x02026;<italic>N</italic>) that can be &#x0201C;active&#x0201D; (<italic>s</italic><sub><italic>i</italic></sub> &#x0003D; 1) or &#x0201C;quiescent&#x0201D; (<italic>s</italic><sub><italic>i</italic></sub> &#x0003D; &#x02212;1). The configuration of the whole network is given by the vector <inline-formula><mml:math id="M1"><mml:mrow><mml:mover accent='true'><mml:mi>s</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>&#x02261;</mml:mo><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mn>..</mml:mn><mml:msub><mml:mi>s</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>&#x0007D;</mml:mo></mml:mrow></mml:math></inline-formula> and its temporal evolution follows the parallel non-linear dynamics:
<disp-formula id="E1"><label>(1)</label><mml:math id="M89"><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x00394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo><mml:mo>&#x02250;</mml:mo><mml:mtext>sign</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mstyle><mml:msub><mml:mi>s</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
where <bold>J</bold> &#x0003D; {<italic>J</italic><sub><italic>ij</italic></sub>} is the connection matrix. We set external inputs to be equal to 0. We assume a symmetric bimodal distribution for the synaptic polarities in the wiring matrix <bold>J</bold>, so 50% of the connections are excitatory and 50% inhibitory. After a transient time related to the finite value of <italic>N</italic>, the network reaches a fixed point, <italic>s</italic><sub><italic>i</italic></sub> &#x0003D; <italic>E</italic>[<italic>s</italic><sub><italic>i</italic></sub>], or a limit cycle of length <italic>L</italic>, s<sub>i</sub> &#x0003D; <italic>E</italic><sup>(<italic>L</italic>)</sup>[<italic>s<sub>i</sub></italic>].</p>
</sec>
<sec>
<title>2.2. The hebbian rule and the storage memory</title>
<p>Previous work has studied the cycle length and transient time distribution as a function of the properties of <bold>J</bold> (Gutfreundt et al., <xref ref-type="bibr" rid="B12">1988</xref>; Sompolinsky et al., <xref ref-type="bibr" rid="B22">1988</xref>; Derrida, <xref ref-type="bibr" rid="B10">1989</xref>; Bastolla et al., <xref ref-type="bibr" rid="B5">1997</xref>). In order to work as an associative memory, the matrix <bold>J</bold> must be tailored in such a way that one or more patterns of neurons are fixed points of the dynamics in Equation (1), i.e., they are the &#x0201C;memory elements&#x0201D; stored in the network. To store one pattern <inline-formula><mml:math id="M3"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, the connection matrix is simply the dyadic form given by <italic>J</italic><sub><italic>ij</italic></sub> &#x0003D; &#x003BE;<sub><italic>i</italic></sub>&#x003BE;<sub><italic>j</italic></sub> <xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> , while to store a generic number <italic>P</italic> of patterns <inline-formula><mml:math id="M4"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (&#x003BC; &#x0003D; 1&#x02026;<italic>P</italic>) one follows the storage prescription of Cooper (<xref ref-type="bibr" rid="B7">1973</xref>) and Cooper et al. (<xref ref-type="bibr" rid="B8">1979</xref>), who exploited an old idea which goes back to Hebb (<xref ref-type="bibr" rid="B14">1949</xref>) and Eccles (<xref ref-type="bibr" rid="B11">1953</xref>) and which states that the change in synaptic transmission is proportional to the product of the signals of pre and post-synaptic neurons. The process for which each matrix element is appropriately determined is called <italic>learning</italic>. Specifically, the &#x0201C;Hebbian&#x0201D; rule results in the following expression for the connectivity matrix,
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>P</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>P</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>&#x003BE;</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>&#x003BE;</mml:mi><mml:mi>j</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
The set of vectors <inline-formula><mml:math id="M6"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mi>u</mml:mi></mml:math></inline-formula> is known as &#x0201C;training set.&#x0201D; In this case, it is not guaranteed that each <inline-formula><mml:math id="M7"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a fixed point. In other words, <inline-formula><mml:math id="M8"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is stable in probabilistic sense. Further, the probability for <inline-formula><mml:math id="M9"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> to be a fixed point depends on the values of <italic>P</italic> and <italic>N</italic>. This dependence has been first studied by Hopfield (<xref ref-type="bibr" rid="B15">1982</xref>); Hopfield et al. (<xref ref-type="bibr" rid="B17">1983</xref>); Hopfield (<xref ref-type="bibr" rid="B16">1984</xref>) who concluded that the retrieval of the memory stored in the Hebbian matrices is guaranteed up to a <italic>P</italic>-value which is a critical fraction on the number of network nodes <italic>N</italic> of the order of 10&#x02013;20%. Above this value, the associative memory quickly degrades. Following these studies, Amit et al. (<xref ref-type="bibr" rid="B3">1985a</xref>,<xref ref-type="bibr" rid="B4">b</xref>), who noticed the similarity between the Hopfield model for the associative memory and the spin glasses, developed a statistical theory for the determination of the critical <italic>P</italic>/<italic>N</italic> ratio, that turned out to be &#x02248; 0.14, in good agreement with the previous Hopfield estimation. Above <italic>P</italic>=0.14<italic>N</italic> the number of errors is so large that the network based on the Hebbian matrix is no longer capable to work as an associative memory. All these studies assumed a modified form of Equation (2): the diagonal elements of <bold>J</bold> are forced to be zero.</p>
</sec>
<sec>
<title>2.3. Numerical simulations and data analysis</title>
<p>In order to demonstrate the validity of our analytic results (see Section 3), we perform numerical simulations by evolving the network model as described in Equation (1). We design the default network by fixing the NxN recurrent connections as given in Equation (2), by randomly assigning the value &#x000B1;1 to <inline-formula><mml:math id="M10"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and retaining the diagonal elements. So, the <italic>N</italic>(<italic>N</italic>&#x02212;1) connections are 50% excitatory and 50% inhibitory and the <italic>N</italic> neurons can form self-connections. We then run simulations by varying the size of the network, <italic>N</italic> &#x0003D; 50, .., 200 and the number of stored memories in Equation (2), <italic>P</italic> &#x0003D; 1, &#x02026;, 2000. Finally, for each pair of <italic>N</italic> and <italic>P</italic>, we perform 1000 different random realizations.</p>
<p>All <italic>P</italic> patterns introduced in Equation (2) are given as input to the network and their dynamics is followed until the network reaches the equilibrium state. The initial patterns are chosen among those that were stored in the adjacency matrix and that have been randomly chosen in the designing of the network. Evolved patterns were recorded at each time step and compared with the initial one. Then, if the evolved pattern is different from the initial state, we calculated the temporal evolution of the distortion (number of wrong bits) and determined the probability that one of the bits was wrong, the probability that the whole vector was exactly recovered, and the number of memory patterns that could not be recovered, as a function of <italic>N</italic> and <italic>P</italic>. Basically, to calculate the storage capacity, it is sufficient to determine all these quantities by using the distortion between the stimulus (the stored memory) and its first evolved pattern.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. The probability of recovery</title>
<p>In order to investigate the maximum storage memory of our model, we calculate the one-step dynamical evolution. We give as input a vector of the training set and we calculate a single step of the dynamical evolution according to Equation (1). Then, we compare the output with the input. We aim to look whether or not a vector, <inline-formula><mml:math id="M11"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, belonging to the training set, is truly a fixed point. If <inline-formula><mml:math id="M12"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a fixed point, the output coincides with the input, and the recovery has been successful. If <inline-formula><mml:math id="M13"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is not a fixed point, the two vectors differ for at least a single bit. We now derive an analytical expression for the probability that the recovery of a stored pattern was not successful. The first step is to find the probability <italic>p</italic><sub><italic>B</italic></sub> that -given the matrix <bold>J</bold> of Equation (2)- a single element of the vector (a &#x0201C;bit&#x0201D;) was wrong, i.e., the probability that <inline-formula><mml:math id="M14"><mml:mi>E</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02260;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Basically, we need to evolve a vector <inline-formula><mml:math id="M15"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mi>u</mml:mi></mml:math></inline-formula> (from the training set) for one step and count how many bits of its time evolution are different from the bits of <inline-formula><mml:math id="M16"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mi>u</mml:mi></mml:math></inline-formula> itself. Obviously, if <inline-formula><mml:math id="M17"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mi>u</mml:mi></mml:math></inline-formula> actually is a fixed point, this distance vanishes. On the contrary, <inline-formula><mml:math id="M18"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mi>u</mml:mi></mml:math></inline-formula> is not a fixed point, the network has made a recovery error. Thus, <italic>p</italic><sub><italic>B</italic></sub> (or better, <italic>p</italic><sub><italic>V</italic></sub>, as we see in the next paragraph) measures &#x0201C;how many&#x0201D; training set vectors are not fixed points. The argument of the <italic>sign</italic> function in Equation (1) is <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, this contains <italic>NP</italic> terms among which there are <italic>N</italic>&#x0002B;<italic>P</italic>&#x02212;1 terms (those with <italic>j</italic>=<italic>i</italic> <xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> and those with &#x003BD;=&#x003BC;) where two out of the three &#x003BE; of the product are equals to each other <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and the third is <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Thus <inline-formula><mml:math id="M22"><mml:mrow><mml:msubsup><mml:mi>A</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:msubsup><mml:mi>&#x003BE;</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>T</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, with <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. The first term is the &#x0201C;coherent&#x0201D; one, its sign is identical to <inline-formula><mml:math id="M24"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and it will -if dominant- guarantee that <inline-formula><mml:math id="M25"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a fixed point of the dynamics. The second term <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> , on the contrary, is &#x0201C;noise&#x0201D; and its presence can either reinforce or weaken the stability of <inline-formula><mml:math id="M27"><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> as fixed point. Specifically, if <inline-formula><mml:math id="M28"><mml:mrow><mml:mo>&#x0007C;</mml:mo><mml:msubsup><mml:mi>T</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup><mml:mo>&#x0007C;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mo>&#x0003E;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> and sign<inline-formula><mml:math id="M29"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02260;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, then the <italic>i</italic>-<italic>th</italic> bit of the vector <inline-formula><mml:math id="M30"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> will turn out to be wrong. The quantity <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the sum of (<italic>N</italic>&#x02212;1)(<italic>P</italic>&#x02212;1) statistically independent terms, each one being &#x0002B;1 or &#x02212;1. Therefore, for large enough <italic>P</italic> and <italic>N</italic>, its distribution <italic>N</italic>(<italic>T</italic>) can be approximated by a gaussian with zero mean and standard deviation <inline-formula><mml:math id="M32"><mml:msqrt><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:math></inline-formula>:
<disp-formula id="E3"><label>(3)</label><mml:math id="M33"><mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
It is now straightforward to determine the probability that <inline-formula><mml:math id="M34"><mml:mrow><mml:mo>&#x0007C;</mml:mo><mml:msubsup><mml:mi>T</mml:mi><mml:mi>i</mml:mi><mml:mi>&#x003BC;</mml:mi></mml:msubsup><mml:mo>&#x0007C;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mo>&#x0003E;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> and sign<inline-formula><mml:math id="M35"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02260;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, thus that one of the bits of <inline-formula><mml:math id="M36"><mml:mi>E</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> was wrong, as <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>P</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:mrow></mml:msubsup><mml:mi>d</mml:mi><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. In conclusion:
<disp-formula id="E4"><label>(4)</label><mml:math id="M38"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
It is worth to note that this expression is symmetric under the exchange of <italic>P</italic> with <italic>N</italic>, and that for large <italic>P</italic> and <italic>N</italic>, with <italic>P</italic>/<italic>N</italic> &#x0003D; 1, it tends to <inline-formula><mml:math id="M39"><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x02248;</mml:mo></mml:mrow></mml:math></inline-formula>0.02275 which corresponds to the maximum of probability in a wrong recovery of a single bit (see Figures <xref ref-type="fig" rid="F1">1</xref>, <xref ref-type="fig" rid="F2">2</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Comparison of the results of the numerical simulation (<italic><bold>p</bold></italic><sub><italic><bold>B</bold></italic></sub>, full dots A</bold>; <italic>p</italic><sub><italic>V</italic></sub>, full squares, <bold>B</bold>; <italic>N</italic><sub><italic>V</italic></sub>, full diamonds, <bold>C</bold>) with the corresponding theoretical function (<italic>p</italic><sub><italic>B</italic></sub>, Equation 4; <italic>p</italic><sub><italic>V</italic></sub>, Equation 5; <italic>N</italic><sub><italic>V</italic></sub>, Equation 6) reported as full lines. The three quantities are reported as a function of <italic>P</italic> for fixed <italic>N</italic>. The values of <italic>N</italic> are 50 (black), 100 (blue), 150 (green), and 200 (red). The <italic>P</italic> range in <bold>(C)</bold> is extended with respect to <bold>(A,B)</bold>.</p></caption>
<graphic xlink:href="fncom-10-00144-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Theoretical curves for the three quantities <italic><bold>p</bold></italic><sub><italic><bold>B</bold></italic></sub>, <italic><bold>p</bold></italic><sub><italic><bold>V</bold></italic></sub>, and <italic><bold>N</bold></italic><sub><italic><bold>V</bold></italic></sub> (<italic><bold>p</bold></italic><sub><italic><bold>B</bold></italic></sub>, Equation 4, A</bold>; <italic>p</italic><sub><italic>V</italic></sub>, Equation 5, <bold>B</bold>; <italic>N</italic><sub><italic>V</italic></sub>, Equation 6, <bold>C</bold>) reported as full lines. The three quantities are reported in linear scale as a function of &#x003B1; &#x0003D; <italic>P</italic>/<italic>N</italic> for fixed <italic>N</italic>. The values of <italic>N</italic> are 50 (black), 100 (blue), and 200 (red). The dotted lines in <bold>(C)</bold> represent <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; <italic>P</italic>.</p></caption>
<graphic xlink:href="fncom-10-00144-g0002.tif"/>
</fig>
<p>The second step is the determination of the probability <italic>p</italic><sub><italic>V</italic></sub> that one of the P vectors encoded into the connection matrix (the training set) turns out not be a fixed point. If only a single bit of the vector is wrong, the whole vector is considered &#x0201C;wrong.&#x0201D; Since there are <italic>N</italic> bits that can be wrong, the probability <italic>p</italic><sub><italic>V</italic></sub> will be much higher than <italic>p</italic><sub><italic>B</italic></sub>. The calculation is straightforward, in order not to be wrong, all the bits of the vector <inline-formula><mml:math id="M40"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> must be right, thus <italic>p</italic><sub><italic>V</italic></sub> &#x0003D; 1&#x02212;<inline-formula><mml:math id="M41"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, therefore:
<disp-formula id="E5"><label>(5)</label><mml:math id="M42"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mi>N</mml:mi></mml:msup><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
Finally, the number, <italic>N</italic><sub><italic>V</italic></sub>, of memory vectors that are not recovered, i.e., that are not true fixed points of the dynamics is given by <italic>Pp</italic><sub><italic>V</italic></sub>, that is:
<disp-formula id="E6"><label>(6)</label><mml:math id="M43"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mi>N</mml:mi></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec>
<title>3.2. The asymptotical approximation</title>
<p>Equations (4), (5) and, in particular, Equation (6) represent the main result of this work. Before showing their validity, via a comparison with numerical simulations, and discussing their relevance in the framework of artificial neural networks, it is important to present the asymptotical approximation for <italic>N</italic><sub><italic>V</italic></sub>. The argument of the error function, for either <italic>P</italic>&#x0226B;<italic>N</italic> or <italic>P</italic>&#x0226A;<italic>N</italic>, is large, and can be expanded as erf(<italic>x</italic>) &#x02248; 1&#x02212;exp<inline-formula><mml:math id="M44"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:msqrt><mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Furthermore, as <italic>p</italic><sub><italic>B</italic></sub> is exponentially small with <italic>N</italic> (or <italic>P</italic>) for large <italic>N</italic> (<italic>P</italic>), we use, <inline-formula><mml:math id="M45"><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02248;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Thus, for large <italic>N</italic> or large <italic>P</italic>:
<disp-formula id="E7"><label>(7)</label><mml:math id="M46"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M47"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
We note that, while in the exact expression for <italic>N</italic><sub><italic>V</italic></sub> (Equation 6) the <italic>P</italic>&#x02194;<italic>N</italic> exchange symmetry is lost, in the approximate form the symmetry is recovered.</p>
<p>For sake of comparison with the previous literature, it is also useful to express the main results as a function of &#x003B1;&#x02250;<italic>P</italic>/<italic>N</italic>. Equations (4) (for large <italic>N</italic>) and (8) read:
<disp-formula id="E9"><label>(9)</label><mml:math id="M48"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M49"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:mi>N</mml:mi><mml:mi>P</mml:mi><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mfrac><mml:mrow><mml:msqrt><mml:mi>&#x003B1;</mml:mi></mml:msqrt></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
While <italic>p</italic><sub><italic>B</italic></sub> only depends on &#x003B1;, <italic>N</italic><sub><italic>V</italic></sub> clearly is an extensive observable, being proportional to <italic>P</italic> and <italic>N</italic>. Furthermore, both expressions keep their symmetry with respect to the exchange of <italic>P</italic> and <italic>N</italic>, thus to the exchange of &#x003B1; with 1/&#x003B1;. The last observation anticipates that there must exists a region at large &#x003B1;-values where the same features are observed as at small values of &#x003B1;.</p>
</sec>
<sec>
<title>3.3. Numerical results</title>
<p>To check the predictions of our network model, we have simulated the Model (1) and studied the dynamics for several values of <italic>N</italic> and <italic>P</italic>, in the range of few hundred, see Section 2.3 for details. In the numerical analysis, the <italic>P</italic> memory vectors have been randomly chosen and used to construct the connection matrix <bold>J</bold>. Next, we tested whether or not the stored memories were fixed points of the dynamics. The values of <italic>p</italic><sub><italic>B</italic></sub>, <italic>p</italic><sub><italic>V</italic></sub> and <italic>N</italic><sub><italic>V</italic></sub> were calculated by averaging over (up to) 1000 different random realizations of <inline-formula><mml:math id="M50"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The results of the numerical simulations are reported (dots) in Figure <xref ref-type="fig" rid="F1">1</xref>, together with the analytical Expressions (4)&#x02013;(6) (lines). The three panels refer to the three quantities <italic>p</italic><sub><italic>B</italic></sub> (Figure <xref ref-type="fig" rid="F1">1A</xref>), <italic>p</italic><sub><italic>V</italic></sub> (Figure <xref ref-type="fig" rid="F1">1B</xref>), and <italic>N</italic><sub><italic>V</italic></sub> (Figure <xref ref-type="fig" rid="F1">1C</xref>) as a function of <italic>P</italic> for the selected values of <italic>N</italic>, as reported in the legend. From Figure <xref ref-type="fig" rid="F1">1</xref>, we observe that on increasing <italic>P</italic>, at fixed <italic>N</italic>, both the single bit probability error, the probability of recovery error (<italic>P</italic><sub><italic>V</italic></sub>), and the number of wrong recoveries <italic>N</italic><sub><italic>V</italic></sub>, after a first fast increase, reach a maximum (equal to 0.02275 for <italic>p</italic><sub><italic>B</italic></sub>, close to one for <italic>p</italic><sub><italic>V</italic></sub>, and larger than <italic>N</italic> for <italic>N</italic><sub><italic>V</italic></sub>) then start to decrease, tending to zero for very large <italic>P</italic>-values.</p>
<p>To better emphasize this behavior, the same quantities are reported (analytic results only) as a function of &#x003B1; in Figure <xref ref-type="fig" rid="F2">2</xref> (linear scale) and in Figure <xref ref-type="fig" rid="F3">3</xref> (log scale) for selected <italic>N</italic>. The dotted lines in panels C of both figures represent <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; <italic>P</italic>, i.e., indicate the case of &#x0201C;totally wrong recovery.&#x0201D; Due to the already observed &#x003B1;&#x02194;1/&#x003B1; symmetry, the asymptotic curve in Figure <xref ref-type="fig" rid="F3">3A</xref> appears with a left-right symmetry around &#x003B1; &#x0003D; 1. From Figures <xref ref-type="fig" rid="F2">2</xref>, <xref ref-type="fig" rid="F3">3</xref>, we can clearly identify two regions of high recovery efficiency. The low &#x003B1; region, already studied many years ago by Hopfield (<xref ref-type="bibr" rid="B15">1982</xref>); Hopfield et al. (<xref ref-type="bibr" rid="B17">1983</xref>); Hopfield (<xref ref-type="bibr" rid="B16">1984</xref>) and Amit et al. (<xref ref-type="bibr" rid="B3">1985a</xref>,<xref ref-type="bibr" rid="B4">b</xref>), shows the existence of a quick transition toward &#x0201C;loss of memory recover&#x0201D; on increasing &#x003B1; around &#x003B1; &#x02248; 0.14. The second region at large &#x003B1;-values is not yet explored.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Theoretical curves for the three quantities <italic>pB, pV</italic>, and <italic>NV</italic> (<italic>pB</italic>, Equation 4, A</bold>; <italic>pV</italic>, Equation 5, <bold>B</bold>; <italic>NV</italic>, Equation 6, <bold>C</bold>) reported in Log-Log scale as full lines. The three quantities are reported in linear scale as a function of &#x003B1; &#x0003D; <italic>P/N</italic> for fixed <italic>N</italic>. The values of <italic>N</italic> are 50 (black), 100 (blue), and 200 (red). The dotted lines in <bold>(C)</bold> represent <italic>NV</italic> &#x0003D; <italic>P</italic>.</p></caption>
<graphic xlink:href="fncom-10-00144-g0003.tif"/>
</fig>
<p>Although the value &#x003B1; &#x0003D; 1 (<italic>P</italic> &#x0003D; <italic>N</italic>) represents traditionally a sort of limit in the computation of the storable memories in a RNN, there is no reason why not to store more than <italic>N</italic> memory elements in a network of <italic>N</italic> neurons, that by construction allows 2<sup><italic>N</italic></sup> possible patterns. Indeed, the number of fixed points in a (random) symmetric matrix is known to be, for fully connected symmetric matrices as in our case, exponentially large with <italic>N</italic> (Tanaka and Edwards, <xref ref-type="bibr" rid="B23">1980</xref>). Specifically, the number of fixed points <italic>P</italic><sub><italic>o</italic></sub> is equal to <italic>P</italic><sub><italic>o</italic></sub> &#x0003D; exp(&#x003B3;<italic>N</italic>), with &#x003B3; &#x02248; 0.2. <italic>P</italic><sub><italic>o</italic></sub>, much larger than <italic>N</italic>, can be considered a natural limit for <italic>P</italic>.</p>
<p>The recovery efficiency increases for large <italic>P</italic>. In fact, the coherent term in the argument of the sign function increases linearly with <italic>P</italic> and the noise increases as <italic>P</italic><sup>1/2</sup>. For large <italic>P</italic>, the relative weight of the noise decreases as <italic>P</italic><sup>&#x02212;1/2</sup>, this allows to store a large number of memories in a relatively small neural network.</p>
<p>For practical purposes, as for example in the design of an artificial neural network with high efficiency (large storage capacity) and effectiveness (low recovery error rates), it is important to study (Equation 6, and its approximation in Equation 8) and, in particular, to find the conditions for which the network shows &#x0201C;perfect recovery.&#x0201D; Let&#x00027;s define perfect recovery as the state where the number of retrieval errors <italic>N</italic><sub><italic>V</italic></sub> is smaller than one.</p>
<p>In Figure <xref ref-type="fig" rid="F4">4</xref> we show the contour plot of the (decimal) logarithm of <italic>N</italic><sub><italic>V</italic></sub>, from Equations (6) and (8), in the <italic>P</italic>-<italic>N</italic> range [0&#x02013;100]. The full lines are the loci of the points where <italic>log</italic><sub>10</sub>(<italic>N</italic><sub><italic>V</italic></sub>) equals 0, 0.4, 0.8, 1.2, and 1.6, as indicated on the right side of the figure. The dashed lines are the same level lines for the (logarithm of the) approximate form of <italic>N</italic><sub><italic>V</italic></sub> reported in Equation (8). As can be observed, for <italic>N</italic><sub><italic>V</italic></sub> &#x02248; 1, the approximation (Equation 8) for <italic>N</italic><sub><italic>V</italic></sub> is highly accurate, indicating that this approximation can be safely applied to find the &#x0201C;perfect recovery&#x0201D; condition.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Contour plot of log<sub><bold>10</bold></sub>(<italic><bold>N</bold></italic><sub><italic><bold>V</bold></italic></sub>) from Equations (6) (full lines) and (8) (dashed lines), in the <italic><bold>P</bold></italic> and <italic><bold>N</bold></italic> range 0&#x02013;100</bold>. The lines are the loci of the points where <italic>log</italic><sub>10</sub>(<italic>N</italic><sub><italic>V</italic></sub>) equals 0.0 (red), 0.4, 0.8, 1.2, and 1.6 (black), as indicated on the right side of the figure. The blue line represents <italic>P</italic> &#x0003D; 0.14<italic>N</italic>, while the black dotted line is the bisectrix <italic>N</italic> &#x0003D; <italic>P</italic>, plotted to emphasize the symmetry of the contour lines.</p></caption>
<graphic xlink:href="fncom-10-00144-g0004.tif"/>
</fig>
<p>In the <italic>P</italic>-<italic>N</italic> plane the existence of two regions (small and large &#x003B1;) where the perfect recovery (<italic>N</italic><sub><italic>V</italic></sub> &#x0003D; 1, red lines) takes place can be easily observed and the result is symmetric under the exchange of <italic>P</italic> and <italic>N</italic>. In the already explored small &#x003B1; region, we also show (full blue line) the <italic>P</italic> &#x0003D; 0.14<italic>N</italic> condition. Similar to the high &#x003B1; region, it is important to find a simple relation between <italic>N</italic> and <italic>P</italic> identifying the <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; 1 condition. We aim, therefore, to obtain a function <italic>P</italic>(<italic>N</italic>) which returns, at given <italic>N</italic>, the <italic>P</italic>-value such that <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; 1. We write the prefactor <italic>NP</italic> in Equation (11) as &#x003B1;<italic>N</italic><sup>2</sup> and exploit the &#x003B1;&#x0226B;1 limit, so to obtain <inline-formula><mml:math id="M51"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> exp<inline-formula><mml:math id="M52"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula>. The equation <italic>N</italic><sup>2</sup>&#x003B1;<sup>1/2</sup>exp<inline-formula><mml:math id="M53"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> can be squared, &#x003B1;exp(&#x02212;&#x003B1;) &#x0003D; 2&#x003C0;/<italic>N</italic><sup>4</sup>, and solved with respect to &#x003B1;, to give <inline-formula><mml:math id="M54"><mml:mi>&#x003B1;</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mo>/</mml:mo><mml:msup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>W</italic><sub>&#x02212;1</sub>(<italic>x</italic>) is the second real branch of the Lambert function (Olver et al., <xref ref-type="bibr" rid="B19">2010</xref>). In conclusion, the &#x0201C;perfect recovery condition&#x0201D; is satisfied -for each <italic>N</italic>-value- if we chose to store a number of memories <italic>larger</italic> than <italic>P</italic>(<italic>N</italic>) given by:
<disp-formula id="E11"><label>(11)</label><mml:math id="M55"><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>N</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mo>/</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mn>4</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
For practical purposes, for large enough <italic>N</italic>, we can use the small-argument expansion of the Lambert function &#x02212;<italic>W</italic><sub>&#x02212;1</sub>(&#x02212;<italic>x</italic>) &#x02248; &#x02212;ln(<italic>x</italic>)&#x0002B;ln(&#x02212;ln(<italic>x</italic>)) (Corless et al., <xref ref-type="bibr" rid="B9">1996</xref>), to have:
<disp-formula id="E12"><label>(12)</label><mml:math id="M56"><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>ln</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mn>4</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mtext>ln</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>ln</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mn>4</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
The results for <italic>P</italic>(<italic>N</italic>) are shown in Figure <xref ref-type="fig" rid="F5">5</xref> as a function of <italic>N</italic> in the range 1&#x02013;1000. The black line represents the exact, numerical, solution to <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; 1, with <italic>N</italic><sub><italic>V</italic></sub> in Equation (6), the blue line is the expression for <italic>P</italic>(<italic>N</italic>) in Equation (11), while the red line is those in Equation (12).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>The quantity <italic><bold>P</bold></italic>(<italic><bold>N</bold></italic>), i.e., the <italic><bold>P</bold></italic>-value where the perfect recovery is guaranteed, is shown as a function of <italic><bold>N</bold></italic></bold>. The blue line is the numerical solution of <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; 1 from Equation (6), the blue line is the plot of Equation (12) and the red line is the plot of Equation (13).</p></caption>
<graphic xlink:href="fncom-10-00144-g0005.tif"/>
</fig>
<p>It is important to note that the presence of a decrease of the retrieval error probability at high <italic>P</italic>, or &#x003B1;, values is due to the presence of non-zero diagonal elements in the <italic>J</italic> matrix that creates a coherent term of weight <italic>P</italic>. Indeed, repeating the rationale leading to Equation (4) with the assumption that <italic>J</italic><sub><italic>ii</italic></sub> &#x0003D; 0, would give rise to the same (Equations 4&#x02013;6) but with the numerator of the argument of the error functions equal to <italic>N</italic> &#x02212; 1 instead of to <italic>N</italic> &#x0002B; <italic>P</italic> &#x02212; 1. This is shown graphically in Figure <xref ref-type="fig" rid="F6">6</xref> where we compare for <italic>N</italic> &#x0003D; 50, both theoretically (full line) and numerically (full dots), the quantities <italic>p</italic><sub><italic>B</italic></sub>, <italic>p</italic><sub><italic>V</italic></sub>, and <italic>N</italic><sub><italic>V</italic></sub> as a function of <italic>P</italic> in the two cases: diagonal elements in Equation (2) (black) and diagonal elements forced to vanish (orange).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>The upper panel (A) reports for a given <italic><bold>N</bold></italic>-value (<italic><bold>N</bold></italic> &#x0003D; 50), as a function of P, the probability <italic><bold>p</bold></italic><sub><italic><bold>B</bold></italic></sub> that, stimulating the network with a vector inside the training set, there is one bit wrong in the network response</bold>. The middle panel <bold>(B)</bold> reports <italic>p</italic><sub><italic>V</italic></sub>, the probability that, stimulating the network with a vector inside the training set, the vector obtained after one dynamical step is not the stimulating vector. The lower panel <bold>(C)</bold> reports <italic>N</italic><sub><italic>V</italic></sub> &#x0003D; <italic>Pp</italic><sub><italic>V</italic></sub>. The black symbols/lines refer to the case where the diagonal elements are as determined in Equation (2), while the oranges ones to diagonal elements forced to vanish. The full lines are the theoretical prediction, the full dots are the results of the numerical simulation.</p></caption>
<graphic xlink:href="fncom-10-00144-g0006.tif"/>
</fig>
<p>The stabilization of the fixed points <inline-formula><mml:math id="M57"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in the high storage region arises from the presence of the non-zero diagonal elements. Asymptotically, on increasing <italic>P</italic>, the diagonal elements growth coherently and the <bold>J</bold> matrix tends to become the unit matrix. However, the dynamics (see Equation 1) dictated by the matrix J does not tend to the dynamics dictated by the unit matrix. In the latter case, indeed, all the 2<sup><italic>N</italic></sup> state vectors should become fixed points and the network should loose on important feature: the capability to distinguish between the stored memories (the vectors <inline-formula><mml:math id="M58"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, for &#x003BC; &#x0003D; 1&#x02026;<italic>P</italic>) and the spurious fixed points, all the vectors <inline-formula><mml:math id="M59"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B6;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula> not belonging to the set <inline-formula><mml:math id="M60"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> but such that <inline-formula><mml:math id="M61"><mml:mi>E</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B6;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> &#x0003D; <inline-formula><mml:math id="M62"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B6;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>. To study this property, we have calculated the probability that a (randomly chosen) vector <inline-formula><mml:math id="M63"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B6;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula> (different from all the <inline-formula><mml:math id="M64"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> used to build the <italic>J</italic> matrix) was recognized as a &#x0201C;memory&#x0201D; from the network dynamics. To be consistent with the previous notation (where we called <italic>p</italic><sub><italic>B</italic></sub> and <italic>p</italic><sub><italic>V</italic></sub> the probability of errors, not that of correct retrieval of the memory states) we define <inline-formula><mml:math id="M65"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (<inline-formula><mml:math id="M66"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) as the probability of correctly not retrieving a vector not belonging to the training set. More specifically, the quantity <inline-formula><mml:math id="M67"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the probability that one dynamical step after presenting a vector &#x003B6; not belonging to the training set to the network, the output a vector is different from &#x003B6;.&#x0201D; More specifically, the quantity <inline-formula><mml:math id="M68"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the probability that presenting a vector &#x003B6; not belonging to the training set to the network, after one dynamical step we found as output a vector different from &#x003B6;. Similarly for <inline-formula><mml:math id="M69"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. It turns out that <xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>:
<disp-formula id="E13"><label>(13)</label><mml:math id="M70"><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mi>B</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mi>P</mml:mi><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="E14"><label>(14)</label><mml:math id="M71"><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mi>V</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mtext>erf</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mi>P</mml:mi><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x0200B;</mml:mtext><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>N</mml:mi></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula>
In Figure <xref ref-type="fig" rid="F7">7</xref> we report the comparison of the <italic>P</italic> dependence of <italic>p</italic><sub><italic>B</italic></sub> and <inline-formula><mml:math id="M72"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Figure <xref ref-type="fig" rid="F7">7A</xref>) and that of <italic>p</italic><sub><italic>V</italic></sub> and <inline-formula><mml:math id="M73"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Figure <xref ref-type="fig" rid="F7">7B</xref>). As usual, full lines are the theoretical results, while the full dots are the outcome of the numerical simulation. Black data are for the &#x0201C;memory states,&#x0201D; while the green ones are for the &#x0201C;spurious state.&#x0201D; As can be seen, the spurious state becomes more and more &#x0201C;present&#x0201D; in the set of memories stored by the network as <italic>P</italic> increases. It seems however that also at high <italic>P</italic>-values the retrieval of the memory states is reasonably good and that of the spurious states reasonably bad.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>(A)</bold> The upper panel reports for a given <italic>N</italic>-value (<italic>N</italic> &#x0003D; 50), as a function of P, the probability that, stimulating the network with a vector inside (<italic>p</italic><sub><italic>B</italic></sub>, black) or outside (<inline-formula><mml:math id="M74"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, green) the training set, there is one bit differing between the input and the output vector. <bold>(B)</bold> The lower panel reports the probability that, stimulating the network with a vector inside (<italic>p</italic><sub><italic>V</italic></sub>, black) or outside (<inline-formula><mml:math id="M75"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, green) the training set, the vector obtained after one dynamical step is not the stimulating vector. The full lines are the theoretical prediction, the full dots are the results of the numerical simulation.</p></caption>
<graphic xlink:href="fncom-10-00144-g0007.tif"/>
</fig>
<p>To be quantitative on this point, we rewrite Equation (14) in its large <italic>N</italic> limit:
<disp-formula id="E15"><label>(15)</label><mml:math id="M76"><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mi>V</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mtext>&#x02009;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mfrac><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:math></disp-formula>
and compare it with Equation (7). In particular, is interesting to calculate the ratio, &#x003C1;, between the probability of wrong retrieval of a spurious state and that of a memory state: &#x003C1; &#x0003D; <inline-formula><mml:math id="M77"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. From Equations (7) and (15) it turns out:
<disp-formula id="E16"><label>(16)</label><mml:math id="M78"><mml:mi>&#x003C1;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi></mml:mrow><mml:mi>P</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x02009;&#x02009;</mml:mtext><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula>
This quantity only depends on &#x003B1;:
<disp-formula id="E17"><label>(17)</label><mml:math id="M79"><mml:mi>&#x003C1;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x02009;&#x02009;</mml:mtext><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></disp-formula>
and &#x003C1; has a finite high &#x003B1; limit:
<disp-formula id="E18"><label>(18)</label><mml:math id="M80"><mml:munder><mml:mrow><mml:mi>lim</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow></mml:munder><mml:mi>&#x003C1;</mml:mi><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula>
In other words, although the number of spurious attractors tends to increase for <italic>P</italic> &#x0226B; <italic>N</italic>, the vectors encoded into the system through the connection matrix are retrieved with an efficiency almost three times better than for the spurious states.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>In this work we have developed a simple theoretical approach to investigate the computational properties and the storage capacity of feed-forward networks with self-connections. We have worked out an exact expression which gives the probability <italic>p</italic><sub><italic>B</italic></sub> of having a wrong bit in the recovery of a memory element from a Hebbian <italic>N</italic>-node neural network, where <italic>P</italic> memory elements are stored. In disagreement with previous studies we have investigated the case in which the diagonal elements were not forced to vanish. Studying the storage capacity, and deriving the related probability <italic>p</italic><sub><italic>V</italic></sub> and number <italic>N</italic><sub><italic>V</italic></sub> of having a wrongly recovered memory element, we discovered that besides the well know <italic>P</italic>&#x0226A;<italic>N</italic> region, there is another region, at <italic>P</italic>&#x0226B;<italic>N</italic>, where the recovery is highly effective. When <italic>P</italic>&#x0226B;<italic>N</italic>, the efficiency of recall for a large number of encoded vectors in the <italic>J</italic> matrix is related to the presence of non-zero diagonal elements of the matrix. Basically, the higher storage performance of the network depends on the number of &#x0201C;coherent&#x0201D; terms (the signal) in the quantity <inline-formula><mml:math id="M81"><mml:msubsup><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (see Section 3.1) with respect to the &#x0201C;incoherent&#x0201D; ones (the noise). The larger the ratio between coherent to incoherent terms, the lower the probability of a wrong recovery. The number of coherent terms is (<italic>N</italic> &#x0002B; <italic>P</italic> &#x02212; 1) in the case of autapses, it is (<italic>N</italic>&#x02212;1) in the case of no autapses. Indeed, the <italic>P</italic> terms disappear if the diagonal is forced to be zero as in the standard Hopfield model. It is clear that, apart from a transient regime at <italic>P</italic> &#x0007E; <italic>N</italic>, increasing <italic>P</italic> &#x0226B; <italic>N</italic> strongly reinforces the signal-to-noise ratio and induces a much larger storage capacity. In addition to the vectors encoded into the system, other unwanted memories also appear in the network. These are the spurious states, fixed points which do not belong to the training set. The presence of spurious states is not a feature specific to the present model, it is a typical characteristic of the standard Hopfield network and its successive improvements. Indeed, as shown by Tanaka and Edwards (<xref ref-type="bibr" rid="B23">1980</xref>), a random <italic>N</italic> &#x000D7; <italic>N</italic> matrix has 2<sup>&#x003B3;<italic>N</italic></sup> fixed points (&#x003B3; &#x02248; 2). As an example, if <italic>N</italic> &#x0003D; 100, the number of fixed points is about one million. A Hebbian 100 &#x000D7; 100 matrix storing <italic>P</italic> &#x0003D; 1000 patterns, besides the &#x0201C;good&#x0201D; <italic>P</italic> fixed points have also an overwhelming number of spurious fixed points (or &#x0201C;false memories&#x0201D;). The interest of our approach does not rely in &#x0201C;how many&#x0201D; spurious (i.e., not belonging to the training set) states are present but rather in how the recognition of a vector belonging to the training set is as a &#x0201C;good&#x0201D; one. Obviously, the argument of Tanaka-Edwards applies only to random matrices. The Hebbian form, with or without autapses, is not fully random (there exists correlation among the matrix elements), but we expect a number of fixed points similar to that of a random matrix. It would be interesting to determine such a number, but this is beyond the scope of the present paper. In spite of the overwhelming majority of spurious fixed points, the network&#x02014;even at very large <italic>P</italic>-values, maintains the capacity of discriminate between &#x0201C;good&#x0201D; state (belonging to the training set) and &#x0201C;wrong&#x0201D; ones (not belonging to the training set). More specifically, looking at the one-step dynamical evolution and comparing the input vector with the output one, we have posed to the network the question: &#x0201C;is the input vector belonging to the training set&#x0201D;? We have demonstrated that, when the input vector actually belongs to the training set, at large <italic>P</italic> (similarly to low <italic>P</italic>) the probability of having a wrong response (&#x0201C;no, it does not belong to the training set) goes to zero. Furthermore, we have demonstrated that when the input vector does not belong to the training set the probability of a wrong response (&#x0201C;yes, it is a fixed point&#x0201D;) is much less that in the previous case, asymptotically 2.7 time worst.</p>
<p>In order to identify whether or not a vector belonging to the training set was a fixed point we propose to the system a vector of the training set as input. Then we perform a one-step dynamic evolution of this input state. If after one step the output vector is equal to the input one, this is a fixed point. On the contrary, if after one step the output vector is not equal to the input one, it could be possible that further dynamical steps lead to the input vector. From this point on, as the dynamic is deterministic, the system enters a limit cycle (of length greater than one). Since it is not clear whether or not a limit cycle can be considered a &#x0201C;right recognition,&#x0201D; we have excluded this possibility from the counts of the right recognition. Only fixed point are considered &#x0201C;good.&#x0201D; For this reason, to determine the probability of &#x0201C;right recognition&#x0201D; one dynamical step is enough. We have also not considered the possibility that, using as input a vector not belonging to the training set, it converges to one of the training vectors. The probability of right recognition reported here is an underestimation of the network capability. A further quantity that it would be interesting to evaluate is the size of the attraction basin of a given fixed point, i.e., how many non-training vectors converge to a given training vector fixed point. The basins size would be an important measure of the network performance, their determination is however difficult to achieve analytically, and is behind the goal of the present paper.</p>
<p>One important finding is summarized in Equation (18). It states that for <italic>P</italic> &#x0226B;<italic>N</italic>, when the connection matrix is dominated by the diagonal term and is still different from the unity matrix (this is due to the great number of off-diagonal elements with zero average and RMS of the order of <inline-formula><mml:math id="M82"><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula>), the network retains its capacity of give more &#x0201C;good&#x0201D; than &#x0201C;wrong&#x0201D; answers. This property, the fact that the limit in Equation (18) is <italic>e</italic> and not &#x0201C;1,&#x0201D; can be ascribed to the observation that, although the matrix <bold>J</bold> tends to the unit matrix for large <italic>P</italic>, the dynamics (see Equation 1) dictated by the matrix J does not tend to the dynamics dictated by the unit matrix. This finding opens the way to a much more efficient use of the artificial Hebbian neural network for information storage. In the first region, as well known since 40 years, the storage capacity is limited as the number of encoded vectors becomes of the order <italic>N</italic>. Indeed, in the high &#x003B1; region, the number of elements is basically unlimited <xref ref-type="fn" rid="fn0004"><sup>4</sup></xref>, when the number of stored elements is taken <italic>larger</italic> than &#x02248; 4<italic>N</italic>ln(<italic>N</italic>).</p>
</sec>
<sec id="s5">
<title>Author contributions</title>
<p>GR designed research. VF, ML, and GR performed numerical simulations, analyzed data and wrote the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abu-Mostafa</surname> <given-names>Y. S.</given-names></name> <name><surname>St. Jacques</surname> <given-names>J.-M.</given-names></name></person-group> (<year>1985</year>). <article-title>Information capacity of the Hopfield model</article-title>. <source>IEEE Trans. Inf. Theory</source> <volume>IT-31</volume>, <fpage>461</fpage>. <pub-id pub-id-type="doi">10.1109/tit.1985.1057069</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Amit</surname> <given-names>D. J.</given-names></name></person-group> (<year>1989</year>). <source>Modelling Brain Function: The World of Attractor Neural Networks</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amit</surname> <given-names>D. J.</given-names></name> <name><surname>Gutfreund</surname> <given-names>H.</given-names></name> <name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>1985a</year>). <article-title>Storing infinite numbers of patterns in a spin-glass model of neural networks</article-title>. <source>Phys. Rev. Lett.</source> <volume>55</volume>:<fpage>1530</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.55.1530</pub-id><pub-id pub-id-type="pmid">10031847</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amit</surname> <given-names>D. J.</given-names></name> <name><surname>Gutfreund</surname> <given-names>H.</given-names></name> <name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>1985b</year>). <article-title>Spin-glass models of neural networks</article-title>. <source>Phys. Rev. A</source> 32:1007. <pub-id pub-id-type="doi">10.1103/physreva.32.1007</pub-id><pub-id pub-id-type="pmid">9896156</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bastolla</surname> <given-names>U.</given-names></name> <name><surname>Parisi</surname> <given-names>G.</given-names></name></person-group> (<year>1997</year>). <article-title>Attractors in fully asymmetric neural networks</article-title>. <source>J. Phys. A Math. Gen</source>. <volume>30</volume>:<fpage>5613</fpage>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunel</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>Is cortical connectivity optimized for storing information?</article-title> <source>Nat. Neurosci</source>. <volume>19</volume>:<fpage>749</fpage>. <pub-id pub-id-type="doi">10.1038/nn.4286</pub-id><pub-id pub-id-type="pmid">27065365</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>L. N.</given-names></name></person-group> (<year>1973</year>). <article-title>A possible organization of animal memory and learning</article-title>, in <source>Proceedings of the Nobel Symposium on Collective Properties of Physical Systems,</source> eds <person-group person-group-type="editor"><name><surname>Lundquist</surname> <given-names>B.</given-names></name> <name><surname>Lundquist</surname> <given-names>S.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>252</fpage>&#x02013;<lpage>264</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>L. N.</given-names></name> <name><surname>Liberman</surname> <given-names>F.</given-names></name> <name><surname>Oja</surname> <given-names>E.</given-names></name></person-group> (<year>1979</year>). <article-title>A theory for the acquisition and loss of neuron specificity in visual cortex</article-title>. <source>Biol. Cybern.</source> <volume>33</volume>:<fpage>9</fpage>. <pub-id pub-id-type="doi">10.1007/BF00337414</pub-id><pub-id pub-id-type="pmid">222355</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corless</surname> <given-names>R. M.</given-names></name> <name><surname>Gonnet</surname> <given-names>G. H.</given-names></name> <name><surname>Hare</surname> <given-names>D. E. G.</given-names></name> <name><surname>Jeffrey</surname> <given-names>D. J.</given-names></name> <name><surname>Knuth</surname> <given-names>D. E.</given-names></name></person-group> (<year>1996</year>). <article-title>On the LambertW function</article-title>. <source>Adv. Comput. Math.</source> <volume>5</volume>:<fpage>329</fpage>. <pub-id pub-id-type="doi">10.1007/BF02124750</pub-id><pub-id pub-id-type="pmid">28043612</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Derrida</surname> <given-names>B.</given-names></name></person-group> (<year>1989</year>). <article-title>Distribution of the activities in a diluted neural network</article-title>. <source>J. Phys. A Math. Gen.</source> <volume>22</volume>:<fpage>2069</fpage>. <pub-id pub-id-type="doi">10.1088/0305-4470/22/12/012</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eccles</surname> <given-names>J. G.</given-names></name></person-group> (<year>1953</year>). <source>The Neurophysiological Basis of Mind</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Clarendon</publisher-name>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gutfreundt</surname> <given-names>H.</given-names></name> <name><surname>Regert</surname> <given-names>J. D.</given-names></name> <name><surname>Young</surname> <given-names>A. P.</given-names></name></person-group> (<year>1988</year>). <article-title>The nature of attractors in an asymmetric spin glass with deterministic dynamics</article-title>. <source>J. Phys. A Math. Gen.</source> <volume>21</volume>:<fpage>2775</fpage>. <pub-id pub-id-type="doi">10.1088/0305-4470/21/12/020</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haykin</surname> <given-names>S.</given-names></name></person-group> (<year>1999</year>). <source>Neural Networks: A Comprehensive Foundation</source>. <publisher-name>Prentice Hall</publisher-name>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://lecteurepub-a507a.firebaseapp.com/3RO4lXk7OXbEm12/Neural%20Networks%20A%20Comprehensive%20Foundation%202nd%20Edition%20Ebooks%20Gratuit.pdf">https://lecteurepub-a507a.firebaseapp.com/3RO4lXk7OXbEm12/Neural%20Networks%20A%20Comprehensive%20Foundation%202nd%20Edition%20Ebooks%20Gratuit.pdf</ext-link></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hebb</surname> <given-names>D. O.</given-names></name></person-group> (<year>1949</year>). <source>The Organization of Behavior</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley</publisher-name>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>1982</year>). <article-title>Neural networks and physical systems with emergent collective computational abilities</article-title>. <source>Proc. Nat. Acad. Sci. U.S.A.</source> <volume>79</volume>:<fpage>2554</fpage>. <pub-id pub-id-type="doi">10.1073/pnas.79.8.2554</pub-id><pub-id pub-id-type="pmid">6953413</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name></person-group> (<year>1984</year>). <article-title>Neurons with graded response have collective computational properties like those of two-state neurons</article-title>. <source>Proc. Nat. Acad. Sci. U.S.A.</source> <volume>81</volume>:<fpage>3088</fpage>. <pub-id pub-id-type="doi">10.1073/pnas.81.10.3088</pub-id><pub-id pub-id-type="pmid">6587342</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopfield</surname> <given-names>J. J.</given-names></name> <name><surname>Feinstein</surname> <given-names>D. I.</given-names></name> <name><surname>Palmer</surname> <given-names>R. G.</given-names></name></person-group> (<year>1983</year>). <article-title>&#x02018;Unlearning&#x02019; has a stabilizing effect in collective memories</article-title> <source>Nature</source> <volume>304</volume>, <fpage>158</fpage>. <pub-id pub-id-type="pmid">6866109</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mc Eliece</surname> <given-names>R. J.</given-names></name> <name><surname>Posner</surname> <given-names>E. C.</given-names></name> <name><surname>Rodemich</surname> <given-names>E. R.</given-names></name> <name><surname>Venkatesh</surname> <given-names>S. S.</given-names></name></person-group> (<year>1987</year>). <article-title>The capacity of the hopfield associative memory</article-title>. <source>IEEE Trans. Inf. Theory</source> <volume>IT-33</volume>, <fpage>461</fpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Olver</surname> <given-names>F. W. J.</given-names></name> <name><surname>Lozier</surname> <given-names>D. W.</given-names></name> <name><surname>Boisvert</surname> <given-names>R. F.</given-names></name> <name><surname>Clark</surname> <given-names>C. W.</given-names></name></person-group> (eds.). (<year>2010</year>). <source>Handbook of Mathematical Functions</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rojas</surname> <given-names>R.</given-names></name></person-group> (<year>1996</year>). <source>Neural Networks</source>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sollacher</surname> <given-names>R.</given-names></name> <name><surname>Gao</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Towards real-world applications of online learning spiral recurrent neural networks</article-title>. <source>J. Intell. Learn. Syst. Appl.</source> <volume>1</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.4236/jilsa.2009.11001</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sompolinsky</surname> <given-names>H.</given-names></name> <name><surname>Crisanti</surname> <given-names>A.</given-names></name> <name><surname>Sommers</surname> <given-names>H.</given-names></name></person-group> (<year>1988</year>). <article-title>Chaos in random neural networks</article-title>. <source>Phys. Rev. Lett.</source> <volume>61</volume>:<fpage>259</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.61.259</pub-id><pub-id pub-id-type="pmid">10039285</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>F.</given-names></name> <name><surname>Edwards</surname> <given-names>S. F.</given-names></name></person-group> (<year>1980</year>). <article-title>Analytic theory of the ground state properties of a spin glass. I. Ising spin glass</article-title>. <source>J. Physi. F Met. Phys.</source> <volume>10</volume>:<fpage>2769</fpage>. <pub-id pub-id-type="doi">10.1088/0305-4608/10/12/017</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wainrib</surname> <given-names>G.</given-names></name> <name><surname>Touboul</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Topological and dynamical complexity of random neural networks</article-title>. <source>Phys. Rev. Lett.</source> <volume>118</volume>:<fpage>101259</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.110.118101</pub-id></citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>The evolution <italic>E</italic>[&#x003BE;<sub><italic>i</italic></sub>] always return &#x003BE;<sub><italic>i</italic></sub> since the sum <inline-formula><mml:math id="M83"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>.</p></fn>
<fn id="fn0002"><p><sup>2</sup>These are <italic>P</italic> terms that are present only if the diagonal elements are kept as they are and are not forced to vanish.</p></fn>
<fn id="fn0003"><p><sup>3</sup>The calculation follows the same steps already depicted before, counting the &#x0201C;coherent&#x0201D; terms, that, in this case, only arise from the diagonal elements (<italic>i</italic>=<italic>j</italic>) and not from the &#x003BC;=&#x003BD; terms that now do not exist. The weight of the coherent part is equal to <italic>P</italic> instead of <italic>N</italic> &#x0002B; <italic>P</italic> &#x02212; 1. The rest of the demonstration follows straightforward.</p></fn>
<fn id="fn0004"><p><sup>4</sup>The word unlimited is obviously non-physical. However, the value of <italic>P</italic><sub><italic>o</italic></sub> arising from the Tanaka-Edwards relation (Tanaka and Edwards, <xref ref-type="bibr" rid="B23">1980</xref>), <italic>P</italic><sub><italic>o</italic></sub> &#x0003D; exp(&#x003B3;<italic>N</italic>), with &#x003B3; &#x0003D; 0.2. <italic>P</italic><sub><italic>o</italic></sub>, already at the <italic>N</italic>-values reported in Equation (5) is so large (<inline-formula><mml:math id="M84"><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>1000</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02248;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>87</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) with respect to <italic>P</italic>(<italic>N</italic>) (<italic>P</italic>(<italic>N</italic> &#x0003D; 1000) &#x02248; 210<sup>4</sup> to state that <italic>P</italic><sub><italic>o</italic></sub> is unlimited to any practical purpose.</p></fn>
</fn-group>
</back>
</article>