<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2021.767953</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Impact of Asymmetric Weight Update on Neural Network Training With Tiki-Taka Algorithm</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Lee</surname> <given-names>Chaeun</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1461933/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Noh</surname> <given-names>Kyungmi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Ji</surname> <given-names>Wonjae</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Gokmen</surname> <given-names>Tayfun</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/339406/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Kim</surname> <given-names>Seyoung</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/191254/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Materials Science and Engineering, Pohang University of Science and Technology</institution>, <addr-line>Pohang-si</addr-line>, <country>South Korea</country></aff>
<aff id="aff2"><sup>2</sup><institution>IBM Research AI</institution>, <addr-line>Yorktown Heights, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yimao Cai, Peking University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Hongwu Jiang, Georgia Institute of Technology, United States; Zhong Sun, Peking University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Seyoung Kim <email>kimseyoung&#x00040;postech.ac.kr</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Neural Technology, a section of the journal Frontiers in Neuroscience</p></fn>
<fn fn-type="present-address" id="fn002"><p>&#x02020;Present address: Chaeun Lee, NAVER Clova, Seongnam-si, South Korea</p></fn></author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>767953</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Lee, Noh, Ji, Gokmen and Kim.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Lee, Noh, Ji, Gokmen and Kim</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Recent progress in novel non-volatile memory-based synaptic device technologies and their feasibility for matrix-vector multiplication (MVM) has ignited active research on implementing analog neural network training accelerators with resistive crosspoint arrays. While significant performance boost as well as area- and power-efficiency is theoretically predicted, the realization of such analog accelerators is largely limited by non-ideal switching characteristics of crosspoint elements. One of the most performance-limiting non-idealities is the conductance update asymmetry which is known to distort the actual weight change values away from the calculation by error back-propagation and, therefore, significantly deteriorates the neural network training performance. To address this issue by an algorithmic remedy, Tiki-Taka algorithm was proposed and shown to be effective for neural network training with asymmetric devices. However, a systematic analysis to reveal the required asymmetry specification to guarantee the neural network performance has been unexplored. Here, we quantitatively analyze the impact of update asymmetry on the neural network training performance when trained with Tiki-Taka algorithm by exploring the space of asymmetry and hyper-parameters and measuring the classification accuracy. We discover that the update asymmetry level of the auxiliary array affects the way the optimizer takes the importance of previous gradients, whereas that of main array affects the frequency of accepting those gradients. We propose a novel calibration method to find the optimal operating point in terms of device and network parameters. By searching over the hyper-parameter space of Tiki-Taka algorithm using interpolation and Gaussian filtering, we find the optimal hyper-parameters efficiently and reveal the optimal range of asymmetry, namely the <italic>asymmetry specification</italic>. Finally, we show that the analysis and calibration method be applicable to spiking neural networks.</p></abstract>
<kwd-group>
<kwd>resistive memory</kwd>
<kwd>update asymmetry</kwd>
<kwd>Tiki-Taka algorithm</kwd>
<kwd>neural network</kwd>
<kwd>deep learning accelerator</kwd>
<kwd>analog AI hardware</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="0"/>
<equation-count count="17"/>
<ref-count count="32"/>
<page-count count="13"/>
<word-count count="8003"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Rapid advances in artificial neural network-based machine learning techniques enable significant performance boost in various artificial intelligence (AI) applications exemplified by image classification and natural language processing. As the larger and deeper neural networks trained with more data generally show higher performance in such cognitive tasks, there is an increasing demand for higher computing power and memory bandwidth to support the large number of operations and data processing required for neural network training and inference. Therefore, improving speed and energy-efficiency in AI computing hardware is one of the key challenges to realize more advanced AI applications and to extend the AI application space on low-power systems such as internet of things (IoT) and edge computing devices (Verhelst and Moons, <xref ref-type="bibr" rid="B28">2017</xref>). To address the issue, various optimization techniques such as quantization (Guo, <xref ref-type="bibr" rid="B9">2018</xref>) and compression (Han et al., <xref ref-type="bibr" rid="B11">2015</xref>) are proposed to reduce the size and number of required computations. Along with such techniques for maximizing the efficiency of the existing hardware, there have been efforts to improve the digital hardware architecture (Chen et al., <xref ref-type="bibr" rid="B4">2020</xref>) for energy-efficient AI computing (Zhou et al., <xref ref-type="bibr" rid="B32">2019</xref>).</p>
<p>As an alternative to the existing digital approaches, analog crosspoint array-based neural network computation accelerators have been intensively studied due to the advances in resistive memory device technologies (Haensch et al., <xref ref-type="bibr" rid="B10">2018</xref>; Tsai et al., <xref ref-type="bibr" rid="B26">2018</xref>; Kim et al., <xref ref-type="bibr" rid="B16">2019b</xref>) and their feasibility for various matrix-involved computations such as inversion, eigenvector solving, matrix pseudoinverse (Sun et al., <xref ref-type="bibr" rid="B25">2019</xref>; Wang et al., <xref ref-type="bibr" rid="B29">2020</xref>) and matrix-vector multiplication (MVM). Especially with MVM, by storing weight matrix in the crosspoint array of resistive memory devices and applying voltages corresponding to the input vector values, one can perform fully-parallel neural network computations in analog domain. As the time complexity of MVM with a resistive crosspoint array is approximately <italic>O</italic>(1) (Sun and Huang, <xref ref-type="bibr" rid="B24">2021</xref>), such analog AI hardware is expected to have significant acceleration factor and higher energy efficiency, compared to those of digital counterparts (Agarwal et al., <xref ref-type="bibr" rid="B1">2016</xref>; Gokmen and Vlasov, <xref ref-type="bibr" rid="B8">2016</xref>). While the concept of the resistive crosspoint array-based neural network computation accelerator is promising, the actual implementation of such system has been difficult due to the non-ideal memory device characteristics. One of the major non-idealities which dramatically degrades the system performance is the conductance update asymmetry which indicates the nonidentical amount of up and down conductance changes at a given conductance level (Islam et al., <xref ref-type="bibr" rid="B13">2019</xref>; Xiao et al., <xref ref-type="bibr" rid="B30">2020</xref>). The update asymmetry causes an unexpected dynamic biasing during the training and prevents the weights from reaching the optimum values (Kim et al., <xref ref-type="bibr" rid="B15">2019a</xref>). One obvious direction to resolve this issue is to build a memory device which features symmetric update property. However, the device with an ideal update symmetry is still under development (Lee et al., <xref ref-type="bibr" rid="B20">2020</xref>). The other direction is to develop a special training algorithm which can train the neural network even with asymmetric devices. Tiki-Taka algorithm (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>) resolves the asymmetry issue by adopting an auxiliary array, which stores the information about the history of gradients, along with the main array to store weight values. It is shown that this algorithm minimizes the performance drop caused by update asymmetry and allows robust neural network training even when significant asymmetry presents. However, detailed quantitative study to reveal the relation between degree of update asymmetry and network performance has yet to be explored.</p>
<p>In this work, we dissect the impact of update asymmetry existing in main and auxiliary arrays on neural network performance when Tiki-Taka algorithm is used to train a neural network. We show that update asymmetry in the system is, indeed, a hyper-parameter of the neural network in the optimization point of view, affecting its convergence and performance. Our analysis shows that the auxiliary array is virtually a part of the optimizer to train a neural network; the degree of update asymmetry in the auxiliary array changes the configuration of the optimizer, similar to the case of momentum SGD where a decay factor determines the optimizer (Rumelhart et al., <xref ref-type="bibr" rid="B22">1986</xref>). To further support this point, we compare the performance of different training algorithms on a convex problem and analyze the role of asymmetry in weight optimization process. On the other hand, if device characteristics such as update symmetry work as an hyper-parameter, it is important to fabricate devices with the hardware-side hyper-parameter in the range where the neural network performance is robust against other software-side hyper-parameters. We propose a method to identify the best asymmetry range by introducing a metric, <italic>robustness score</italic>, which provides a concrete rule for comparing the efficiency of training neural networks. By comparing <italic>robustness score</italic> among the different combinations of asymmetry levels in two arrays, it is possible to find the optimal asymmetry range for each array while maximizing the training efficiency and neural network performance. Optimized asymmetry range for each array can relieve the hardware constraints further hence hardware implementation of the algorithm can be accelerated. Finally, our approach can be applied to domains beyond deep neural networks (DNN), including spiking neural networks (SNN) which can be mapped into the resistive crosspoint arrays.</p>
</sec>
<sec id="s2">
<title>2. Preliminaries</title>
<sec>
<title>2.1. Stochastic Gradient Descent With Resistive Crosspoint Array</title>
<p>Stochastic gradient descent (SGD) is a widely-used, first-order optimization method for training neural networks. During forward pass, a neural network infers output and estimates loss by comparing the output with expected values. Based on the loss, gradients of weights in a neural network are calculated during backward pass and later updated during the weight update phase by error back-propagation algorithm. The amount of weight update is proportional to the calculated gradient with the scaling factor called learning rate, &#x003B7;:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable class="aligned"><mml:mtr><mml:mtd columnalign="right"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd columnalign="left"><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>w</italic><sub><italic>ij</italic></sub> indicates a weight from pre-synaptic neuron <italic>i</italic> to post-synaptic neuron <italic>j</italic>, <italic>L</italic> is the loss of a neural network, and &#x02207;<sub><italic>ij, k</italic></sub><italic>L</italic> indicates the derivative of <italic>L</italic> with respect to <italic>w</italic><sub><italic>ij</italic></sub>. <italic>x</italic><sub><italic>i</italic></sub> is the input activation of pre-synaptic neuron <italic>i</italic>, and &#x003B4;<sub><italic>j</italic></sub> is the error of post-synaptic neuron <italic>j</italic> propagated from output neurons.</p>
<p>Analog computation with a resistive cross-point array resembles fixed point or quantized number system unlike the digital computation done with floating point number system. For the forward pass, the input activation, <italic>x</italic><sub><italic>i</italic></sub> is converted to a corresponding discrete voltage, <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, through analog-to-digital converter (ADC). Then, <italic>x</italic><sub><italic>i</italic></sub> is encoded into a pulse, <inline-formula><mml:math id="M3"><mml:mi>P</mml:mi><mml:mi>W</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, by pulse width modulation (PWM). As the pulse trains are applied to the rows, currents flowing through the devices are summed at each column line, and one can obtain the weighted sum, <italic>z</italic><sub><italic>j</italic></sub>:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>P</mml:mi><mml:mi>W</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where operation &#x0201C;&#x0002A;" includes the integration of currents through the time domain and <italic>g</italic><sub><italic>ij</italic></sub> is conductance value of <italic>i</italic><sup><italic>th</italic></sup> column and <italic>j</italic><sup><italic>th</italic></sup> row device. Then, analog-valued <italic>z</italic><sub><italic>j</italic></sub> is again quantized into <inline-formula><mml:math id="M5"><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> by ADC, and <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> propagates to the next layer as inputs. For the backward pass, the error of pre-synaptic neuron <italic>i</italic>, <italic>e</italic><sub><italic>i</italic></sub>, is calculated in a similar fashion to the errors of post-synaptic neurons, <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> :</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>P</mml:mi><mml:mi>W</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the update phase, <inline-formula><mml:math id="M9"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is converted to a pulse train, <inline-formula><mml:math id="M10"><mml:mi>S</mml:mi><mml:mi>P</mml:mi><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, by stochastic pulse generator (SPG) for stochastic update of a resistive memory device. In compliance with the rules in Equation (1), the update of <italic>g</italic><sub><italic>ij</italic></sub> is determined by <italic>x</italic><sub><italic>i</italic></sub> and &#x003B4;<sub><italic>j</italic></sub>.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x003B4;</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy='true'>(</mml:mo></mml:mstyle><mml:mi>S</mml:mi><mml:mi>P</mml:mi><mml:mi>G</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003B7;</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mi>q</mml:mi></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02227;</mml:mo><mml:mi>S</mml:mi><mml:mi>P</mml:mi><mml:mi>G</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003B7;</mml:mi><mml:mi>&#x003B4;</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x003B4;</mml:mi><mml:mi>j</mml:mi><mml:mi>q</mml:mi></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='true'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where SPG function is defined by pulse bit length (BL) and parameters for pulse probability, satisfying &#x003B7; &#x0003D; &#x003B7;<sub><italic>x</italic></sub>&#x003B7;<sub>&#x003B4;</sub>. &#x003B7;<sub><italic>x</italic></sub> and &#x003B7;<sub>&#x003B4;</sub> can be modulated, such that two SPGs will have similar pulse generating probability (Gokmen and Vlasov, <xref ref-type="bibr" rid="B8">2016</xref>; Gokmen et al., <xref ref-type="bibr" rid="B7">2017</xref>). &#x00394;<italic>g</italic><sub><italic>ij</italic></sub>(<italic>g</italic><sub><italic>ij</italic></sub>) is the amount of conductance change of <italic>g</italic><sub><italic>ij</italic></sub> as a function of <italic>g</italic><sub><italic>ij</italic></sub>, which is dependent on the update characteristics of the specific memory device. Here, we assume a linear model for &#x00394;<italic>g</italic><sub><italic>ij</italic></sub>(<italic>g</italic><sub><italic>ij</italic></sub>) which is described in section 2.2.</p>
<p>When mapping a neural network into resistive cross-point arrays for analog operations, we note that each weight can be expressed with a single or multiple memory cells depending on the architecture or weight mapping scheme (Xiao et al., <xref ref-type="bibr" rid="B30">2020</xref>). As an example, a pair of memory cells, <italic>g</italic><sub><italic>ij</italic></sub> and <italic>g</italic><sub><italic>ij, ref</italic></sub>, can be read differentially to represent a single weight: <italic>w</italic><sub><italic>ij</italic></sub> &#x0003D; <italic>K</italic>(<italic>g</italic><sub><italic>ij</italic></sub>&#x02212;<italic>g</italic><sub><italic>ij, ref</italic></sub>), where <italic>K</italic> is a factor that scales the weight into the conductance (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>).</p>
</sec>
<sec>
<title>2.2. Asymmetric Conductance Update in Resistive Memory Devices</title>
<p>Practical resistive memory devices feature various types of non-ideal characteristics such as asymmetric conductance response (or update asymmetry), short retention time, device-to-device variation and limited number of states. Among them, update asymmetry is known to hamper the convergence (Huang et al., <xref ref-type="bibr" rid="B12">2020</xref>) of SGD-based neural network training and causes a significant performance drop. As shown in Equation (4), it is the &#x00394;<italic>g</italic><sub><italic>ij</italic></sub>(<italic>g</italic><sub><italic>ij</italic></sub>) term which distorts the actual update from the ideal update behavior defined in Equation (1). That is, <italic>g</italic><sub><italic>ij</italic></sub> will represent a weight value distorted form ideal software calculation after several cycles of weight update unless &#x00394;<italic>g</italic><sub><italic>ij</italic></sub>(<italic>g</italic><sub><italic>ij</italic></sub>) is unity. For most variants of resistive memory devices, as illustrated in <xref ref-type="fig" rid="F1">Figure 1A</xref>, &#x00394;<italic>g</italic><sub><italic>ij</italic></sub>(<italic>g</italic><sub><italic>ij</italic></sub>) scales up or down the amounts of updates, which results in inconsistent updates of the conductance of devices (or weights). This type of behavior not only misguides the direction of updates, but also reduces the effective number of states in resistive memory devices.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>(A)</bold> Normalized conductance of three resistive memory devices with different AF values as a function of pulse number when 4,000 positive pulses (&#x0002B;<italic>V</italic>) and 4,000 negative pulses (&#x02212;<italic>V</italic>) are applied. The conductance change (&#x00394;<italic>g</italic>) with respect to the normalized conductance (<italic>g</italic>) is shown below. (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 3.1</xref> for the details of experimental environments.) <bold>(B)</bold> Comparison of conductance update process between SGD and Tiki-Taka algorithm. In case of SGD, an array, <italic>W</italic>, is updated with two parameters, <italic>x</italic> and &#x003B4; in the update phase. In Tiki-Taka algorithm, there are two steps: (1) update and (2) occasional transfer. In the update phase, like SGD, the auxiliary array, <italic>W</italic><sup><italic>A</italic></sup>, is updated. In the following transfer phase, its state is occasionally transferred and updates <italic>W</italic><sup><italic>C</italic></sup>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0001.tif"/>
</fig>
<p><xref ref-type="fig" rid="F1">Figure 1A</xref> illustrates the conductance(<italic>g</italic>) and the conductance change behavior (&#x00394;<italic>g</italic>) which can be described by the following equations, Equation (5):</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(6)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo> <mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>+</mml:mo></mml:msubsup></mml:mrow><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>+</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>when&#x000A0;</mml:mtext><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:mtext>0</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>&#x02212;</mml:mo></mml:msubsup></mml:mrow><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>&#x02212;</mml:mo></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x000A0;when&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mtext>0</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow> </mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>g</italic><sub><italic>min, ij</italic></sub> and <italic>g</italic><sub><italic>max, ij</italic></sub> is the minimum and maximum conductance state of resistive device and <inline-formula><mml:math id="M14"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x000B1;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is the amount of &#x00394;<italic>g</italic><sub><italic>ij</italic></sub> when <italic>g</italic><sub><italic>ij</italic></sub> of the device is at <italic>g</italic><sub>0, <italic>ij</italic></sub> &#x0003D; (<italic>g</italic><sub><italic>max, ij</italic></sub>&#x0002B;<italic>g</italic><sub><italic>min, ij</italic></sub>)/2. <italic>s</italic><sub><italic>ij</italic></sub>, namely <italic>slope</italic>, is defined as <inline-formula><mml:math id="M15"><mml:mi>&#x00394;</mml:mi><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x000B1;</mml:mo></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for nonlinear device, which represents the reciprocal of approximate number of states of the device. For ideal device, i.e., perfectly symmetric and linear device, <italic>slope</italic> is defined zero. <italic>d</italic><sub><italic>ij,k</italic></sub> is number of pulse needed to achieve desired amount of <italic>w</italic><sub><italic>ij,k</italic></sub> (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 1.1</xref> for the details).</p>
<p>To embody the update asymmetry of a device, we introduce a parameter called <italic>asymmetry factor</italic>, AF(&#x0003E;0), which is associated with existing parameters: <inline-formula><mml:math id="M16"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">AF</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math id="M17"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a unit amount of &#x00394;<italic>g</italic><sub>0, <italic>ij</italic></sub>. Thus, Equation (6) can be rewritten into Equation (7):</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">AF</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">when&#x000A0;</mml:mtext><mml:mstyle class="math"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn><mml:mtext class="textrm" mathvariant="normal"></mml:mtext></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">AF</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">when&#x000A0;</mml:mtext><mml:mstyle class="math"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mtext class="textrm" mathvariant="normal"></mml:mtext></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Equation (7) shows that symmetry of device is solely defined with <inline-formula><mml:math id="M19"><mml:mi>A</mml:mi><mml:mi>F</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, if <italic>g</italic><sub><italic>min, ij</italic></sub>, <italic>g</italic><sub><italic>max, ij</italic></sub> and <inline-formula><mml:math id="M20"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x000B1;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is fixed. Therefore, we can compare the asymmetry of device within the same range by modulating AF. In the case of ideal device, <italic>slope</italic> equals zero and &#x00394;<italic>g</italic><sub>0, <italic>ij</italic></sub> is constant regardless of <italic>AF</italic><sub><italic>ij</italic></sub>. However, to make the notation consistent, we denote this case as <italic>AF</italic><sub><italic>ij</italic></sub> &#x0003D; 0. For instance, there are three different cases in <xref ref-type="fig" rid="F1">Figure 1A</xref>, namely, <italic>AF</italic><sub><italic>ij</italic></sub> &#x0003D; 0;<italic>AF</italic><sub><italic>ij</italic></sub> &#x0003D; 2.0;and<italic>AF</italic><sub><italic>ij</italic></sub> &#x0003D; 6.0. All three cases have the same <italic>g</italic><sub><italic>min, ij</italic></sub>, <italic>g</italic><sub><italic>max, ij</italic></sub> and <inline-formula><mml:math id="M21"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. We note that many devices show non-zero AF behavior (Brivio et al., <xref ref-type="bibr" rid="B2">2018</xref>; van De Burgt et al., <xref ref-type="bibr" rid="B27">2018</xref>; Islam et al., <xref ref-type="bibr" rid="B13">2019</xref>; Kim et al., <xref ref-type="bibr" rid="B16">2019b</xref>).</p>
</sec>
<sec>
<title>2.3. Tiki-Taka Algorithm</title>
<p>Tiki-Taka algorithm (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>) is proposed to resolve the performance degradation issue in neural network training caused by the update asymmetry in the crosspoint elements. We showed that the asymmetric update prevents devices from being accurately updated by the amount of calculated gradients. <xref ref-type="fig" rid="F1">Figure 1B</xref> illustrates the key difference between SGD and Tiki-Taka algorithm in the weight update phase. In the case of SGD, an array, <italic>W</italic>, stores the weight vectors of a neural network. In the update phase, a weight, <italic>w</italic><sub><italic>ij</italic></sub>, is updated with the gradient &#x02207;<sub><italic>ij</italic></sub><italic>L</italic>(&#x0003D; <italic>x</italic><sub><italic>i</italic></sub>&#x003B4;<sub><italic>j</italic></sub>). On the other hand, Tiki-Taka algorithm requires one additional array, namely <italic>A</italic>, and it stores &#x00394;<italic>W</italic> by accumulating gradient vectors, &#x02207;<italic>L</italic>. The weight vectors stored in the array <italic>A</italic> are denoted as <italic>W</italic><sup><italic>A</italic></sup> and the array, <italic>C</italic>, stores the weight vectors denoted as <italic>W</italic><sup><italic>C</italic></sup>. When compared to SGD, Tiki-Taka algorithm accumulates the gradient, &#x02207;<sub><italic>ij</italic></sub><italic>L</italic>, in <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Then, the accumulated gradient is transferred to update the weight, <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> at every <italic>ns</italic> steps. Tiki-Taka algorithm supports sparse update by adopting cell-selection vector, <bold>u</bold><sub>&#x003C0;</sub>, where <bold>u</bold><sub>&#x003C0;</sub>(<italic>i</italic>) is the <italic>i</italic><sup><italic>th</italic></sup> element of a sparse vector, <bold>u</bold>, such as a one-hot vector and &#x003C0; is a permutation, &#x003C0;:{1, &#x022EF;&#x02009;, <italic>n</italic>} &#x02192; {1, &#x022EF;&#x02009;, <italic>n</italic>}. Finally, Tiki-Taka algorithm allows the array <italic>A</italic> to participate in the forward pass by introducing a parameter, &#x003B3;&#x02208;[0, 1]: <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Therefore, the effective weight <italic>w</italic><sub><italic>ij</italic></sub> depends on &#x003B3;: if &#x003B3; &#x0003D; 1, <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> replace the effective weight vector. Otherwise, if &#x003B3; &#x0003D; 0, then the effective weight is <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. In summary, Tiki-Taka algorithm can be formulated as follows:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02190;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msubsup><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msubsup><mml:mi>L</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E9"><label>(9)</label><mml:math id="M28"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>C</mml:mi></mml:msubsup><mml:mo>&#x02190;</mml:mo><mml:mrow><mml:mo>{</mml:mo> <mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>C</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>A</mml:mi></mml:msubsup><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>u</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003C0;</mml:mi></mml:mstyle></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>,&#x000A0;every&#x000A0;</mml:mtext><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mtext>&#x000A0;step</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>C</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>,&#x000A0;otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow> </mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msubsup><mml:mi>L</mml:mi></mml:math></inline-formula> indicates the derivative of the <italic>L</italic> with respect to <inline-formula><mml:math id="M30"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. &#x003B7; and &#x003BB; indicates SGD learning rate and transfer learning rate, respectively.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Tiki-Taka Algorithm and Update Asymmetry</title>
<sec>
<title>3.1. Mismatch Factor and Weight Update Formulation</title>
<p>Here, we formulate the impact of asymmetry by adopting even-odd function decomposition, with which any function can be represented (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>), and integrate the model into the SGD update rule in Equation (1) (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 1.2</xref> for the details):</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mi>S</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mi>A</mml:mi><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>A</mml:mi><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the function <italic>Sym</italic> represents symmetry of the devices and the function <italic>Asym</italic> corresponds to asymmetry. The weight, <italic>w</italic><sub><italic>ij</italic></sub>, is scaled with a factor <italic>K</italic> from the conductance, <italic>g</italic><sub><italic>ij</italic></sub>: <italic>w</italic><sub><italic>ij</italic></sub> &#x0003D; <italic>K</italic>&#x000B7;(<italic>g</italic><sub><italic>ij</italic></sub>&#x02212;<italic>g</italic><sub><italic>ref, ij</italic></sub>) and <inline-formula><mml:math id="M32"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>K</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <italic>w</italic><sub><italic>max, ij</italic></sub> is a correspond weight to the maximum amount of conductance, <italic>g</italic><sub><italic>max, ij</italic></sub>. To analyze it in the aspect of optimization, Equation (10) is reformulated as follows:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mi>S</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mi>A</mml:mi><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mtext class="textrm" mathvariant="normal">sgn</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>A</mml:mi><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mi>M</mml:mi><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext class="textrm" mathvariant="normal">sgn</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><bold> Definition 3.1 (Mismatch factor (MF))</bold>. <italic>Mismatch factor (MF) is an additional factor appeared in the original SGD weight update rule, originated from the update asymmetry of the weight storage device:</italic></p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mi>A</mml:mi><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy='true'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo stretchy='true'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>MF</italic> is multiplied to the gradient of a weight, &#x02207;<sub><italic>ij</italic></sub><italic>L</italic>, and affects the weight optimization process. As observed in the Definition 3.1, <italic>MF</italic><sub><italic>ij</italic></sub> scales with <italic>AF</italic>, and is dependent on the sign of gradient and the current weight state, <italic>w</italic><sub><italic>ij</italic></sub>, of a device. The dependence on the sign of gradient can create fluctuation of <italic>MF</italic><sub><italic>ij</italic></sub> if sgn(&#x02207;<sub><italic>ij</italic></sub><italic>L</italic>) changes the sign, and <italic>AF</italic><sub><italic>ij</italic></sub> magnifies the fluctuation of <italic>MF</italic><sub><italic>ij</italic></sub>. Therefore, there can be abrupt change in <italic>MF</italic><sub><italic>ij</italic></sub> affecting the neural network training.</p>
</sec>
<sec>
<title>3.2. Impact of Asymmetry: Case Studies</title>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> displays the classification error in MNIST handwritten digit recognition problem as a function of training epoch for various scenarios of training algorithm and <italic>AF</italic> values. We perform the experimental simulation using the open-source toolkit <italic>aihwkit</italic> (Rasch et al., <xref ref-type="bibr" rid="B21">2021</xref>) and further details about the experimental environment can be found in <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 3.3</xref>. First, performing SGD with asymmetric devices results in severe accuracy drop both for the train and test datasets when compared with the software baseline. Since the update asymmetry leads to non-unity <italic>MF</italic> and consequent unfavorable behavior in weight update process, a sharp drop in performance is found during the training. On the other hand, the cases with Tiki-Taka algorithm and finite asymmetry show robust performance reaching to the floating point baseline (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>). As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, when <italic>AF</italic><sup><italic>A</italic></sup> &#x0003D; <italic>AF</italic><sup><italic>C</italic></sup> &#x0003D; 1.0, the train and test error gradually decrease and can reach to the baseline. However, the remaining question is how much asymmetry Tiki-Taka algorithm can tolerate. Other cases shown in <xref ref-type="fig" rid="F2">Figure 2</xref> with various pairs of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> answer the question. The train error does not gradually decrease, which indicates that the weight vectors of neural networks are stuck at bad local optima. Interestingly, the case with <italic>AF</italic><sup><italic>A</italic></sup> &#x0003D; <italic>AF</italic><sup><italic>C</italic></sup> &#x0003D; 0 shows higher train and test error than those of SGD with asymmetric non-linear devices. This result indicates that there exist optimal pairs of <italic>AF</italic>, and a reliable method to determine the optimal pair of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> is necessary. To find the method, it is necessary to quantify and understand how <italic>MF</italic> affects the optimizer, Tiki-Taka algorithm, during the neural network training.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The train error <bold>(A)</bold> and test error <bold>(B)</bold> on handwritten digits images (MNIST) dataset with MLP. The baseline indicates the averaged error of the last 3 epochs with symmetric linear device, and the neural network is trained with SGD. In the case of SGD with asymmetric non-linear device, the error is averaged over various <italic>AF</italic> &#x0003E; 0. In the case of Tiki-Taka algorithm, the error is averaged over various pair of SGD and transfer learning rates. The shaded area indicates standard deviation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0002.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4. Optimization based on History of Gradients</title>
<sec>
<title>4.1. General Formulation for Gradient History-Based SGD Variants</title>
<p>Recent variants of SGD introduce the concept of history-dependent update and utilize gradient information at each step, including momentum and adaptive gradient, to overcome the issues of vanilla SGD (Kandel et al., <xref ref-type="bibr" rid="B14">2020</xref>). In such methods, the weight update is determined not only by the current gradient calculation, but also by the history of previous gradients and their importance at the current step. For instance, if the current weight update includes only the first-order gradient, then the history dependent update is composed of first-order gradient vectors and their respective importance for each step.</p>
<p><bold> Definition 4.1 (Gradient Scheduler)</bold>. <italic>History-dependent optimization algorithm is generalized by the</italic> <italic>governing rules</italic>  <italic>including</italic> <italic>n</italic><sup><italic>th</italic></sup>  <italic>order gradients:</italic></p>
<disp-formula id="E14"><label>(13)</label><mml:math id="M36"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B7;</mml:mi><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>t</mml:mi></mml:mstyle></mml:msub><mml:mo stretchy='false'>[</mml:mo><mml:mtext>sgn</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mo>&#x02207;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>n</mml:mi></mml:msup><mml:mo stretchy='false'>]</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><bold>t</bold> &#x0003D; <bold>t</bold>(<italic>k</italic>) (<italic>k</italic> &#x0003D; 1, &#x022EF;&#x02009;, <italic>T</italic>) is a <italic>gradient scheduler</italic> or <italic>scheduler</italic>, <italic>which contains information on the importance of the</italic> <italic>k</italic><sup><italic>th</italic></sup> <italic>gradients</italic>, <inline-formula><mml:math id="M37"><mml:mstyle class="text"><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x02207;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>L</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>Therefore, the scheduler, <bold>t</bold>(<italic>k</italic>), defines how the gradients of previous steps for <italic>k</italic> &#x0003C; <italic>T</italic> affect the update of current step at <italic>k</italic> &#x0003D; <italic>T</italic>, since <bold>t</bold>(<italic>k</italic>) contains information on weighting of the gradients dependent on the step, <italic>k</italic>.</p>
</sec>
<sec>
<title>4.2. Gradient Scheduler in Tiki-Taka Algorithm</title>
<p>From the Definition 4.1, we can find the <bold>t</bold>(<italic>k</italic>) for Tiki-Taka algorithm as follows (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 2.2</xref> for the detailed derivation):</p>
<disp-formula id="E15"><label>(14)</label><mml:math id="M38"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02261;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>u</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <bold>1</bold><sub><italic>T</italic><sub>&#x02261;</sub><sub><italic>ns</italic></sub>0</sub> is an indicator function which returns 1 every ns steps, <bold>t</bold>(<italic>k</italic>) is controlled by three hyper-parameters (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>): &#x003B3;, <italic>ns</italic>, and <bold>u</bold><sub>&#x003C0;</sub>. These parameters cause the weight matrix <italic>W</italic> to be updated sparsely and asynchronously.</p>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> compares the gradient schedulers in three algorithms: vanilla SGD, momentum-based optimizer, and Tiki-Taka algorithm. The parameters of Tiki-Taka used are <italic>ns</italic> &#x0003D; 1, <bold>u</bold> &#x0003D; <bold>1</bold>, and &#x003B3; &#x0003D; 0, and the ideal, symmetric and linear, switching characteristics are assumed for the devices in both arrays (i.e., <inline-formula><mml:math id="M39"><mml:mi>A</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>). With ideal devices, Tiki-Taka algorithm memorizes all the previous gradients with equal importance. Therefore, it resembles the momentum-based optimizer if &#x003B2; &#x02192; 1. However, it should not be confused with the momentum-based optimizer, because the scheduler of Tiki-Taka algorithm with asymmetric devices depends not only on the step, but also on the history of gradients and slopes of devices unlike the momentum-based optimizer.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>(A)</bold> Comparison of gradient scheduler, <bold>t</bold>(<italic>k</italic>), of SGD, momentum-based optimizer, and Tiki-Taka algorithm with respect to steps, k. <bold>t</bold>(<italic>k</italic>) is normalized by the value of <bold>t</bold>(<italic>T</italic>). <bold>(B)</bold> Summarized table of gradient scheduler of optimizers from <bold>(A)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0003.tif"/>
</fig>
<p>By including the update asymmetry, we can reformulate the scheduler of Tiki-Taka algorithm, Equation (14), as follows (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 2.2</xref>) for the detailed derivation):</p>
<disp-formula id="E16"><label>(15)</label><mml:math id="M40"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mrow><mml:mo>&#x02261;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>u</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x000B7;</mml:mo><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Equation (15) indicates that both <italic>MF</italic><sup><italic>A</italic></sup> and <italic>MF</italic><sup><italic>C</italic></sup> play an critical role in weight optimization. <italic>MF</italic><sup><italic>A</italic></sup> kicks in every step of <italic>k</italic>, and <italic>MF</italic><sup><italic>C</italic></sup> provides a factor until the last step of <italic>T</italic> in the scheduler of Tiki-Taka algorithm. Therefore, the optimizer of Tiki-Taka algorithm is uniquely determined by two <italic>MF</italic>&#x00027;s. Considering that <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is determined by the history of gradients, <inline-formula><mml:math id="M42"><mml:mo>&#x02207;</mml:mo><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:mo>&#x02207;</mml:mo><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B3;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, the scheduler of Tiki-Taka algorithm is parameterized by the step and the history of the first-order gradients. <italic>MF</italic><sup><italic>C</italic></sup> affects the optimization process for different <italic>T, T</italic>&#x0002B;1, &#x022EF;&#x02009; steps in common by scaling all the ideal <bold>t</bold>(<italic>k</italic>). Therefore, <italic>MF</italic><sup><italic>C</italic></sup> is one of key components in building the optimizer, and <italic>MF</italic><sup><italic>A</italic></sup> determines the optimizer of Tiki-Taka algorithm.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Linear Regression Analysis</title>
<p>To analyze the impact of update asymmetry in neural network training, we perform a series of linear regression experiments with different algorithms and display the results in <xref ref-type="fig" rid="F4">Figure 4</xref>. With the presence of update asymmetry, the weight vectors of a neural network are stuck at sub-optima (Kim et al., <xref ref-type="bibr" rid="B15">2019a</xref>) as shown in <xref ref-type="fig" rid="F4">Figure 4B</xref>, and the linear regression is unable to approach the global optimum in this convex problem. Since the expected change of a weight around the sub-optima region over the given distribution of data <italic>D</italic>, <bold>E</bold><sub><italic>X</italic>&#x0007E;<italic>D</italic></sub>[&#x00394;<italic>w</italic><sub><italic>ij</italic></sub>], is 0 on average, there is no room for the weight to progress further (Gokmen and Haensch, <xref ref-type="bibr" rid="B6">2020</xref>). If the asymmetry factor of a crosspoint element, <italic>AF</italic>, is large, then the weight vector will converge to the region further away from the global optimum. This is because larger <italic>AF</italic> values cause larger variations in <italic>MF</italic> and equilibrium region to be formed away from the global optimum.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>(A)</bold> An example of linear regression model and sampled data: <italic>y</italic> &#x0003D; <italic>ax</italic> &#x0002B; <italic>b</italic> &#x0002B; &#x003F5;, where <italic>a</italic> &#x02208; <italic>R</italic> and &#x003F5;&#x0007E;<italic>N</italic>(0, &#x003C3; <sup>2</sup>). In this example, <italic>a</italic> &#x0003D; 0.5, <italic>b</italic> &#x0003D; 0.5, <italic>and</italic> &#x003C3; &#x0003D; 0.05. <bold>(B)</bold> Trajectories of weights in the log-scaled loss surface. A &#x0201C;target&#x0201D; point indicates a closed-form solution of linear regression model. i.e., (<italic>a b</italic>)<sup><italic>T</italic></sup>= (<sup><italic>X</italic><sup><italic>T</italic></sup><italic>X</italic>)&#x02212;1</sup><italic>X</italic><sup><italic>T</italic></sup><italic>y</italic>, where <italic>X</italic> is the <italic>N</italic> &#x000D7; 2 input matrix and <italic>y</italic> is the <italic>N</italic> &#x000D7; 1 output vector. The &#x0201C;ideal&#x0201D; trajectory indicates a trajectory of a neural network optimized by vanilla SGD using ideal device. We denote it as &#x0201C;ideal&#x0201D; only in reference to the problem of convex optimization (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 3.2</xref> for the details of experimental environments).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0004.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F5">Figure 5</xref>, we repeat the linear regression experiments with Tiki-Taka algorithm and different combinations of <italic>AF</italic> values and observe the impact of update asymmetry on the linear regression results. Unlike the previous results with vanilla SGD, neural networks trained by Tiki-Taka algorithm with certain combinations of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> are able to converge to the global optimum, even with the update asymmetry. Since Tiki-Taka algorithm can accumulate gradients of previous steps in the auxilary array and utilize the information for optimization, the weight vector can escape from the sub-optima dug by <italic>MF</italic> and converge to the global optimum.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Trajectories of weights in the log-scaled loss surface for convex problem where the data is sampled from the same distribution as described in <xref ref-type="fig" rid="F4">Figure 4A</xref> and the normalized schedulers, <bold>t</bold>(<italic>k</italic>), of a variable <italic>b</italic> described in <xref ref-type="fig" rid="F4">Figure 4A</xref>. <bold>(A)</bold> <italic>AF</italic><sup><italic>C</italic></sup> is fixed with 2.0 and <bold>(B)</bold> <italic>AF</italic><sup><italic>A</italic></sup> is fixed with 2.0. The points of each trajectory are marked for every epoch.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0005.tif"/>
</fig>
<p>To analyze the impact of update asymmetry in the array A and C independently, we modulate <italic>AF</italic><sup><italic>A</italic></sup> (<xref ref-type="fig" rid="F5">Figure 5A</xref>) and <italic>AF</italic><sup><italic>C</italic></sup> (<xref ref-type="fig" rid="F5">Figure 5B</xref>) in an alternating fashion and compare the neural network training performance. The corresponding <bold>t</bold>(<italic>K</italic>) and <italic>W</italic><sup><italic>A</italic></sup> value as a function of training step for each case are visualized in <xref ref-type="fig" rid="F6">Figure 6</xref>. <xref ref-type="fig" rid="F5">Figure 5A</xref> shows the trajectories of weight vectors in three cases where <italic>AF</italic><sup><italic>A</italic></sup> values are 0.0, 2.0 and 4.0, respectively, and <italic>AF</italic><sup><italic>C</italic></sup> is fixed at 2.0. First, if <italic>AF</italic><sup><italic>A</italic></sup> &#x0003D; 0, then <inline-formula><mml:math id="M43"><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02243;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> at all steps <italic>k</italic>, which indicates that <inline-formula><mml:math id="M44"><mml:mstyle mathvariant="bold"><mml:mtext>t</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02243;</mml:mo><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. In this case, the large oscillation and deviation from the path by SGD are observed in the weight trajectory since the optimizer memorizes all the history of gradients with equal importance in the array <italic>A</italic> and the previous gradients <inline-formula><mml:math id="M45"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> can only be changed slowly. On the other hand, when <italic>AF</italic><sup><italic>A</italic></sup> is 2.0 or 4.0, <bold>t</bold>(<italic>k</italic>) abruptly changes curvature near the 150<sup><italic>th</italic></sup> step point as illustrated in <xref ref-type="fig" rid="F6">Figure 6</xref>. If &#x02207;<sub><italic>ij</italic></sub><italic>L</italic> &#x0003C; 0 for <italic>T</italic>&#x02212;1 steps, then <inline-formula><mml:math id="M46"><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> decreases as <inline-formula><mml:math id="M47"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> increases as discussed in Definition 3.1. After the <italic>T</italic> steps, <inline-formula><mml:math id="M48"><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02265;</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> abruptly increases as sgn(&#x02207;<sub><italic>ij</italic></sub><italic>L</italic>) is inverted. The implication of the &#x0201C;abruptness" is that the optimizer is likely to forget the gradients of previous steps right before sgn(&#x02207;<sub><italic>ij</italic></sub><italic>L</italic>) is inverted, and accept newly updated gradients. The &#x0201C;abruptness" amplifies with devices with larger <italic>AF</italic><sup><italic>A</italic></sup>. In other words, the optimizer has an &#x0201C;easy come, easy go" memory that stores the history of gradients, and <bold>t</bold>(<italic>k</italic>) determines the speed of accumulating and forgetting the gradients. This behavior resembles that of the first order momentum optimizer. However, Tiki-Taka algorithm decides to decay the gradients by accumulating the successive gradients of same sgn(&#x02207;<sub><italic>ij</italic></sub><italic>L</italic>) while the first order momentum optimizer merely decays the gradients far from the current step. In <xref ref-type="fig" rid="F5">Figure 5</xref>, the trajectories with large <italic>AF</italic><sup><italic>A</italic></sup> are less inclined to oscillate and deviate as the optimizer forgets the information right before sgn(&#x02207;<sub><italic>ij</italic></sub><italic>L</italic>) is inverted. It helps the weight vector to converge to the global optimum. However, it also fades out the regularization effect. Thus, there exists optimal range of <italic>AF</italic><sup><italic>A</italic></sup> that guarantees convergence and regularization.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Comparisons of approximated <bold>t</bold>(<italic>k</italic>) and <italic>W</italic><sup><italic>A</italic></sup> of three cases in <bold>(A)</bold> <xref ref-type="fig" rid="F5">Figure 5A</xref> and <bold>(B)</bold> <xref ref-type="fig" rid="F5">Figure 5B</xref>. Raw data of <bold>t</bold>(<italic>k</italic>) and <italic>W</italic><sup><italic>A</italic></sup> was averaged using exponential moving average and approximated with natural smoothing spline method.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0006.tif"/>
</fig>
<p>Modulating <italic>AF</italic><sup><italic>C</italic></sup> impacts differently on the weight vector trajectories as shown in <xref ref-type="fig" rid="F5">Figure 5B</xref>, where the trajectories of weight vectors with <italic>AF</italic><sup><italic>C</italic></sup> = 0.0, 2.0, and 4.0 are displayed while <italic>AF</italic><sup><italic>A</italic></sup> is fixed at 2.0. As <italic>AF</italic><sup><italic>C</italic></sup> increases, the weight update becomes less frequent with larger update steps, and, consequently, the frequency and magnitude of the oscillation in the weight vector becomes larger. This observation is expected from the fact that <inline-formula><mml:math id="M49"><mml:mstyle mathvariant="bold"><mml:mtext>V</mml:mtext></mml:mstyle><mml:mi>a</mml:mi><mml:mstyle mathvariant="bold"><mml:mtext>r</mml:mtext></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>A</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M50"><mml:mstyle mathvariant="bold"><mml:mtext>E</mml:mtext></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:mi>A</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. With large <italic>AF</italic><sup><italic>C</italic></sup>, the array <italic>A</italic> value fluctuates by frequently changing signs as shown in <xref ref-type="fig" rid="F6">Figure 6</xref> and cannot drive <italic>W</italic><sup><italic>C</italic></sup> in a consistent manner, which leads to a large oscillation in <italic>W</italic><sup><italic>C</italic></sup>. In that sense, array <italic>A</italic> can capture the gradient history better when <italic>AF</italic><sup><italic>C</italic></sup> is smaller. However, with larger <italic>AF</italic><sup><italic>C</italic></sup>, <inline-formula><mml:math id="M51"><mml:mi>M</mml:mi><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> gradually decreases as a function of step <italic>k</italic>, if the signs of <inline-formula><mml:math id="M52"><mml:msubsup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are the same. This observation shows the similar effects of the adaptive gradient (Duchi, <xref ref-type="bibr" rid="B5">2011</xref>; Zeiler, <xref ref-type="bibr" rid="B31">2012</xref>; Kingma and Ba, <xref ref-type="bibr" rid="B18">2014</xref>), which can help the weight vector to converge by decreasing the amount of updates around the global optimum when <italic>AF</italic><sup><italic>C</italic></sup> is set within the proper range. Therefore, the impact of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> becomes highly entangled during optimization process, and it is a challenging task to find the optimal <italic>AF</italic> value for each array. A systematic method to find the optimal values of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup>, which are highly dependent on the specific dataset and neural network architecture, is required for the convergence and performance of the neural network training.</p>
</sec>
<sec id="s6">
<title>6. Impact of Update Asymmetries on Neural Network Performance</title>
<sec>
<title>6.1. Experimental Results</title>
<p>As we discussed in the previous sections, the update asymmetry in array <italic>A</italic> and <italic>C</italic> plays an important role in neural network training with Tiki-Taka algorithm as an optimizer. To experimentally explore the impact of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> on training neural networks, we perform a series of experiments by varying <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> values and training a multi-layer perceptron (MLP) with MNIST dataset. For this experiment, we set (<italic>w</italic><sub><italic>min</italic></sub>, <italic>w</italic><sub><italic>max</italic></sub>) = (&#x02013;1,1) and <inline-formula><mml:math id="M53"><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>001</mml:mn></mml:math></inline-formula> for all devices in the array as mentioned in section 2.2, which illustrates the device reaching <italic>w</italic><sub><italic>max</italic></sub> &#x0003D; 1 from initial weight value of <italic>w</italic><sub><italic>min</italic></sub> &#x0003D; &#x02212;1 within 600 updates if (<inline-formula><mml:math id="M54"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M55"><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>) = (1, &#x02013;1) is given during update phase (see <xref ref-type="supplementary-material" rid="SM1">Supplementary Materials section 3.2</xref> for the details of experimental environments).</p>
<p><xref ref-type="fig" rid="F7">Figure 7</xref> shows the heat map of train (<xref ref-type="fig" rid="F7">Figure 7A</xref>) and test accuracies of MLP as a function of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> values and with two different set of SGD and transfer learning rates (<xref ref-type="fig" rid="F7">Figures 7B,C</xref>). The results show the consistent observations as analyzed in section 5. First, when <italic>AF</italic><sup><italic>C</italic></sup> is fixed, the train and test accuracies peak at <italic>AF</italic><sup><italic>A</italic></sup> &#x0003D; 1.78 and decrease as it becomes larger or smaller. Smaller <italic>AF</italic><sup><italic>A</italic></sup> causes the optimizer to accumulate the history of gradient with equal importance, which hampers the convergence of the weight vectors due to oscillation. However, this effect is reduced with larger <italic>AF</italic><sup><italic>C</italic></sup> thanks to the adaptive gradient impact. In contrary, larger <italic>AF</italic><sup><italic>A</italic></sup> values disallow accumulation of the full history of gradients, which results in reduction of the regularization effect. This tendency is more significant for larger <italic>AF</italic><sup><italic>C</italic></sup> values, where the variance and expectation of updates for the array <italic>C</italic> is large enough not to store the entire history of gradients. Considering that larger <italic>AF</italic><sup><italic>A</italic></sup> causes memorization of short-term gradients, the regularization effect becomes weakened.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The <bold>(A)</bold> train and <bold>(B,C)</bold> test accuracy of a multi-layer perceptron as a function of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup>. The displayed values are averaged over the last three epochs of the training and also three different initialization conditions. The mini-batch size is 1. The transfer learning rate and SGD learning rate is (0.01 and 0.02) for <bold>(B)</bold> and (0.02 and 0.01) for <bold>(C)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0007.tif"/>
</fig>
</sec>
<sec>
<title>6.2. Impact of Learning Rates and Robustness Score</title>
<p>The optimal combination of <italic>AF</italic> values, which provide the best accuracy, can be changed when the combination of learning rates vary. For example, in <xref ref-type="fig" rid="F7">Figure 7B</xref>, the test accuracy peaks at the combinations, (<italic>AF</italic><sup><italic>A</italic></sup>, <italic>AF</italic><sup><italic>C</italic></sup>) &#x02208; {(1.0, 1.0), (1.78, 1.0), (1.0, 1.78)} when the transfer learning rate and SGD learning rate is (0.01, 0.02). However, as shown in <xref ref-type="fig" rid="F7">Figure 7C</xref>, the combination of <italic>AF</italic> is optimal when (<italic>AF</italic><sup><italic>A</italic></sup>, <italic>AF</italic><sup><italic>C</italic></sup>)&#x02208;{(3.16, 1.0), (5.62, 1.0), (10.0, 1.0)} if the neural network is trained with different pair of learning rates, (0.02, 0.01). Considering that <italic>AF</italic> of each device is typically pre-determined at the moment of device fabrication, it is crucial to find the robust <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> values over the learning rate space which do not degrade the accuracy. From this perspective, finding the robust region of <italic>AF</italic> without searching the best pair of learning rates can be guaranteed by introducing the thresholds of certain measurements such as test accuracy. The following <italic>robustness score</italic> provides the hard-bound threshold for finding the optimal combination of <italic>AF</italic> over the space of learning rates.</p>
<p><bold> Definition 6.1 (Robustness score)</bold>. Suppose that <italic>m</italic> is a neural network model, and <italic>Meas</italic>(&#x000B7;) is a measurement such as accuracy or loss of the model <italic>m</italic> after training for a given dataset <italic>D</italic>. <inline-formula><mml:math id="M56"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> is a hyper-parmeter space for software training, <inline-formula><mml:math id="M57"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B7;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>, where &#x003B7; &#x0003D; SGD learning rate and&#x003BB; &#x0003D; transfer learning rate}. A <italic>Robustness score</italic>, <italic>RS</italic>(<italic>m</italic>), for the given threshold, <italic>th</italic>, is defined as follows:</p>
<disp-formula id="E17"><label>(16)</label><mml:math id="M58"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>M</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003E;</mml:mo><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><xref ref-type="fig" rid="F8">Figure 8</xref> illustrates the robustness score, <italic>RS</italic>(<italic>m</italic>), over a learning rate space, <inline-formula><mml:math id="M59"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B7;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, where &#x003B7;&#x02208;{0.01, 0.02, 0.04} and &#x003BB;&#x02208;{0.01, 0.02, 0.04}. As discussed in section 3.2, <italic>AF</italic> of each array affects the test accuracy of the network using Tiki-Taka algorithm. Test accuracies do not reach sufficient value when AF of each array is too small (<italic>AF</italic><sup><italic>A, C</italic></sup> &#x0003D; 0) or large (<italic>AF</italic><sup><italic>A, C</italic></sup> &#x0003D; 10). Therefore, optimal <italic>AF</italic> value pairs which maximize the robustness score exist near <italic>AF</italic><sup><italic>A, C</italic></sup> &#x0003D; 1 as shown in <xref ref-type="fig" rid="F8">Figure 8</xref>. The score map has two implications. First of all, if the minimum requirement of test accuracy (threshold) is determined, the region whose robustness scores are over a certain criteria could be used to define the minimum specifications of update asymmetry. Therefore, one can obtain the range of <italic>AF</italic> value for each array which gurantees the minimum test accuracy. Second, by scanning over various pairs of learning rates, one can identify the region which provides the <italic>AF</italic> values which are most independent on the learing rates. To find such region, the score is post-processed with interpolation and gaussian filtering. The robustness score is approximated by piece-wise linear interpolation in the domain of <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup>. The interpolated data is smoothed by gaussian filtering (<italic>g</italic>(<italic>AF</italic><sup><italic>A</italic></sup>, <italic>AF</italic><sup><italic>C</italic></sup>))) to obtain the robust region in the domain: <italic>g</italic>(<italic>AF</italic><sup><italic>A</italic></sup>, <italic>AF</italic><sup><italic>C</italic></sup>) &#x0003D; (2&#x003C0;<sup>&#x003C3;<sup>2</sup>)&#x02212;1/2</sup>&#x000B7;<italic>exp</italic>(&#x02212;(((<italic>A</italic><sup><italic>F</italic><sup><italic>A</italic></sup>)2</sup>&#x0002B;(<italic>A</italic><sup><italic>F</italic><sup><italic>C</italic></sup>)2</sup>)/(2&#x003C3;<sup>2</sup>)), where we define the degree of &#x0201C;robustness" of the score map by modulating &#x003C3; of a gaussian filter. Finding the robust region with a gaussian filter is consistent with the definition of variations in <italic>AF</italic> of devices. In <xref ref-type="fig" rid="F8">Figure 8C</xref>, the robust region is around <italic>AF</italic><sup><italic>A</italic></sup> &#x0003D; 1.2 and <italic>AF</italic><sup><italic>C</italic></sup> &#x0003D; 1.0 with the score greater than 8.95. As we experiment with a total of 9 scenarios of learning rates, the region with the score greater than 8.95 displays the required test accuracy within the given range of learning rate pairs. When compared to the region with <italic>RS</italic>(<italic>m</italic>) &#x0003D; 4.5, the region with <italic>RS</italic>(<italic>m</italic>)&#x02265;8.95 shows almost doubled efficiency during on-device training. In other words, it requires half of the resources to find the optimal pair of learning rates. This calibration method to find the best combination of <italic>AF</italic> is applicable to other types of hyper-parameters if the hyper-parameters space <inline-formula><mml:math id="M60"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> can be expanded to include them.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>(A)</bold> The robustness score based on the test accuracy with a given threshold, 97.0%, where the number indicates how many cases with the test accuracy over the threshold. <italic>dw</italic><sub><italic>min</italic></sub> is set at 0.001. <bold>(B)</bold> The post-processed robustness score, where the step size of each domain in <italic>AF</italic> is 0.2 and &#x003C3; &#x0003D; 0.2. <bold>(C)</bold> The regions with high robustness score are highlighted to visualize the area where the test accuracy is universally high as the learning rates change.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0008.tif"/>
</fig>
</sec>
<sec>
<title>6.3. Inspection on Optimal Range of Slope for Different Thresholds</title>
<p>In practical applications, the target performance of the network, such as the test accuracy of the MNIST classification task using MLP, is one of the most significant design choices as some tasks require the maximum accuracy while others do not. Considering that the robustness score, <italic>RS</italic>(<italic>m</italic>), varies as target accuracy changes, studying the relationship between <italic>RS</italic>(<italic>m</italic>) and the target threshold value can provide insights on the choice of <italic>AF</italic> for each array. <xref ref-type="fig" rid="F9">Figure 9</xref> illustrates the robustness score maps with two different threshold values, 97.0 and 97.5%. When the threshold is set to 97.0%, the score is robust against the change in <italic>AF</italic><sup><italic>C</italic></sup> value and the efficiency of searching over the hyper-parameter space, <inline-formula><mml:math id="M61"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> with <italic>AF</italic><sup><italic>A</italic></sup> is not degraded in the proper region. However, in the case of the threshold being 97.5%, the best choice is to fix <italic>AF</italic><sup><italic>C</italic></sup> around 1.0. Then the asymmetry requirement for the array A is lifted. Although the robustness scores of the case with the threshold being 97.5% are relatively lower than those of 97.0%, selecting <italic>AF</italic><sup><italic>A</italic></sup> and <italic>AF</italic><sup><italic>C</italic></sup> of the region has comparative advantages over the other regions.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Post-processed test robustness score on various thresholds: <bold>(A)</bold> 97.0% and <bold>(B)</bold> 97.5%. The green-colored regions are highlighted with different light and shade. Red cross point indicates the pair of asymmetric factors for baseline model used in original Tiki-Taka algorithm (Gokmen and Haensch,2020). Boxes on the left side of <bold>(A)</bold> show the asymmetry of device response in accordance with each asymmetric factor.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-767953-g0009.tif"/>
</fig>
<p>Consistent with the analysis in sections 5 and 6.1, <xref ref-type="fig" rid="F9">Figure 9A</xref> describes that the neural networks are able to find the optimum with a wide range of <italic>AF</italic><sup><italic>C</italic></sup> values over <inline-formula><mml:math id="M62"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">H</mml:mi></mml:mrow></mml:math></inline-formula> if <italic>AF</italic><sup><italic>A</italic></sup> is in the proper region. This observation is due to the adaptive gradient effect that <italic>AF</italic><sup><italic>C</italic></sup> brings, which helps the convergence around the nearest local optimum. In this case, the neural network converges to the local optimum instead of further exploring the loss surface. This results in the high robustness score region along the wide range of <italic>AF</italic><sup><italic>C</italic></sup>, though it does not guarantee higher accuracy. In <xref ref-type="fig" rid="F9">Figure 9B</xref>, on the other hand, the region satisfying over 97.5% accuracy is spread along the axis of <italic>AF</italic><sup><italic>A</italic></sup>, which indicates that the specification for <italic>AF</italic><sup><italic>C</italic></sup> value is strict. If <italic>AF</italic><sup><italic>C</italic></sup> value is within the specification, then the robustness score depends on the choice of <italic>AF</italic><sup><italic>A</italic></sup>. With smaller <italic>AF</italic><sup><italic>A</italic></sup> values, the optimizer allows the weights of neural networks to explore over loss surface, which enables to find better optimum.</p>
</sec>
<sec>
<title>6.4. Application to SNN</title>
<p>Recent research on SNN and its implementations using resistive memory devices have shown that the devices are designed explicitly to be embedded with the learning algorithm for SNN such as Spike-timing-dependent plasticity (STDP) and equilibrium propagation (Scellier and Bengio, <xref ref-type="bibr" rid="B23">2017</xref>). Accordingly, the update asymmetry of devices is more likely to affect the process of training neural networks (Kwon et al., <xref ref-type="bibr" rid="B19">2020</xref>). Inspired by this motivation, they demonstrated that asymmetric non-linear devices are more powerful than symmetric linear devices (Brivio et al., <xref ref-type="bibr" rid="B2">2018</xref>, <xref ref-type="bibr" rid="B3">2021</xref>; Kim et al., <xref ref-type="bibr" rid="B17">2021</xref>). Our analytical tools and calibration method can be applied to illustrate the relationship between the update asymmetry of devices and its impact on training in detail. Therefore, it is worthwhile to pursue further studies on the relationship in more extensive range of cases of learning algorithms.</p>
</sec>
</sec>
<sec id="s7">
<title>7. Conclusion and Discussion</title>
<p>In this work, the impact of update asymmetry in array <italic>A</italic> and <italic>C</italic> and learning rates on network performance when training neural networks with Tiki-Taka algorithm is quantitatively analyzed. We introduced the concept of <italic>gradient scheduler</italic>, which defines the weights of previous gradient values, to explain how the update asymmetry impacts on solving a convex problem. With update asymmetry, we derived that asymmetry factor involved in <italic>gradient scheduler</italic> that the training of neural networks is affected. As we inspected on the gradient schedulers, larger asymmetry factor of <italic>A</italic> accelerates accumulating and forgetting the history of gradients while that of <italic>C</italic> brings adaptive gradient effects. We showed that the update asymmetry levels in the main and auxiliary arrays impact differently on training neural networks, indicating that requirements for asymmetry on each array are different. We also proposed a novel calibration metric, <italic>Robustness Score</italic>, to quantify and visualize the performance of the network as a function of device asymmetry and network parameters with respect to the target accuracy threshold. By searching over the hyper-parameter space of Tiki-Taka algorithm, we quantified the specification of device asymmetry for <italic>A</italic> and <italic>C</italic> arrays depending on the target accuracy value of choice. Our analysis shed light on further relaxing device specifications for the realization of analog neural network accelerators in hardware and provide a guideline to realize the resistive switching devices with robust training performance over the space of hyper parameters.</p>
</sec>
<sec sec-type="data-availability" id="s8">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>SK conceived the original idea. CL formulated the theory. CL, KN, and WJ developed methodology and conducted experiments. CL, TG, and SK analyzed and interpreted results. All authors drafted and revised the manuscript.</p>
</sec>
<sec sec-type="funding-information" id="s10">
<title>Funding</title>
<p>This work was supported by Samsung Science &#x00026; Technology Foundation (grant no. SRFC-IT2001-06).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>CL is employed by NAVER Clova, South Korea. TG is employed by IBM Research, USA. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<ack><p>We thank Malte J. Rasch and Diego Moreda for many useful discussions and technical supports.</p>
</ack><sec sec-type="supplementary-material" id="s12">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fnins.2021.767953/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fnins.2021.767953/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Agarwal</surname> <given-names>S.</given-names></name> <name><surname>Quach</surname> <given-names>T.-T.</given-names></name> <name><surname>Parekh</surname> <given-names>O.</given-names></name> <name><surname>Hsia</surname> <given-names>A. H.</given-names></name> <name><surname>DeBenedictis</surname> <given-names>E. P.</given-names></name> <name><surname>James</surname> <given-names>C. D.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Energy scaling advantages of resistive memory crossbar based computation and its application to sparse coding</article-title>. <source>Front. Neurosci</source>. <volume>9</volume>:<fpage>484</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2015.00484</pub-id><pub-id pub-id-type="pmid">26778946</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brivio</surname> <given-names>S.</given-names></name> <name><surname>Conti</surname> <given-names>D.</given-names></name> <name><surname>Nair</surname> <given-names>M. V.</given-names></name> <name><surname>Frascaroli</surname> <given-names>J.</given-names></name> <name><surname>Covi</surname> <given-names>E.</given-names></name> <name><surname>Ricciardi</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Extended memory lifetime in spiking neural networks employing memristive synapses with nonlinear conductance dynamics</article-title>. <source>Nanotechnology</source> <volume>30</volume>, <fpage>015102</fpage>. <pub-id pub-id-type="doi">10.1088/1361-6528/aae81c</pub-id><pub-id pub-id-type="pmid">30378572</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brivio</surname> <given-names>S.</given-names></name> <name><surname>Ly</surname> <given-names>D. R.</given-names></name> <name><surname>Vianello</surname> <given-names>E.</given-names></name> <name><surname>Spiga</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Nonlinear memristive synaptic dynamics for efficient unsupervised learning in spiking neural networks</article-title>. <source>Front. Neurosci</source>. <volume>15</volume>:<fpage>27</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2021.580909</pub-id><pub-id pub-id-type="pmid">33633531</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>F.</given-names></name> <name><surname>Tang</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of accelerator architectures for deep neural networks</article-title>. <source>Engineering</source> <volume>6</volume>, <fpage>264</fpage>&#x02013;<lpage>274</lpage>. <pub-id pub-id-type="doi">10.1016/j.eng.2020.01.007</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duchi</surname> <given-names>J.</given-names></name> <name><surname>Hazan</surname> <given-names>E.</given-names></name> <name><surname>Singer</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). <article-title>Adaptive subgradient methods for online learning and stochastic optimization</article-title>. <source>J. Mach. Learn. Res.</source> <volume>12</volume>, <fpage>2121</fpage>&#x02013;<lpage>2159</lpage>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Haensch</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Algorithm for training neural networks on resistive device arrays</article-title>. <source>Front. Neurosci</source>. <volume>14</volume>:<fpage>103</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2020.00103</pub-id><pub-id pub-id-type="pmid">32174807</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Onen</surname> <given-names>M.</given-names></name> <name><surname>Haensch</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Training deep convolutional neural networks with resistive cross-point devices</article-title>. <source>Front. Neurosci</source>. <volume>11</volume>:<fpage>538</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2017.00538</pub-id><pub-id pub-id-type="pmid">29066942</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Vlasov</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Acceleration of deep neural network training with resistive cross-point devices: design considerations</article-title>. <source>Front. Neurosci</source>. <volume>10</volume>:<fpage>333</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2016.00333</pub-id><pub-id pub-id-type="pmid">27493624</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Y..</given-names></name></person-group> (<year>2018</year>). <article-title>A survey on methods and theories of quantized neural networks</article-title>. <source>arXiv preprint arXiv:1808.04752</source>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haensch</surname> <given-names>W.</given-names></name> <name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Puri</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>The next generation of deep learning hardware: analog computing</article-title>. <source>Proc. IEEE</source> <volume>107</volume>, <fpage>108</fpage>&#x02013;<lpage>122</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2018.2871057</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Mao</surname> <given-names>H.</given-names></name> <name><surname>Dally</surname> <given-names>W. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding</article-title>. <source>arXiv preprint arXiv:1510.00149</source>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Peng</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>H.</given-names></name> <name><surname>Yu</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Overcoming challenges for achieving high in-situ training accuracy with emerging memories,</article-title> in <source>2020 Design, Automation &#x00026;Test in Europe Conference &#x00026;Exhibition (DATE)</source> (<publisher-loc>Grenoble</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1025</fpage>&#x02013;<lpage>1030</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Islam</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>P.-Y.</given-names></name> <name><surname>Wan</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>H.-Y.</given-names></name> <name><surname>Gao</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Device and materials requirements for neuromorphic computing</article-title>. <source>J. Phys. D Appl. Phys</source>. 52, 113001. <pub-id pub-id-type="doi">10.1088/1361-6463/aaf784</pub-id><pub-id pub-id-type="pmid">33837601</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kandel</surname> <given-names>I.</given-names></name> <name><surname>Castelli</surname> <given-names>M.</given-names></name> <name><surname>Popovi&#x0010D;</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Comparative study of first order optimizers for image classification using convolutional neural networks on histopathology images</article-title>. <source>J. Imaging</source> <volume>6</volume>, <fpage>92</fpage>. <pub-id pub-id-type="doi">10.3390/jimaging6090092</pub-id><pub-id pub-id-type="pmid">34460749</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Rasch</surname> <given-names>M.</given-names></name> <name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Ando</surname> <given-names>T.</given-names></name> <name><surname>Miyazoe</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>J.-J.</given-names></name> <etal/></person-group>. (<year>2019a</year>). <article-title>Zero-shifting technique for deep neural network training on resistive cross-point arrays</article-title>. <source>arXiv preprint arXiv:1907.10228</source>.</citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Todorov</surname> <given-names>T.</given-names></name> <name><surname>Onen</surname> <given-names>M.</given-names></name> <name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Bishop</surname> <given-names>D.</given-names></name> <name><surname>Solomon</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2019b</year>). <article-title>Metal-oxide based, cmos-compatible ecram for deep learning accelerator,</article-title> in <source>2019 IEEE International Electron Devices Meeting (IEDM)</source> (<publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>35</fpage>&#x02013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>T.</given-names></name> <name><surname>Hu</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Kwak</surname> <given-names>J. Y.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Spiking neural network (snn) with memristor synapses having non-linear weight update</article-title>. <source>Front. Comput. Neurosci</source>. <volume>15</volume>:<fpage>22</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2021.646125</pub-id><pub-id pub-id-type="pmid">33776676</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>arXiv preprint arXiv:1412.6980</source>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kwon</surname> <given-names>D.</given-names></name> <name><surname>Lim</surname> <given-names>S.</given-names></name> <name><surname>Bae</surname> <given-names>J.-H.</given-names></name> <name><surname>Lee</surname> <given-names>S.-T.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Seo</surname> <given-names>Y.-T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>On-chip training spiking neural networks using approximated backpropagation with analog synaptic devices</article-title>. <source>Front. Neurosci</source>. <volume>14</volume>:<fpage>423</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2020.00423</pub-id><pub-id pub-id-type="pmid">32733180</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>C.</given-names></name> <name><surname>Rajput</surname> <given-names>K. G.</given-names></name> <name><surname>Choi</surname> <given-names>W.</given-names></name> <name><surname>Kwak</surname> <given-names>M.</given-names></name> <name><surname>Nikam</surname> <given-names>R. D.</given-names></name> <name><surname>Kim</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Pr 0.7 ca 0.3 mno 3-based three-terminal synapse for neuromorphic computing</article-title>. <source>IEEE Electr. Device Lett</source>. <volume>41</volume>, <fpage>1500</fpage>&#x02013;<lpage>1503</lpage>. <pub-id pub-id-type="doi">10.1109/LED.2020.3019938</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rasch</surname> <given-names>M. J.</given-names></name> <name><surname>Moreda</surname> <given-names>D.</given-names></name> <name><surname>Gokmen</surname> <given-names>T.</given-names></name> <name><surname>Gallo</surname> <given-names>M. L.</given-names></name> <name><surname>Carta</surname> <given-names>F.</given-names></name> <name><surname>Goldberg</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>A flexible and fast pytorch toolkit for simulating training and inference on analog crossbar arrays</article-title>. <source>arXiv preprint arXiv:2104.02184</source>. <pub-id pub-id-type="doi">10.1109/AICAS51828.2021.9458494</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Williams</surname> <given-names>R. J.</given-names></name></person-group> (<year>1986</year>). <article-title>Learning representations by back-propagating errors</article-title>. <source>Nature</source> <volume>323</volume>, <fpage>533</fpage>&#x02013;<lpage>536</lpage>. <pub-id pub-id-type="doi">10.1038/323533a0</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scellier</surname> <given-names>B.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Equilibrium propagation: bridging the gap between energy-based models and backpropagation</article-title>. <source>Front. Comput. Neurosci</source>. <volume>11</volume>:<fpage>24</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2017.00024</pub-id><pub-id pub-id-type="pmid">28522969</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Time complexity of in memory matrix vector multiplication</article-title>. <source>IEEE Trans. Circ. Syst. II Express Briefs</source> <volume>68</volume>, <fpage>2785</fpage>&#x02013;<lpage>2789</lpage>. <pub-id pub-id-type="doi">10.1109/TCSII.2021.3068764</pub-id><pub-id pub-id-type="pmid">34413731</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Z.</given-names></name> <name><surname>Pedretti</surname> <given-names>G.</given-names></name> <name><surname>Ambrosi</surname> <given-names>E.</given-names></name> <name><surname>Bricalli</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Ielmini</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Solving matrix equations in one step with cross-point resistive arrays</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>116</volume>, <fpage>4123</fpage>&#x02013;<lpage>4128</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1815682116</pub-id><pub-id pub-id-type="pmid">30782810</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tsai</surname> <given-names>H.</given-names></name> <name><surname>Ambrogio</surname> <given-names>S.</given-names></name> <name><surname>Narayanan</surname> <given-names>P.</given-names></name> <name><surname>Shelby</surname> <given-names>R. M.</given-names></name> <name><surname>Burr</surname> <given-names>G. W.</given-names></name></person-group> (<year>2018</year>). <article-title>Recent progress in analog memory-based accelerators for deep learning</article-title>. <source>J. Phys. D Appl. Phys</source>. 51, 283001. <pub-id pub-id-type="doi">10.1088/1361-6463/aac8a5</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van De Burgt</surname> <given-names>Y.</given-names></name> <name><surname>Melianas</surname> <given-names>A.</given-names></name> <name><surname>Keene</surname> <given-names>S. T.</given-names></name> <name><surname>Malliaras</surname> <given-names>G.</given-names></name> <name><surname>Salleo</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Organic electronics for neuromorphic computing</article-title>. <source>Nat. Electron</source>. <volume>1</volume>, <fpage>386</fpage>&#x02013;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1038/s41928-018-0103-3</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verhelst</surname> <given-names>M.</given-names></name> <name><surname>Moons</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>Embedded deep neural network processing: algorithmic and processor techniques bring deep learning to iot and edge devices</article-title>. <source>IEEE Solid State Circ. Mag</source>. <volume>9</volume>, <fpage>55</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1109/MSSC.2017.2745818</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>H.</given-names></name> <name><surname>Burr</surname> <given-names>G. W.</given-names></name> <name><surname>Hwang</surname> <given-names>C. S.</given-names></name> <name><surname>Wang</surname> <given-names>K. L.</given-names></name> <name><surname>Xia</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Resistive switching materials for information processing</article-title>. <source>Nat. Rev. Mater</source>. <volume>5</volume>, <fpage>173</fpage>&#x02013;<lpage>195</lpage>. <pub-id pub-id-type="doi">10.1038/s41578-019-0159-3</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>T. P.</given-names></name> <name><surname>Bennett</surname> <given-names>C. H.</given-names></name> <name><surname>Feinberg</surname> <given-names>B.</given-names></name> <name><surname>Agarwal</surname> <given-names>S.</given-names></name> <name><surname>Marinella</surname> <given-names>M. J.</given-names></name></person-group> (<year>2020</year>). <article-title>Analog architectures for neural network acceleration based on non-volatile memory</article-title>. <source>Appl. Phys. Rev</source>. 7, 031301. <pub-id pub-id-type="doi">10.1063/1.5143815</pub-id><pub-id pub-id-type="pmid">31553955</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zeiler</surname> <given-names>M. D..</given-names></name></person-group> (<year>2012</year>). <article-title>Adadelta: an adaptive learning rate method</article-title>. <source>arXiv preprint arXiv:1212.5701</source>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>E.</given-names></name> <name><surname>Zeng</surname> <given-names>L.</given-names></name> <name><surname>Luo</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Edge intelligence: Paving the last mile of artificial intelligence with edge computing</article-title>. <source>Proc. IEEE</source> <volume>107</volume>, <fpage>1738</fpage>&#x02013;<lpage>1762</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2019.2918951</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
</ref-list> 
</back>
</article> 