<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Oncol.</journal-id>
<journal-title>Frontiers in Oncology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Oncol.</abbrev-journal-title>
<issn pub-type="epub">2234-943X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fonc.2023.1101225</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Oncology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Comparison of initial learning algorithms for long short-term memory method on real-time respiratory signal prediction</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Sun</surname>
<given-names>Wenzheng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1417151"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dang</surname>
<given-names>Jun</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/847156"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Lei</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wei</surname>
<given-names>Qichun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/796817"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Radiation Oncology, The Second Affiliated Hospital, School of Medicine, Zhejiang University</institution>, <addr-line>Hangzhou, Zhejiang</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Radiation Oncology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital and Shenzhen Hospital, Chinese Academy of Medical Science and Peking Union Medical College</institution>, <addr-line>Shenzhen, Guangdong</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Department of Oncology, The First Affiliated Hospital of Chongqing Medical University</institution>, <addr-line>Chongqing</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Graduate Program of Medical Physics and Data Science Research Center, Duke Kunshan University</institution>, <addr-line>Kunshan, Jiangsu</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Humberto Rocha, University of Coimbra, Portugal</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Xiaokun Liang, Shenzhen Institutes of Advanced Technology (CAS), China; Ahmad Esmaili Torshabi, Graduate University of Advanced Technology, Iran</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Qichun Wei, <email xlink:href="mailto:qichun_wei@zju.edu.cn">qichun_wei@zju.edu.cn</email>
</p>
</fn>
<fn fn-type="other" id="fn002">
<p>This article was submitted to Radiation Oncology, a section of the journal Frontiers in Oncology</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>13</volume>
<elocation-id>1101225</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Sun, Dang, Zhang and Wei</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Sun, Dang, Zhang and Wei</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Aim</title>
<p>This study aimed to examine the effect of the weight initializers on the respiratory signal prediction performance using the long short-term memory (LSTM) model.</p>
</sec>
<sec>
<title>Methods</title>
<p>Respiratory signals collected with the CyberKnife Synchrony device during 304 breathing motion traces were used in this study. The effectiveness of four weight initializers (Glorot, He, Orthogonal, and Narrow-normal) on the prediction performance of the LSTM model was investigated. The prediction performance was evaluated by the normalized root mean square error (NRMSE) between the ground truth and predicted respiratory signal.</p>
</sec>
<sec>
<title>Results</title>
<p>Among the four initializers, the He initializer showed the best performance. The mean NRMSE with 385-ms ahead time using the He initializer was superior by 7.5%, 8.3%, and 11.3% as compared to that using the Glorot, Orthogonal, and Narrow-normal initializer, respectively. The confidence interval of NRMSE using Glorot, He, Orthogonal, and Narrow-normal initializer were [0.099, 0.175], [0.097, 0.147], [0.101, 0.176], and [0.107, 0.178], respectively.</p>
</sec>
<sec>
<title>Conclusions</title>
<p>The experiment results in this study indicated that He could be a valuable initializer in the LSTM model for the respiratory signal prediction.</p>
</sec>
</abstract>
<kwd-group>
<kwd>respiratory signals prediction</kwd>
<kwd>initializer</kwd>
<kwd>long short-term memory</kwd>
<kwd>radiation therapy</kwd>
<kwd>He initializer</kwd>
<kwd>Glorot initializer</kwd>
<kwd>orthogonal initializer</kwd>
<kwd>narrow-normal initializer</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="15"/>
<ref-count count="29"/>
<page-count count="7"/>
<word-count count="3266"/>
</counts>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>During radiation therapy treatment delivery process, tumor in certain organs, such as the lung, would be subject to substantial motion due to patient respiration (<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B5">5</xref>). This motion may lead to the leakage of radiation dose from the tumor target to nearby normal tissues, which would sharply degrade the accuracy and quality of the radiation therapy treatment. Respiratory motion could be measured and monitored by several mature techniques. However, the real-time adaptation to motion during radiotherapy treatment is challenging, and latencies in hundreds of milliseconds may still exist (<xref ref-type="bibr" rid="B6">6</xref>&#x2013;<xref ref-type="bibr" rid="B10">10</xref>). Hence, prediction of the tumor motion in advance could help reduce these latencies and improve the quality of radiotherapy treatment in mobile cancers.</p>
<p>Many machine learning models have been proposed to predict the respiratory motion. Putra et&#xa0;al. investigated the prediction performance of the Kalman filter (KF) for a short latency (<xref ref-type="bibr" rid="B11">11</xref>). Recent studies have demonstrated the merit of artificial neural network (ANN) models on respiratory signal prediction, especially for the nonlinear signals (<xref ref-type="bibr" rid="B3">3</xref>, <xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B13">13</xref>). Sharp et&#xa0;al. indicated that the ANN models could have better prediction performance as compared to the KF method (<xref ref-type="bibr" rid="B6">6</xref>). Sun et&#xa0;al. proposed a revamped multilayer perceptron neural network (MLP-NN) called Adaboost MLP-NN (ADMLP-NN), which showed more accurate predictions than the MLP-NN (<xref ref-type="bibr" rid="B3">3</xref>). One of the main limitations of the ANN models was that they generally ignore the temporal dependence of the previous inputs. The recurrent neural network (RNN) was introduced to include the consideration of temporal information. However, the gradient disappearance and explosion problems restricted the application of RNN on long-term memory prediction. A special RNN model known as long short-term memory (LSTM) (<xref ref-type="bibr" rid="B14">14</xref>) had been proposed to overcome the above weakness of RNN (gradient disappearance and explosion problems). Various studies have demonstrated the superior performance of the LSTM model in different time-series prediction tasks (<xref ref-type="bibr" rid="B15">15</xref>&#x2013;<xref ref-type="bibr" rid="B20">20</xref>), including the respiratory signal prediction. Wang et&#xa0;al. showed that the prediction performance of a suitable LSTM model could be three-fold higher than that of the ADMLP-NN model (<xref ref-type="bibr" rid="B18">18</xref>). Another two studies also showed the potential of the LSTM model in the respiratory signal prediction (<xref ref-type="bibr" rid="B19">19</xref>, <xref ref-type="bibr" rid="B20">20</xref>).</p>
<p>Neural network models are sensitive to their initial weights (<xref ref-type="bibr" rid="B21">21</xref>, <xref ref-type="bibr" rid="B22">22</xref>). When the neural network was successfully proposed initially, the Narrow-normal initializer was usually used to generate the initial value of weight from a predefined normal sampling distribution. The weights between layers are initialized using a fixed variance distribution, which may cause the problem of gradient disappearance or gradient explosion (<xref ref-type="bibr" rid="B21">21</xref>, <xref ref-type="bibr" rid="B22">22</xref>). This defect hampered the extensive use of the Narrow-normal initializer. The Glorot initializer (<xref ref-type="bibr" rid="B23">23</xref>) provided a normalized initialization, which could maintain the activation and the back-propagated gradient variances during training based on the linear activation. Then, He et&#xa0;al. took the rectifier nonlinearity into account and proposed the He initializer (<xref ref-type="bibr" rid="B24">24</xref>). Sachs et&#xa0;al. illustrated that if the initial weight obeyed an orthogonal matrix, the initial conditions could keep the error vector norm through the deep neural network during back propagation process while generating depth independent learning times. Based on this finding, the orthogonal initializer (<xref ref-type="bibr" rid="B25">25</xref>) was proposed.</p>
<p>While the critical importance of the proper weight initializers on time-series prediction performance has often been noted in previous works (<xref ref-type="bibr" rid="B21">21</xref>&#x2013;<xref ref-type="bibr" rid="B26">26</xref>), researchers often focused on improving the model architecture. Therefore, there is a lack of knowledge on the effect of initializer on respiratory signal prediction using the LSTM model. The aim of this study was to examine the effect of different initializer on the performance of LSTM model in the patient respiratory signal prediction. The primary contributions of this study were concluded as follows:</p>
<p>1. To the best of our knowledge, this is the first study to investigate the effectiveness of the weight initializers on the respiratory prediction problem using the LSTM model. In this study, we investigated the influence of the four common weight initializers discussed in the literature (<xref ref-type="bibr" rid="B22">22</xref>, <xref ref-type="bibr" rid="B27">27</xref>) on the prediction performance of LSTM model using 304 breathing motion cases from an open-access database collected by the CyberKnife Synchrony tracking system (Accuray, Sunnyvale, CA) with a 26-Hz sampling rate.</p>
<p>2. We further investigated the effect of the irregular breathing patterns on the prediction performance for each initializer and demonstrated the advantage of using the He initializer on irregular respiratory pattern patients.</p>
<p>The results illustrated that the initial weight algorithms in the LSTM model would exert substantial effect on the respiratory signal prediction performance. The He initializer could be an optimal choice for the respiratory signal prediction, especially for the irregular respiratory pattern patients.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Prediction process</title>
<p>The general workflow for the prediction process used in this study is outlined in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. Each respiratory signal was divided into two segments separated by time A. Respiratory signals prior to time A were used as training data, and those after time A were used for testing. Among the training data, the signals before and after time C were used as the input and prediction outputs, respectively. LSTM neural network was implemented as the prediction model. The testing data (positions after the time A) were used to evaluate the developed LSTM prediction model. The testing data were divided into two segments separated by time D. The signals before time D were defined as the testing input, and those after time D were defined as the testing outputs or ground truth. The trained LSTM prediction model was applied to the testing inputs to generate the prediction signal P&#x2019;, which was compared to the testing outputs or ground truth P for evaluation.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Flow chart of the prediction algorithm.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-13-1101225-g001.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Long short-term memory neural network</title>
<p>A bidirectional architecture LSTM layer was developed to test the performance of all the four initializers in this study (<xref ref-type="bibr" rid="B17">17</xref>, <xref ref-type="bibr" rid="B18">18</xref>). The formulas of the LSTM layer are illustrated in the <bold>Equations 1&#x2013;8</bold>.</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>h</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>b</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>f</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>b</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>b</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>f</mml:mi>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>b</mml:mi>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>y</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here, <italic>i<sub>t</sub>
</italic>, <italic>f<sub>t</sub>
</italic>, <italic>c<sub>t</sub>
</italic>, <italic>o<sub>t</sub>
</italic>, <italic>h<sub>t</sub>
</italic>, <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>f</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>b</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, and <italic>y<sub>t</sub> </italic> refer to the input gate, forget gate, memory cell vectors, output gate, hidden vector sequence, forward hidden vector sequence, backward hidden vector sequence, and output, respectively. The tanh and &#x3c3; stand for the two activation functions as given by the <bold>Equations 9</bold> and <bold>10</bold>.</p>
<disp-formula>
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mi>tanh</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<bold>
<italic>W</italic>
</bold>
<italic>
<sub>xi</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>hi</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>ci</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>xf</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>hf</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>cf</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>xc</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>hc</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>xo</sub>
</italic>, <bold>
<italic>W</italic>
</bold>
<italic>
<sub>ho</sub>
</italic>, and <bold>
<italic>W</italic>
</bold>
<italic>
<sub>co</sub>
</italic> denote the weighted parameters, while the <bold>
<italic>b</italic>
</bold>
<italic>
<sub>i</sub>
</italic>, <bold>
<italic>b</italic>
</bold>
<italic>
<sub>f</sub>
</italic>, <bold>
<italic>b</italic>
</bold>
<italic>
<sub>c</sub>
</italic>, and <bold>
<italic>b</italic>
</bold>
<italic>
<sub>o</sub>
</italic> represent the intercepts. A plurality of LSTM layers can be stacked into a deeper neural network, which can fit the complicated functions between the inputs and targets.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Initializers</title>
<p>For the Glorot initializers (also called the Xavier initializer), the weights <italic>W<sub>ij</sub>
</italic> of each layer were initialized with the heuristic as follows (<xref ref-type="bibr" rid="B23">23</xref>):</p>
<disp-formula>
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>~</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi mathvariant="script">U</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo> <mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>6</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi>&#x2110;</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi mathvariant="script">O</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
<mml:mo>,</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>6</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi>&#x2110;</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi mathvariant="script">O</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
</mml:mrow> <mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here, <italic>U</italic>[ &#x2212;&#x3b4;,&#xa0;&#x3b4; ] was sampled independently from a zero-mean uniform distribution in the bounds [ &#x2212;&#x3b4;,&#xa0;&#x3b4; ] . <italic>N<sub>0</sub>
</italic> was four times of the hidden unit number, and <italic>N<sub>J</sub>
</italic> was the input channel number.</p>
<p>For the He initializer (<xref ref-type="bibr" rid="B24">24</xref>), the weights <italic>W<sub>ij</sub>
</italic> were sampled according to a zero-mean normal (Gaussian) distribution with the following standard deviation (SD):</p>
<disp-formula>
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>~</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi mathvariant="script">G</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo> <mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi>&#x2110;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
<mml:mo>,</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="script">N</mml:mi>
<mml:mi>&#x2110;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
</mml:mrow> <mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here, the size of <italic>N<sub>J</sub>
</italic> was the input channel number for the input weight and the hidden unit number for the recurrent weight.</p>
<p>For the Orthogonal initializer (<xref ref-type="bibr" rid="B25">25</xref>), the orthogonal matrix <bold>
<italic>Q</italic>
</bold> was first established by producing the Gaussian matrices randomly and then computed by the <bold>
<italic>QR</italic>
</bold> decomposition using the formula <bold>
<italic>Z</italic>
</bold> = <bold>
<italic>QR</italic>
</bold>, where <bold>
<italic>Z</italic>
</bold> obeyed unit normal distribution.</p>
<p>For the Narrow-normal initializer, the weights were obtained by sampling from a normal distribution with 0 mean and 0.01 SD independently.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Prediction performance evaluation</title>
<p>A total of 304 breathing motion traces collected by the CyberKnife Synchrony (Accuray, Sunnyvale, CA) tracking system and procured from an open dataset (<xref ref-type="bibr" rid="B18">18</xref>, <xref ref-type="bibr" rid="B28">28</xref>) were used in this study. The detail information of the dataset is illustrated in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. The first 1-min signal was used to train the LSTM model with different initial methods, while the following 30 s was applied to evaluate the effectiveness of each initial method. The ahead time was about 385 ms (10 samples). All the evaluation metrics were based on the following default hyper-parameters in this study: three LSTM layers, 0.001 initial learning rate, 50 time lags, and 300 hidden units.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Detail information of the dataset used in this study.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Items</th>
<th valign="middle" align="center">Result</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">Sampling Rate</td>
<td valign="middle" align="left">26 Hz</td>
</tr>
<tr>
<td valign="middle" align="left">Collected Equipment</td>
<td valign="middle" align="left">CyberKnife Synchrony</td>
</tr>
<tr>
<td valign="middle" align="left">Treatment Place</td>
<td valign="middle" align="left">Georgetown University Hospital</td>
</tr>
<tr>
<td valign="middle" align="left">Trace Number</td>
<td valign="middle" align="left">304</td>
</tr>
<tr>
<td valign="middle" align="left">Treatment Fraction</td>
<td valign="middle" align="left">102</td>
</tr>
<tr>
<td valign="middle" align="left">Patient Number</td>
<td valign="middle" align="left">31</td>
</tr>
<tr>
<td valign="middle" align="left">Private Information</td>
<td valign="middle" align="left">None</td>
</tr>
<tr>
<td valign="middle" align="left">Respiratory Signal Record Way</td>
<td valign="middle" align="left">Fiducial marker on patient&#x2019;s chest</td>
</tr>
<tr>
<td valign="middle" align="left">Datasets Range in Duration</td>
<td valign="middle" align="left">80&#x2013;158 min</td>
</tr>
<tr>
<td valign="middle" align="left">Imaging Data</td>
<td valign="middle" align="left">None</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>RMSE is one of the main metrics used for respiratory signal prediction evaluation (<xref ref-type="bibr" rid="B2">2</xref>&#x2013;<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B8">8</xref>&#x2013;<xref ref-type="bibr" rid="B11">11</xref>, <xref ref-type="bibr" rid="B19">19</xref>, <xref ref-type="bibr" rid="B20">20</xref>). However, the root mean square error (RMSE) was not a dimensionless and normalized metric (<xref ref-type="bibr" rid="B29">29</xref>), and it would not be suitable for comparing the performance of respiratory signal prediction across different patients before normalization (<xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B18">18</xref>). Hence, in order to minimize the effect of signal amplitude on different cases, the NRMSE instead of the RMSE between the real and predicted signal were used to evaluate the predication performance for all initializers (<xref ref-type="bibr" rid="B29">29</xref>). The NRMSEs used in this study are illustrated by <bold>Equations 13</bold> and <bold>14</bold>.</p>
<disp-formula>
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mi>R</mml:mi>
<mml:mi>M</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mtext>D</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mtext>D</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>P</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>P</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
<mml:mo>)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>P</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
<mml:mo>)</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>P</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
<mml:mo>)</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>(</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here, <italic>P</italic>(<italic>t</italic>) and <italic>P'</italic>(<italic>t</italic>) were the ground truth and predicted signal at the time fame <italic>t</italic>, respectively.</p>
<p>Breathing irregularity was examined to test prediction performance of the LSTM model using the four initializers on different breathing patterns. As illustrated in the <bold>Equation 15</bold>, the breathing irregularity was defined as the average of the standard deviation (SD) of the maximum (i.e., peak) and minimum (i.e., valley) amplitudes. Patients were split into two groups by the median value of the irregularity (<italic>r</italic>= 0.22).</p>
<disp-formula>
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>y</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results</title>
<p>
<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref> shows the NRMSEs using four initializers (Glorot, He, Orthogonal, and Narrow-normal) with a 385-ms ahead time. The He initializer showed the best prediction performance. The mean of NRMSE using the He initializer was lower by 7.5%, 8.3%, and 11.3% compared to the Glorot, Orthogonal, and Narrow-normal initializer, respectively. The He initializer was more robust than the other three initializers, as it achieved the narrowest confidence interval (CI) among all the four initializers.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Prediction performance of the four initializers.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-13-1101225-g002.tif"/>
</fig>
<p>The prediction performance of the four initializers with different ahead time is shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>. The prediction performance using all of the four initializers decreased when the ahead time increases. The prediction performance using the He initializer for all the ahead time examined in this study were higher than the other three initializers. The average prediction performance gap between the He initializer and other three initializers increased when the ahead time increases.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Prediction performance of the four initializers with different ahead times.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-13-1101225-g003.tif"/>
</fig>
<p>The prediction performance of the regular and irregular breathing groups, which are divided by the median irregularity is shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. For all the four initializers, the performance of the irregular group was inferior to the regular group. The prediction performances of Glorot, He, and Orthogonal were similar in the regular group. However, the mean of NRMSE in the irregular group using the He initializer was superior by 11.0% (Glorot), 11.6% (Orthogonal), and 14.1% (Narrow-normal), respectively. The upper limits of the 95% confidence interval (CI) of the NRMSE for the irregular group patients were lowest for He (0.147) and were 0.175, 0.176, and 0.178 for Glorot, Orthogonal, and Narrow-normal, respectively.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Prediction performance of the two groups divided by the median irregularity. The green and red bars represented 95% CI of the NRMSE for the RE and IR groups, respectively.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-13-1101225-g004.tif"/>
</fig>
<p>The effect of the four important hyper-parameters on the prediction performance was explored and shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>. Four different values (1, 2, 3, and 4) were selected to evaluate the effect of the <bold>
<italic>N<sub>l</sub>
</italic>
</bold> on the prediction performance (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5A</bold>
</xref>). The LSTM model showed the similar prediction performance with one to three LSTM layers for the Glorot and Orthogonal initializers. For He and Narrow-normal initializers, the LSTM model with two or three LSTM layers showed similar prediction performance. The effect of <bold>
<italic>T<sub>l</sub>
</italic>
</bold> on the prediction performance was examined by five selected values (1, 5, 10, 25 and 50) (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5B</bold>
</xref>). Time lag of 10 showed the best performance for the He and Narrow-normal initializers, while 5 showed the best performance for the other two initializers. The influence of <bold>
<italic>H<sub>u</sub>
</italic>
</bold> was explored on five values (30, 50, 100, 300, and 500) (<xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5C</bold>
</xref>). The prediction performance kept improving as the <bold>
<italic>H<sub>u</sub>
</italic>
</bold> increased from 30 to 500 for all the four initializers. Four <bold>
<italic>L<sub>r</sub>
</italic>
</bold> values (0.0001, 0.001, 0.01, and 0.1) were investigated. The NRMSE using the first three initializers (Glorot, He, and Orthogonal) was lowest when <bold>
<italic>L<sub>r</sub>
</italic>
</bold> was 0.001. Narrow-normal initializer showed the best prediction performance when <bold>
<italic>L<sub>r</sub>
</italic>
</bold> was 0.01.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Effect of the hyper-parameters on the prediction performance. <bold>(A)</bold> Number of the LSTM layers (<italic>N<sub>l</sub>
</italic>). <bold>(B)</bold> Time lags (<italic>T<sub>l</sub>
</italic>). <bold>(C)</bold> Hidden units (<italic>H<sub>u</sub>
</italic>). <bold>(D)</bold> Learning rate (<italic>L<sub>r</sub>
</italic>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-13-1101225-g005.tif"/>
</fig>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>The effect of four common initializers on the performance of respiratory signal prediction using the LSTM model was examined in this study. The results illustrated that the He initializer outperformed the other three initializers for its higher respiratory prediction performance. A suitable initializer would substantially improve the prediction performance.</p>
<p>The prediction performance using all the four initializers became lower when the ahead time increases. This was probably because the relationship between the training and predicted respiratory signals would diminish when the ahead time increases. However, the prediction performance deterioration using the He initializer was slower than the other three methods. This may suggest that the He initializer could enhance LSTM&#x2019;s ability to capture longer ahead time information.</p>
<p>The prediction performance of the irregular breathing group was lower than the regular breathing group for all the four initializers. This may be contributed by the factor that the relationship between the prior and future signals in the irregular group was more difficult to capture than that in the regular group. The prediction performance of Glorot, He, and Orthogonal was similar in the regular group. However, the mean of NRMSE using the He initializer in the irregular group was lower compared to other three initializers. This suggested the superior ability of He initializer in capturing connection between prior and future signals. The upper limit of the 95% CI of the NRMSE using the He initializer was lower than that using other three initializers, suggesting that the He initializer might improve the general performance of the LSTM model in breathing signal prediction.</p>
<p>We also investigated the influence of hyper-parameters setting on the prediction performance for each initializer. A total of four important hyper-parameters was examined in this study. A large numerical value of the first three hyper-parameters (<bold>
<italic>N<sub>l,</sub> T<sub>l</sub>
</italic>
</bold>, and <bold>
<italic>H<sub>u</sub>
</italic>
</bold>) represented a complex network, which would fit a more complicated function but easy to overfit. The prediction performance became better as the first three hyper-parameters increased initially. However, as these three hyper-parameters continue to increase, the prediction performance improved only slightly and even could deteriorate. This may be because when these three hyper-parameters were too small, the LSTM model would not fit the respiratory curve well. Hence, the increase in these three hyper-parameters could improve the prediction performance initially. However, too large <bold>
<italic>N<sub>l</sub>
</italic>
</bold>, <bold>
<italic>T<sub>l</sub>
</italic>
</bold>, and <bold>
<italic>H<sub>u</sub>
</italic>
</bold> would raise the risk of overfitting for the LSTM model and potentially degrade the prediction accuracy. The <bold>
<italic>L<sub>r</sub>
</italic>
</bold> scales and updates the magnitude of the LSTM model weights to minimize the loss function. If <bold>
<italic>L<sub>r</sub>
</italic>
</bold> was too small, the converge time would be long, and the risk of trapping in undesirable local minimum increases. On the other hand, if <bold>
<italic>L<sub>r</sub>
</italic>
</bold> was too large, a suboptimal result may be obtained. Finally, 0.001 and 0.01 achieved the best performance for Narrow-normal and the other three initializers, respectively.</p>
<p>One of the limitations of this study was that all the respiratory signals from this database were originally detected by the fiducial marker placed on patient&#x2019;s chest. These external signals may be different from the real internal tumor motion. Besides, the dataset was collected from a single center. In the future, we would further evaluate the effect of the initializers on actual tumor motion signals or internal respiratory signals ideally from multi-centers.</p>
</sec>
<sec id="s5" sec-type="conclusion">
<label>5</label>
<title>Conclusion</title>
<p>The influence of the four weight initializers on the performance of respiratory signal prediction using the LSTM model was investigated in this study. The results suggested that the weight initialization methods would exert substantial effect on the respiratory signal prediction performance. The He initializer could be an optimal initializer for the respiratory signal prediction using the LSTM model, especially for the irregular respiratory pattern patients.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>WS and QW designed the methodology. WS and JD wrote the program. WS performed data analysis, interpretation and drafted the manuscript. LZ, JD, and QW revised this manuscript critically for important intellectual content and contributed to the interpretation of data. All authors reviewed the manuscript. All authors read and approved the final manuscript.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="funding-information">
<title>Funding</title>
<p>This work was partially supported by the National Natural Science Foundation of China (62103366), the General Project of Chongqing Natural Science Foundation (grant cstc2020jcyj-msxm2928), and Duke Kunshan University Education Development Foundation (22KDKUF032).</p>
</sec>
<sec id="s9" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Verma</surname> <given-names>P</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>H</given-names>
</name>
<name>
<surname>Langer</surname> <given-names>M</given-names>
</name>
<name>
<surname>Das</surname> <given-names>I</given-names>
</name>
<name>
<surname>Sandison</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>Survey: Real-time tumor motion prediction for image-guided radiation treatment</article-title>. <source>Computing Sci Eng</source> (<year>2011</year>) <volume>13</volume>(<issue>5</issue>):<fpage>24</fpage>&#x2013;<lpage>35</lpage>.</citation>
</ref>
<ref id="B2">
<label>2</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chang</surname> <given-names>PC</given-names>
</name>
<name>
<surname>Dang</surname> <given-names>J</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>JR</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>WZ</given-names>
</name>
</person-group>. <article-title>Real-time respiratory tumor motion prediction based on a temporal convolutional neural network: Prediction model development study</article-title>. <source>J Med Internet Res</source> (<year>2021</year>) <volume>23</volume>(<issue>8</issue>):<fpage>0</fpage>&#x2013;<lpage>e27235</lpage>.</citation>
</ref>
<ref id="B3">
<label>3</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>WZ</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>MY</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>L</given-names>
</name>
<name>
<surname>Dang</surname> <given-names>J</given-names>
</name>
<name>
<surname>You</surname> <given-names>T</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>FF</given-names>
</name>
<etal/>
</person-group>. <article-title>Respiratory signal prediction based on adaptive boosting and multi-layer perceptron neural network</article-title>. <source>Phys Med Biol</source> (<year>2017</year>) <volume>62</volume>(<issue>17</issue>):<page-range>6822&#x2013;35</page-range>.</citation>
</ref>
<ref id="B4">
<label>4</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bukhari</surname> <given-names>W</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>SM</given-names>
</name>
</person-group>. <article-title>Real-time prediction and gating of respiratory motion using an extended kalman filter and Gaussian process regression</article-title>. <source>Phys Med Biol</source> (<year>2015</year>) <volume>60</volume>(<issue>1</issue>):<page-range>233&#x2013;52</page-range>.</citation>
</ref>
<ref id="B5">
<label>5</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vergalasova</surname> <given-names>I</given-names>
</name>
<name>
<surname>Cai J, Yin</surname> <given-names>F-F</given-names>
</name>
</person-group>. <article-title>A novel technique for markerless, self-sorted 4d-cbct: Feasibility study</article-title>. <source>Med Phys</source> (<year>2012</year>) <volume>39</volume>(<issue>3</issue>):<page-range>1442&#x2013;51</page-range>.</citation>
</ref>
<ref id="B6">
<label>6</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sharp</surname> <given-names>GC</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>SB</given-names>
</name>
<name>
<surname>Shimizu S, Shirato</surname> <given-names>H</given-names>
</name>
</person-group>. <article-title>Prediction of respiratory tumour motion for real-time image guided radiotherapy</article-title>. <source>Phys Med Biol</source> (<year>2004</year>) <volume>49</volume>:<page-range>425&#x2013;40</page-range>.</citation>
</ref>
<ref id="B7">
<label>7</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ernst</surname> <given-names>F</given-names>
</name>
<name>
<surname>Schlaefer</surname> <given-names>A</given-names>
</name>
<name>
<surname>Schweikard</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Predicting the outcome of respiratory motion prediction</article-title>. <source>Med Phys</source> (<year>2011</year>) <volume>38</volume>(<issue>10</issue>):<page-range>5569&#x2013;81</page-range>.</citation>
</ref>
<ref id="B8">
<label>8</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>WZ</given-names>
</name>
<name>
<surname>Chun</surname> <given-names>QW</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>L</given-names>
</name>
<name>
<surname>Dang J, Yin</surname> <given-names>FF</given-names>
</name>
</person-group>. <article-title>Adaptive respiratory signal prediction using dual multi-layer perceptron neural networks (MLP-NNs)</article-title>. <source>Phys Med Biol</source> (<year>2020</year>) <volume>65</volume>(<issue>18</issue>):<fpage>185005</fpage>.</citation>
</ref>
<ref id="B9">
<label>9</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Torshabi</surname> <given-names>AE</given-names>
</name>
<name>
<surname>Riboldi</surname> <given-names>M</given-names>
</name>
<name>
<surname>Fooladi</surname> <given-names>AAI</given-names>
</name>
<name>
<surname>Mosalla</surname> <given-names>M</given-names>
</name>
<name>
<surname>Borani</surname> <given-names>G</given-names>
</name>
<etal/>
</person-group>. <article-title>An adaptive fuzzy prediction model for real time tumor tracking in radiotherapy <italic>via</italic> external surrogates</article-title>. <source>J Appl Clin Med Phys</source> (<year>2013</year>) <volume>14</volume>(<issue>1</issue>):<page-range>102&#x2013;14</page-range>.</citation>
</ref>
<ref id="B10">
<label>10</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Nishioka</surname> <given-names>S</given-names>
</name>
<name>
<surname>Shirato H, Berbeco</surname> <given-names>RI</given-names>
</name>
</person-group>. <article-title>Adaptive prediction of respiratory motion for motion compensation radiotherapy</article-title>. <source>Phys Med Biol</source> (<year>2007</year>) <volume>52</volume>(<issue>22</issue>):<page-range>6651&#x2013;61</page-range>.</citation>
</ref>
<ref id="B11">
<label>11</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Putra</surname> <given-names>D</given-names>
</name>
<name>
<surname>Haas</surname> <given-names>OCL</given-names>
</name>
<name>
<surname>Mills JA, Bumham</surname> <given-names>KJ</given-names>
</name>
</person-group>. <article-title>Prediction of tumour motion using interacting multiple model filter</article-title>. In: <italic>2006 IET 3rd international conference on advances in medical, signal and information processing - MEDSIP 2006</italic>. <conf-loc>Glasgow, UK</conf-loc> (<year>2016</year>) pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>, doi:&#xa0;<pub-id pub-id-type="doi">10.1049/cp:20060350</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Isaksson</surname> <given-names>M</given-names>
</name>
<name>
<surname>Jalden J, Murphy</surname> <given-names>MJ</given-names>
</name>
</person-group>. <article-title>On using an adaptive neural network to predict lung tumor motion during respiration for radiotherapy applications</article-title>. <source>Med Phys</source> (<year>2005</year>) <volume>32</volume>:<page-range>3801&#x2013;9</page-range>.</citation>
</ref>
<ref id="B13">
<label>13</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsai</surname> <given-names>TI</given-names>
</name>
<name>
<surname>Li</surname> <given-names>DC</given-names>
</name>
</person-group>. <article-title>Approximate modeling for high order non-linear functions using small sample sets</article-title>. <source>Expert Syst Appl</source> (<year>2008</year>) <volume>34</volume>(<issue>1</issue>):<page-range>564&#x2013;9</page-range>.</citation>
</ref>
<ref id="B14">
<label>14</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hochreiter</surname> <given-names>S</given-names>
</name>
<name>
<surname>Schmidhuber</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>Long short-term memory</article-title>. <source>Neural Comput</source> (<year>1997</year>) <volume>9</volume>:<page-range>1735&#x2013;80</page-range>.</citation>
</ref>
<ref id="B15">
<label>15</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sutskever</surname> <given-names>I</given-names>
</name>
<name>
<surname>Vinyals O, Le</surname> <given-names>QV</given-names>
</name>
</person-group>. <article-title>Sequence to sequence learning with neural networks</article-title>. <source>Adv Neural Inf Process Syst</source> (<year>2014</year>) <volume>2</volume>:<page-range>3104&#x2013;3112</page-range>.</citation>
</ref>
<ref id="B16">
<label>16</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cho</surname> <given-names>K</given-names>
</name>
<name>
<surname>Van Merri&#xeb;nboer</surname> <given-names>B</given-names>
</name>
<name>
<surname>Gulcehre</surname> <given-names>C</given-names>
</name>
<name>
<surname>Bahdanau</surname> <given-names>D</given-names>
</name>
<name>
<surname>Bougares</surname> <given-names>F</given-names>
</name>
<name>
<surname>Schwenk</surname> <given-names>H</given-names>
</name>
<etal/>
</person-group>. <article-title>Learning phrase representations using RNN encoder-decoder for statistical machine translation</article-title>, in: <conf-name>Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing</conf-name> <conf-loc>(EMNLP). Doha, Qatar</conf-loc> (<year>2014</year>) pp. <page-range>1724&#x2013;34</page-range>.</citation>
</ref>
<ref id="B17">
<label>17</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rahman</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Watanobe Y, Nakamura</surname> <given-names>K</given-names>
</name>
</person-group>. <article-title>A bidirectional lstm language model for code evaluation and repair</article-title>. <source>Symmetry</source> (<year>2021</year>) <volume>13</volume>(<issue>2</issue>):<fpage>247</fpage>.</citation>
</ref>
<ref id="B18">
<label>18</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>R</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>X</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>Y</given-names>
</name>
</person-group>. <article-title>A feasibility of respiration prediction based on deep bi-LSTM for real-time tumor tracking</article-title>, in: <conf-name>IEEE Access</conf-name>. (<year>2018</year>) Vol. <volume>6</volume>. pp. <page-range>51262&#x2013;8</page-range>.</citation>
</ref>
<ref id="B19">
<label>19</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>H</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>CY</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>B</given-names>
</name>
<name>
<surname>Chan</surname> <given-names>MF</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>W</given-names>
</name>
</person-group>. <article-title>Towards real-time respiratory motion prediction based on long short-term memory neural networks</article-title>. <source>Phys Med Biol</source> (<year>2019</year>) <volume>64</volume>(<issue>8</issue>):<fpage>085010</fpage>.</citation>
</ref>
<ref id="B20">
<label>20</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>G</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Li</surname> <given-names>G</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>G</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Bai</surname> <given-names>L</given-names>
</name>
<etal/>
</person-group>. <article-title>Real-time liver tracking algorithm based on LSTM and SVR networks for use in surface-guided radiation therapy</article-title>. <source>Radiat Oncol</source> (<year>2021</year>) <volume>16</volume>(<issue>1</issue>):<fpage>13</fpage>.</citation>
</ref>
<ref id="B21">
<label>21</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname> <given-names>AY</given-names>
</name>
<name>
<surname>Koh</surname> <given-names>YS</given-names>
</name>
<name>
<surname>Riddle</surname> <given-names>P</given-names>
</name>
<name>
<surname>Pfahringer</surname> <given-names>B</given-names>
</name>
</person-group>. <article-title>Using supervised pretraining to improve generalization of neural networks on binary classification problems</article-title>. <source>Joint Eur Conf Mach Learn Knowledge Discovery Database</source> (<year>2018</year>) <volume>11051</volume>, <page-range>410&#x2013;425</page-range>.</citation>
</ref>
<ref id="B22">
<label>22</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H</given-names>
</name>
<name>
<surname>Krek M, Perin</surname> <given-names>G</given-names>
</name>
</person-group>. <article-title>A comparison of weight initializers in deep learning-based side-channel analysis</article-title>. In: <source>Applied cryptography and network security workshops</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2020</year>) <volume>11051</volume>:<page-range>410&#x2013;425</page-range>.</citation>
</ref>
<ref id="B23">
<label>23</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Glorot</surname> <given-names>X</given-names>
</name>
<name>
<surname>Bengio</surname> <given-names>Y</given-names>
</name>
</person-group>. <article-title>Understanding the difficulty of training deep feedforward neural networks</article-title>. <source>JMLR Workshop Conf Proc</source> (<year>2010</year>) <volume>9</volume>:<page-range>249&#x2013;256</page-range>.</citation>
</ref>
<ref id="B24">
<label>24</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>S</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>Developing deep into rectifiers: Surpassing human-level performance on ImageNet classification</article-title>, in: <conf-name>2015 IEEE International Conference on Computer Vision</conf-name> <conf-loc>(ICCV). Santiago, Chile</conf-loc> (<year>2015</year>) pp. <fpage>1026-34</fpage>, doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV.2015.123</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saxe</surname> <given-names>AM</given-names>
</name>
<name>
<surname>Mcclelland</surname> <given-names>JL</given-names>
</name>
<name>
<surname>Ganguli</surname> <given-names>S</given-names>
</name>
</person-group>. <article-title>Exact solutions to the nonlinear dynamics of learning in deep linear neural networks</article-title>. <source>Int Conf Learn Representations</source> (<year>2014</year>).</citation>
</ref>
<ref id="B26">
<label>26</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>HC</given-names>
</name>
<name>
<surname>Kang</surname> <given-names>MJ</given-names>
</name>
</person-group>. <article-title>Comparison of weight initialization techniques for deep neural networks</article-title>. <source>Int J Advanced Culture Technol</source> (<year>2019</year>) <volume>7</volume>(<issue>4</issue>):<page-range>283&#x2013;8</page-range>.</citation>
</ref>
<ref id="B27">
<label>27</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qi</surname> <given-names>ZA</given-names>
</name>
<name>
<surname>Ro</surname> <given-names>B</given-names>
</name>
</person-group>. <article-title>Performance of neural network for indoor airflow prediction: Sensitivity towards weight initialization</article-title>. <source>Energy Buildings</source> (<year>2021</year>) <volume>246</volume>:<fpage>111106</fpage>.</citation>
</ref>
<ref id="B28">
<label>28</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Floris</surname> <given-names>E</given-names>
</name>
</person-group>. <source>Compensating for quasi-periodic motion in robotic radiosurgery</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2012</year>). doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-1-4614-1912-9</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Debaditya</surname> <given-names>C</given-names>
</name>
<name>
<surname>Hazem</surname> <given-names>E</given-names>
</name>
</person-group>. <article-title>Performance testing of energy models: Are we using the right statistical metrics</article-title>? <source>J Building Perform Simulation</source> (<year>2018</year>) <volume>11</volume>(<issue>4</issue>):<page-range>433&#x2013;48</page-range>.</citation>
</ref>
</ref-list>
</back>
</article>