<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1210703</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2023.1210703</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Study of diagnosis for rotating machinery in advanced nuclear reactor based on deep learning model</article-title>
<alt-title alt-title-type="left-running-head">Sun and Wang</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fenrg.2023.1210703">10.3389/fenrg.2023.1210703</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Sun</surname>
<given-names>Yuanli</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2055441/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Hang</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Tsinghua University Nuclear Research Institute</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Nuclear Science and Technology of Harbin Engineering University</institution>, <addr-line>Harbin</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1880228/overview">Yaoli Zhang</ext-link>, Xiamen University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/839224/overview">Haochun Zhang</ext-link>, Harbin Institute of Technology, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/874360/overview">Deqi Chen</ext-link>, Chongqing University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence:Yuanli Sun, <email>syl850122@126.com</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>18</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1210703</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Sun and Wang.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Sun and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Many types of rotating mechanical equipment, such as the primary pump, turbine, and fans, are key components of fourth-generation (Gen IV) advanced reactors. Given that these machines operate in challenging environments with high temperatures and liquid metal corrosion, accurate problem identification and health management are essential for keeping these machines in good working order. This study proposes a deep learning (DL)-based intelligent diagnosis model for the rotating machinery used in fast reactors. The diagnosis model is tested by identifying the faults of bearings and gears. Normalization, augmentation, and splitting of data are applied to prepare the datasets for classification of faults. Multiple diagnosis models containing the multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN), and residual network (RESNET) are compared and investigated with the Case Western Reserve University datasets. An improved Transformer model is proposed, and an enhanced embeddings generator is designed to combine the strengths of the CNN and transformer. The effects of the size of the training samples and the domain of data preprocessing, such as the time domain, frequency domain, time-frequency domain, and wavelet domain, are investigated, and it is found that the time-frequency domain is most effective, and the improved Transformer model is appropriate for the fault diagnosis of rotating mechanical equipment. Because of the low probability of the occurrence of a fault, the imbalanced learning method should be improved in future studies.</p>
</abstract>
<kwd-group>
<kwd>fault diagnosis model</kwd>
<kwd>deep learning</kwd>
<kwd>rotating machine</kwd>
<kwd>advanced nuclear reactor</kwd>
<kwd>improved transformer model</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Nuclear Energy</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Commercial pressure water reactors (PWRs) are undergoing life extensions from 40 years of operation to 60 years of operation. The safe and economic operation of these plants is considered in anticipation of a second round of license extensions. The technical issues are similar to the life extension of advanced reactors. In fact, many small modular reactors (SMRs) and designed advanced reactors have increased their operating cycles (typically ten more years after the designed 40 years). As key mechanical components both in PWRs and advanced reactors, rotating mechanical equipment unusually operates under harsh environments of elevated temperature, cyclic loading, and even corrosion. Potential flaws could lead to catastrophic incidents with significant financial losses anddeaths; prognostics health management (PHM) has become indispensable to rotating machinery in advanced nuclear reactors. PHM, which is one of the crucial systems in a rotating machine, uses an intelligent diagnosis method as a crucial component to monitor and diagnose issues efficiently (<xref ref-type="bibr" rid="B7">Hamadache et al., 2019</xref>). Traditional diagnosis methods mainly apply signal processing methods to identify the features of the recorded signals and then infer possible faults from the identified features (<xref ref-type="bibr" rid="B13">Li et al., 2018</xref>; <xref ref-type="bibr" rid="B16">Sun et al., 2018</xref>; <xref ref-type="bibr" rid="B20">Zhao et al., 2018</xref>). However, signal processing and feature extraction are time consuming and empirical in nature because both operations depend on prior knowledge of the massive and heterogeneous data. Thus, the precise and efficient performance of the diagnosis model is still a challenging problem.</p>
<p>Artificial intelligence-based diagnosis models have recently become a suitable method in many areas, including computer vision (CV), natural language processing (NLP), state diagnosis, and other fields. Typical methods include artificial neural networks (ANN), k-nearest neighbors (KNN), naive Bayes, and deep learning (DL) (<xref ref-type="bibr" rid="B11">LeCun et al., 2015</xref>; <xref ref-type="bibr" rid="B14">Liu et al., 2018</xref>). Early fault detection techniques for gears, bearings, and rotors were categorized as signal processing techniques and AI-based techniques (<xref ref-type="bibr" rid="B17">Wei et al., 2019</xref>). Data-driven health monitoring by DL-based methods has been summarized and tested by <xref ref-type="bibr" rid="B19">Zhao et al. (2019)</xref> with regard to monitoring the operation condition of rotating machinery. <xref ref-type="bibr" rid="B4">Duan et al. (2018)</xref> reviewed the diagnosis and prognosis of mechanical equipment based on DL algorithms such as the dynamic Bayesian network (DBN) and convolutional neural network (CNN). <xref ref-type="bibr" rid="B12">Lei et al. (2020)</xref> performed a comprehensive review of developing intelligent diagnosis models by using machine learning methods. In addition; <xref ref-type="bibr" rid="B5">Ellefsen et al. (2019)</xref> introduced and reviewed four well-known DL algorithms&#x2014;Auto-encoder (AE), CNN, extended short-term memory network (LSTM), and DBN&#x2014;In practical PHM applications. For smart manufacturing and manufacturing diagnostics, AI-based methods and applications (smart sensors, intelligent manufacturing, PHM, and cyber-physical systems) were also reviewed by <xref ref-type="bibr" rid="B2">Chang et al. (2018)</xref>.</p>
<p>This article mainly focuses on the deep learning-based diagnosis models of the rotating machinery used in advanced nuclear reactors. An improved Transformer model is proposed, and an enhanced embeddings generator is designed with a CNN-based positional information extractor. A comparison and analysis of the different diagnosis models, including the MLP, CNN, RNN, RESNET, and Transformer, are performed in <xref ref-type="sec" rid="s2">Section 2</xref>. Because nuclear rotating mechanical equipment adapts the vibration signals to diagnose the current status, this article uses the typical healthy bearings datasets from the Case Western Reserve University (CWRU) to train and investigate DL-based diagnosis models. The datasets for the network training and the data preprocessing technique are also introduced in this section. Major findings are presented in <xref ref-type="sec" rid="s3">Section 3</xref>, and a conclusion is drawn in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
</sec>
<sec id="s2">
<title>2 Methods and data</title>
<p>Multiple deep learning algorithms are available for constructing diagnosis models. Their features are firstly introduced in <xref ref-type="sec" rid="s2-1">Section 2.1</xref>. The labeled datasets are crucial for training the neural networks, and the datasets used in the present study are described in <xref ref-type="sec" rid="s2-2">Section 2.2</xref>. The data processing techniques are introduced in <xref ref-type="sec" rid="s2-3">Section 2.3</xref>.</p>
<sec id="s2-1">
<title>2.1 Introduction of DL models</title>
<sec id="s2-1-1">
<title>2.1.1 Convolutional neural network</title>
<p>A type of feedforward neural network with deep structure and convolution computing is known as the &#x201c;convolutional neural network&#x201d; (CNN). It has been used extensively in image processing and natural language processing since it was initially presented in 1997 (<xref ref-type="bibr" rid="B10">LeCun and Bengio, 1995</xref>). The basic components of the CNN include convolutional layers, activation functions, pooling layers, and fully connected layers. This particular type of neural network has demonstrated dominance in image processing (AlexNet won the ImageNet competition in 2012) and classification (ResNet&#x2019;s accuracy surpassed that of humans in 2016) (<xref ref-type="bibr" rid="B8">He et al., 2016</xref>). The RESNET is a deep CNN architecture designed to address the issues of vanishing gradients and degradation in training deep networks. Residual blocks employ skip connections to directly pass the input signal to subsequent layers while learning the residual to represent the change in the network layers. This design allows the model to directly learn the residual, making it easier to optimize the network. In this study, we create five CNN layers for input data and also customize the well-known RESNET (ResNet18) for four different input data types. As a feed-forward neural network with convolution computation and deep design, the convolutional neural network (CNN) is frequently employed in image and natural language processing. A convolutional and pooling layer are included in each hidden layer of a CNN. The convolutional layer converts the local signal of the preceding layer to the ones of the following layer by employing a filter with common weights to derive characteristics from the input signal.</p>
<p>It is challenging to train CNN models with acceptable accuracy of fault diagnosis since the volume of the labeled samples used in fault diagnosis is comparably modest to that of the images in ImageNet. However, deep CNN models can perform well on small data (<xref ref-type="bibr" rid="B18">Yosinski et al., 2014</xref>) by combining with transfer learning (<xref ref-type="bibr" rid="B3">Donahue et al., 2014</xref>). Since the RESNET may boost accuracy with higher network depth, we transfer the RESNET trained on ImageNet to the fault diagnostic field in this study. The RESNET is naturally applied in the layers of feature extraction to applications for defect diagnosis since it performs well in picture classification and feature extraction.</p>
<p>We have presented a model (<xref ref-type="fig" rid="F1">Figure 1</xref>) that contains four convolutional layers, batch normalization layers, ReLU layers, and a pooling layer for the purpose of fault diagnostics, which is shown in <xref ref-type="table" rid="T1">Table 1</xref>. Moreover, <xref ref-type="table" rid="T2">Table 2</xref> presents the parameters of the RESNET model we used. A three-layered basic block structure with 64, 128, 256, 512, and N neurons, where N is the total number of defect categories, makes up the classification module. The retrieved characteristics are sent into the categorization module.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Diagramof the CNN model for fault diagnosis of rotating machine.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g001.tif"/>
</fig>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Parameters of the CNN model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No.</th>
<th align="left">Layers</th>
<th align="left">Hyperparameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Input signal</td>
<td align="left">Channel &#x3d; 1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Data preprocessing method</td>
<td align="left">Time Domain, Frequency Domain, Time&#x2013;Frequency Domain, or Wavelet Domain</td>
</tr>
<tr>
<td rowspan="6" align="left">3</td>
<td rowspan="6" align="left">CNN model</td>
<td align="left">Conv (16,15), BN, ReLU</td>
</tr>
<tr>
<td align="left">Conv (32,3), BN, ReLU, Max-pooling (2,2 &#xd7; 1)</td>
</tr>
<tr>
<td align="left">Conv (64,3), BN, ReLU</td>
</tr>
<tr>
<td align="left">Conv (128,3), BN, ReLU, Adaptive Max-pooling (4)</td>
</tr>
<tr>
<td align="left">Linear (128,256), ReLU</td>
</tr>
<tr>
<td align="left">Linear (256,64), ReLU</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Fault classifier</td>
<td align="left">Linear (class number)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Parameters of the RESNET model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No.</th>
<th align="left">Layers</th>
<th align="left">Hyperparameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Input signal</td>
<td align="left">Channel &#x3d; 1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Data preprocessing method</td>
<td align="left">Time Domain, Frequency Domain, Time-Frequency Domain, or Wavelet Domain</td>
</tr>
<tr>
<td rowspan="6" align="left">3</td>
<td rowspan="6" align="left">RESNET model</td>
<td align="left">Conv (64,7,2,3), BN, ReLU, Max-pooling (3,2 &#xd7; 1)</td>
</tr>
<tr>
<td align="left">Basic Block [Conv1x1 (64,1)]</td>
</tr>
<tr>
<td align="left">Basic Block [Conv1x1 (128,2)]</td>
</tr>
<tr>
<td align="left">Basic Block [Conv1x1 (256,2)]</td>
</tr>
<tr>
<td align="left">Basic Block [Conv1x1 (512,3)]</td>
</tr>
<tr>
<td align="left">Adaptive Max-pooling (1)</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Fault classifier</td>
<td align="left">Linear (class number)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-1-2">
<title>2.1.2 Short-term memory network</title>
<p>Due to the CNN&#x2019;s inability to obtain the long-term dependence characteristics of time series data and the complex features of non-linear data, the RNN is well suited for dealing with time series and has the ability to characterize temporal dynamic behavior. However, there are certain drawbacks to the RNN, such as the potential for gradient explosion or disappearance during backpropagation. For processing continuous input streams, LSTM was introduced in 1997 as a solution to these issues. Bidirectional LSTM (BiLSTM) is able to selectively recall and forget information and can capture bidirectional relationships across extended distances (<xref ref-type="bibr" rid="B9">Hochreiter and Schmidhuber, 1997</xref>). For the classification problem, we use BiLSTM to cope with two different forms of input data (<xref ref-type="table" rid="T3">Table 3</xref>). The RNN is constructed with three layers: the input layer, hidden layer, and output layer. Each layer&#x2019;s features are represented by the notations <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msubsup>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, where the superscript and subscript denote the sample number and time step, respectively. <italic>U</italic>, <italic>V</italic>, and <italic>W</italic> are the weight matrix. The forward propagation of a standard RNN is defined as follows:<disp-formula id="e1">
<mml:math id="m4">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>U</mml:mi>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>u</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m5">
<mml:mrow>
<mml:msubsup>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>v</mml:mi>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>v</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>b</italic>
<sub>
<italic>u</italic>
</sub> and <italic>b</italic>
<sub>
<italic>v</italic>
</sub> are the bias vectors, and <italic>f</italic>( ) is the non-linear activation function (<xref ref-type="bibr" rid="B6">Gers et al., 2000</xref>). The standard RNN is often limited by the long-term dependencies and becomes unable as the sequence grows (<xref ref-type="bibr" rid="B1">Bengio et al., 1994</xref>). A common solution is using the architecture of the LSTM network (<xref ref-type="fig" rid="F2">Figure 2</xref>).</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Parameters of the BiLSTM model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No.</th>
<th align="left">Layers</th>
<th align="left">Hyperparameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Input signal</td>
<td align="left">Channel &#x3d; 1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Data preprocessing method</td>
<td align="left">Time Domain, Frequency Domain, Time-Frequency Domain, or Wavelet Domain</td>
</tr>
<tr>
<td rowspan="2" align="left">3</td>
<td rowspan="2" align="left">Embeddings generator</td>
<td align="left">Conv (16,3,1), BN, ReLU, Max-pooling (2,2 &#xd7; 1)</td>
</tr>
<tr>
<td align="left">Conv (32,3,1), BN, ReLU, Adaptive Max-pooling (25)</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">BiLSTM model</td>
<td align="left">T &#x3d; 64, layers &#x3d; 2</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Fault classifier</td>
<td align="left">Linear (class number)</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>LSTM neuron internal structure.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g002.tif"/>
</fig>
<p>The enhanced classic RNN, or LSTM, can record the entire history of the input data. By combining the forgetting gates, output gates, and input gates, LSTM addresses these issues. The LSTM neuron&#x2019;s internal organization is depicted in <xref ref-type="fig" rid="F2">Figure 2</xref>. The primary concept is that a number of gates regulate how information flow along the time axis is updated. To decide whether or not the input <italic>x</italic>
<sub>
<italic>t</italic>
</sub> and hidden state of the preceding layer <italic>h</italic>(<italic>t&#x2212;1</italic>) should be added to the current cell, we introduce the input state. The forget gate controls whether or not the cell value should be preserved in relation to the current input and the prior concealed state. The current layer&#x2019;s output is based on both the current input and the preceding layer&#x2019;s output. A set of state responses that incorporate both the most recent input and output data will be produced by LSTM neurons. The memory cell makes sure that the gradient can be transferred to several stages without disappearing or exploding. The update of the input, forget, and output gates is listed as follows:<disp-formula id="e3">
<mml:math id="m6">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="italic">f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
<disp-formula id="e4">
<mml:math id="m7">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:msubsup>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m8">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:msubsup>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m9">
<mml:mrow>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>tan</mml:mi>
<mml:mspace width=".17em"/>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
<disp-formula id="e7">
<mml:math id="m10">
<mml:mrow>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x26ac;</mml:mo>
</mml:mrow>
<mml:mspace width="0.17em"/>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:msubsup>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x26ac;</mml:mo>
</mml:mrow>
<mml:mspace width="0.17em"/>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m11">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:msubsup>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="0.17em"/>
<mml:mo>&#x26ac;</mml:mo>
</mml:mrow>
<mml:mspace width=".17em"/>
<mml:mi>tan</mml:mi>
<mml:mspace width=".17em"/>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>W</italic> and <italic>V</italic> are the input and hidden state weights, respectively, and <italic>b</italic> stands for the biases. In the <italic>t</italic>-th update step, the input gate, forget gate, output gate, and cell state are updated by the input <italic>x</italic> and the hidden state of the (<italic>n-</italic>1)-th step. The hyperparameters of the BiLSTM model presents in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
</sec>
<sec id="s2-1-3">
<title>2.1.3 Multi-layer perceptron (MLP)</title>
<p>The MLP, which is a fully connected network with numerous hidden layers, was put forth as the ANN&#x2019;s model in 1987 (<xref ref-type="bibr" rid="B15">Rumelhart et al., 1985</xref>). With such a basic framework, the MLP is capable of performing some basic categorization tasks. However, when the task gets more difficult, the MLP is challenging to train due to the vast number of factors. For the one-dimensional (1D) input data in the current study, an MLP with five fully linked layers and five batch normalization layers is utilized (<xref ref-type="table" rid="T4">Table 4</xref>). CE loss refers to the softmax cross-entropy loss, BN refers to the batch normalization layer, and FC refers to the fully connected layer.</p>
<p>We use the CNN and MLP neural network topologies in this article. We will assess how well they perform in applications involving fault diagnostics and transfer learning. One-dimensional vibration signals are the type of data used in this work. In this case, a one-dimensional CNN structure is employed. Let<italic>x</italic> and <italic>y</italic> represent the neural networks&#x2019; input and output vectors, respectively. An input layer, an output layer, and a number of hidden layers make up the MLP. The letter <italic>h</italic>
<sub>
<italic>i</italic>
</sub> represents the hidden vector. Then, the forward propagation process of the MLP is as follows: where <bold>
<italic>b</italic>
</bold>
<sub>
<italic>i</italic>
</sub> is the bias vector, <bold>
<italic>w</italic>
</bold>
<sub>
<italic>i</italic>
</sub> is the weight matrix, <italic>j</italic> (&#xd7;) is the activation function, and <italic>softmax</italic>(&#xd7;) represents the softmax classifier. The forward propagation process of the CNN is as follows:<disp-formula id="e9">
<mml:math id="m12">
<mml:mrow>
<mml:mtable class="aligned" columnalign="left">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2a;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mo>&#x22ef;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>z</mml:mi>
<mml:mi>i</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2a;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>y</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where <italic>down</italic> (&#xd7;) is the pooling operation and &#x2a; stands for the convolution operation. For both the MLP and CNN, the loss function is a cross-entropy one as follows:<disp-formula id="e10">
<mml:math id="m13">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>where <bold>
<italic>y</italic>
</bold>
<sub>
<italic>i</italic>
</sub> is one of the labels, <inline-formula id="inf4">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is one of the predictions, <italic>n</italic> is the number of the training samples, and <italic>n</italic> is the number of the batch samples. The training method of both the MLP and CNN is back-propagation (BP) based on the gradient descent strategy. The parameters of the MLP model are shown in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Parameters of the MLP model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No.</th>
<th align="left">Layers</th>
<th align="left">Hyperparameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Input signal</td>
<td align="left">Channel &#x3d; 1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Data preprocessing method</td>
<td align="left">Time Domain, Frequency Domain, Time-Frequency Domain, or Wavelet Domain</td>
</tr>
<tr>
<td rowspan="5" align="left">3</td>
<td rowspan="5" align="left">BiLSTM model</td>
<td align="left">Linear (1024), BN, ReLU</td>
</tr>
<tr>
<td align="left">Linear (512), BN, ReLU</td>
</tr>
<tr>
<td align="left">Linear (256), BN, ReLU</td>
</tr>
<tr>
<td align="left">Linear (128), BN, ReLU</td>
</tr>
<tr>
<td align="left">Linear (64), BN, ReLU</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Fault classifier</td>
<td align="left">Linear (class number)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-1-4">
<title>2.1.4 Improved transformer</title>
<p>Transformer is a deep learning model widely used for natural language processing and other sequence processing tasks. It was developed by Google based on the encoder-decoder framework with an attention mechanism. The Transformer model can handle variable-length sequence data and has better parallelization capability and faster training speed compared to traditional recurrent neural networks (RNNs). The Transformer is composed of multiple stacked encoders and decoders, each consisting of multiple attention sub-layers and fully connected neural network sub-layers. During training, the Transformer uses the self-attention mechanism to capture the relationships between sequences, which effectively handles sequence data.</p>
<p>This study proposes a novel Transformer model for diagnosing faults in rotating machinery by integrating the strengths of the CNN and transformer. The architecture of the proposed model is presented in <xref ref-type="fig" rid="F3">Figure 3</xref> and consists of three major components: an enhanced embeddings generator, a transformer encoder, and a fault classifier. The parameters of the transformer architecture are listed in <xref ref-type="table" rid="T5">Table 5</xref>. The enhanced embeddings generator incorporates a CNN-based positional information extractor to convert the preprocessed 1-D data into a sequence of token embeddings. Since the input of the transformer is 1-D sequence tokens, a three-layer 1-D CNN architecture has been designed as an enhanced embeddings generator to ensure that the token embeddings have the same spatial arrangement information as the preprocessed data and that the transformer can access the inductive bias while having local feature extraction ability. The transformer encoder is connected after the input embeddings, which can reduce the model complexity to avoid gradient disappearance during the training process. Because the final output of the transformer is a sequence of the embeddings, an FC layer is used in the fault classifier.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Overall architecture of the improved Transformer model.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g003.tif"/>
</fig>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Parameters of the proposed Transformer model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No.</th>
<th align="left">Layers</th>
<th align="left">Hyperparameters</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Input signal</td>
<td align="left">Channel &#x3d; 1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Data preprocessing method</td>
<td align="left">Time Domain, Frequency Domain, Time&#x2013;Frequency Domain, or Wavelet Domain</td>
</tr>
<tr>
<td rowspan="3" align="left">3</td>
<td rowspan="3" align="left">Enhanced embeddings generator</td>
<td align="left">Conv (16,15,7), BN, ReLU, Max-pooling (3,2 &#xd7; 1)</td>
</tr>
<tr>
<td align="left">Conv (32,7,3), BN, ReLU, Max-pooling (3,2 &#xd7; 1)</td>
</tr>
<tr>
<td align="left">Conv (64,3,1), BN, ReLU</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Input embeddings</td>
<td align="left">Linear</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Transformer encoder</td>
<td align="left">T &#x3d; 32, Heads &#x3d; 8, drop rate &#x3d; 0.5</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Fault classifier</td>
<td align="left">Linear</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s2-2">
<title>2.2 Datasets</title>
<p>Many types of rotating mechanical equipment, such as the primary pump, turbine, and fans, are key components of Gen IV advanced reactors. Normally, industrial vibration detectors are used to monitor and diagnose the rotating machinery state. Although many different failure modes exist in the nuclear industry, the similar fault diagnosis process makes it feasible to investigate the DL-based diagnosis models. We use the typical healthy bearings datasets from CWRU as benchmark testing datasets in order to assess the efficacy of the suggested diagnosis models. Datasets from the Case Western Reserve University (CWRU) have been procured through the Bearing Data Center. Under four different motor loads, vibration signals were recorded at 12 or 48 kHz for healthy bearings and damaged bearings with single-point faults. Single-point defects were added for each operating condition on the rolling element, inner ring, and outer ring, with fault widths of 0.007, 0.014, and 0.021 inches, respectively. The data used in this study were gathered at the drive end, and the sampling frequency employed was 12 kHz.</p>
<p>As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the acceleration data were collected both close to and far from the motor bearings. Using electro-discharge machining, defects were introduced into the motor bearings (EDM). The inner raceway, rolling component (the ball), and outer raceway all experienced faults that ranged in dimension from 0.007 inches to 0.040 inches. The test motor&#x2019;s defective bearings were replaced, and vibration data were collected at loads ranging from 0 to 3 horsepower (the rotating speed ranges from 1797 to 1720 RPM).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Experiment platform for rolling bearing fault sensing.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g004.tif"/>
</fig>
</sec>
<sec id="s2-3">
<title>2.3 Data preprocessing</title>
<sec id="s2-3-1">
<title>2.3.1 Time domain</title>
<p>The time domain input denotes the usage of vibration signals without any prior processing as the input. Each sample in this study has a length of 1024. A total of 20% of all the samples are used as the testing set, while 80% of all samples are used as the training set.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Frequency domain</title>
<p>The term &#x201c;frequency domain&#x201d; refers to the process of employing the fast Fourier transform (FFT) to convert a sample into the frequency domain. The data length is cut in half as a result of this procedure, and the new sample is described as:<disp-formula id="e11B">
<mml:math id="m16">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>where the operator <bold>
<italic>FFT</italic>
</bold>(&#xb7;) represents transforming <bold>
<italic>x</italic>
</bold>
<sub>
<bold>
<italic>i</italic>
</bold>
</sub> into the frequency domain and taking the first half of the result.</p>
</sec>
<sec id="s2-3-3">
<title>2.3.3 Time-frequency domain</title>
<p>This kind of input is generated by applying the short-time Fourier transform (STFT) to the samples. The Hanning window is used, and the window length is 64. After this operation, the time&#x2013;frequency representation (a 33 &#xd7; 33 image) is generated as:<disp-formula id="e12">
<mml:math id="m17">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>T</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>where the operator <bold>
<italic>STFT</italic>
</bold>(&#xb7;) represents transforming <bold>
<italic>x</italic>
</bold>
<sub>
<bold>
<italic>i</italic>
</bold>
</sub> into the time&#x2013;frequency domain.</p>
</sec>
<sec id="s2-3-4">
<title>2.3.4 Wavelet domain</title>
<p>Continuous wavelet transform (CWT) is used to obtain the wavelet domain representation for the wavelet domain input, as illustrated in Eq. 16 The length of each sample is set at 100 because the CWT requires a lot of time. Following this procedure, the (a 100 &#xd7; 100 image) wavelet coefficients are produced as:<disp-formula id="e13">
<mml:math id="m18">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>W</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">C</mml:mi>
<mml:mi mathvariant="bold-italic">W</mml:mi>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>where the operator <bold>
<italic>CWT</italic>
</bold>(&#xb7;) represents transforming <bold>
<italic>x</italic>
</bold>
<sub>
<bold>
<italic>i</italic>
</bold>
</sub> into the wavelet domain.</p>
<p>The performance of the DL models is significantly influenced by the type of input data and method of normalization. The complexity of feature extraction is determined by the types of input data, and the difficulty of calculation is determined by normalization techniques and evaluation. This article&#x2019;s normalizing technique can be used by:<disp-formula id="e14">
<mml:math id="m19">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>z</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>min</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
</sec>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and discussion</title>
<p>The bearing dataset from the CWRU&#x2019;s bearing data center have been compared in this study. Four groups of data from the normal bearing, the bearing with the fault on the inner race, the bearing with the defect on the outer race, and the bearing with the fault on the ball make up the dataset. In these tests, four spinning speeds are used, namely, 1797 rpm, 1772 rpm, 1750 rpm, and 1730 rpm. <xref ref-type="table" rid="T6">Table 6</xref> contains a list of the label data.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Label of the modes of fault.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Label</th>
<th align="center">Fault mode</th>
<th align="center">Description</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">1</td>
<td align="center">Health State</td>
<td align="center">The normal bearing at 1791 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">2</td>
<td align="center">Inner ring 1</td>
<td align="center">0.007-inch inner ring fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">3</td>
<td align="center">Inner ring 2</td>
<td align="center">0.014-inch inner ring fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">4</td>
<td align="center">Inner ring 3</td>
<td align="center">0.021-inch inner ring fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">5</td>
<td align="center">Rolling Element 1</td>
<td align="center">0.007-inch rolling element fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">6</td>
<td align="center">Rolling Element 2</td>
<td align="center">0.014-inch rolling element fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">7</td>
<td align="center">Rolling Element 3</td>
<td align="center">0.021-inch rolling element fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">8</td>
<td align="center">Outer ring 1</td>
<td align="center">0.007-inch outer ring fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">9</td>
<td align="center">Outer ring 2</td>
<td align="center">0.014-inch outer ring fault at 1797 rpm and 0 HP</td>
</tr>
<tr>
<td align="center">10</td>
<td align="center">Outer ring 3</td>
<td align="center">0.021-inch outer ring fault at 1797 rpm and 0 HP</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="F4">Figure 4</xref> displays the bearing&#x2019;s vibration signals at a speed of 1797 revolutions per minute. The captured signals were divided into 100 samples, each of which has 1024 points. The testing set and training set can be selected at random. We used 5, 10, 15, and 20 examples of each fault mode for training and the remaining 40 samples for testing in order to compare the performances of the classifiers with various sizes of the training set.</p>
<p>The typical raw vibration signals for ten modes are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. It is apparent that it is challenging to determine the kind of failure from the raw data without data processing. The time-frequency spectrums with the time-frequency domain and wavelet domain for the ten modes are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. The proposed diagnosis models have been evaluated against these spectrums.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Vibration signals in time domain (left) and frequency domain (right) of the 10 fault modes.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g005.tif"/>
</fig>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The time-frequency spectrum with time-frequency domain (left) and wavelet domain (right) of the 10 fault modes.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g006.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, the accuracies of prediction of the CNN, BiLSTM, MPL, RESNET, and Transformer models with different data preprocessing methods are compared. A total of 10 samples are chosen to train the diagnosis models. The CNN, BiLSTM, RESNET, and Transformer models can reach a stable and accurate state after 25 epochs without overfitting. The four data processing techniques are suitable for the three diagnosis models. However, the frequency domain input and time&#x2013;frequency domain input are appropriate for the MPL as they can achieve the precise model The MPL with the time domain input and wavelet domain input are unable to distinguish the fault modes. In <xref ref-type="fig" rid="F7">Figure 7</xref>, the CNN, BiLSTM, and RESNET have a good accuracy with the time domain and wavelet domain. However, the MLP is unavailable with these two data preprocessing methods. The frequency domain input and time&#x2013;frequency domain are compatible with these diagnosis models.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The accuracy of the CNN, BiLSTM, MLP, RESNET, and Transformer with epoch number.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g007.tif"/>
</fig>
<p>In addition to data preprocessing, the train samples number has a great effect on the performance of the diagnosis models (<xref ref-type="fig" rid="F8">Figure 8</xref>). The CNN, BiLSTM, and RESNET methods have a good accuracy with the four data preprocessing methods, with 10 samples for each fault mode. However, the time domain and wavelet domain method show the shortage for the MPL method. The time-frequency domain and frequency method are available for the MPL method, with 5 samples for each fault mode.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The accuracy with the CNN, BiLSTM, MLP, RESNET, and Transformer models with different training sample numbers under different data preprocessing methods: <bold>(A)</bold> time domain; <bold>(B)</bold> frequency domain; <bold>(C)</bold> time-frequency domain; and <bold>(D)</bold> wavelet domain.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g008.tif"/>
</fig>
<p>As it is known, signal noise also has great influence on diagnostic accuracy. To test the robustness of the proposed models, white Gaussian noise is added to the training and the testing samples with different signal-to-noise ratios (SNRs). The trend of the different diagnosis models accuracies is shown in <xref ref-type="fig" rid="F9">Figure 9</xref>, with different SNRs using the time-frequency domain input data. For the BiLSTM, MLP, RESNET, and Transformer, the satisfied models are obtained when the SNR is more than 10 dB. The BiSLTM and Transformer show strong anti-noise ability and performs significantly well. Meanwhile, the other models show a shortage in this range of SNR due to overlooking long-term dependencies in the time-series data. BiLSTM models can handle long-term dependencies in time-series data and capture local features in the sequence. Transformer models are based on self-attention mechanisms that effectively handle long-term dependencies between sequence data, which can cause the robustness of the diagnosis model in case of signal noise.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>The accuracy of the CNN, BiLSTM, MLP, RESNET, and Transformer models with different SNRs.</p>
</caption>
<graphic xlink:href="fenrg-11-1210703-g009.tif"/>
</fig>
<p>The key conditions for deep learning diagnosis model processing are sufficient and representative data, proper data preprocessing, and hyperparameter tuning. Deep learning models require sufficient high-quality data to learn and generalize effectively. The dataset should be diverse, balanced, and representative of the various fault types. Sufficient data ensure that the model can capture complex patterns and make accurate predictions. Data preprocessing plays a crucial role in deep learning. It involves tasks such as cleaning, normalization, feature scaling, and handling missing data. Preprocessing ensures that the data are in a suitable format for the model and removes any noise or biases that could hinder learning. Deep learning models have various hyperparameters, such as drop rate, layer number, and regularization parameters, that need to be tuned. Proper hyperparameter tuning can significantly impact the model&#x2019;s performance, convergence speed, and generalization ability.</p>
</sec>
<sec sec-type="conclusion" id="s4">
<title>4 Conclusion</title>
<p>In the present study, we propose the five DL-based diagnosis models, i.e., the MLP, CNN, RNN, RESNET, and improved Transformer model. The performance of the five models has been evaluated and compared. It is shown that the CNN, BiLSTM, and RESNET can achieve good accuracy of prediction with four types of data preprocessing, even with only 10 samples for training. The MLP method can also yield satisfactory accuracy if the input is provided in the time&#x2013;frequency domain. The improved Transformer model with the input provided in the time domain and frequency domain is most suitable for the fault diagnosis of rotating machinery.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary materials, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>YS contributed to the conception and design of the study. HW established the fault diagnosis model, organized the database, carried out statistical analysis, and wrote the first draft. YS wrote sections of the manuscript. HW partially revised the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<ack>
<p>Thanks are extended to Tsinghua University and Harbin Engineering University for their help in this research.</p>
</ack>
<sec sec-type="COI-statement" id="s7">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Simard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Frasconi</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>Learning long-term dependencies with gradient descent is difficult</article-title>. <source>IEEE Trans. neural Netw.</source>
<volume>5</volume>, <fpage>157</fpage>&#x2013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1109/72.279181</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chang</surname>
<given-names>C.-W.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>H.-W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C.-H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A review of artificial intelligence algorithms used for smart machine tools</article-title>. <source>Inventions</source>
<volume>3</volume>, <fpage>41</fpage>. <pub-id pub-id-type="doi">10.3390/inventions3030041</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Donahue</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vinyals</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Hoffman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Tzeng</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <source>Decaf: A deep convolutional activation feature for generic visual recognition</source>. <publisher-loc>PMLR</publisher-loc>: <publisher-name>International conference on machine learning</publisher-name>,<fpage>647</fpage>&#x2013;<lpage>655</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep learning enabled intelligent fault diagnosis: Overview and applications</article-title>. <source>J. Intelligent Fuzzy Syst.</source>
<volume>35</volume>, <fpage>5771</fpage>&#x2013;<lpage>5784</lpage>. <pub-id pub-id-type="doi">10.3233/jifs-17938</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ellefsen</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>&#xc6;s&#xf8;y</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ushakov</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A comprehensive survey of prognostics and health management based on deep learning for autonomous ships</article-title>. <source>IEEE Trans. Reliab.</source>
<volume>68</volume>, <fpage>720</fpage>&#x2013;<lpage>740</lpage>. <pub-id pub-id-type="doi">10.1109/tr.2019.2907402</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gers</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cummins</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Learning to forget: Continual prediction with LSTM</article-title>. <source>Neural Comput.</source>
<volume>12</volume>, <fpage>2451</fpage>&#x2013;<lpage>2471</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300015015</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hamadache</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jung</surname>
<given-names>J. H.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Youn</surname>
<given-names>B. D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A comprehensive review of artificial intelligence-based approaches for rolling element bearing PHM: Shallow and deep learning</article-title>. <source>JMST Adv.</source>
<volume>1</volume>, <fpage>125</fpage>&#x2013;<lpage>151</lpage>. <pub-id pub-id-type="doi">10.1007/s42791-019-0016-y</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. <source>Proc. IEEE Conf. Comput. Vis. pattern Recognit.</source>, <fpage>770</fpage>&#x2013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Long short-term memory Neural computation</article-title>. <volume>9</volume>, <source>Nov</source>, <volume>15</volume>, <fpage>8</fpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-24797-2</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LeCun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Convolutional networks for images, speech, and time series</article-title>. <source>Handb. Brain theory neural Netw.</source>
<volume>3361</volume>, <fpage>1995</fpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LeCun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>nature</source>
<volume>521</volume>, <fpage>436</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lei</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Nandi</surname>
<given-names>A. K.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Applications of machine learning to machine fault diagnosis: A review and roadmap</article-title>. <source>Mech. Syst. Signal Process.</source>
<volume>138</volume>, <fpage>106587</fpage>. <pub-id pub-id-type="doi">10.1016/j.ymssp.2019.106587</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>De Oliveira</surname>
<given-names>J. V.</given-names>
</name>
<name>
<surname>Cerrada</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cabrera</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>S&#xe1;nchez</surname>
<given-names>R. V.</given-names>
</name>
<name>
<surname>Zurita</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A systematic review of fuzzy formalisms for bearing fault diagnosis</article-title>. <source>IEEE Trans. Fuzzy Syst.</source>
<volume>27</volume>, <fpage>1362</fpage>&#x2013;<lpage>1382</lpage>. <pub-id pub-id-type="doi">10.1109/tfuzz.2018.2878200</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zio</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Artificial intelligence for fault diagnosis of rotating machinery: A review</article-title>. <source>Mech. Syst. Signal Process.</source>
<volume>108</volume>, <fpage>33</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1016/j.ymssp.2018.02.016</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rumelhart</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>R. J.</given-names>
</name>
</person-group> (<year>1985</year>). <source>Learning internal representations by error propagation</source>. <publisher-loc>USA</publisher-loc>: <publisher-name>California Univ San Diego La Jolla Inst for Cognitive Science</publisher-name>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Sparse deep stacking network for fault diagnosis of motor</article-title>. <source>IEEE Trans. Industrial Inf.</source>
<volume>14</volume>, <fpage>3261</fpage>&#x2013;<lpage>3270</lpage>. <pub-id pub-id-type="doi">10.1109/tii.2018.2819674</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A review of early fault diagnosis approaches and their applications in rotating machinery</article-title>. <source>Entropy</source>
<volume>21</volume>, <fpage>409</fpage>. <pub-id pub-id-type="doi">10.3390/e21040409</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yosinski</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Clune</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lipson</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>How transferable are features in deep neural networks?</article-title>
<source>Adv. neural Inf. Process. Syst.</source>
<volume>27</volume>, <fpage>1</fpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Mao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>R. X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Deep learning and its applications to machine health monitoring</article-title>. <source>Mech. Syst. Signal Process.</source>
<volume>115</volume>, <fpage>213</fpage>&#x2013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1016/j.ymssp.2018.05.050</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Qiao</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Enhanced sparse period-group lasso for bearing fault diagnosis</article-title>. <source>IEEE Trans. Industrial Electron.</source>
<volume>66</volume>, <fpage>2143</fpage>&#x2013;<lpage>2153</lpage>. <pub-id pub-id-type="doi">10.1109/tie.2018.2838070</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>