<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2023.1096083</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A multi-scale pooling convolutional neural network for accurate steel surface defects classification</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Fu</surname> <given-names>Guizhong</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2035115/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Zengguang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Le</surname> <given-names>Wenwu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2112842/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Jinbin</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhu</surname> <given-names>Qixin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Niu</surname> <given-names>Fuzhou</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1379363/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Hao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1488592/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sun</surname> <given-names>Fangyuan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Shen</surname> <given-names>Yehu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1488604/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Mechanical Engineering, Suzhou University of Science and Technology</institution>, <addr-line>Suzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>College of Mechanical and Electrical Engineering, Shihezi University</institution>, <addr-line>Shihezi</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Feihu Zhang, Northwestern Polytechnical University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Horacio Rostro Gonzalez, University of Guanajuato, Mexico; Shuqiang Wang, Shenzhen Institutes of Advanced Technology (CAS), China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Guizhong Fu &#x02709; <email>fuguizhongchina&#x00040;163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1096083</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Fu, Zhang, Le, Li, Zhu, Niu, Chen, Sun and Shen.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Fu, Zhang, Le, Li, Zhu, Niu, Chen, Sun and Shen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Surface defect detection is an important technique to realize product quality inspection. In this study, we develop an innovative multi-scale pooling convolutional neural network to accomplish high-accuracy steel surface defect classification. The model was built based on SqueezeNet, and experiments were carried out on the NEU noise-free and noisy testing set. Class activation map visualization proves that the multi-scale pooling model can accurately capture the defect location at multiple scales, and the defect feature information at different scales can complement and reinforce each other to obtain more robust results. Through T-SNE visualization analysis, it is found that the classification results of this model have large inter-class distance and small intra-class distance, indicating that this model has high reliability and strong generalization ability. In addition, the model is small in size (3MB) and runs at up to 130FPS on an NVIDIA 1080Ti GPU, making it suitable for applications with high real-time requirements.</p></abstract>
<kwd-group>
<kwd>convolutional neural network</kwd>
<kwd>multi-scale</kwd>
<kwd>defect classification</kwd>
<kwd>class activation map</kwd>
<kwd>feature visualization</kwd>
</kwd-group>
<contract-num rid="cn001">51875380</contract-num>
<contract-num rid="cn001">51975394</contract-num>
<contract-num rid="cn001">52105526</contract-num>
<contract-num rid="cn001">61903269</contract-num>
<contract-num rid="cn002">2020M671383</contract-num>
<contract-num rid="cn002">2020M681517</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<contract-sponsor id="cn002">China Postdoctoral Science Foundation<named-content content-type="fundref-id">10.13039/501100002858</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="7"/>
<equation-count count="12"/>
<ref-count count="39"/>
<page-count count="12"/>
<word-count count="8319"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Surface defect detection is one of the most important processes that affects the quality of the products (Ravikumar et al., <xref ref-type="bibr" rid="B28">2011</xref>). Some surface defects will not only affect the appearance of the product surface but also endanger the user&#x00027;s property and life safety of users. In the beginning, surface defect detection is realized by manual inspection, hindering the improvement of productivity. It is vital to develop competent defect detection systems to replace manual work and satisfy the growing demands for automated inspection in the manufacturing sector (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>; Neogi et al., <xref ref-type="bibr" rid="B27">2014</xref>).</p>
<p>With the development of machine vision technology, defect detection task has attracted extensive attention from researchers in the industry. A typical visual inspection system includes hardware and defects identification algorithms. These algorithms use different kinds of approaches to implement defect detection, template-based (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>), morphological filter (Mak et al., <xref ref-type="bibr" rid="B24">2009</xref>), Fourier transforms (Zori&#x00107; et al., <xref ref-type="bibr" rid="B39">2022</xref>), Gabor filters (Bissi et al., <xref ref-type="bibr" rid="B3">2013</xref>), wavelet (Li and Tsai, <xref ref-type="bibr" rid="B19">2012</xref>; Li et al., <xref ref-type="bibr" rid="B18">2015</xref>), Markov random field (Dogand&#x0017E;i&#x00107; et al., <xref ref-type="bibr" rid="B6">2005</xref>), sparse dictionary reconstruction (Kang and Zhang, <xref ref-type="bibr" rid="B14">2020</xref>), decision tree (Aghdam et al., <xref ref-type="bibr" rid="B1">2012</xref>), random forest (Zhang et al., <xref ref-type="bibr" rid="B37">2019</xref>), and support vector machines (Chu et al., <xref ref-type="bibr" rid="B5">2017</xref>). These methods achieve the representation of defect features through manually designed feature extractors, which are highly subjective, and their defect recognition performance is affected by the designer. The detection performance will be somewhat compromised as the defect&#x00027;s morphology changes and the generalization ability of these detection approaches are limited.</p>
<p>In recent years, the theory of artificial neural networks and graphic processing unit (GPU) has been developed rapidly. Convolutional neural network (CNN) brings a new solution to vision-based tasks such as object classification and detection. LeNet is proposed by LeCun et al. (<xref ref-type="bibr" rid="B17">1998</xref>), it is the first CNN that can be applied to handwriting character recognition on letters. AlexNet is the next influential CNN model, it is the first CNN deployed on GPU, which can greatly improve the speed of training and testing, and provides a research basis for the subsequent extensive application of the CNN model (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>). The next representative model is VGG (Simonyan and Zisserman, <xref ref-type="bibr" rid="B30">2015</xref>), it incorporates 3 &#x000D7; 3 kernel to reduce parameters, and it is deeper and better than the previous model. The VGG16 model achieves 92.7% top-five test accuracy in ImageNet. He Kaiming et al. proposed ResNet (He et al., <xref ref-type="bibr" rid="B10">2016</xref>), a very deep neural network with hundreds of layers, and skip connections are used to jump over some layers to enhance gradient backpropagation and restrain gradient vanish. Afore-mentioned models tend to use more stacked layers to obtain higher classification accuracy, which leads to an increasing number of network parameters, reducing the computation efficiency. To solve this problem, some researchers have proposed a CNN model with fewer parameters and high accuracy. Iandola et al. proposed SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>), which incorporates 1 &#x000D7; 1 and 3 &#x000D7; 3 kernels to build the model, it is about 1/50 parameters of AlexNet. Howard et al. proposed a lightweight deep neural network-MobileNet (Howard et al., <xref ref-type="bibr" rid="B11">2017</xref>) which could be applied to mobile and embedded vision applications.</p>
<p>With the emergence of constantly updated image classification models, CNN has been applied to various fields, such as medical image processing (Wang et al., <xref ref-type="bibr" rid="B33">2018</xref>, <xref ref-type="bibr" rid="B34">2020</xref>). Surface defect classification of industrial products is also an important application of CNN. Khumaidi et al. proposed a CNN model to obtain welding defect classification (Khumaidi et al., <xref ref-type="bibr" rid="B15">2017</xref>). Li et al. proposed an end-to-end surface defects recognition system that incorporates a defect saliency map and convolutional neural network (Li et al., <xref ref-type="bibr" rid="B20">2016</xref>). Fu et al. proposed a deep-learning-based model, which emphasizes the training of low-level features and incorporates multiple receptive fields (Fu et al., <xref ref-type="bibr" rid="B7">2019</xref>). Ren et al. presented a CNN model to perform surface defects inspection task, and feature transferring from pre-trained models is used in the model (Ren et al., <xref ref-type="bibr" rid="B29">2018</xref>). At present, in most of the research articles on defect detection, the feature distribution of the testing set and training set is relatively consistent, which cannot effectively test the generalization ability of the model. In the actual model deployment process, there are differences between the collected images and testing set, and the generalization ability of the model should also be considered an important aspect of model performance evaluation. In addition, the contents of current defect detection research papers mainly focus on the accuracy and performance comparison, and there are few defects features and rules studied by researchers, which are insufficient to establish clear corresponding relationship between the internal features of neural networks and defect detection tasks.</p>
<p>To settle the two problems, we propose a lightweight CNN model to achieve precise and efficient steel surface defect classification. Our CNN model is constructed on the SqueezeNet pre-trained model, an innovative multi-scale pooling (MSP) module which is proposed to learn semantic features at different scales. These features pooled on the three dimensions have been jointly considered to predict an optimal classification result. To explore the hyperparameter characteristics and their distribution rules obtained through training the model, the class activation map is used to visualize the activated features. Furthermore, to analyze the distribution information of the high-dimensional features in the CNN model, T-distribution and stochastic neighbor embedding (T-SNE) are used to reduce the data dimension to obtain a more comprehensible class-specific feature distribution rule.</p>
<p>Our proposed model runs about 130 FPS on a single NVIDIA 1080Ti GPU (12G memory), which can meet the needs of manufacturing enterprises for defect identification efficiency. Overall, the main contributions of this study are summarized as follows:</p>
<list list-type="bullet">
<list-item><p>We propose a lightweight CNN model, which is constructed based on SqueezeNet. An architecture that integrates class-specific defect cues after pooling rich features at multiple scales is proposed.</p></list-item>
<list-item><p>The proposed model is initialized using the SqueezeNet pre-trained model, it is then fine-tuned by transferring learning across the NEU dataset. The model is tested on the noise dataset to confirm the generalization ability of the proposed model.</p></list-item>
<list-item><p>The distribution characteristic of the defect at multi-scale learned by the neural network is analyzed using the class activation map, which sheds insight into the position of the defect features that are crucial for classifying the defect type. T-SNE is used to analyze the classification feature vectors in neural networks to further demonstrate the generalization ability of the proposed model.</p></list-item>
</list>
<p>The rest of the article is organized as follows. In Section 2, the SqueezeNet model and an optimization module are introduced. The performance and experimental comparisons of our proposed model are presented in Section 3. Finally, Section 4 concludes the article.</p></sec>
<sec id="s2">
<title>2. Proposed method</title>
<p>In this section, the details of pre-trained model and a multi-scale pooling module are presented. The complete structure of the proposed model is shown in <bold>Figure 2</bold>.</p>
<sec>
<title>2.1. SqueezeNet-based defect classification</title>
<p>In the past decade, much of the research on convolution neural networks (CNNs) has focused on image classification task. Researchers have found that deepening the network depth can effectively improve the classification accuracy (He et al., <xref ref-type="bibr" rid="B10">2016</xref>), as a result, increasing the CNN depth has become an important technique to improve performance. With the improvement of the network depth, the parameters in the network and the computational burden are also increasing. However, some industrial product inspection task requires real-time speed, so it is extremely important to develop a model that can accurately and quickly identify the image. The lightweight SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>) can achieve high accuracy without loss of efficiency. The whole structure schematic diagram of the SqueezeNet model is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The architecture of the pre-trained SqueezeNet model (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0001.tif"/>
</fig>
<p>In previous models, such as AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>) and VGG (Simonyan and Zisserman, <xref ref-type="bibr" rid="B30">2015</xref>), the convolutional layers are constructed entirely using 3 &#x000D7; 3 filters. Compared to AlexNet and VGG, SqueezeNet has two advantages (1) Replace of 3 &#x000D7; 3 filters with 1 &#x000D7; 1 filters, (2) Decrease the number of input channels of 3 &#x000D7; 3 filters.</p>
<p>We plan to build our model based on the SqueezeNet pre-trained model. Pre-trained model is helpful for improving model performance in machine/computer vision-related tasks. Traditional non-deep-learning-based (machine learning-based or statistical) methods use hand-crafted features (Mak et al., <xref ref-type="bibr" rid="B24">2009</xref>; Bissi et al., <xref ref-type="bibr" rid="B3">2013</xref>; Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>; Zori&#x00107; et al., <xref ref-type="bibr" rid="B39">2022</xref>). Researchers can observe the dominant feature regularity in a small amount of data to design feature extractors. However, for deep-learning-based models, the model learns the characteristic distribution of the features from a large amount of data. In general, the higher the complexity of the model, the more data are required. Inadequate data may lead to problems such as overfitting and weak generalization ability (Belkin et al., <xref ref-type="bibr" rid="B2">2019</xref>; Xu et al., <xref ref-type="bibr" rid="B35">2020</xref>). In the industrial defect inspection task, it is difficult to obtain a large number of images for there are limited specimens. However, transfer learning provides a feasible scheme to alleviate this problem, where a model is trained on ImageNet at first and then fine-tuned on the target dataset. The effectiveness of transfer learning methods based on pre-trained models has been demonstrated on a large number of machine/computer vision-related tasks (Lu et al., <xref ref-type="bibr" rid="B23">2020</xref>; Bouaafia et al., <xref ref-type="bibr" rid="B4">2021</xref>). The surface defect inspection task is an important application of machine/computer vision in the industrial field, so it is reasonable and effective to apply the pre-trained model to our task. In addition, the pre-trained models like VGG, ResNet, and SqueezeNet are publicly available (Simonyan and Zisserman, <xref ref-type="bibr" rid="B30">2015</xref>; He et al., <xref ref-type="bibr" rid="B10">2016</xref>). There is no need for the researcher to train the model on ImageNet again. Therefore, we build the proposed convolutional neural network based on the SqueezeNet pre-training model.</p>
<p>There are two individual convolutional layers ( Conv1 and Conv10), three pooling layers, and nine fire modules. Each fire module is comprised of a squeeze layer (namely Fire <italic>i</italic>/ Squeeze) and two expand layers (namely Fire <italic>i</italic>/ Expand 1 &#x000D7; 1 and Fire <italic>i</italic>/ Expand 3 &#x000D7; 3), among which <italic>i</italic> is the sequence number of fire module. The detailed configurations of nine fire modules and pooling layers are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Detailed parameter configurations of layers in the SqueezeNet model.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>Layer name</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold><italic>i</italic></bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold><italic>S</italic><sub><italic>i</italic></sub></bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold><italic>E</italic><sub><italic>i</italic></sub></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Fire 2</td>
<td valign="top" align="center"><italic>i</italic> = 2</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">64</td>
</tr> <tr>
<td valign="top" align="left">Fire 3</td>
<td valign="top" align="center"><italic>i</italic> = 3</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">64</td>
</tr> <tr>
<td valign="top" align="left">Fire 4</td>
<td valign="top" align="center"><italic>i</italic> = 4</td>
<td valign="top" align="center">32</td>
<td valign="top" align="center">128</td>
</tr> <tr>
<td valign="top" align="left">Pool 4</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">Fire 5</td>
<td valign="top" align="center"><italic>i</italic> = 5</td>
<td valign="top" align="center">32</td>
<td valign="top" align="center">128</td>
</tr> <tr>
<td valign="top" align="left">Fire 6</td>
<td valign="top" align="center"><italic>i</italic> = 6</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">192</td>
</tr> <tr>
<td valign="top" align="left">Fire 7</td>
<td valign="top" align="center"><italic>i</italic> = 7</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">192</td>
</tr> <tr>
<td valign="top" align="left">Fire 8</td>
<td valign="top" align="center"><italic>i</italic> = 8</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">256</td>
</tr> <tr>
<td valign="top" align="left">Pool 8</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">Fire 9</td>
<td valign="top" align="center"><italic>i</italic> = 9</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">256</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The stride of pooling layers Pool 1, Pool 4, and Pool 8 is set to 2. It can be observed that the third pooling layer, Pool 8, appeared in the deep location of SqueezeNet, ensuring the feature maps (Fire 5 &#x0007E; Fire 8) have high resolution. In each fire module, the channel number of the squeeze layer is <italic>S</italic><sub><italic>i</italic></sub>, expand layer <italic>E</italic><sub><italic>i</italic></sub>, <italic>S</italic><sub><italic>i</italic></sub> is set to a quarter of <italic>E</italic><sub><italic>i</italic></sub> in all fire modules to reduce the number of input channels to 3 &#x000D7; 3 expand layer. The pre-trained model is trained in ImageNet; <italic>n</italic> is the channel number of &#x0201C;Conv 10&#x0201D; as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, which is set to 1,000 originally. To apply SqueezeNet to our task, <italic>n</italic> of &#x0201C;Conv 10&#x0201D; is modified to 6 because there are six types of defects in the NEU datasets (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>). A global average pooling (GAP) layer &#x0201C;Pool 10&#x0201D; is then connected to Conv10. GAP computes the average value of the input feature map. GAP could partially retain the spatial structure information of the input feature. Furthermore, compared to the traditional fully connected layer, the GAP does not generate additional parameters thus reducing the computational burden. The final layer is the Softmax loss layer, which is used to compute the cross-entropy between network output and label. Specifically, the loss function is defined as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>n</italic> &#x0003D; 6, <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> when the label of input image is <italic>i</italic>, otherwise <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. <italic>f</italic>(<italic>z</italic><sub><italic>i</italic></sub>) is the confidence score, which is calculate by softmax, defined as</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>z</italic><sub><italic>i</italic></sub> is the number <italic>i</italic> output value of &#x0201C;Pool 10.&#x0201D;</p>
</sec>
<sec>
<title>2.2. Model Optimization</title>
<p>Based on the constructed Squeezenet model, this study proposes a multi-scale pooling (MSP) and multi-scale feature fusion structure to improve the accuracy of defect classification. The whole architecture of the proposed model is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. All the activation functions used in our model are Rectified Linear Unit (Relu). All the activation functions are omitted for brevity, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The architecture of the proposed multi-scale pooling convolutional neural network.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0002.tif"/>
</fig>
<p>The layers from Conv1 to fire 9 in the SqueezeNet model are used as a feature extractor to extract defect features of an input image. The input image size is 256 &#x000D7; 256. After feature extraction based on SqueezeNet, the output feature dimension is <italic>n</italic>&#x000D7;16 &#x000D7; 16, where <italic>n</italic> is as defined earlier. In the multi-scale pooling module, three different convolutional layers are connected to the feature extractor separately. The detailed parameters of three convolutional layers: (1) Conv11-1 uses <italic>k</italic><sub>1</sub>&#x000D7;1 &#x000D7; 1 filter, stride &#x0003D; 1; (2) Conv11-2 uses <italic>k</italic><sub>2</sub>&#x000D7;1 &#x000D7; 1 filter, stride &#x0003D; 2; (3) Conv11-3 uses <italic>k</italic><sub>3</sub>&#x000D7;1 &#x000D7; 1 filter, stride &#x0003D; 4. After convoluted by Conv11, the dimension of output feature maps at three scales are <italic>k</italic><sub>1</sub>&#x000D7;16 &#x000D7; 16, <italic>k</italic><sub>2</sub>&#x000D7;8 &#x000D7; 8, and <italic>k</italic><sub>3</sub>&#x000D7;4 &#x000D7; 4. We set <italic>k</italic><sub>1</sub> &#x0003D; 6, <italic>k</italic><sub>2</sub> &#x0003D; 6, and <italic>k</italic><sub>3</sub> &#x0003D; 6 for six classes of defect, which is helpful to promote the convolutional layer <italic>Conv</italic>11 to learn better class-specific features. Considering there are different types of defects, the defects differ in pattern size and morphology. The proposed multi-scale module in our model could effectively capture semantic defect features at different scales. The detailed visual comparison of the defect recognition effect of the model at different scales will be shown in Section 3.5. The afore-mentioned process is described as</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mi>x</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>x</italic> is the output of SqueezeNet feature extractor, <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M38"><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> the weight and bias of the <italic>m</italic><sub><italic>i</italic></sub>(<italic>j</italic>)<italic>th</italic> channel filter in Conv11&#x02212;<italic>j</italic> layer (<italic>j</italic> is the scale number, and <italic>j</italic> &#x0003D; 1, 2, 3), <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <italic>m</italic>(<italic>j</italic>)<italic>th</italic> channel filter output after processed by Conv11&#x02212;<italic>i</italic>, <italic>m</italic><sub><italic>i</italic></sub>(<italic>j</italic>) &#x0003D; 1&#x0007E;<italic>k</italic><sub><italic>j</italic></sub>.</p>
<p>Global average pooling layer calculate the average value of input tensors (<inline-formula><mml:math id="M9"><mml:msup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>) across <italic>k</italic><sub><italic>j</italic></sub> channels, each channel generates a class-related value (Lin et al., <xref ref-type="bibr" rid="B21">2013</xref>). The operation is described as</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02208;</mml:mo><mml:mi>R</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>p</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <italic>m</italic><sub><italic>i</italic></sub>(<italic>j</italic>)<italic>th</italic> feature map output value in scale <italic>j</italic>, <italic>R</italic> is the element total number of <italic>m</italic><sub><italic>i</italic></sub>(<italic>j</italic>)<italic>th</italic> feature map, and <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>p</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the element value at (<italic>p, q</italic>) in region <italic>R</italic>.</p>
<p>In the multi-scale feature fusion stage, the output of three global average pooling layers are stacked together by a concat layer, which defined as</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>Y</italic><sub><italic>j</italic></sub> is the full set of <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, defined as <inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, <italic>f</italic><sup><italic>cat</italic></sup> represents the concatenation operation stacking pooled values <italic>Y</italic><sub>1</sub>, <italic>Y</italic><sub>2</sub>, and <italic>Y</italic><sub>3</sub> together. <italic>F</italic><sub><italic>y</italic></sub> is the fused tensor, due to <italic>m</italic><sub><italic>i</italic></sub>(<italic>j</italic>) &#x0003D; 1&#x0007E;<italic>k</italic><sub><italic>j</italic></sub>, total length of <italic>F</italic><sub><italic>y</italic></sub> is <italic>k</italic><sub>1</sub>&#x0002B;<italic>k</italic><sub>2</sub>&#x0002B;<italic>k</italic><sub>3</sub>. <italic>F</italic><sub><italic>y</italic></sub> is finally convoluted by a convolution layer, using six channel number and 1 &#x000D7; 1 kernel size. The parameters in feature extractor and new-added layer are updated by cross-entropy loss, which is already introduced in Section 2.1.</p>
<p>Liu et al. proposed a lightweight model with multi-scale features for steel surface defect classification (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>). Their model is named ConCNN, which is a concurrent CNN including input of two different image scales. Specifically, the 200 &#x000D7; 200 and 400 &#x000D7; 400 pixel images are input into two independent sub-networks with the same structure, and the output feature vectors of both sub-networks are fused to obtain the final classification output. Our proposed model is quite different from the ConCNN, and the difference includes two significant aspects. First, the multi-scale features in our model are achieved by constructing different pooling layers at the high level of the network; ConCNN inputs two scales of images so that two individual sub-networks learn features at two scales. Second, our model uses a pre-trained model to initialize the parameters in the network, while the parameters of ConCNN are randomly initialized. Yu et al. (<xref ref-type="bibr" rid="B36">2021</xref>) applied the high-order pooling for Alzheimer&#x00027;s disease assessment. The high-order pooling module is incorporated into the classifier to make full use of the correlation within feature maps along the channel axis to capture more discriminative CNN features. Different from high-order pooling, the proposed multi-scale pooling is constructed using first-order pooling at different scales. The focus of the two types of pooling is different.</p>
<p>In our proposed model, the output of the SqueezeNet feature extractor passes through the convolution layer with large step size; the most significant defect features will be retained, while small-size defect features will be suppressed in this process. More detailed features will be preserved when processed by the convolution layer with a small step size. The suggested model has a good detection performance because it can collect defect characteristics of both large and small sizes at a high level by combining the defect feature information from both sources. The effectiveness of the proposed multi-scale pooling structure is systematically evaluated in Section 3.3.</p></sec></sec>
<sec id="s3">
<title>3. Experiments</title>
<sec>
<title>3.1. Training and testing database</title>
<p>The NEU steel surface defect dataset (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>), which is freely accessible online, is the basis for our experiment. There are six classes of defects in the dataset, Crazing (Cr), Inclusion (In), Patches (Pa), Pitted surface (Ps), Rolled-in scale (Rs), and Scratches (Sc). Some sample images of different defect classes are shown in the top three rows of <xref ref-type="fig" rid="F3">Figure 3</xref>. It can be seen that the morphology of different defect types can show great differences, including color, texture features, and defect size, bringing challenges for defect classification. For each defect class, there are 300 images. The dataset is divided into the training set and the testing set. A total of 80% of samples is selected as the training set and 20% as the testing set. We do not make use of any data augmentation techniques in the training and testing sets.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Samples images of six classes of surface defects in the NEU surface defect dataset and a corresponding noise testing dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0003.tif"/>
</fig>
<p>Image noise, which is typically brought on by electronic noise or environmental variables, is a random variation of color values in acquired photographs. A typical type of picture noise, Gaussian noise, is applied to the testing set to examine the effectiveness and generalization ability of the suggested strategy in the case of potential noise. Comparative samples of NEU testing set and corresponding Gaussian noise images are shown in the bottom two rows of <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<p>Gaussian noise is statistical noise in which probability density function is Gaussian distribution. The Gaussian distribution is defined as</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi>&#x003C3;</mml:mi></mml:mrow></mml:mfrac><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BC; is average value and &#x003C3; is standard deviation. To control the intensity of noise, &#x003BC; and &#x003C3; are determined by a signal-to-noise ratio (SNR) value, which is defined as</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mrow><mml:mi>S</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:msub><mml:mrow><mml:mi>log</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>H</mml:mi></mml:msubsup><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>W</mml:mi></mml:msubsup><mml:mrow><mml:mi>s</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>H</mml:mi></mml:msubsup><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>W</mml:mi></mml:msubsup><mml:mrow><mml:mi>n</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>H</italic> and <italic>W</italic> indicate the height and width of the input image; <italic>s</italic>(<italic>i, j</italic>) and <italic>n</italic>(<italic>i, j</italic>) are the pixel values of signal and noise at pixel location (<italic>i, j</italic>). The SNR value is set to 20 dB. Each noise image are generated five times, and add up to 1,800 images (60 &#x000D7; 6 &#x000D7; 5). With the addition of noise, defect morphologies of six types have changed, which brings difficulties to the defect identification task.</p>
</sec>
<sec>
<title>3.2. Implementation Details</title>
<p>Caffe is one of the widely used deep-learning frameworks that are publicly available (Jia et al., <xref ref-type="bibr" rid="B13">2014</xref>). All the experiments are implemented in Caffe. Stochastic gradient descent policy is used to train the CNN models with a weight decay of 10<sup>&#x02212;4</sup> and a momentum of 0.9. The batch size is set to 32, which means 32 sample images are computed per iteration. The basic learning rate is 0.01 and after every 800 iterations, the learning rate becomes one-tenth of the original. NVIDIA 1080Ti GPU(12GB) is used in experiments to realize parallel computing and achieve good performance. We use Xavier&#x00027;s initialization (Glorot and Bengio, <xref ref-type="bibr" rid="B8">2010</xref>) for convolutional layers in our proposed model.</p>
</sec>
<sec>
<title>3.3. Comparisons with other models</title>
<p>In order to verify the effectiveness of the proposed MSP module and model, we compare our model with other defect classification models, including machine learning (ML)-based and CNN-based approaches. Among these approaches, the ML-based classifiers include support vector machine (SVM), nearest neighbor clustering (NNC), and multiple linear regression (MLR). Three feature extractors&#x02013;Gray level co-occurrence matrix (GLCM) (Haralick et al., <xref ref-type="bibr" rid="B9">1973</xref>), adaptive extended local ternary pattern (AELTP) (Mohamed and Yampolskiy, <xref ref-type="bibr" rid="B25">2013</xref>), and adjacent evaluation completed local binary patterns (AECLBP) (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>)&#x02014;are used. Nine classification methods can be obtained by combining different feature extractors and classifiers. Moreover, several CNN-based approaches are also compared, including ETE (Li et al., <xref ref-type="bibr" rid="B20">2016</xref>), DECAF&#x0002B;MLR (Ren et al., <xref ref-type="bibr" rid="B29">2018</xref>), AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>), ConCNN (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>), and SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>; Fu et al., <xref ref-type="bibr" rid="B7">2019</xref>). For a fair comparison, all the approaches are trained and tested on the same training and testing set, respectively. The training set includes 1,440 sample images (240 &#x000D7; 6), and the testing set includes 360 sample images without noise (60 &#x000D7; 6) and 1,800 with noise (60 &#x000D7; 6 &#x000D7; 5).</p>
<p>The comparative results are shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>The classification accuracy (%) of various steel surface defect classification approaches in both the NEU datasets without/with noise.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>Method</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>NEU</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>NEU with noise</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">GLCM&#x0002B;SVM (Haralick et al., <xref ref-type="bibr" rid="B9">1973</xref>)</td>
<td valign="top" align="center">88.1</td>
<td valign="top" align="center">67.8</td>
</tr> <tr>
<td valign="top" align="left">GLCM&#x0002B;NNC (Haralick et al., <xref ref-type="bibr" rid="B9">1973</xref>)</td>
<td valign="top" align="center">89.7</td>
<td valign="top" align="center">71.0</td>
</tr> <tr>
<td valign="top" align="left">GLCM&#x0002B;MLR (Haralick et al., <xref ref-type="bibr" rid="B9">1973</xref>)</td>
<td valign="top" align="center">94.7</td>
<td valign="top" align="center">52.3</td>
</tr> <tr>
<td valign="top" align="left">AELTP&#x0002B;SVM (Mohamed and Yampolskiy, <xref ref-type="bibr" rid="B25">2013</xref>)</td>
<td valign="top" align="center">76.1</td>
<td valign="top" align="center">44.6</td>
</tr> <tr>
<td valign="top" align="left">AELTP&#x0002B;NNC (Mohamed and Yampolskiy, <xref ref-type="bibr" rid="B25">2013</xref>)</td>
<td valign="top" align="center">96.4</td>
<td valign="top" align="center">64.4</td>
</tr> <tr>
<td valign="top" align="left">AELTP&#x0002B;MLR (Mohamed and Yampolskiy, <xref ref-type="bibr" rid="B25">2013</xref>)</td>
<td valign="top" align="center">98.6</td>
<td valign="top" align="center">48.8</td>
</tr> <tr>
<td valign="top" align="left">AECLBP&#x0002B;SVM (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>)</td>
<td valign="top" align="center">98.9</td>
<td valign="top" align="center">39.9</td>
</tr> <tr>
<td valign="top" align="left">AECLBP&#x0002B;NNC (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>)</td>
<td valign="top" align="center">98.3</td>
<td valign="top" align="center">42.7</td>
</tr> <tr>
<td valign="top" align="left">AECLBP&#x0002B;MLR (Song and Yan, <xref ref-type="bibr" rid="B31">2013</xref>)</td>
<td valign="top" align="center">98.3</td>
<td valign="top" align="center">43.7</td>
</tr> <tr>
<td valign="top" align="left">ETE (Li et al., <xref ref-type="bibr" rid="B20">2016</xref>)</td>
<td valign="top" align="center">95.8</td>
<td valign="top" align="center">47.6</td>
</tr> <tr>
<td valign="top" align="left">DECAF&#x0002B;MLR (Ren et al., <xref ref-type="bibr" rid="B29">2018</xref>)</td>
<td valign="top" align="center">99.7</td>
<td valign="top" align="center">89.7</td>
</tr> <tr>
<td valign="top" align="left">AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>)</td>
<td valign="top" align="center">91.4</td>
<td valign="top" align="center">83.4</td>
</tr> <tr>
<td valign="top" align="left">ConCNN (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>)</td>
<td valign="top" align="center">99.6</td>
<td valign="top" align="center">84.5</td>
</tr> <tr>
<td valign="top" align="left">SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>; Fu et al., <xref ref-type="bibr" rid="B7">2019</xref>)</td>
<td valign="top" align="center">99.7</td>
<td valign="top" align="center">82.9</td>
</tr> <tr>
<td valign="top" align="left">SqueezeNet&#x0002B;MSP(Proposed)</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">94.6</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It is noted that the classification accuracy of most machine-learning approaches is higher than 85% on the NEU dataset without noise. The defect feature characteristic of the testing set and the training set is relatively consistent. After the feature extractor obtains the defect features on the training set, the model trained by the machine learning classifier is also applicable to the testing set. The proposed model achieves 100% accuracy, better than the rest of CNN approaches (100% vs. 95.8%, 99.7%, 91.4%, 99.6% and 99.7%). For the testing set NEU with noise, the accuracy of most approaches has dropped dramatically, especially ML-based approaches&#x02014;GLCM/AELTP/AECLBP&#x0002B;SVM/NNC/MLR. With the addition of noise, defect feature characteristics of the sample images change accordingly. As a result, those models with poor generalization ability could not identify the defect class accurately. Among the CNN-based models, our proposed model still achieved the highest accuracy (94.6%), which is much higher than DECAF&#x0002B;MLR (Ren et al., <xref ref-type="bibr" rid="B29">2018</xref>) (89.7%), AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>) (83.4%), and ConCNN (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>) (84.5%). Among these methods, ConCNN (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>) could achieve a close performance to ours in the NEU dataset, but the accuracy drops to 84.5% in the noisy dataset. The reason behind this is that ConCNN constructed by inputting two scale images encountered difficulties in recognizing defect features with noise. According to the previous discussions, it can be concluded that by adding the MSP module to SqueezeNet, the accuracy improves from 82.9% to 94.6%, which proves that our proposed MSP is effective and robust.</p>
<p>In order to obtain more detailed defect classification information, the confusion matrixes of SqueezeNet and the proposed model on two testing sets are shown in <xref ref-type="table" rid="T3">Table 3</xref>. Due to six classes of defect type, each confusion matrix has six columns and six rows. Each column of the confusion matrix represents the prediction class, and the total number of each column represents the number of data predicted for this class. Each row represents the real class of defect image and the total number of data in each row represents the number of image samples of that class. For example, the number in the third row (Pa) and the second column (In) represents the total number of samples whose real class is Pa, which is predicted to be In. <xref ref-type="table" rid="T3">Table 3A</xref> shows the classification results of SqueezeNet in the NEU testing set, it can be seen that the only one In defect sample is wrongly identified as PS. <xref ref-type="table" rid="T3">Table 3B</xref> shows the results in the noisy set; it can be seen that the accuracy of identifying In and PS is low, indicating that they are easy to be confused with other defect types. From <xref ref-type="table" rid="T3">Table 3C</xref>, it can be seen from the experimental results that all defect classes are correctly identified on the NEU testing set without noise. <xref ref-type="table" rid="T3">Table 3D</xref> shows the results of the NEU testing set with noise, the defect samples Cr, Pa, and RS are 100% accurately identified, In is easily confused with PS and Sc and Ps is easily confused with Cr. The experimental results prove that the proposed MSP module is helpful to improve the accuracy of defect classification.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>The confusion matrix of SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>; Fu et al., <xref ref-type="bibr" rid="B7">2019</xref>) and the proposed model on NEU steel surface defect dataset without/with noise.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496" colspan="7"><bold>(A)</bold></th>
<th valign="top" align="center" style="background-color:#8f9496" colspan="7"><bold>(B)</bold></th>
</tr>
</thead>
<tbody>
 <tr>
<td style="background-color:#8f9496"/>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Cr</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>In</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Pa</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>PS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>RS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Sc</bold></td>
<td style="background-color:#8f9496"/>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Cr</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>In</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Pa</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>PS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>RS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Sc</bold></td>
</tr> <tr>
<td valign="top" align="left">Cr</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">Cr</td>
<td valign="top" align="center">300</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">In</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">59</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">In</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">133</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">72</td>
<td valign="top" align="center">52</td>
</tr> <tr>
<td valign="top" align="left">Pa</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">Pa</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">299</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">PS</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">PS</td>
<td valign="top" align="center">106</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">169</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">RS</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">RS</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">296</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">Sc</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">Sc</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">296</td>
</tr> <tr>
<td valign="top" align="left" colspan="7" style="background-color:#8f9496"><bold>(C)</bold></td>
<td valign="top" align="center" colspan="7" style="background-color:#8f9496"><bold>(D)</bold></td>
</tr>
 <tr>
<td style="background-color:#8f9496"/>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Cr</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>In</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Pa</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>PS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>RS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Sc</bold></td>
<td style="background-color:#8f9496"/>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Cr</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>In</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Pa</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>PS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>RS</bold></td>
<td valign="top" align="center" style="background-color:#8f9496"><bold>Sc</bold></td>
</tr> <tr>
<td valign="top" align="left">Cr</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">Cr</td>
<td valign="top" align="center">300</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">In</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">In</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">267</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">20</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">13</td>
</tr> <tr>
<td valign="top" align="left">Pa</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">Pa</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">300</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">PS</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">PS</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">239</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">RS</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">RS</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">300</td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="top" align="left">Sc</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">60</td>
<td valign="top" align="center">Sc</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">298</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><bold>(A)</bold> SqueezeNet in NEU testing set; <bold>(B)</bold> SqueezeNet in NEU testing set with noise; <bold>(C)</bold> proposed model in NEU testing set; and <bold>(D)</bold> proposed model in NEU testing set with noise.</p>
</table-wrap-foot>
</table-wrap>
<p>In order to comprehensively analyze and propose the model, the running speed and model size of different CNN-based defect classification methods are also compared. The comparison results are shown in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>The running time and model size of several CNN-based methods.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>Method</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>Running time(s)</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>Model size (MB)</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>Accuracy in noisy testing set (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ETE (Li et al., <xref ref-type="bibr" rid="B20">2016</xref>)</td>
<td valign="top" align="center">0.005</td>
<td valign="top" align="center">1.9</td>
<td valign="top" align="center">47.6</td>
</tr> <tr>
<td valign="top" align="left">DECAF&#x0002B;MLR (Ren et al., <xref ref-type="bibr" rid="B29">2018</xref>)</td>
<td valign="top" align="center">0.015</td>
<td valign="top" align="center">244</td>
<td valign="top" align="center">89.7</td>
</tr> <tr>
<td valign="top" align="left">AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B16">2012</xref>)</td>
<td valign="top" align="center">0.085</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">83.4</td>
</tr> <tr>
<td valign="top" align="left">SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B12">2016</xref>)</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">3.0</td>
<td valign="top" align="center">82.9</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">0.007</td>
<td valign="top" align="center">3.0</td>
<td valign="top" align="center">94.6</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The running speed is the time for the model to process an image sample, and three repeated experiments are used to calculate the average speed. It is observed that the file size of the proposed model is 3 MB, which is easy to deploy on mobile devices. Although ETE&#x00027;s model runs faster, it is less accurate than the proposed model. The running speed of the proposed model reaches 130FPS on 1080TI GPU, which is able to fully satisfy the demand for fast detection in industrial scenarios.</p>
</sec>
<sec>
<title>3.4. Ablation study</title>
<p>In the proposed model, the scale parameter settings of MSP will impact the classification performance. To determine the optimal scale parameters, the following comparison experiments are conducted. In a pooling layer, the value of the stride is usually set to 2<sup><italic>n</italic></sup>, that is, 1, 2, 4, 8, etc. Considering the output feature width of the SqueezeNet feature extractor is 16. As a result, all the optional strides are 1, 2, 4, 8, and 16. The MSP module contains five different implementation ways, as shown in <xref ref-type="table" rid="T5">Table 5</xref>.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>The classification accuracy (%) of MSP using different scales in the NEU dataset without/with noise.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>The strides in MSP</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>NEU</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>NEU with noise</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.871</td>
</tr> <tr>
<td valign="top" align="left">1, 2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.898</td>
</tr> <tr>
<td valign="top" align="left">1, 2, 4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.946</td>
</tr> <tr>
<td valign="top" align="left">1, 2, 4, 8</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.932</td>
</tr> <tr>
<td valign="top" align="left">1, 2, 4, 8, 16</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.930</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The experimental results show that the combination of the three scales achieves the best classification accuracy. When the scale increases from one to three, the accuracy increases, and the classification accuracy declines when the number of scales increases further. This phenomenon of decreased accuracy is due to the fact that the pooling layer with large strides will lose too much location information, which is detrimental to defect feature recognition. Therefore, the proposed MSP in our study is composed of three pooling layers using 1, 2, and 4 strides, respectively.</p>
</sec>
<sec>
<title>3.5. Class activation map analysis</title>
<p>In order to analyze the defect features learned in neural networks and which defect features are the key to judge the defect type, the class activation map (CAM) is used for feature analysis in neural networks (Zhou et al., <xref ref-type="bibr" rid="B38">2016</xref>). In the proposed multi-scale pooling module, conv11&#x02212;<italic>i</italic> are the last convolutional layers, and defect features are learned at three scales. By multiplying the feature map in conv11&#x02212;<italic>i</italic> and the output value of the corresponding global average pooling layer, summing up all products, the CAM is obtained. The process is calculated as</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> are the feature map and corresponding activated score at <italic>jth</italic> scale, detailed description is given in Section 2.2. <italic>M</italic><sub><italic>j</italic></sub> is the class activation map at scale <italic>j</italic>. To explain the calculation process more intuitively, two types of defect sample images are selected as examples. The detailed calculation process of CAM at scale 1 is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, where two types of defects were selected for analysis.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>The detailed calculation process of class activation map at scale 1. <bold>(A)</bold> Test sample Crazing; <bold>(B)</bold> Test sample Patches.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0004.tif"/>
</fig>
<p><xref ref-type="fig" rid="F4">Figure 4A</xref> shows the calculation process of using Crazing as a testing image. In the Conv11 &#x02212; 1 convolution layer, there are six feature channels, namely <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x0007E;</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>. The six feature maps are feed into the GAP to get six activation scores, namely <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x0007E;</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>. Specifically, the six scores are 10.588, 0.005, 49.979, 5.6444, 1.007, and 0.049, which are calculated by the intensity of pixels in the feature image. The CAM at this scale is obtained by multiplying <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>.</p>
<p>It is worth noting that some channels have richer feature composition(<inline-formula><mml:math id="M24"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>), while others have fewer(<inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M27"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>). When another defect type Pa is selected as the input image, the corresponding feature maps are shown in <xref ref-type="fig" rid="F4">Figure 4B</xref>. Different from <xref ref-type="fig" rid="F4">Figure 4A</xref>, <inline-formula><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M30"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> have richer feature composition, while <inline-formula><mml:math id="M32"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> has fewer. The explanation is that different types of defect features will be activated on different channels.</p>
<p>For a comprehensive analysis, the class activation map visualization results of six defect types at three scales in the multi-scale pooling module are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. For better visual effects, the pixels of all CAM results are adjusted to 200 &#x000D7; 200, and the grayscale images are converted to heat map mode. Two samples of each defect type in NEU with noise testing set were selected for analysis, namely test sample A and B. To study the CAM characteristic at three different scales and their differences, the CAM of three scales in the proposed MSP module are given. To verify the effectiveness of the proposed MSP module, the MSP of SqueezeNet model is also shown for comparison.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The class activation map visualization results of convolutional features in MSP module and SqueezeNet.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0005.tif"/>
</fig>
<p>It can be observed that the CAM at three scales can highlight the defect regions, but the focus is different. In general, the results of scale 1 have a higher resolution and can retain more detailed defect cues; scale 3 highlights the main defect regions and ignores some small-scale defect cues, while scale 2 is characterized by a synthesis of scales 1 and 3. The focused regions at three scales complement each other, such as the test sample A of Crazing, Inclusion, Patches, Pitted surface, Rolled-in scale, and Scratches. Taking sample A of Crazing as an example, it can be seen that the highlighted areas of scale 1 are relatively scattered; the highlighted areas of scale 2 are compact and concentrated in the lower right and upper right corners; highlighted areas are concentrated in the upper right corner at scale 3.</p>
<p>In addition to being complementary, the highlighted areas at multiple scales may show very high consistency, such as test sample B of six defect classes. Consistency of highlight areas at multiple scales could strengthen the identification of defects. Take sample B of Inclusion as an example; the CAM at all three scales focus on long-striped defect features. In addition, comparing the CAM of the proposed model with SqueezeNet, SqueezeNet does not properly focus on the defect area, such as sample A/B of Crazing, Pitted surface, and Rolled-in scale. Missing the correct defect region will result in false detection of the defect class. The afore-mentioned comparison results could confirm the validity of the newly added MSP module. In conclusion, the CAM of the proposed model could accurately locate defect locations at multiple scales, and the highlighted areas at multiple scales can complement and reinforce each other. In the subsequent feature fusion layer, the neural network can adaptively learn the relationship between features of different scales according to the characteristics of defects to obtain more reliable defect classification cues.</p>
</sec>
<sec>
<title>3.6. T-SNE dimension reduction visualization Analysis</title>
<p>To study the class-related information learned in the hidden layer of the neural network, the t-distributed stochastic neighbor embedding (T-SNE) (Van der Maaten and Hinton, <xref ref-type="bibr" rid="B32">2008</xref>) dimension reduction method is used to visualize the neural network parameters. Specifically, the six numerical parameters in Conv-12 (shown in <xref ref-type="fig" rid="F2">Figure 2</xref>) are used as raw data. T-SNE shows the representation of six-dimensional data in two-dimensional, intuitively, the degree of aggregation between different defect image samples. The distance of two samples in raw data (six dimensions) is calculated as</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>|</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext class="textrm" mathvariant="normal">exp</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mtext class="textrm" mathvariant="normal">exp</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mo>&#x02225;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>x</italic><sub><italic>i</italic></sub>, <italic>x</italic><sub><italic>j</italic></sub> is the CNN features of two samples, and <inline-formula><mml:math id="M34"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is variance. Then, the joint distribution <italic>P</italic><sub><italic>ij</italic></sub> is calculated as</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02223;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02223;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> is the total number of samples in testing set. The distance of projection points in two-dimension is calculated as</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M36"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>y</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>j</italic></sub> is the projection points of two samples in two-dimension. Kullback&#x02013;Leibler (KL) divergence is used to measure the similarity between points at high and low dimensions to ensure that points with high similarity at high dimension also have high similarity at low dimension. KL-divergence is calculated as</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M37"><mml:mtable class="eqnarray" columnalign="center"><mml:mtr><mml:mtd><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The T-SNE dimension reduction visualization results of AlexNet, SqueezeNet, and the proposed model are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The T-SNE feature dimension reduction visualization results comparison of different CNN model classifiers in NEU dataset without/with noise. <bold>(A)</bold> AlexNet in NEU testing set; <bold>(B)</bold> AlexNet in NEU testing set with noise; <bold>(C)</bold> SqueezeNet in NEU testing set; <bold>(D)</bold> SqueezeNet in NEU testing set with noise; <bold>(E)</bold> proposed model in NEU testing set; <bold>(F)</bold> proposed model in NEU testing set with noise.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0006.tif"/>
</fig>
<p><xref ref-type="fig" rid="F6">Figure 6A</xref> shows the result of AlexNet in the NEU testing set, it can be seen that the samples Cr, Pa, RS, and Sc aggregate together and have a clear separation, while In and PS mix together. <xref ref-type="fig" rid="F6">Figures 6C</xref>, <xref ref-type="fig" rid="F6">E</xref> shows the results of SquuezeNet and the proposed model in the NEU testing set, respectively; six classes of the defect image samples are completely separated in both. It is worth noting that the PS cluster in the proposed model is no longer in the middle of multiple classes, thus increasing the inter-class distance, which is helpful in obtaining more reliable classification results. <xref ref-type="fig" rid="F6">Figures 6B</xref>, <xref ref-type="fig" rid="F6">D</xref> shows the results of AlexNet and SqueezeNet in the NEU testing set with noise, respectively. It can be seen that many defect class clusters show dispersion and mix with each other, such as Pa, PS, PS, and Sc. <xref ref-type="fig" rid="F6">Figure 6F</xref> shows the results of the proposed model in the NEU testing set with noise, and these clusters have a clear separation; only a few samples are mixed into other classes. To sum up, the clustering results of the proposed model have a small intra-class distance and a large inter-class distance, proving that our model has better performance and strong generalization ability.</p>
</sec>
<sec>
<title>3.7. Generalization ability verification</title>
<p>To verify the generalization ability of the MSP module, we conduct comparative experiments in the field of medical image processing. We chose the cell Malaria image classification task (Narayanan et al., <xref ref-type="bibr" rid="B26">2019</xref>), which is provided by Kaggle. The Malaria dataset includes two classes, parasitic and uninfected, and the typical image samples are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. It is observed that both cell color and shape are highly variable. The training set and the testing set have been partitioned in the original dataset and the partition of the dataset is shown in <xref ref-type="table" rid="T6">Table 6</xref>. The training set contains 220 parasitic and 196 uninfected samples, and the testing set contains 91 parasitic and 43 uninfected samples.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Sample images of parasitic and uninfected cell image samples in the Malaria dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1096083-g0007.tif"/>
</fig>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>The number of parasitic and uninfected cell image samples in the Malaria training and testing dataset (Narayanan et al., <xref ref-type="bibr" rid="B26">2019</xref>).</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>Malaria class</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>Parasite</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>Uninfected</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Training set</td>
<td valign="top" align="center">220</td>
<td valign="top" align="center">196</td>
</tr> <tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">43</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The Malaria training set is trained on SqueezeNet and SqueezeNet&#x0002B;MSP model with the same training parameters, and the prediction accuracy is shown in <xref ref-type="table" rid="T7">Table 7</xref>. It can be seen that the accuracy of the SqueezeNet reaches 97.7%, and the accuracy is further improved by adding the MSP module, reaching 98.5%. This proves that the proposed MSP is effective in improving the classification accuracy of Malaria images. It can be concluded that MSP is not only applicable in the field of industrial defect detection but also in medical image processing.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>The classification accuracy (%) of SqueezeNet and proposed model in the Malaria dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" style="background-color:#8f9496"><bold>Method</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>SqueezeNet</bold></th>
<th valign="top" align="center" style="background-color:#8f9496"><bold>SqueezeNet&#x0002B;MSP (Proposed)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Accuracy (%)</td>
<td valign="top" align="center">97.7</td>
<td valign="top" align="center">98.5</td>
</tr>
</tbody>
</table>
</table-wrap></sec></sec>
<sec sec-type="conclusions" id="s4">
<title>4. Conclusion</title>
<p>For the classification of steel surface defects, we propose a multi-scale pooling convolutional neural network in this research. Our model is based on SqueezeNet, and to capture defect features at various scales, we propose an innovative multi-scale pooling module. In the module, the multi-scale features are combined to produce more reliable defect cues. It is demonstrated that our model has higher accuracy by contrasting it with other defect classification models in a noisy NEU testing set. The MSP module is able to locate the defect location accurately, according to class activation map analysis, and the highlighted areas at different scales could complement and reinforce one another to produce more reliable results. According to the visualization results of T-SNE, the suggested model has a small intra-class spacing and a big inter-class spacing, which suggests a strong generalization capacity in handling noise. Furthermore, our model&#x00027;s 3MB size and 130 FPS performance on a single NVIDIA 1080Ti GPU, which could be applied to scenarios where device computation power is constrained and detection speed is required. In future, we intend to develop a model that can perform well on various surface defect dataset classification tasks, saving time on model fine-tuning.</p></sec>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="http://faculty.neu.edu.cn/songkechen/zh_CN/zhym/263269/list/index.htm">http://faculty.neu.edu.cn/songkechen/zh_CN/zhym/263269/list/index.htm</ext-link>.</p></sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This research was supported by the National Natural Science Foundation of China (Nos. 52105526, 51975394, 61903269, and 51875380) and the China Postdoctoral Science Foundation (Nos. 2020M671383 and 2020M681517).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aghdam</surname> <given-names>S. R.</given-names></name> <name><surname>Amid</surname> <given-names>E.</given-names></name> <name><surname>Imani</surname> <given-names>M. F.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;A fast method of steel surface defect detection using decision trees applied to lbp based features,&#x0201D;</article-title> in <source>2012 7th IEEE Conference on Industrial Electronics and Applications (ICIEA)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1447</fpage>&#x02013;<lpage>1452</lpage>. <pub-id pub-id-type="doi">10.1109/ICIEA.2012.6360951</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belkin</surname> <given-names>M.</given-names></name> <name><surname>Hsu</surname> <given-names>D.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <name><surname>Mandal</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Reconciling modern machine-learning practice and the classical bias-variance trade-off</article-title>. <source>Proc. Nat. Acad. Sci. U.S.A</source>. <volume>116</volume>, <fpage>15849</fpage>&#x02013;<lpage>15854</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1903070116</pub-id><pub-id pub-id-type="pmid">31341078</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bissi</surname> <given-names>L.</given-names></name> <name><surname>Baruffa</surname> <given-names>G.</given-names></name> <name><surname>Placidi</surname> <given-names>P.</given-names></name> <name><surname>Ricci</surname> <given-names>E.</given-names></name> <name><surname>Scorzoni</surname> <given-names>A.</given-names></name> <name><surname>Valigi</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>Automated defect detection in uniform and structured fabrics using gabor filters and pca</article-title>. <source>J. Vis. Commun. Image Represent</source>. <volume>24</volume>, <fpage>838</fpage>&#x02013;<lpage>845</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2013.05.011</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bouaafia</surname> <given-names>S.</given-names></name> <name><surname>Messaoud</surname> <given-names>S.</given-names></name> <name><surname>Maraoui</surname> <given-names>A.</given-names></name> <name><surname>Ammari</surname> <given-names>A. C.</given-names></name> <name><surname>Khriji</surname> <given-names>L.</given-names></name> <name><surname>Machhout</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Deep pre-trained models for computer vision applications: traffic sign recognition,&#x0201D;</article-title> in <source>2021 18th International Multi-Conference on Systems, Signals and Devices (SSD)</source> (<publisher-loc>Monastir</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>23</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1109/SSD52085.2021.9429420</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chu</surname> <given-names>M.</given-names></name> <name><surname>Gong</surname> <given-names>R.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Steel surface defects recognition based on multi-type statistical features and enhanced twin support vector machine</article-title>. <source>Chemometr. Intellig. Lab. Syst</source>. <volume>171</volume>, <fpage>140</fpage>&#x02013;<lpage>150</lpage>. <pub-id pub-id-type="doi">10.1016/j.chemolab.2017.10.020</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dogand&#x0017E;i&#x00107;</surname> <given-names>A.</given-names></name> <name><surname>Eua-anant</surname> <given-names>N.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name></person-group> (<year>2005</year>). <article-title>Defect detection using hidden markov random fields</article-title>. <source>AIP Conf. Proc</source>. <volume>760</volume>, <fpage>704</fpage>&#x02013;<lpage>711</lpage>. <pub-id pub-id-type="doi">10.1063/1.1916744</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>G.</given-names></name> <name><surname>Sun</surname> <given-names>P.</given-names></name> <name><surname>Zhu</surname> <given-names>W.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>M. Y.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>A deep-learning-based approach for fast and robust steel surface defects classification</article-title>. <source>Opt. Lasers Eng</source>. <volume>121</volume>, <fpage>397</fpage>&#x02013;<lpage>405</lpage>. <pub-id pub-id-type="doi">10.1016/j.optlaseng.2019.05.005</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glorot</surname> <given-names>X.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Understanding the difficulty of training deep feedforward neural networks,&#x0201D;</article-title> in <source>Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics</source> (<publisher-loc>Sardinia</publisher-loc>), <fpage>249</fpage>&#x02013;<lpage>256</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haralick</surname> <given-names>R. M.</given-names></name> <name><surname>Shanmugam</surname> <given-names>K.</given-names></name> <name><surname>Dinstein</surname> <given-names>I. H.</given-names></name></person-group> (<year>1973</year>). <article-title>Textural features for image classification</article-title>. <source>IEEE Trans. Syst. Man Cybern</source>. <volume>6</volume>, <fpage>610</fpage>&#x02013;<lpage>621</lpage>. <pub-id pub-id-type="doi">10.1109/TSMC.1973.4309314</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep residual learning for image recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id><pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howard</surname> <given-names>A. G.</given-names></name> <name><surname>Zhu</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Kalenichenko</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Weyand</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Mobilenets: efficient convolutional neural networks for mobile vision applications</article-title>. <source>arXiv preprint</source> arXiv:1704.04861.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iandola</surname> <given-names>F. N.</given-names></name> <name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Moskewicz</surname> <given-names>M. W.</given-names></name> <name><surname>Ashraf</surname> <given-names>K.</given-names></name> <name><surname>Dally</surname> <given-names>W. J.</given-names></name> <name><surname>Keutzer</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Squeezenet: alexnet-level accuracy with 50x fewer parameters and &#x0003C; 0.5mb model size</article-title>. <source>arXiv preprint</source> arXiv:1602.07360.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Shelhamer</surname> <given-names>E.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Karayev</surname> <given-names>S.</given-names></name> <name><surname>Long</surname> <given-names>J.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>&#x0201C;Caffe: convolutional architecture for fast feature embedding,&#x0201D;</article-title> in <source>Proceedings of the 22nd ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>675</fpage>&#x02013;<lpage>678</lpage>. <pub-id pub-id-type="doi">10.1145/2647868.2654889</pub-id><pub-id pub-id-type="pmid">32210685</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>E.</given-names></name></person-group> (<year>2020</year>). <article-title>A universal and adaptive fabric defect detection algorithm based on sparse dictionary learning</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>221808</fpage>&#x02013;<lpage>221830</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3041849</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Khumaidi</surname> <given-names>A.</given-names></name> <name><surname>Yuniarno</surname> <given-names>E. M.</given-names></name> <name><surname>Purnomo</surname> <given-names>M. H.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Welding defect classification based on convolution neural network (cnn) and gaussian kernel,&#x0201D;</article-title> in <source>2017 International Seminar on Intelligent Technology and Its Applications (ISITIA)</source> (<publisher-loc>Surabaya</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>261</fpage>&#x02013;<lpage>265</lpage>. <pub-id pub-id-type="doi">10.1109/ISITIA.2017.8124091</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Imagenet classification with deep convolutional neural networks,&#x0201D;</article-title> in <source>International Conference on Neural Information Processing Systems</source> (<publisher-loc>Lake Tahoe, NV</publisher-loc>), <fpage>1097</fpage>&#x02013;<lpage>1105</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Haffner</surname> <given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x02013;<lpage>2324</lpage>. <pub-id pub-id-type="doi">10.1109/5.726791</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>P.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Jing</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Fabric defect detection based on multi-scale wavelet transform and gaussian mixture model method</article-title>. <source>J. Textile Inst</source>. <volume>106</volume>, <fpage>587</fpage>&#x02013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.1080/00405000.2014.929790</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.-C.</given-names></name> <name><surname>Tsai</surname> <given-names>D.-M.</given-names></name></person-group> (<year>2012</year>). <article-title>Wavelet-based defect detection in solar wafer images with inhomogeneous texture</article-title>. <source>Pattern Recognit</source>. <volume>45</volume>, <fpage>742</fpage>&#x02013;<lpage>756</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2011.07.025</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Jiang</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>An end-to-end steel strip surface defects recognition system based on convolutional neural networks</article-title>. <source>Steel Res. Int</source>. 88, 1600068. <pub-id pub-id-type="doi">10.1002/srin.201600068</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>Q.</given-names></name> <name><surname>Yan</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Network in network</article-title>. <source>arXiv preprint</source> arXiv:1312.4400.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Balta</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>A light-weight deep-learning model with multi-scale features for steel surface defect classification</article-title>. <source>Materials</source> <volume>13</volume>, <fpage>4629</fpage>. <pub-id pub-id-type="doi">10.3390/ma13204629</pub-id><pub-id pub-id-type="pmid">33081388</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Goswami</surname> <given-names>V.</given-names></name> <name><surname>Rohrbach</surname> <given-names>M.</given-names></name> <name><surname>Parikh</surname> <given-names>D.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;12-in-1: Multi-task vision and language representation learning,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>10437</fpage>&#x02013;<lpage>10446</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01045</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mak</surname> <given-names>K.-L.</given-names></name> <name><surname>Peng</surname> <given-names>P.</given-names></name> <name><surname>Yiu</surname> <given-names>K. F. C.</given-names></name></person-group> (<year>2009</year>). <article-title>Fabric defect detection using morphological filters</article-title>. <source>Image Vis. Comput</source>. <volume>27</volume>, <fpage>1585</fpage>&#x02013;<lpage>1592</lpage>. <pub-id pub-id-type="doi">10.1016/j.imavis.2009.03.007</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohamed</surname> <given-names>A. A.</given-names></name> <name><surname>Yampolskiy</surname> <given-names>R. V.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Adaptive extended local ternary pattern (aeltp) for recognizing avatar faces,&#x0201D;</article-title> in <source>International Conference on Machine Learning and Applications</source> (<publisher-loc>Boca Raton, FL</publisher-loc>). <pub-id pub-id-type="doi">10.1109/ICMLA.2012.19</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Narayanan</surname> <given-names>B. N.</given-names></name> <name><surname>Ali</surname> <given-names>R.</given-names></name> <name><surname>Hardie</surname> <given-names>R. C.</given-names></name></person-group> (<year>2019</year>). <article-title>Performance analysis of machine learning and deep learning architectures for malaria detection on cell images</article-title>. <source>Appl. Mach. Learn</source>. <volume>11139</volume>, <fpage>240</fpage>&#x02013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1117/12.2524681</pub-id><pub-id pub-id-type="pmid">33465520</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neogi</surname> <given-names>N.</given-names></name> <name><surname>Mohanta</surname> <given-names>D. K.</given-names></name> <name><surname>Dutta</surname> <given-names>P. K.</given-names></name></person-group> (<year>2014</year>). <article-title>Review of vision-based steel surface inspection systems</article-title>. <source>EURASIP J. Image Video Process</source>. 2014, 50. <pub-id pub-id-type="doi">10.1186/1687-5281-2014-50</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ravikumar</surname> <given-names>S.</given-names></name> <name><surname>Ramachandran</surname> <given-names>K.</given-names></name> <name><surname>Sugumaran</surname> <given-names>V.</given-names></name></person-group> (<year>2011</year>). <article-title>Machine learning approach for automated visual inspection of machine components</article-title>. <source>Expert Syst. Appl</source>. <volume>38</volume>, <fpage>3260</fpage>&#x02013;<lpage>3266</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2010.09.012</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>R.</given-names></name> <name><surname>Hung</surname> <given-names>T.</given-names></name> <name><surname>Tan</surname> <given-names>K. C.</given-names></name></person-group> (<year>2018</year>). <article-title>A generic deep-learning-based approach for automated surface inspection</article-title>. <source>IEEE Trans. Cybern</source>. <volume>48</volume>, <fpage>929</fpage>&#x02013;<lpage>940</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2017.2668395</pub-id><pub-id pub-id-type="pmid">28252414</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <source>Very Deep Convolutional Networks for Large-Scale Image Recognition</source>. (San Diego, CA: ICLR).</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>K.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>A noise robust method based on completed local binary patterns for hot-rolled steel strip surface defects</article-title>. <source>Appl. Surf. Sci</source>. <volume>285</volume>, <fpage>858</fpage>&#x02013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1016/j.apsusc.2013.09.002</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Visualizing data using t-sne</article-title>. <source>J. Mach. Learn. Res</source>. <volume>9</volume>, <fpage>2579</fpage>&#x02013;<lpage>2605</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Automatic recognition of mild cognitive impairment and alzheimers disease using ensemble based 3d densely connected convolutional networks,&#x0201D;</article-title> in <source>2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA)</source> (<publisher-loc>Orlando, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>517</fpage>&#x02013;<lpage>523</lpage>. <pub-id pub-id-type="doi">10.1109/ICMLA.2018.00083</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>He</surname> <given-names>B.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Cheung</surname> <given-names>P. W.-H.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>An ensemble-based densely-connected deep learning system for assessment of skeletal maturity</article-title>. <source>IEEE Trans. Syst. Man Cybern. Syst</source>. <volume>52</volume>, <fpage>426</fpage>&#x02013;<lpage>437</lpage>. <pub-id pub-id-type="doi">10.1109/TSMC.2020.2997852</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Du</surname> <given-names>S. S.</given-names></name> <name><surname>Kawarabayashi</surname> <given-names>K.-,i.</given-names></name> <name><surname>Jegelka</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>How neural networks extrapolate: from feedforward to graph neural networks</article-title>. <source>arXiv preprint</source> arXiv:2009.11848.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>W.</given-names></name> <name><surname>Lei</surname> <given-names>B.</given-names></name> <name><surname>Ng</surname> <given-names>M. K.</given-names></name> <name><surname>Cheung</surname> <given-names>A. C.</given-names></name> <name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Tensorizing gan with high-order pooling for Alzheimer&#x00027;s disease assessment</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst</source>. <volume>33</volume>, <fpage>4945</fpage>&#x02013;<lpage>4959</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2021.3063516</pub-id><pub-id pub-id-type="pmid">33729958</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Ren</surname> <given-names>W.</given-names></name> <name><surname>Wen</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>Random forest-based real-time defect detection of al alloy in robotic arc welding using optical spectrum</article-title>. <source>J. Manuf. Process</source>. <volume>42</volume>, <fpage>51</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmapro.2019.04.023</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Lapedriza</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Learning deep features for discriminative localization,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>2921</fpage>&#x02013;<lpage>2929</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.319</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zori&#x00107;</surname> <given-names>B.</given-names></name> <name><surname>Mati&#x00107;</surname> <given-names>T.</given-names></name> <name><surname>Hocenski</surname> <given-names>&#x0017D;.</given-names></name></person-group> (<year>2022</year>). <article-title>Classification of biscuit tiles for defect detection using fourier transform features</article-title>. <source>ISA Trans</source>. <volume>125</volume>, <fpage>400</fpage>&#x02013;<lpage>414</lpage>. <pub-id pub-id-type="doi">10.1016/j.isatra.2021.06.025</pub-id><pub-id pub-id-type="pmid">34217499</pub-id></citation></ref>
</ref-list> 
</back>
</article>