<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2022.889308</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Recognition and Classification of Ship Images Based on SMS-PCNN Model</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Fengxiang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1705726/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Liang</surname> <given-names>Huang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Yalun</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Qingxia</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zong</surname> <given-names>Ruirui</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>College of Electronic Engineering, Naval University of Engineering</institution>, <addr-line>Wuhan</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Institute of Noise &#x00026; Vibration, Naval University of Engineering</institution>, <addr-line>Wuhan</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>College of International Studies, National University of Defense Technology</institution>, <addr-line>Changsha</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Marco Leo, National Research Council (CNR), Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Divya Sardana, Nike, United States; Tyler Highlander, Wittenberg University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Huang Liang <email>jwf20&#x00040;mails.tsinghua.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>889308</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>03</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Wang, Liang, Zhang, Xu and Zong.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Wang, Liang, Zhang, Xu and Zong</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>In the field of ship image recognition and classification, traditional algorithms lack attention to the differences between the grain of ship images. The differences in the hull structure of different categories of ships are reflected in the coarse-grain, whereas the differences in the ship equipment and superstructures of different ships of the same category are reflected in the fine-grain. To extract the ship features of different scales, the multi-scale paralleling CNN oriented on ships images (SMS-PCNN) model is proposed in this paper. This model has three characteristics. (1) Extracting image features of different sizes by parallelizing convolutional branches with different receptive fields. (2) The number of channels of the model is adjusted two times to extract features and eliminate redundant information. (3) The residual connection network is used to extend the network depth and mitigate the gradient disappearance. In this paper, we collected open-source images on the Internet to form an experimental dataset and conduct performance tests. The results show that the SMS-PCNN model proposed in this paper achieves 84.79% accuracy on the dataset, which is better than the existing four state-of-the-art approaches. By the ablation experiments, the effectiveness of the optimization tricks used in the model is verified.</p></abstract>
<kwd-group>
<kwd>image classification</kwd>
<kwd>multi-scale</kwd>
<kwd>CNN</kwd>
<kwd>ship images</kwd>
<kwd>ResNet</kwd>
</kwd-group>
<counts>
<fig-count count="13"/>
<table-count count="4"/>
<equation-count count="25"/>
<ref-count count="48"/>
<page-count count="19"/>
<word-count count="9350"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>In the military field, ship image classification is used to conduct precise strikes against hostile targets and important to carry out counter-terrorism missions. In the civilian field, ship image classification can assist relevant departments in maritime traffic control, search and rescue, and anti-smuggling activities. Therefore, ship image classification has broad applications and technical requirements in both military and civilian fields.</p>
<p>Currently, there are four types of maritime target images: radar images, remote sensing images, infrared images, and visible light images. Radar (Jiang et al., <xref ref-type="bibr" rid="B16">2021</xref>; Tang et al., <xref ref-type="bibr" rid="B38">2021</xref>) image recognition is all-weather and daylong, which means that it is not easily affected by light and weather. Its mainstream approach is extracting and classifying the features of radar echo signals, so as to achieve autonomous deep feature extraction of the data. Remote sensing (Yang et al., <xref ref-type="bibr" rid="B46">2014</xref>, <xref ref-type="bibr" rid="B45">2017</xref>) image recognition is extracting geometric features such as length and contour of targets in high-resolution SAR remote sensing images, so as to enhance SAR image recognition capability. Infrared image recognition can work in a long distance, which can penetrate thick fog and work day and night.</p>
<p>For visible images, traditional image recognition and classification use techniques including pixel-level edge detection, genetic algorithm, and support vector machine (SVM) classification. Atsuto and Kazuhiro (<xref ref-type="bibr" rid="B3">2004</xref>) proposed a multi-frame image processing algorithm to extract contours as basic features of targets for vector analysis, achieving good recognition performance. Xu et al. (<xref ref-type="bibr" rid="B43">2017</xref>) designed a multi-level discrimination method based on the improved entropy and pixel distribution with multi-scale and multidirectional decomposition of high-frequency coefficients, which can effectively resist background interferences and improve recognition accuracy and efficiency. Yang and Kim (<xref ref-type="bibr" rid="B44">2012</xref>) integrated SAR and automatic identification system (AIS) datasets as one system that can display the position, size, and classification of ships on SAR images. Enriquez de Luna et al. (<xref ref-type="bibr" rid="B10">2005</xref>) proposed a silhouette-based decision support system for ship image classification, using an evolved version of the Curvature Scale space (CSS) to improve recognition accuracy.</p>
<p>In the recent years, convolutional neural networks (CNNs) have gradually been widely applied in visible light image classification and recognition. A series of classic CNN models, including AlexNet, VGG, GoogLeNet, ResNet, DenseNet, and so on, stand out in the ImageNet Large Scale Visual Recognition Challenge (ILSVRC), which covers various fields such as image classification and target detection.</p>
<p>The AlexNet (Alex et al., <xref ref-type="bibr" rid="B2">2017</xref>) model is the first multi-layer CNN with five convolutional layers, three fully connected layers, and a max-pooling layer. This state-of-the-art model is a landmark of CNN model. The proposal of the VGG (Simonyan and Zisserman, <xref ref-type="bibr" rid="B34">2014</xref>) model made the 3 &#x000D7; 3 convolution filters mainstream and improved the accuracy based on the AlexNet, achieving the state-of-the-art results and accelerating further research on the use of deep visual representations in computer vision. Since 2014, the GoogLeNet series has begun to emerge. GoogLeNet (Christian et al., <xref ref-type="bibr" rid="B6">2014</xref>) proposed the inception convolutional neural network that started application of 1 &#x000D7; 1 convolution, reducing the amounts of computation and improving the utilization of computational resources. Its improved model GoogLeNet-V2 (Sergey and Christian, <xref ref-type="bibr" rid="B32">2015</xref>) replaces the 5 &#x000D7; 5 convolution with two layers of 3 &#x000D7; 3 convolution, and the batch normalization proposed in this paper is widely used in deep neural networks. The Inception-V3 model proposed by GoogLeNet-V3 (Christian et al., <xref ref-type="bibr" rid="B7">2015</xref>) achieved the state-of-the-art in the 2015 ILSVRC classification challenge. This model summarized four guidelines for network model design and three optimization tricks to effectively reduce the number of parameters and improve computational efficiency. GoogleNet-V4 (Szegedy et al., <xref ref-type="bibr" rid="B37">2017</xref>) proposed the Inception-V4 model to summarize and integrate. It designed the Inception-ResNet network, which introduces residual connections into the GoogleNet series. The ResNet (He et al., <xref ref-type="bibr" rid="B12">2016</xref>) model is designed with residual connection module, which is easier to optimize and has lower computational complexity and higher accuracy. The ResNeXt (Xie et al., <xref ref-type="bibr" rid="B41">2017</xref>) model improved based on the ResNet model which integrated the ideas of VGG, ResNet, and Inception series. ResNeXt proposed cardinality index to measure complexity, winning the runner-up in the 2016 ILSVRC classification challenge. The DenseNet (Huang et al., <xref ref-type="bibr" rid="B14">2017</xref>) model can effectively alleviate gradient disappearance and strengthen feature propagation, and therefore, it can outperform ResNet in all aspects on Cifar-10, SVHN, and Cifar-100 standard datasets. Besides, it can reduce the number of parameters by half with the same accuracy. The Senet (Jie et al., <xref ref-type="bibr" rid="B17">2017</xref>) model focuses on the channel relationship and introduces attention mechanism into the convolutional neural network, which can adaptively recalibrate channel-wise feature to improve performance.</p>
<p>Convolutional neural networks are also widely applied in speech recognition (Partha et al., <xref ref-type="bibr" rid="B26">2020</xref>; Pradeep and Nirmaladevi, <xref ref-type="bibr" rid="B27">2021</xref>; Yang et al., <xref ref-type="bibr" rid="B47">2021</xref>), medical diagnosis (Seo and Kim, <xref ref-type="bibr" rid="B31">2020</xref>; Toktam et al., <xref ref-type="bibr" rid="B39">2020</xref>; Mustaqeem, <xref ref-type="bibr" rid="B25">2021</xref>), biometrics (Alay, <xref ref-type="bibr" rid="B1">2020</xref>; Sadasivan et al., <xref ref-type="bibr" rid="B29">2020</xref>; Mekruksavanich and Jitpattanakul, <xref ref-type="bibr" rid="B23">2021</xref>; Mohaghegh and Payne, <xref ref-type="bibr" rid="B24">2021</xref>), and other fields. The exploration of convolutional neural networks in other fields provides a reference for us to use CNN for ship image recognition (Cazzato et al., <xref ref-type="bibr" rid="B4">2020</xref>). Jeon and Yang (<xref ref-type="bibr" rid="B15">2021</xref>) proposed a classification model integrating CNN and KNN, which classifies from dual-polarization data with 10-m pixel distance to improve the efficiency of ship classification. Li et al. (<xref ref-type="bibr" rid="B20">2021</xref>) contribute to intelligent ship vision system by traditionally integrating image processing with machine learning and using target detection method based on CNN. Julianto et al. (<xref ref-type="bibr" rid="B18">2020</xref>) proposed a modified method using CNN to classify a patrol vessel dataset and achieved great results. Chen et al. (<xref ref-type="bibr" rid="B5">2020</xref>) combined the improved generative adversarial network (GAN) and convolutional neural network, proposing a small ship detection method, which significantly improved the accuracy and robustness of the results. Zhao et al. (<xref ref-type="bibr" rid="B48">2020</xref>) used the improved AlexNet convolutional neural network for deep feature extraction of ship images, which significantly improved the classification performance. Xu et al. (<xref ref-type="bibr" rid="B42">2020</xref>) proposed a detection method combining visual saliency and a convolutional neural network cascade, which can effectively improve ship detection accuracy and efficiency. Based on the pixel-level fusion of visible and infrared bispectral images, Gao (<xref ref-type="bibr" rid="B11">2020</xref>) integrated the algorithms that include image preprocessing, image smoothing, and anti-cloud interference to achieve the detection of ship targets in complex land and sea backgrounds. Ren et al. (<xref ref-type="bibr" rid="B28">2019</xref>) proposed a CNN framework with fewer layers and parameter. This method used the Softmax function to classify different ship types and achieved good results. Li et al. (<xref ref-type="bibr" rid="B21">2019</xref>) proposed a method to learn discriminative features through supervised learning and good classification performance and built two small optical ship image datasets to verify that results of this method. Hou et al. (<xref ref-type="bibr" rid="B13">2019</xref>) proposed a ship detection method based on the visual attention enhanced network to process optical remote sensing images, which improved the recognition accuracy and the detection performance. Endang and Agfianto (<xref ref-type="bibr" rid="B9">2019</xref>) combined the CNN-ZFNet architecture and the random forest method to extract the so-called best features defined in the paper and improved the ship image classification accuracy. Shao et al. (<xref ref-type="bibr" rid="B33">2019</xref>) proposed a saliency-aware CNN framework for ship detection, comprising comprehensive ship discriminative features. Makantasis et al. (<xref ref-type="bibr" rid="B22">2016</xref>) proposed that a methodology fuses a visual attention method that exploits low-level image features appropriately selected for maritime environment. Dong et al. (<xref ref-type="bibr" rid="B8">2018</xref>) proposed a hierarchical detection method that finishes the automatic ship detection well. Sobral et al. (<xref ref-type="bibr" rid="B35">2015</xref>) proposed a double-constrained robust principal component analysis (RPCA) to improve the object foreground detection in maritime scenes.</p>
<p>By now, various methods to solve the classification and recognition problem of ship targets by convolutional neural networks do not focus on the differences between the grain of ship images. We found that the differences between categories of ships are the hull structure, and the differences between ships of the same category are the equipment and superstructures. Therefore, for the differences in the hull structure between different categories of ships, we should use convolution kernels with larger receptive fields, which is conducive to extracting features such as hull structures that occupy large pixel areas. At the same time, for small differences between diverse ships of the same category, we should use convolution kernels with smaller receptive fields to extract fine features such as ship equipment and superstructures.</p>
<p>Based on the above ideas, we proposed the SMS-PCNN model to classify ship images. The model extracts the feature of images in different sizes by parallelizing convolution branches with different receptive fields. To enhance the speed of converge and avoid overfitting, three optimization tricks are introduced including enhancing specific ship dataset applications, improving training learning rate, and avoiding overfitting. Then, we collected 5,062 ship images of 12 categories on the Internet for the experiments to test the algorithm performance of the SMS-PCNN model and the effectiveness of three optimization tricks.</p>
<p>An outline of the paper is as follows. Section Introduction briefly introduces the research background, technical status quo, faced challenges, and the innovations of this paper. Section SMS-PCNN Model gives a detailed introduction to SMS-PCNN model, which uses three parallel CNN branches to extract the features from different receptive fields, and then classifies the feature maps based on the multilayer perceptron neural networks. Moreover, three optimization tricks of SMS-PCNN model to optimize network performance are also introduced in this chapter. Section Experiment tests the performance of the proposed SMS-PCNN model using the ship image dataset. Besides, a series of comparison experiments were also conducted to test the performance of algorithms using other state-of-the-art approaches as references. Besides, a series of ablation experiments are also conducted to verify the effectiveness of three optimization tricks used in the model. In addition, we conduct experiments on Branch1-CNN model only using 3 &#x000D7; 3 convolution to demonstrate the necessity of designing a multi-branch parallel structure. By comparing the accuracy of both experiments, we demonstrate that designing of a multi-branch parallel structure has better feature extraction capabilities for ship images than other. Section Conclusion summarizes the content of the paper.</p></sec>
<sec id="s2">
<title>SMS-PCNN Model</title>
<p>This chapter presents the SMS-PCNN model proposed for the ship dataset and optimization tricks used by the model. SMS-PCNN uses three parallel CNN branches to extract feature maps from different receptive fields and then classifies the feature maps based on a multilayer perceptron neural network. This chapter is divided into two sections. The first section introduces the network architecture of the SMS-PCNN model in two levels. It introduces the basic single-branch architecture design for implementing the feature map extraction function in the SMS-PCNN model and then presents the SMS-PCNN model architecture design for integrating three single-branch CNNs with different receptive fields. The second section focuses on three optimization tricks for optimizing the network performance, including enhancing specific ship dataset applications, improving training learning rate, and avoiding overfitting.</p>
<sec>
<title>Model Architecture</title>
<p>This section is divided into two parts. The first part introduces the basic single-branch architecture design of SMS-PCNN. In this part, basic single-branch in parallel CNNs is used as an example, and the data processing procedure of RGB images after being inputting to each branch is introduced in a systematic way. In the second part, the overall network architecture of SMS-PCNN is presented. This part details the process of fusing feature maps from different single-branch CNNs and classifying the input images based on multilayer perceptron neural networks.</p>
<sec>
<title>Basic Single-Branch Architecture of SMS-PCNN Model</title>
<p>The SMS-PCNN model consists of three parallel CNN branches with different receptive fields. The three CNN branches have similar structures, so the first branch in the single-branch benchmark architecture network is taken as an example to illustrate the data flow after the image is inputted to the network. The first branch uses a 3 &#x000D7; 3 convolutional kernel with a relatively small receptive field. The schematic structure of this branch is shown as <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Network architecture of <italic>Branch1</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0001.tif"/>
</fig>
<p>The single-branch network design is inspired by the Resnet network (Xie et al., <xref ref-type="bibr" rid="B41">2017</xref>) and consists of three parts. The first part is the CNN <italic>Head</italic>, which includes the convolutional layer, (group normalization (GN) layer, ReLu activation layer, and max pooling layer sequentially. The second and third parts are the two layers with similar structure, each containing two blocks. Through two layers, the feature maps will be inputted to the global average pooling layer and then be flattened into a one-dimensional feature vector.</p>
<p>The convolutional layer of CNN <italic>Head</italic> uses 64 convolution kernels of size 3 &#x000D7; 3 with depth 3 to filter the image of size 512 &#x000D7; 512 &#x000D7; 3. The stride is 2 pixels and the padding is 3 pixels. Then, the image goes through the GN layer and ReLu activation layer. Then, max pooling is performed on the feature maps. The size of pooling is 3 &#x000D7; 3, the stride is 2 pixels, and the padding is 1 pixel.</p>
<p>The basic architecture of the block in <italic>Layer</italic> is shown as follows.</p>
<p>A formula can express the output of the block
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>x</italic> and <italic>y</italic> are, respectively, the input and output of the block. <italic>F</italic>(<italic>x</italic>) is the residual mapping which the network needs to learn, and <italic>x</italic>&#x02032; is the output from the side road. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the block contains two roads, the main road and the side road. After entering the block, <italic>x</italic> will be inputted to the main road and the side road for data processing, respectively. In the main road, feature map <italic>x</italic> goes through the first convolution layer, the first batch normalization (BN) layer, the first ReLu activation layer, the second convolution layer, and the second BN layer and then becomes feature map <italic>F</italic>(<italic>x</italic>). To complete the summation in Eqn (1), the dimensions of <italic>x</italic> and <italic>F</italic>(<italic>x</italic>) must be equal. In the s side road, if the dimensions of <italic>F</italic>(<italic>x</italic>) are the same of <italic>x</italic>, no further data processing will be performed, which is called <italic>x</italic>&#x02032; &#x0003D; <italic>x</italic>. However, if the dimensions of the feature map output by the main road decrease, x needs to be processed by 1 &#x000D7; 1 convolution layer. The stride of 1 &#x000D7; 1 convolution kernel in <xref ref-type="fig" rid="F2">Figure 2</xref> is 2 pixels, and the padding is 1 pixel. The number of channels of the 1 &#x000D7; 1 convolutional layer depends on the number of channels outputting <italic>F</italic>(<italic>x</italic>) from the main road in the block. <italic>F</italic>(<italic>x</italic>) from the main road and <italic>x</italic>&#x02032; from the side road are activated by ReLu function after the summation of the pixels at the corresponding positions. The feature map y is the final output of this block.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Architecture of <italic>Block</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0002.tif"/>
</fig>
<p>There are two cascaded blocks in each layer with the same structure. The difference between the two blocks is that the resolution of the output <italic>F</italic>(<italic>x</italic>) in <italic>Block1</italic> decreases and the number of channels changes. In <italic>Block2</italic>, the stride of the first convolution layer is 1 pixel. Therefore, the size of feature map and the number of channels do not change, and a 1 &#x000D7; 1 convolution operation is not required in <italic>Block2</italic>. <italic>Block1</italic> uses a convolution kernel with the stride of 2 pixels instead of a pooling layer to reduce the resolution of the feature map. Springenberg et al. (<xref ref-type="bibr" rid="B36">2014</xref>) show that max pooling layer can be replaced by a convolution layer with a stride bigger than 1 pixel without loss of accuracy. Moreover, this can reduce the error to some degree in vast majority of networks. Therefore, we use stride of 2 pixels to compress the size of the feature map to ensure that the network not only reduces the dimension and removes redundant information, but also improves the discriminative accuracy.</p>
<p><italic>Layer2</italic> has nearly the same design as of <italic>Layer1</italic>, except that <italic>Layer1</italic> has 128 convolutional channels and <italic>Layer2</italic> has 64 convolutional channels. Moreover, when the image is put into the single-branch benchmark network, it passes through the CNN <italic>Head</italic> and 2 layers one by one. Then, the network uses global average pooling to further reduce the resolution of the image and flatten the high-dimensional feature map into a one-dimensional feature vector, pending further classification processing.</p>
<p>Based on the above single-branch architecture network of <italic>Branch1</italic>, other two branches with similar structures can be designed, which are <italic>Branch2</italic> and <italic>Branch3</italic>. The main difference is that the size of the convolutional kernel used in <italic>Branch2</italic> is 5 &#x000D7; 5, and the size of the convolutional kernel used in <italic>Branch3</italic> is 7 &#x000D7; 7. Moreover, to make the feature map of each branch consistent in size for comparison and calculation, the padding of different branches is slightly different. <italic>Branch2</italic> uses 5 &#x000D7; 5 convolutional kernels, and the padding of the first block in each layer is 2 pixels. <italic>Branch3</italic> uses 7 &#x000D7; 7 convolutional kernels, the padding of the <italic>head</italic> convolutional layer is 4, and the padding of the first block in each layer is 3 pixels.</p></sec>
<sec>
<title>Overall Architecture of SMS-PCNN Model</title>
<p>Generally, ship images are natural scene images, so the input of SMS-PCNN model is RGB images. Analyzing the ship dataset, we found that the difference in appearance between different categories of ships mainly lies in the hull structure of ships, while the difference in appearance between ships of the same category lies in the superstructures and equipment of ships.</p>
<p>Therefore, for different categories of ships, we should pay more attention to discriminating architectural differences and use convolutional kernels with large receptive fields, which are more conducive to extracting features such as hull structures that occupy large pixel areas. For light differences between ships of the same category, we should use convolutional kernels with small receptive fields to extract features such as ship equipment and superstructures that occupy small pixel areas. For example, there are huge differences between the hull of the Ariake class destroyer and landing ship, which are two different categories of ships, so the convolution kernel with large receptive field is more advantageous. While aircraft carriers contain many categories of ships such as Queen Elizabeth class, the Nimitz class, the main difference between them is the bridge and other special equipment that occupy small pixel areas. In this case, the convolution kernel with smaller receptive field is more advantageous.</p>
<p>Combining the above two considerations, the SMS-PCNN model uses different sizes of receptive fields to extract features of ship images at different levels, so as to achieve a high classification accuracy. The SMS-PCNN model has three parallel branches with different sizes of convolutional kernels that is 3 &#x000D7; 3, 5 &#x000D7; 5, and 7 &#x000D7; 7. The receptive fields of the branches increase, so that the features of the ship images can be extracted and classified in a multi-scale manner. The network structure of SMS-PCNN model is shown in the <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>SMS-PCNN model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0003.tif"/>
</fig>
<p>The number of channels in the SMS-PCNN model has been adjusted two times. In the CNN <italic>Head</italic>, the convolution layer contains 64 channels. In <italic>Layer1</italic>, this number is expanded to 128. The reason is that the feature map shrinks and part of information loses during the process of extracting features from low-dimensional to high-dimensional features. To expand the amount of data flow between layers as much as possible, we expand the number of channels. However, the number of channels is reduced to 64 in <italic>Layer2</italic>. That is because the redundant information increases in the process of further abstraction of high-dimensional features. So, as to further improve the accuracy and reduce the redundant feature information, we decrease the number of channels.</p>
<p>After generated from three parallel branches with different sizes of receptive fields, feature vectors are linearly concatenated in the SMS-PCNN model.
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>F</mml:mi><mml:mi>u</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x000D7;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>q</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>q</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mn>3</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x000D7;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>q</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>]</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>p</italic> and <italic>q</italic> denote the dimensions of feature vectors. <italic>p</italic> is 192 and <italic>q</italic> is 64. <italic>Branch1</italic> denotes the branch which uses 3 &#x000D7; 3 convolution, <italic>Branch2</italic> denotes the branch which uses 5 &#x000D7; 5 convolution, and <italic>Branch3</italic> denotes the branch which uses 7 &#x000D7; 7 convolution.</p>
<p>These fused feature vectors are processed by multi-layer perceptron neural networks (MLP neural networks), mapping feature vectors with different sizes of receptive fields into different classifications of ship images by decreasing the dimension layer by layer.</p>
<p>Suppose after flatten layer, MLP with <italic>L</italic> layer is used for feature classification, then the layer <italic>l</italic> has <italic>m</italic><sup><italic>l</italic></sup> neurons and the input vector is
<disp-formula id="E4"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><italic>a</italic><sup><italic>l</italic>&#x02212;1</sup> is called the feature vector from layer (<italic>l</italic> &#x02212; 1). The matrix of weight connected between neurons in layer (<italic>l</italic> &#x02212; 1) and <italic>l</italic> is:
<disp-formula id="E5"><label>(4)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mtext>&#x02009;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>21</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>12</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd><mml:mo>&#x022EF;</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>22</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd><mml:mo>&#x022EF;</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:msup><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mn>1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:msup><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:mn>2</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mo>&#x022F1;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x022EF;</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mo>&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:msup><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msup><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mi>l</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the weights connected between neuron <italic>k</italic> in layer (<italic>l</italic> &#x02212; 1) and neuron <italic>j</italic> in layer <italic>l</italic>.</p>
<p>The bias vector of neurons in layer <italic>l</italic> is
<disp-formula id="E6"><label>(5)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
then the feature vector <italic>a</italic><sup><italic>l</italic></sup> from neurons in layer <italic>l</italic> is
<disp-formula id="E7"><label>(6)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02022;</mml:mo><mml:msup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where the <italic>f</italic><sup><italic>l</italic></sup>( ) is the activation function used by neurons in layer <italic>l</italic>. After the signal enters the feedforward neuron network, it passes layer by layer and gets the output <italic>a</italic><sup><italic>L</italic></sup> in the end. The whole process in the feedforward neural network is actually inputting the feature vectors from PCNN into the composite function <italic>F</italic>(.; <italic>W, b</italic>). Finally, the output of the network <italic>a</italic><sup><italic>L</italic></sup> is as follows.
<disp-formula id="E8"><label>(7)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>;</mml:mo><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02022;</mml:mo><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x02026;</mml:mo><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02022;</mml:mo><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02022;</mml:mo><mml:mi>x</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x02026;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
After the multilayer perceptron neural network, the SMS-PCNN model uses softmax function to normalize the vector <italic>a</italic><sup><italic>L</italic></sup>. Let the i-th element of the vector <italic>a</italic><sup><italic>L</italic></sup> be <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, then the result obtained after <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> being processed by softmax function is
<disp-formula id="E10"><label>(8)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
After softmax function, <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is mapped to the <italic>S</italic><sub><italic>i</italic></sub>.<italic>S</italic><sub><italic>i</italic></sub> satisfies the properties of the probability distribution, <italic>S</italic><sub><italic>i</italic></sub> &#x02208; (0, 1) and <inline-formula><mml:math id="M15"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Therefore, <italic>S</italic><sub><italic>i</italic></sub> can represent the prediction probability. By selecting the node with the largest prediction probability, we can obtain the final classification category.</p>
<p>The predicted probability <italic>S</italic><sub><italic>i</italic></sub> will also be inputted to loss function. In the SMS-PCNN model, we use cross-entropy loss function used. It is
<disp-formula id="E11"><label>(9)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>K</italic> is the number of categories, and <italic>p</italic><sub><italic>i</italic></sub> denotes the indicator variable (takes the value of 0 or 1). If the predicted category of the sample <italic>i</italic> is the same as its true category, <italic>p</italic><sub><italic>i</italic></sub> takes the value of 1 and 0 otherwise. Cross-entropy loss function can measure the prediction performance of the model.</p>
<p>To address the problem of too many parameters in the SMS-PCNN model, the SMS-PCNN model uses L2 regularization to mitigate overfitting. Regularization is a method to reduce the complexity of the model by introducing penalty terms. The L2 regularization is to add the L2 norm of the parameters after the loss function:
<disp-formula id="E12"><label>(10)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x003A9;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M18"><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:math></inline-formula> is cross-entropy loss function with L2 regularization, and <italic>w</italic> is the weight coefficient matrix, and the parameter &#x003B1; controls the strength of regularization. &#x003A9; function is the L2 norm, namely, the sum of squares of the weight coefficients:
<disp-formula id="E13"><label>(11)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mo>&#x003A9;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>w</mml:mi><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></sec></sec>
<sec>
<title>Tricks for Model Optimization</title>
<sec>
<title>Optimization for Specific Ship Dataset Applications</title>
<p>Generally, the distinguishing features of some ships are mainly concentrated in parts that account for a small proportion of the images, such as the bridge of aircraft carriers and the radar of frigates. Therefore, to improve the accuracy of the model, we need to retain as much information of images as possible and use images with large size to contain more information as much as possible. For images with large size, the batch size is reduced to fit the training. However, relatively low batch size will lead to bad performance of the batch normalization (BN) used in the SMS-PCNN model. In this case, Wu and Kaiming (<xref ref-type="bibr" rid="B40">2020</xref>) proposed group normalization (GN) to improve the dependence of BN on batch size and proved that when the batch size is &#x0003C;8, the error of using GN is smaller than BN.</p>
<p>The GN layer calculates the mean and variance based on the input data and then uses these two values to normalize the input data. Specifically, GN divides the channels of the layer into groups and uses the mean and variance of the data within the groups.
<disp-formula id="E14"><label>(12)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><inline-formula><mml:math id="M21"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> is the feature map of the layer, and <italic>i</italic> &#x0003D; (<italic>i</italic><sub><italic>N</italic></sub>, <italic>i</italic><sub><italic>C</italic></sub>, <italic>i</italic><sub><italic>H</italic></sub>, <italic>i</italic><sub><italic>w</italic></sub>). <italic>N</italic> is the batch axis, <italic>C</italic> is the channel axis, <italic>H</italic> and <italic>W</italic> are the spatial height and width axis. &#x003BC; and &#x003C3; here can be calculated by the following equations
<disp-formula id="E15"><label>(13)</label><mml:math id="M22"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E16"><label>(14)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003F5;</mml:mi></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where &#x003B5; is a small constant. <italic>S</italic><sub><italic>i</italic></sub> is a feature collection with a size of <italic>m</italic> used to calculate &#x003BC; and &#x003C3;. The solution of <italic>S</italic><sub><italic>i</italic></sub> in GN is the following equation.
<disp-formula id="E17"><label>(15)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x0230A;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfrac><mml:mo>&#x0230B;</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo>&#x0230A;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfrac><mml:mo>&#x0230B;</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><italic>G</italic> is the number of groups. In the SMS-PCNN model, we pre-set <italic>G</italic> = 2. <italic>C</italic> is the number of channels, and <italic>C/G</italic> is the number of channels in each group. &#x0201C;<inline-formula><mml:math id="M25"><mml:mrow><mml:mo>&#x0230A;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfrac><mml:mo>&#x0230B;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>&#x0230A;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfrac><mml:mo>&#x0230B;</mml:mo></mml:mrow></mml:math></inline-formula>&#x0201D; indicates that <italic>i</italic> and <italic>k</italic> are in the same channel group.</p>
<p>When we used GN in all the normalization layers of the SMS-PCNN models, gradient explosion appeared. Therefore, we only use GN to normalize the output of the head convolution layer of each branch whereas BN is used in other layers, which achieved good results.</p></sec>
<sec>
<title>Learning Rate Optimization</title>
<p>The model uses gradient descent to minimize the value of loss function and thus converge to the optimal solution. During the iterations performed by the gradient descent, the learning rate controls the learning progress of the model. To optimize the model learning rate, we use cosine annealing and Warmup (He et al., <xref ref-type="bibr" rid="B12">2016</xref>) technique. Warmup is used to train with a small learning rate at the beginning of to make the network familiar with the data. As the training continues, the learning rate slowly increases and reaches the initial learning rate within a pre-set iteration. Then, we use cosine annealing to adjust learning rate.</p>
<p>The equation of decreasing learning rate is as follows:
<disp-formula id="E18"><label>(16)</label><mml:math id="M26"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">cos</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mfrac><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where &#x003B7;<sub><italic>t</italic></sub> is the learning rate of the <italic>t</italic>-th iteration. &#x003B7;<sub>min</sub> is the termination learning rate. &#x003B7;<sub>max</sub> is the initial learning rate. <italic>T</italic><sub><italic>cur</italic></sub> is the current iteration. <italic>T</italic> is the total iteration. The key of the formula is <inline-formula><mml:math id="M27"><mml:mo class="qopname">cos</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mfrac><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. As iteration increases, the value of <inline-formula><mml:math id="M28"><mml:mo class="qopname">cos</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mfrac><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> decreases from 1 to &#x02212;1, so that the learning rate &#x003B7;<sub>max</sub> decreases to &#x003B7;<sub>min</sub>.</p>
<p>The learning rate adjustment for the whole model training is shown in the <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Curve of learning rate adjustment.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0004.tif"/>
</fig>
<p>The learning rate is adjusted with iterations. In the first 2,709 iterations (3 epochs), warmup is used. As iteration gradually increases, the learning rate grows to 0.1, and then, this number decreases to 10<sup>&#x02212;4</sup> using cosine annealing.</p></sec>
<sec>
<title>Optimization of Overfitting of the Ship Dataset</title>
<p>On the Internet, ship images are limited in amounts and in low quality, so the ship dataset we compiled from open-source data is limited in size, which leads to overfitting to some degrees when we use the SMS-PCNN model for training. For this reason, we use two methods, label smoothing and dropout, to mitigate overfitting in training.</p>
<p>Label smoothing is a technique proposed by Sergey and Christian (<xref ref-type="bibr" rid="B32">2015</xref>) in the model of inception-V2 to mitigate overfitting. The traditional one-hot coding suffers from overconfidence, which leads to overfitting. To address overconfidence, label smoothing decays the item with probability 1 in one-hot, which reduces the weight of category of the ground-truth label in the calculation of loss function. The confidence of the decayed items in one-hot is equally divided into other items, so that each category has a certain amount of confidence, which can suppress overfitting. In the SMS-PCNN model, we add the label smoothing to cross-entropy loss function, and loss function is changed as follows.
<disp-formula id="E19"><label>(17)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>K</italic> denotes the number of categories, and <italic>p</italic><sub><italic>i</italic></sub> denotes the indicator variable (takes the value of 0 or 1). <italic>q</italic><sub><italic>i</italic></sub> is the predicted probability value.
<disp-formula id="E21"><label>(18)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B5;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B5;</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where &#x003B5; is a smoothing parameter, which is a pre-set hyperparameter. <italic>u</italic>(<italic>k</italic>) is a probability distribution and uniform distribution is used here, so <inline-formula><mml:math id="M32"><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula>. The result of cross-entropy loss function through label smoothing <inline-formula><mml:math id="M33"><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is as follows.
<disp-formula id="E22"><label>(19)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B5;</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B5;</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
In addition to label smoothing, we also use dropout to prevent the model from overfitting. In dropout (Hou et al., <xref ref-type="bibr" rid="B13">2019</xref>), the training process randomly drops out some neurons in hidden layers and maintains the number of neurons in input and output layers. Then, we inputted data through the modified network to propagate forward, and the loss which is calculated by the network is propagated back. Suppose the multilayer perceptron neural network has L layers, and <italic>l</italic> &#x02208; {1 &#x02026;, <italic>L</italic>} denotes the neuron in layer <italic>l</italic>. Let <italic>z</italic><sup><italic>l</italic></sup> denotes the input of the neuron in layer <italic>l</italic>, <italic>y</italic><sup><italic>l</italic></sup> denotes the output of the neuron in layer <italic>l</italic>, and <italic>W</italic><sup><italic>l</italic></sup> and <italic>b</italic><sup><italic>l</italic></sup> denote the weight coefficients and bias of the neuron in layer <italic>l</italic>. Then, the feedforward neural network can be expressed by the following equation.
<disp-formula id="E23"><label>(20)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E24"><label>(21)</label><mml:math id="M36"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>f</italic> denotes the activation function. When we use dropout during the model training, the feedforward neural network becomes of the following form.
<disp-formula id="E25"><label>(22)</label><mml:math id="M37"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0007E;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E26"><label>(23)</label><mml:math id="M38"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>&#x01EF9;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>*</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E27"><label>(24)</label><mml:math id="M39"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mi>&#x01EF9;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E28"><label>(25)</label><mml:math id="M40"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where the probability <italic>p is a</italic> pre-set hyperparameter. <italic>r</italic><sup><italic>l</italic></sup> is a random variable that obeys the Bernoulli distribution and is designed to randomly generate a vector of 0 or 1 with probability <italic>p</italic>. Probability <italic>p</italic> denotes the proportion of the elements of this vector taking the value of 0. <italic>r</italic><sup><italic>l</italic></sup> is to sample each layer and multiply <italic>y</italic><sup>(<italic>l</italic>)</sup> which is the output of that layer element by element. <italic>y</italic><sup>(<italic>l</italic>&#x0002B;1)</sup> denotes the final result of Dropout.</p>
<p>During the prediction of the model, the weight parameter of each neural unit is multiplied by the probability <italic>p</italic>, <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>p</mml:mi><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>In the SMS-PCNN model, we used the dropout before two fully connected layers and set the probability <italic>p</italic> to be 0.3, which means that 30% of the neurons were dropped out randomly, and the experiment achieved a good accuracy improvement.</p></sec></sec></sec>
<sec id="s3">
<title>Experiment</title>
<sec>
<title>Dataset</title>
<p>On the Internet, ship images are limited in amounts and in low quality. Therefore, there are no large-scale standard open-source ship datasets. We collected open-source images from the Internet and compiled a dataset containing 12 categories of ship targets. Part of the ship images is shown in <xref ref-type="fig" rid="F5">Figure 5</xref>.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Ship dataset. <bold>(A)</bold> Type 054 frigate, <bold>(B)</bold> Type 055 destroyer, <bold>(C)</bold> Sovremenny class destroyer, <bold>(D)</bold> Chinese aircraft carrier Liaoning, <bold>(E)</bold> Nimitz class aircraft carrier, <bold>(F)</bold> Queen Elizabeth-class aircraft carrier, <bold>(G)</bold> Arleigh Burke class destroyer, <bold>(H)</bold> Landing ship, <bold>(I)</bold> Oil tanker, <bold>(J)</bold> Surveillance boat, <bold>(K)</bold> Missile boat, and <bold>(L)</bold> Motorboat.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0005.tif"/>
</fig>
<p>The number of each sample of 12 categories applicable to the image classification task is shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Distribution of ship dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Ship Category</bold></th>
<th valign="top" align="center"><bold>Number of images</bold></th>
<th valign="top" align="left"><bold>Ship category</bold></th>
<th valign="top" align="center"><bold>Number of images</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Surveillance boat</td>
<td valign="top" align="center">397</td>
<td valign="top" align="left">Type 055 destroyer</td>
<td valign="top" align="center">409</td>
</tr>
<tr>
<td valign="top" align="left">Motorboat</td>
<td valign="top" align="center">474</td>
<td valign="top" align="left">Arleigh Burke class destroyer</td>
<td valign="top" align="center">445</td>
</tr>
<tr>
<td valign="top" align="left">Landing ship</td>
<td valign="top" align="center">495</td>
<td valign="top" align="left">Type 054 frigate</td>
<td valign="top" align="center">404</td>
</tr>
<tr>
<td valign="top" align="left">Missile boat</td>
<td valign="top" align="center">403</td>
<td valign="top" align="left">Chinese aircraft carrier Liaoning</td>
<td valign="top" align="center">378</td>
</tr>
<tr>
<td valign="top" align="left">Oil tanker</td>
<td valign="top" align="center">436</td>
<td valign="top" align="left">Nimitz-class aircraft carriers</td>
<td valign="top" align="center">502</td>
</tr>
<tr>
<td valign="top" align="left">Sovremenny class destroyer</td>
<td valign="top" align="center">420</td>
<td valign="top" align="left">Queen Elizabeth-class aircraft carriers</td>
<td valign="top" align="center">299</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the ship dataset shown in <xref ref-type="table" rid="T1">Table 1</xref>, we divide the training set and valid set according to the ratio of 7:3.</p></sec>
<sec>
<title>Experiment and Result</title>
<sec>
<title>Experiment and Result of SMS-PCNN Model</title>
<p>We use the SMS-PCNN model optimized by three tricks to conduct experiments. Samples of 12 categories were trained iteratively using 175 epochs with the batch size of 4. The model was trained iteratively using the stochastic gradient descent (Sebastian, <xref ref-type="bibr" rid="B30">2016</xref>). Cosine annealing and warmup optimization techniques were used on the learning rate. After warmup, the learning rate is initially set to 0.1 and finally decreased to 10<sup>&#x02212;4</sup>. The weight of label smoothing is set to 10<sup>&#x02212;3</sup>. The SMS-PCNN model uses L2 regularization and its weight is set to 10<sup>&#x02212;3</sup>. MLP uses two dropouts, both set to discarding 30% of the neurons. The batch normalization, group normalization, and MLP in the model all use the correlation functions that come with PyTorch. After being preprocessed, the natural scene images with the adjusted resolution of 512 &#x000D7; 512 are inputted to the SMS-PCNN model for training.</p>
<p><xref ref-type="table" rid="T2">Table 2</xref> shows the architecture of SMS-PCNN model for a ship dataset with resolution of 512 &#x000D7; 512. Among them, the images are processed by <italic>Branch1, Branch2</italic>, and <italic>Branch3</italic> in parallel and then jointly processed by average pooling. Finally, the feature vectors from three branches are concatenated together. The concatenated feature vectors with 192 dimensions sequentially go through the input layer, the hidden layer with 64 neurons and the output layer with 12 neurons of the MLP neural network. Then, these vectors pass through the softmax function, and the results of classification are outputted.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Architectures of SMS-PCNN model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Layer name</bold></th>
<th valign="top" align="center"><bold>Output size</bold></th>
<th valign="top" align="center"><bold>Branch1</bold></th>
<th valign="top" align="center"><bold>Branch2</bold></th>
<th valign="top" align="center"><bold>Branch3</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Conv1</td>
<td valign="top" align="center">256 &#x000D7; 256</td>
<td valign="top" align="center">3 &#x000D7; 3, 64, stride = 2<break/> padding = 3</td>
<td valign="top" align="center">5 &#x000D7; 5, 64, stride = 2<break/> padding = 3</td>
<td valign="top" align="center">7 &#x000D7; 7, 64, stride = 2<break/> padding = 4</td>
</tr>
<tr>
<td valign="top" align="left">Pooling</td>
<td valign="top" align="center">129 &#x000D7; 129</td>
<td valign="top" align="center" colspan="3">3 &#x000D7; 3 max pooling, stride = 2, padding = 1</td>
</tr>
<tr>
<td valign="top" align="left">Layer1</td>
<td valign="top" align="center">65 &#x000D7; 65</td>
<td valign="top" align="center"><inline-formula><mml:math id="M42"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>*</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>*</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M43"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M44"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Layer2</td>
<td valign="top" align="center">33 &#x000D7; 33</td>
<td valign="top" align="center"><inline-formula><mml:math id="M45"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>*</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mo>*</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>128</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M46"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M47"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>5</mml:mn><mml:mo>*</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>64</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo class="qopname">padding</mml:mo><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">1 &#x000D7; 1</td>
<td valign="top" align="center" colspan="3">Average pooling</td>
</tr>
<tr>
<td valign="top" align="left" colspan="2">Fully connected layer1</td>
<td valign="top" align="center" colspan="3">(192, 64), dropout = 0.3</td>
</tr>
<tr>
<td valign="top" align="left" colspan="2">Fully connected layer2</td>
<td valign="top" align="center" colspan="3">(64, 12), dropout = 0.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The &#x0201C;64&#x0201D; in Conv1 denotes the number of convolutional kernels. The &#x0201C;128&#x0201D; in <italic>Layer1</italic> and 64 in <italic>Layer2</italic> both denote the number of convolutional kernels. The &#x0201C;&#x000D7;2&#x0201D; of each Branch in <italic>Layer1</italic> and <italic>Layer2</italic> denotes that the two cascaded blocks using the same parameters. Take the parameters of fully connected layer1 (192, 64) as an example, &#x0201C;192&#x0201D; indicates the dimension of the input vector and &#x0201C;64&#x0201D; the dimension of the output vector.</p>
<p>Generally, the indicator to evaluate the classification effectiveness of a classifier is its accuracy. The classification accuracy is the proportion of correctly predicted results to the total sample.</p>
<p>As seen in <xref ref-type="fig" rid="F6">Figure 6</xref>, the SMS-PCNN model, which combines all tricks, converges faster in training and has nearly reached the highest accuracy rate at the 100th epochs. With the increase of epochs, the learning rate gradually decreases and the training gradually tends to be stable, and the final model optimal results reach 84.79% accuracy.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Accuracy curve.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0006.tif"/>
</fig>
<p>The loss function in the SMS-PCNN model uses cross entropy loss function with label smoothing and is shown in the <xref ref-type="fig" rid="F7">Figure 7</xref>. The improved loss function converges faster and more stably in the training set. At the end of the training, the model training finally converges. The decrease of the loss function on the valid set tends to be stable as the learning rate decreases and reaches the optimal result around 150th epochs. The confusion matrix of the SMS-PCNN model is shown in the following figure.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Loss curve.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0007.tif"/>
</fig>
<p>From the confusion matrix of the SMS-PCNN model on the valid set as shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, three points can be concluded. (1) The darkest areas of the confusion matrix are concentrated in the diagonal line, showing great classification performance. (2) Among all categories, landing ships, type 055 destroyers, and oil tankers have the highest classification accuracy. (3) Sovremenny class destroyers and Nimitz class aircraft carriers cannot be classified well. The reason why these two categories have a slightly worse classification performance than others is that there are diverse ships of these two categories with minor differences, resulting in the lack of obvious different features between categories.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Confusion matrix of the valid set of SMS-PCNN model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0008.tif"/>
</fig>
<p>T-sne (Laurens, <xref ref-type="bibr" rid="B19">2014</xref>) clustering analysis of the SMS-PCNN model for the valid data is plotted in the <xref ref-type="fig" rid="F9">Figure 9</xref>.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>SMS-PCNN model t-sne clustering analysis.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0009.tif"/>
</fig>
<p>From the t-sne clustering analysis graph, we can find that the samples of each category are basically concentrated together, and the SMS-PCNN model can classify the samples well.</p></sec>
<sec>
<title>Performance Comparison</title>
<p>This section is divided into two parts for performance comparison. The first part is the performance comparison of the SMS-PCNN model with classical model frameworks and the model with single branch. The second part is the validation of the tricks of the SMS-PCNN model.</p>
<list list-type="simple">
<list-item><p>(1) Comparison to state-of-the-art approaches.</p></list-item>
</list>
<p>We compared our SMS-PCNN model with 4 state-of-the-art approaches: Resnet-50 (He et al., <xref ref-type="bibr" rid="B12">2016</xref>), Alexnet (Alex et al., <xref ref-type="bibr" rid="B2">2017</xref>), VGG-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B34">2014</xref>), and Resnet-18 (He et al., <xref ref-type="bibr" rid="B12">2016</xref>). All the methods were compared on the valid set of our ship dataset. The accuracy comparison is presented in <xref ref-type="table" rid="T3">Table 3</xref>. The results in <xref ref-type="table" rid="T3">Table 3</xref> show that the proposed SMS-PCNN model achieved the best results among all algorithms. All the state-of-the-art comparisons on the ship dataset are fine-tuned well. These comparisons use the same hyperparameters as the SMS-PCNN model we proposed.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Accuracy comparison with state-of-the-art approaches.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Resnet-50</td>
<td valign="top" align="center">63.82</td>
</tr>
<tr>
<td valign="top" align="left">Alexnet</td>
<td valign="top" align="center">64.23</td>
</tr>
<tr>
<td valign="top" align="left">VGG-16</td>
<td valign="top" align="center">67.52</td>
</tr>
<tr>
<td valign="top" align="left">Resnet-18</td>
<td valign="top" align="center">70.09</td>
</tr>
<tr>
<td valign="top" align="left">Branch1-CNN (3 &#x000D7; 3 conv)</td>
<td valign="top" align="center">75.53</td>
</tr>
<tr>
<td valign="top" align="left">SMS-PCNN model</td>
<td valign="top" align="center">84.79</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, we conducted experiments on network architecture with single branch to analyze the classification of ships. The experiments showed that the Branch1-CNN with small receptive field of 3 &#x000D7; 3 convolutional kernel achieved an accuracy of 75.53%, which is lower than the accuracy of the SMS-PCNN model. This demonstrates the necessity of designing a multi-branch structure.</p>
<p>Comparing with the confusion matrix and t-sne clustering analysis graph in <xref ref-type="fig" rid="F10">Figures 10</xref>, <xref ref-type="fig" rid="F11">11</xref>, it can be clearly concluded that compared with the SMS-PCNN model, other models do not have well classification effect and have limitations in generalization performance. The confusion matrix shows that other models have low accuracy in discriminating the Queen Elizabeth-class aircraft carrier and the Chinese aircraft carrier Liaoning. The t-sne clustering analysis shows that the point representing Queen Elizabeth-class aircraft carrier and the Chinese aircraft carrier Liaoning are scattered and disorganized, showing a bad performance on clustering. The results of confusion matrix and t-sne clustering analysis are consistent and unified.</p>
<list list-type="simple">
<list-item><p>(2) Validation of tricks of SMS-PCNN model.</p></list-item>
</list>
<p>Basic SMS-PCNN means the SMS-PCNN model we used to perform an experiment without trick1, trick2, and trick3. These tricks are used for optimizing the network, including enhancing specific ship dataset applications, improving training learning rate and avoiding overfitting, respectively.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Confusion matrix of state-of-the-art approaches and Branch1-CNN in valid dataset. <bold>(A)</bold> Resnet-50, <bold>(B)</bold> Alexnet, <bold>(C)</bold> VGG-16, <bold>(D)</bold> Resnet-18, <bold>(E)</bold> Branch1-CNN (3 &#x000D7; 3 conv), <bold>(F)</bold> SMS-PCNN model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0010.tif"/>
</fig>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>T-sne clustering analysis of state-of-the-art approaches in valid dataset. <bold>(A)</bold> Resnet-50, <bold>(B)</bold> Alexnet, <bold>(C)</bold> VGG-16, <bold>(D)</bold> Resnet-18, <bold>(E)</bold> Branch1-CNN (3 &#x000D7; 3 conv), <bold>(F)</bold> SMS-PCNN model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0011.tif"/>
</fig>
<p>We designed the ablation experiments. The result is shown in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Ablation experiments of SMS-PCNN model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Basic SMS-PCNN</td>
<td valign="top" align="center">77.28</td>
</tr>
<tr>
<td valign="top" align="left">Basic SMS-PCNN(add trick1)</td>
<td valign="top" align="center">79.44</td>
</tr>
<tr>
<td valign="top" align="left">Basic SMS-PCNN(add trick1 and trick2)</td>
<td valign="top" align="center">83.24</td>
</tr>
<tr>
<td valign="top" align="left">SMS-PCNN(add trick1-3)</td>
<td valign="top" align="center">84.79</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It demonstrates that the accuracy of Basic SMS-PCNN model without any tricks is 77.28%, which is still higher than the other state-of-the-art approaches and Branch1-CNN in valid dataset mentioned above. This demonstrates the superiority of the model design principles and the basic framework structure. Meanwhile, the ablation experiments examined that the accuracy of the Basic SMS-PCNN model with trick1 increases by 2.16&#x02013;79.44%. For the Basic SMS-PCNN (add trick1 and trick2), the accuracy increases by 5.96&#x02013;83.79%. When using trick1, trick2, and trick3, the model accuracy increases by 7.51&#x02013;84.79%. It can be seen that the use of three tricks did improve the accuracy of SMS-PCNN model and enhance the classification effect.</p></sec>
<sec>
<title>Mechanistic Analysis of the Network</title>
<p>To further analyze the mechanism of the SMS-PCNN model and verify the effectiveness of parallelizing multi-scale convolutional branch and adjusting the number of channels, we visualize the convolutional kernels and the feature maps of tanker image through each convolutional module. The head convolution layer of each Branch in the SMS-PCNN model is a convolutional kernel with depth 3 designed for three channels of RGB. When visualizing, we combine these convolutional kernels into single convolutional kernel with color. The visualization of the 64 combined convolutional kernels of the head convolutional layer of each branch of the SMS-PCNN model is shown in the following figure.</p>
<p>From the visualization of the convolution kernels in <xref ref-type="fig" rid="F12">Figure 12</xref>, we can find that the head convolution layer of <italic>Branch3</italic> with 7 &#x000D7; 7 size convolution kernels mainly focuses on the edge features of images, so it can better extract features such as the hull of ships. The head convolution layers of <italic>Branch2</italic> and <italic>Branch1</italic> with 5 &#x000D7; 5 and 3 &#x000D7; 3 size convolution kernels focus on the fine-grained features of images, so it can better extract features such as ship superstructures and equipment.</p>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p>Convolutional kernel visualization. <bold>(A)</bold> Branch1_Head(3 &#x000D7; 3conv), <bold>(B)</bold> Branch2_Head(5 &#x000D7; 5 conv), <bold>(C)</bold> Branch3_Head(7 &#x000D7; 7 conv).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0012.tif"/>
</fig>
<p>We take the tanker image as an example to visualize feature maps that pass through each branch convolution module. In each branch, the tanker image passes through the CNN <italic>Head, Layer1</italic>, and <italic>Layer2</italic>, and the feature maps are shown in the following figure.</p>
<p>From <xref ref-type="fig" rid="F13">Figures 13A&#x02013;C</xref>, we can see that after the head convolution layer with different convolution kernels in three branches, the feature maps of the tanker image have different attention. The feature map from Branch3_Head with a larger receptive field is obviously coarser than that of Branch2_Head and Branch1_Head, which have smaller receptive fields. It shows that the feature map from Branch3_Head is more concerned with coarse-grained feature extraction, whereas Branch1_Head with the smallest receptive field is clearer in detail and is obviously more concerned with the fine-grained feature extraction.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>Feature maps of oil tanker image after different convolution layers. <bold>(A)</bold> Branch1_Head(3 &#x000D7; 3 conv), <bold>(B)</bold> Branch2_Head (5 &#x000D7; 5 conv), <bold>(C)</bold> Branch3_Head(7 &#x000D7; 7 conv), <bold>(D)</bold> Branch1_layer1 (3 &#x000D7; 3 conv), <bold>(E)</bold> Branch2_layer1 (5 &#x000D7; 5 conv), <bold>(F)</bold> Branch3_layer1(7 &#x000D7; 7 conv), <bold>(G)</bold> Branch1_layer2(3 &#x000D7; 3 conv), <bold>(H)</bold> Branch2_layer2 (5 &#x000D7; 5 conv), <bold>(I)</bold> Branch3_layer2(7 &#x000D7; 7 conv).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-889308-g0013.tif"/>
</fig>
<p><xref ref-type="fig" rid="F13">Figures 13D&#x02013;F</xref> show the feature maps of the tanker image after processing by the layer1 of each branch. Branch1_layer1 obviously retains more details and pays more attention to the fine-grained features of the ship. Branch1_layer2 has fewer details than Branch1_layer1 and focuses more on the moderate scale features such as the superstructures of the tanker. Branch1_layer3 focuses on the contours of the ship and the coarse-grained features of the hull structure of ships.</p>
<p><xref ref-type="fig" rid="F13">Figures 13G&#x02013;I</xref> shows the feature maps of the tanker image after the layer2 of each branch. It can be clearly seen that there are huge differences between the feature maps after all the convolution layers of each branch. The feature map of <italic>Branch1</italic> is clear and focuses on fine-grained features such as ship equipment. The feature map of <italic>Branch2</italic> focuses on the superstructures of the tanker. The feature map of <italic>Branch3</italic> pays attention to contours and the hull structure of the tanker.</p>
<p>By comparing the feature maps of different convolution of the three branches, we demonstrate the effectiveness of using multi-scale convolutional branch parallelization for ship images.</p>
<p>We also compared the feature maps of each convolutional layer of single branch. Taking <xref ref-type="fig" rid="F13">Figures 13A,D,G</xref> as an example, after the processing of CNN <italic>Head, Layer1</italic>, and <italic>Layer2</italic>, feature maps are gradually abstracted from specific low dimension to high dimension, which is effective for image feature extraction. Besides, from <xref ref-type="fig" rid="F13">Figures 13D&#x02013;F</xref>, we observe that there is redundant clutter in part of the output from <italic>Layer1</italic> of each branch. Therefore, the reduced numbers of channels in <italic>Layer2</italic> can filter redundant feature maps from the previous layer. This verified the effectiveness of changing numbers of channels in the network.</p></sec></sec></sec>
<sec sec-type="conclusions" id="s4">
<title>Conclusion</title>
<p>We found that the differences between different categories of ships are mainly focused on the hull structures, whereas the differences between different ships of the same category are concentrated in parts such as ship equipment and superstructures. Therefore, we used convolutional kernels with larger receptive fields to extract coarse-grained features such as the hull structure of different categories of ships and used convolutional kernels with smaller receptive fields to extract fine-grained features of different ships of the same category. In this paper, we proposed the SMS-PCNN model for the recognition and classification of ship images. The SMS-PCNN model extracted features of image in different sizes by parallelizing convolutional branches with different receptive fields. In addition, the results of experiments showed that the SMS-PCNN model achieved 84.79% accuracy on a ship dataset containing samples of 12 ship categories. We also conducted a series of comparison experiments to test the performance of algorithms using other state-of-the-art approaches. The results showed that comparing with the Resnet-50, Alexnet, VGG-16, and Resnet-18 models, the SMS-PCNN model improved the accuracy by 14.7&#x02013;20.97%. We also experimented with the Branch1-CNN model, and the results showed that the SMS-PCNN model with parallel structure improved the accuracy by 9.26% compared with the single-branch model, thus proving that the multi-branch parallel structure outperformed the single-branch for the ship image feature extraction.</p>
<p>The SMS-PCNN model uses three tricks to improve the performance, including enhancing specific ship dataset applications, improving training learning rate, and avoiding overfitting. The experiments showed that the model achieves an accuracy of 77.28% without using any tricks, which is 7.19% higher than Resnet-18. This accuracy is the highest among state-of-the-art approaches. It proved the superiority of the model design principles and the basic framework structure. To verify the effect of optimization tricks on the SMS-PCNN model, we performed a series of ablation experiments. When trick1 was added on the basic SMS-PCNN model, the accuracy increased by 2.16%. When adding trick2 to the SMS-PCNN model with trick1, the accuracy increased by 3.8%. Finally, when adding trick3 to the SMS-PCNN model with trick1 and trick2, the accuracy increased by 1.55%. The ablation experiments validated the effectiveness of the three tricks we used for the ship dataset.</p></sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p></sec>
<sec id="s6">
<title>Author Contributions</title>
<p>FW designed research, performed research, analyzed data, and wrote the paper. HL supervised the entire research process. YZ was involved in writing part of the manuscript. QX and RZ were involved in collected the data. All authors contributed to the article and approved the submitted version.</p></sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This study was funded by Naval University of Engineering, College of Electric Engineering.</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alay</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>AlBaity HH. Deep learning approach for multimodal biometric recognition system based on fusion of Iris, Face, and Finger Vein Traits</article-title>. <source>Sensors</source>. <volume>20</volume>, <fpage>5523</fpage>. <pub-id pub-id-type="doi">10.3390/s20195523</pub-id><pub-id pub-id-type="pmid">32992524</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alex</surname> <given-names>K.</given-names></name> <name><surname>Ilya</surname> <given-names>S.</given-names></name> <name><surname>Geoffrey</surname> <given-names>E. H.</given-names></name></person-group> (<year>2017</year>). <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Commun. ACM.</source> <volume>60</volume>, <fpage>84</fpage>&#x02013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1145/3065386</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atsuto</surname> <given-names>M.</given-names></name> <name><surname>Kazuhiro</surname> <given-names>F. K.</given-names></name></person-group> (<year>2004</year>). <article-title>Ship identification in sequential ISAR imagery</article-title>. <source>Mach. Vis. Appl.</source> <volume>15</volume>, <fpage>149</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-004-0140-y</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cazzato</surname> <given-names>D.</given-names></name> <name><surname>Cimarelli</surname> <given-names>C.</given-names></name> <name><surname>Sanchez-Lopez</surname> <given-names>J. L.</given-names></name> <name><surname>Voos</surname> <given-names>H.</given-names></name> <name><surname>Leo</surname> <given-names>M. A.</given-names></name></person-group> (<year>2020</year>). <article-title>Survey of Computer Vision Methods for 2D Object Detection from Unmanned Aerial Vehicles</article-title>. <source>J. Imag.</source> <volume>6</volume>, <fpage>78</fpage>. <pub-id pub-id-type="doi">10.3390/jimaging6080078</pub-id><pub-id pub-id-type="pmid">34460693</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Z. J.</given-names></name> <name><surname>Chen</surname> <given-names>D. P.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. S.</given-names></name> <name><surname>Cheng</surname> <given-names>X. Z.</given-names></name> <name><surname>Zhang</surname> <given-names>M. Y.</given-names></name> <name><surname>Wu</surname> <given-names>C. Z.</given-names></name></person-group> <article-title>Deep learning for autonomous ship-oriented small ship detection</article-title>. <source>Saf. Sci</source>. (<year>2020</year>). <pub-id pub-id-type="doi">10.1016./j.ssci.2020.104812</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Christian</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Jia</surname> <given-names>Y. Q.</given-names></name> <name><surname>Pierre</surname> <given-names>S.</given-names></name> <name><surname>Scott</surname> <given-names>E. R.</given-names></name> <name><surname>Dragomir</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Going deeper with convolutions</article-title>. <source>IEICE Trans Fundam Electron Commun Comput Sci</source>, abs/1409,4842.</citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Christian</surname> <given-names>S.</given-names></name> <name><surname>Vincent</surname> <given-names>V.</given-names></name> <name><surname>Sergey</surname> <given-names>I.</given-names></name> <name><surname>Zbigniewm</surname> <given-names>W.</given-names></name></person-group> (<year>2015</year>). <article-title>Rethinking the inception architecture for computer vision</article-title>. <source>IEICE Trans Fundam Electron Commun Comput Sci</source>, abs/1512, 00567.</citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Xu</surname> <given-names>F.</given-names></name></person-group> (<year>2018</year>). <article-title>Ship Detection in Optical remote sensor Images Based on Saliency and a Rotation-Invariant Descriptor</article-title>. <source>Rem. Sens.</source> <volume>18</volume>, <fpage>400</fpage>. <pub-id pub-id-type="doi">10.3390/rs10030400</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Endang</surname> <given-names>A.</given-names></name> <name><surname>Agfianto</surname> <given-names>E. P.</given-names></name></person-group> (<year>2019</year>). <article-title>Ship Identification on Satellite Image Using Convolutional Neural Network and Random Forest</article-title>. <source>IJCCS (Indones. J. Comput. Cybern. Syst).</source> <volume>13</volume>, <fpage>117</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.22146/ijccs.37461</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Enriquez de Luna</surname> <given-names>C.</given-names></name> <name><surname>Otaduy</surname> <given-names>M. D.</given-names></name> <name><surname>Dorronsoro</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title>A decision support system for ship identification based on the curvature scale space representation</article-title>. <source>Proc. SPIE</source>. <volume>5988</volume>, <fpage>171</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1117/12.630532</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>X. B.</given-names></name></person-group> (<year>2020</year>). <article-title>Design and implementation of marine automatic target recognition system based on visible rem sensor images</article-title>. <source>J. Coast Res.</source> <volume>115</volume>, <fpage>277</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.2112/JCR-SI115-088.1</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. In: <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>. p. <fpage>770</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>B. F.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Ship</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Ship detection for optical remote sensor images based on visual attention enhanced network</article-title>. <source>Sensors (Basel)</source>. <volume>19</volume>, <fpage>2271</fpage>. <pub-id pub-id-type="doi">10.3390/s19102271</pub-id><pub-id pub-id-type="pmid">31100909</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2017</year>). <article-title>Densely connected convolutional networks</article-title>. <source>CVPR 4700&#x02013;4708.</source> <volume>4</volume>, <fpage>5</fpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jeon</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>C. S.</given-names></name></person-group> (<year>2021</year>). <article-title>Enhancement of Ship Type Classification from a Combination of CNN and KNN</article-title>. <source>Electronics</source>. <volume>10</volume>, <fpage>1169</fpage>. <pub-id pub-id-type="doi">10.3390/electronics10101169</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J. H.</given-names></name> <name><surname>Fu</surname> <given-names>X. J.</given-names></name> <name><surname>Qin</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>X. Y.</given-names></name> <name><surname>Ma</surname> <given-names>Z. F.</given-names></name></person-group> (<year>2021</year>). <article-title>High-Speed Lightweight Ship Detection Algorithm Based on YOLO-V4 for Three-Channels RGB SAR Image</article-title>. <source>Rem. Sens</source>. <volume>13</volume>, <fpage>1909</fpage>. <pub-id pub-id-type="doi">10.3390/rs13101909</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jie</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Gang</surname> <given-names>S.</given-names></name> <name><surname>Albanie</surname> <given-names>S. S.</given-names></name></person-group> (<year>2017</year>). <article-title>Squeeze-and-excitation networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Julianto</surname> <given-names>E.</given-names></name> <name><surname>Khumaidi</surname> <given-names>A.</given-names></name> <name><surname>Priyonggo</surname> <given-names>P.</given-names></name> <name><surname>Munadhif</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>Object recognition on patrol ship using image processing and convolutional neural network (CNN)</article-title>. <source>J. Phys. Conf. Ser</source>. <volume>1450</volume>, <fpage>012081</fpage>. <pub-id pub-id-type="doi">10.1088./1742-6596/1450/1/012081</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laurens</surname> <given-names>V. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Accelerating t-SNE using tree-based algorithms</article-title>. <source>J. Mach. Learn. Res.</source> <volume>15</volume>, <fpage>3221</fpage>&#x02013;<lpage>45</lpage>.</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J. H.</given-names></name> <name><surname>Ruan</surname> <given-names>J. C.</given-names></name> <name><surname>Xie</surname> <given-names>Y. Q.</given-names></name></person-group> (<year>2021</year>). <article-title>Research on the development of object detection algorithm in the field of ship target recognition[J]</article-title>. <source>Int. Core J. Eng</source>. <volume>7</volume>, <fpage>233</fpage>&#x02013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.6919/ICJE.202101_7(1).0031</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zhao</surname> <given-names>B.</given-names></name> <name><surname>Tang</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Feng</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>Ship classification based on convolutional neural networks</article-title>. <source>J. Eng</source>. <volume>2019</volume>, <fpage>7343</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1049/joe.2019.0422</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Makantasis</surname> <given-names>K.</given-names></name> <name><surname>Protopapadakis</surname> <given-names>E.</given-names></name> <name><surname>Doulamis</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Semi-supervised vision-based maritime surveillance system using fused visual attention maps</article-title>. <source>Multimed Tools Appl.</source> <volume>75</volume>, <fpage>15051</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-015-2512-x</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mekruksavanich</surname> <given-names>S.</given-names></name> <name><surname>Jitpattanakul</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Biometric user identification based on human activity recognition using wearable sensors: an experiment using deep learning models</article-title>. <source>Electronics</source>. <volume>10</volume>, <fpage>308</fpage>. <pub-id pub-id-type="doi">10.3390/electronics10030308</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohaghegh</surname> <given-names>M.</given-names></name> <name><surname>Payne</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Automated biometric identification using dorsal hand images and convolutional neural networks</article-title>. <source>J Phys Conf Ser</source>. <volume>1880</volume>, <fpage>012014</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/1880/1/012014</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mustaqeem</surname> <given-names>K. S.</given-names></name></person-group> (<year>2021</year>). <article-title>MLT-DNet: Speech emotion recognition using 1D dilated CNN based on multi-learning trick approach</article-title>. <source>Exp. Syst. Appl.</source> <fpage>167</fpage>. <pub-id pub-id-type="doi">10.1016./j.eswa.2020.114177</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Partha</surname> <given-names>P. B.</given-names></name> <name><surname>Rappy</surname> <given-names>S.</given-names></name> <name><surname>Ki-Doo</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>An automatic nucleus segmentation and CNN model based classification method of white blood cell</article-title>. <source>Exp. Syst. Appl</source>. <volume>149</volume>, <fpage>113211</fpage>.</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pradeep</surname> <given-names>S.</given-names></name> <name><surname>Nirmaladevi</surname> <given-names>P.</given-names></name></person-group> (<year>2021</year>). <article-title>A review on speckle noise reduction techniques in ultrasound medical images based on spatial domain, transform domain and CNN methods</article-title>. <source>IOP Conf Ser Mat Sci Eng</source>. <volume>1055</volume>, <fpage>012116</fpage>. <pub-id pub-id-type="doi">10.1088/1757-899X/1055/1/012116</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Guo</surname> <given-names>Z.</given-names></name></person-group> (<year>2019</year>). <article-title>Multi-feature fusion with convolutional neural network for ship classification in optical images</article-title>. <source>Appl. Sci</source>. <volume>9</volume>, <fpage>4209</fpage>. <pub-id pub-id-type="doi">10.3390/app9204209</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sadasivan</surname> <given-names>S.</given-names></name> <name><surname>Sivakumar</surname> <given-names>T. T.</given-names></name> <name><surname>Joseph</surname> <given-names>A. P.</given-names></name> <name><surname>Zacharias</surname> <given-names>G. C.</given-names></name> <name><surname>Nair</surname> <given-names>M. S.</given-names></name> <name><surname>Thampi</surname> <given-names>S. M.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Tongue print identification using deep CNN for forensic analysis</article-title>. <source>J Intell Fuzzy Syst</source>. <volume>38</volume>, <fpage>6415</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.3233/JIFS-179722</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sebastian</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>An overview of gradient descent optimization algorithms</article-title>. <source>IEICE Transac Fundam Electron Commun Comput Sci</source>,, abs/1609,04747. <pub-id pub-id-type="pmid">26186171</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seo</surname> <given-names>M.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Fusing visual attention CNN and bag of visual words for cross-corpus speech emotion recognition</article-title>. <source>Sensors</source>. <volume>20</volume>, <fpage>5559</fpage>. <pub-id pub-id-type="doi">10.3390/s20195559</pub-id><pub-id pub-id-type="pmid">32998382</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sergey</surname> <given-names>I.</given-names></name> <name><surname>Christian</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Batch normalization: accelerating deep network training by reducing internal covariate shift</article-title>. <source>IEICE Trans Fundam Electron Commun Comput Sci</source>, abs/1502,03167.</citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shao</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Du</surname> <given-names>W.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>Saliency-aware convolution neural network for ship detection in surveillance video</article-title>. <source>IEEE Transac. Circ. Syst. Video Technol</source>. <volume>30</volume>, <fpage>781</fpage>&#x02013;<lpage>94</lpage>.</citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv preprint arXiv:</source>1409,1556.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sobral</surname> <given-names>A.</given-names></name> <name><surname>Bouwmans</surname> <given-names>T.</given-names></name> <name><surname>Zahzah</surname> <given-names>E. D.</given-names></name></person-group> <article-title>Based onSaliency Maps for Foreground Detection in Automated Maritime Surveillance</article-title>. <source>ISBC2015 Workshop Conjunct AVSS 2015, Karlsruhe, Germany</source>. (<year>2015</year>) <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109./AVSS.2015.7301753</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Springenberg</surname> <given-names>J. T.</given-names></name> <name><surname>Dosovitskiy</surname> <given-names>A.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name> <name><surname>Riedmiller</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412, 6806</article-title>. <italic>and ICLR</italic>. 2015.</citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Vanhoucke</surname> <given-names>V.</given-names></name></person-group> (<year>2017</year>). <article-title>Inception-v4, inception-resnet and the impact of residual connections on learning</article-title>. In: <source>Thirty-First AAAI Conference on Artificial Intelligence, San Francisco</source>, <fpage>4278</fpage>&#x02013;<lpage>84</lpage>.</citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>G.</given-names></name> <name><surname>Zhuge</surname> <given-names>Y. C.</given-names></name> <name><surname>Claramunt</surname> <given-names>C.</given-names></name> <name><surname>Men</surname> <given-names>S. Y.</given-names></name></person-group> (<year>2021</year>). <article-title>N-YOLO: A SAR ship detection using noise-classifying and complete-target extraction</article-title>. <source>Rem. Sens</source>. <volume>13</volume>, <fpage>871</fpage>. <pub-id pub-id-type="doi">10.3390/rs13050871</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toktam</surname> <given-names>Z.</given-names></name> <name><surname>Mohammad</surname> <given-names>M. H.</given-names></name> <name><surname>Mahmood</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Adaptive windows multiple deep residual networks for speech recognition</article-title>. <source>Exp Syst Appl</source>. <volume>139</volume>, <fpage>112840</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2019.112840</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y. X.</given-names></name> <name><surname>Kaiming</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Group Normalization</article-title>. <source>Int. J. Comput. Vis.</source> <volume>128</volume>, <fpage>742</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-019-01198-w</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>S. R.</given-names></name> <name><surname>Girshick</surname> <given-names>P.</given-names></name> <name><surname>Dollar</surname></name> <name><surname>Tu</surname> <given-names>Z</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name></person-group> (<year>2017</year>). <article-title>Aggregated residual transformations for deep neural networks</article-title>. <source>CVPR</source>. <volume>1</volume>, <fpage>3</fpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.634</pub-id><pub-id pub-id-type="pmid">31141794</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Yin</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Han</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Fast ship detection combining visual saliency and a cascade CNN in SAR images</article-title>. <source>IET Radar Sonar Navigation</source>. <volume>14</volume>, <fpage>1879</fpage>&#x02013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1049/iet-rsn.2020.0113</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>F.</given-names></name> <name><surname>Liu</surname> <given-names>J. H.</given-names></name> <name><surname>Dong</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Ship detection in optical remote sensing images based on wavelet transform and multi-level false alarm identification</article-title>. <source>Rem Sens</source>. <volume>9</volume>, <fpage>985</fpage>. <pub-id pub-id-type="doi">10.3390/rs9100985</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Kim</surname> <given-names>T</given-names></name></person-group>. <article-title>Integration of SAR AIS for ship detection identification Ocean Sensing Monitoring</article-title>. <source>Proc. SPIE.</source> (<year>2012</year>). <volume>8372</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage> <pub-id pub-id-type="doi">10.1117./12.920359</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Q. Z.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>Ship detection from optical satellite images based on saliency segmentation and structure-LBP feature</article-title>. <source>IEEE Geosci. Rem. Sens. Lett.</source> <volume>14</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2017.2664118</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Ji</surname> <given-names>S.</given-names></name> <name><surname>Gao</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Q.</given-names></name></person-group> (<year>2014</year>). <article-title>Ship detection from optical satellite images based on sea surface analysis</article-title>. <source>IEEE Geosci. Rem. Sens. Lett.</source> <volume>11</volume>, <fpage>641</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2013.2273552</pub-id><pub-id pub-id-type="pmid">34844147</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>Y. Y.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Two-stage selective ensemble of CNN via deep tree training for medical image classification</article-title>. <source>IEEE Trans. Cybern</source>. <pub-id pub-id-type="doi">10.1109./TCYB.2021.3061147</pub-id><pub-id pub-id-type="pmid">33705343</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Lang</surname></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Optical remote sensor ship image classification based on deep feature combined distance metric learning</article-title>. <source>J. Coast Res</source>. <volume>102</volume>, <fpage>82</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.2112/SI102-011.1</pub-id></citation></ref>
</ref-list>
</back>
</article>