<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Phys.</journal-id>
<journal-title>Frontiers in Physics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Phys.</abbrev-journal-title>
<issn pub-type="epub">2296-424X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1208781</article-id>
<article-id pub-id-type="doi">10.3389/fphy.2023.1208781</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Physics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Multi-level semantic information guided image generation for few-shot steel surface defect classification</article-title>
<alt-title alt-title-type="left-running-head">Hao et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphy.2023.1208781">10.3389/fphy.2023.1208781</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Hao</surname>
<given-names>Liang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2286883/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shen</surname>
<given-names>Pei</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Pan</surname>
<given-names>Zhiwei</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Xu</surname>
<given-names>Yong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Computer Science and Technology</institution>, <institution>Harbin Institute of Technology</institution>, <addr-line>Shenzhen</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>HBIS Digital Technology Co., Ltd.</institution>, <addr-line>Shijiazhuang</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1664211/overview">Guanqiu Qi</ext-link>, Buffalo State College, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2140551/overview">Jian Sun</ext-link>, Southwest University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1773018/overview">Zhiqin Zhu</ext-link>, Chongqing University of Posts and Telecommunications, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2290657/overview">Yong Li</ext-link>, Chongqing University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Yong Xu, <email>yongxu@ymail.com</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>30</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1208781</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>05</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Hao, Shen, Pan and Xu.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Hao, Shen, Pan and Xu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Surface defect classification is one of key points in the field of steel manufacturing. It remains challenging primarily due to the rare occurrence of defect samples and the similarity between different defects. In this paper, a multi-level semantic method based on residual adversarial learning with Wasserstein divergence is proposed to realize sample augmentation and automatic classification of various defects simultaneously. Firstly, the residual module is introduced into model structure of adversarial learning to optimize the network structure and effectively improve the quality of samples generated by model. By substituting original classification layer with multiple convolution layers in the network framework, the feature extraction capability of model is further strengthened, enhancing the classification performance of model. Secondly, in order to better capture different semantic information, we design a multi-level semantic extractor to extract rich and diverse semantic features from real-world images to efficiently guide sample generation. In addition, the Wasserstein divergence is introduced into the loss function to effectively solve the problem of unstable network training. Finally, high-quality defect samples can be generated through adversarial learning, effectively expanding the limited training samples for defect classification. The experimental results substantiate that our proposed method can not only generate high-quality defect samples, but also accurately achieve the classification of defect detection samples.</p>
</abstract>
<kwd-group>
<kwd>few-shot steel surface defect classification</kwd>
<kwd>adversarial learning</kwd>
<kwd>residual module</kwd>
<kwd>multi-level semantic feature extractor</kwd>
<kwd>Wasserstein divergence</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Radiation Detectors and Imaging</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Steel is an essential material for industrial production, with a broad range of uses in areas such as automobile, aerospace and machinery. As the demand for material fitness in various industries increases, the surface quality of steel has become increasingly important. However, during the steel manufacturing process, due to the influence of various unstable factors such as raw materials and production conditions, various types of defects may appear on the surface of steel, which affect the quality of steel to varying degrees and easily lead to serious production accidents, resulting in immeasurable losses to producer and users [<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B2">2</xref>]. Thus, it is of great importance to classify the defects on the surface of steel efficiently for further quality enhancement.</p>
<p>Generally, steel surface defects belonging to the same category meet a large intra-class difference, while those of different categories are highly similar [<xref ref-type="bibr" rid="B3">3</xref>], making the classification of steel surface defects more complicated. To address this problem, various approaches have been studies. For instance, Zaghdoudi et al. [<xref ref-type="bibr" rid="B4">4</xref>] proposed a steel surface defect classification method based on the binary Gabor pattern (BGP) algorithm and support vector machine (SVM). Hu et al. [<xref ref-type="bibr" rid="B5">5</xref>] extracted various visual features such as geometry, texture, and shape of the defect image and fed them to SVM for classification. Despite the fact that these methods do classify different defects, these hand-crafted features are not optimal, making a constraint on the further performance improvement. Fortunately, thanks to the development of deep learning, deep learning based methods have attracted much attention in the field of steel surface defect classification due to it powerful capability in feature extraction. Specifically, Duan et al. [<xref ref-type="bibr" rid="B6">6</xref>] used RGB images and gradient images as inputs to a dual-flow convolutional neural network, and fused multi-source information to recognize aluminum surface defects. Liu et al. [<xref ref-type="bibr" rid="B7">7</xref>] proposed an improved dual CNN model fusion framework, which uses pre-trained VGG16 and AlexNet to extract different features from the input source to classify and identify aluminum surface defects.</p>
<p>Although deep learning based methods enjoy superiority compared with conventional methods, they also meet the limitation on the large scale of training data. However, the number of non-defective samples in actual industrial production environments is far greater than that of defective samples. Moreover, it is difficult to identify and collect defective samples, further leading to an insufficient number of samples [<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B9">9</xref>, <xref ref-type="bibr" rid="B10">10</xref>]. To address this issue of insufficient samples, many researchers have begun to focus on the unsupervised data enhancement algorithm: Generative Adversarial Networks (GANs). Currently, many improved GANs and adversarial learning strategies have been derived, such as Wasserstein GAN (WGAN) [<xref ref-type="bibr" rid="B11">11</xref>], Deep Convolutional GAN (DCGAN) [<xref ref-type="bibr" rid="B12">12</xref>], and ACGAN [<xref ref-type="bibr" rid="B13">13</xref>]. These generative models augment the original data by generating synthetic samples, thereby mitigating the effect of few-shot on the classification performance and improving the accuracy. Dosovitskiy et al. [<xref ref-type="bibr" rid="B14">14</xref>] showed that even with low-fidelity images, the performance can be significantly improved. If the generated images enjoy the high-quality, the over-fitting problem can further be solved [<xref ref-type="bibr" rid="B15">15</xref>]. However, despite the wide application of GANs and its related improved models, there are still some tough difficulties, such as insufficient model feature capture capability, gradient disappearance, and model collapse, etc.</p>
<p>Furthermore, generating high-quality data similar to the original data distribution can solve the over-fitting problem, and enhance the detection accuracy and generalization ability of the model [<xref ref-type="bibr" rid="B15">15</xref>]. Lu and Su [<xref ref-type="bibr" rid="B16">16</xref>] proposed a novel method to eliminate mura patterns from defect images by using conditional generation adversarial networks; Li et al. [<xref ref-type="bibr" rid="B17">17</xref>] studied a cross-domain fault diagnosis method based on deep neural networks, which has a good industrial application prospect; Liu et al. [<xref ref-type="bibr" rid="B18">18</xref>] introduced an attention mechanism into feature extraction, proposed a structural defect detection framework based on GAN-CNN, and achieved satisfactory results. Despite the wide application of GANs and its related improved models, there are still some tough difficulties, such as insufficient model feature capture capability, gradient disappearance, and model collapse, etc.</p>
<p>Aimed at above problems, we propose a steel surface defect classification method based on residual adversarial learning with Wasserstein divergence. First, the residual module is introduced into the network framework of adversarial learning, to enhance the feature extraction ability of the model and improve the quality of generated samples. Subsequently, to extract semantic information from defect samples at different levels, we design a multi-level semantic feature extractor (MSFE), which guides sample generation by extracting the most relevant semantic features from images. Then Wasserstein divergence is used to alleviate gradient disappearance, gradient explosion and mode collapse during model training. Finally, high-quality samples are generated, and few-shot steel surface defect classification is realized by adversarial learning. The experimental results show that the proposed method improves the accuracy of steel surface defect classification, which are superior to many state-of-the-arts.</p>
<p>The main contributions of this paper are as follows:<list list-type="simple">
<list-item>
<p>&#x2022; The residual module is introduced into the network structure of adversarial learning to contribute to the feature extraction. Moreover, multiple convolutional layers are employed in the model architecture to replace the original classification layer, further boosting the classification performance of the model.</p>
</list-item>
<list-item>
<p>&#x2022; A multi-level semantic feature extractor (MSFE) which effectively extracts features at different levels is designed, fully capturing diverse semantic information of images to guide the generator in sample generation and improve the quality of generated samples.</p>
</list-item>
<list-item>
<p>&#x2022; The proposed method can generate high-quality samples to compensate for the deficiencies under few-shot conditions, further improving the classification performance.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Related works and preliminary knowledge</title>
<sec id="s2-1">
<title>2.1 Steel surface defect classification</title>
<p>Steel surface defect classification based on deep learning has gained considerable attention in recent years and achieved remarkable results. Chenon et al. [<xref ref-type="bibr" rid="B19">19</xref>] proposed a defect classification approach based on a single convolutional neural network, which can extract effective features for defect classification without the prior of hand-crafted features. Nakazawa et al. [<xref ref-type="bibr" rid="B20">20</xref>] proposed a method for surface defect classification and image retrieval using convolutional neural networks. The model was trained, validated, and tested using generated data samples, and it was demonstrated that the model trained by synthetic data can be classified efficiently. Zhu et al. [<xref ref-type="bibr" rid="B21">21</xref>] studied an intelligent identification algorithm based on convolutional neural networks and random forest algorithms, which enabled the intelligent identification of weld surface defects. However, obtaining effective defect samples is very challenging in the actual industrial environment, and there is the problem of insufficient samples, which leads to a low performance of the surface defect classification model based on deep learning. Therefore, data augmentation and transfer learning have been proposed by many researchers to address the few-shot problem in this field. Wan et al. [<xref ref-type="bibr" rid="B22">22</xref>] studied an improved VGG19 neural network based on small samples and unbalanced datasets for strip steel defect detection. Through fast image preprocessing algorithms and transfer learning theory, excellent results have been achieved on multiple datasets. Han et al. [<xref ref-type="bibr" rid="B23">23</xref>] proposed a new framework for intelligent fault diagnosis, namely, Deep Transfer Network (DTN), which generalized deep learning models to domain self-adaptation scenarios. By using the discriminative structure associated with the labeled data in the source domain to adapt to the unlabeled data, more accurate distribution matching is ensured. Furthermore, Liu et al. [<xref ref-type="bibr" rid="B24">24</xref>] designed ImDeep, a deep learning model for unbalanced multi-label surface defect classification, which combines three key technologies to improve the classification performance of the model: imbalanced sampler, Fussy-FusionNet, and transfer learning.</p>
<p>Apart from the domain adaptation, some scholars also utilize GAN as a data augmentation technique to address few-shot issue. Goodfellow et al. [<xref ref-type="bibr" rid="B15">15</xref>] first proposed the unsupervised deep learning model GAN in 2014, which was inspired by the two-player zero-sum game in game theory and consists of two components: the generator and the discriminator. The generator is mainly responsible for generating data that is as similar as possible to the original data samples, while the discriminator is tasked with distinguishing between real and fake images. Currently, GAN has been widely applied in various fields, such as image generation, data augmentation, image restoration, and image coloring. Specifically, Jain et al. [<xref ref-type="bibr" rid="B25">25</xref>] trained three GAN architectures to generate synthetic images for data augmentation, which significantly improved the performance of surface defect classification. He et al. [<xref ref-type="bibr" rid="B26">26</xref>] proposed a semi-supervised learning for defect classification based on GAN and ResNet to expand the training samples and exploit the unlabeled images. Zhao et al. [<xref ref-type="bibr" rid="B27">27</xref>] designed a reconstruction network to reconstruct the potential defect areas in the sample image, and determine the final defect area according to the difference between the reconstructed sample and the original sample. Lian et al. [<xref ref-type="bibr" rid="B28">28</xref>] proposed a novel machine vision method for automatic identification of tiny defects in a single image. To effectively achieve pixel-level defect detection on textured surfaces without manual annotation, Tsai et al. [<xref ref-type="bibr" rid="B29">29</xref>] introduced a two-stage deep learning scheme. Particularly, the first stage used CycleGAN to automatically synthesize and annotate the pixels of defect in images. The second stage used the synthesized defect images and their corresponding annotation results as input-output pairs for training the U-Net semantic network.</p>
</sec>
<sec id="s2-2">
<title>2.2 Preliminary knowledge</title>
<p>GAN consists of a generator and a discriminator [<xref ref-type="bibr" rid="B15">15</xref>], as shown in <xref ref-type="fig" rid="F1">Figure 1A</xref>. The input of the generator is a random noise vector z, and the output is the fake sample generated by it. The discriminator uses the fake sample generated by the generator and the real data x as the input, and the output is the discrimination score of the discriminator on the fake sample. GAN&#x2019;s overall objective function is:<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:munder>
<mml:mi>min</mml:mi>
<mml:mi>G</mml:mi>
</mml:munder>
<mml:munder>
<mml:mi>max</mml:mi>
<mml:mi>D</mml:mi>
</mml:munder>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the probability density distribution of the real data <inline-formula id="inf2">
<mml:math id="m3">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>; <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the noise vector randomly sampled from the prior distribution <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>; <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the generator, <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the discriminator, and <inline-formula id="inf7">
<mml:math id="m8">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mo>&#x2219;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the calculated expected value; <inline-formula id="inf8">
<mml:math id="m9">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is a probability distribution, that is, the probability of classifying data <inline-formula id="inf9">
<mml:math id="m10">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> as a real sample, and <inline-formula id="inf10">
<mml:math id="m11">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is derived from a real sample <inline-formula id="inf11">
<mml:math id="m12">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> or a generated sample <inline-formula id="inf12">
<mml:math id="m13">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The model framework of GAN and ACGAN. <bold>(A)</bold> GAN. <bold>(B)</bold> ACGAN.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g001.tif"/>
</fig>
<p>
<xref ref-type="disp-formula" rid="e1">Formula 1</xref> shows that the optimization problem of GAN is same as the max-min optimization problem, which includes the optimization goals of the generator and the discriminator. The main function of the discriminator is to perform binary classification on the input data to determine whether the input data comes from the distribution of the real data or the generated pseudo data. Thus, its objective function is:<disp-formula id="e2">
<mml:math id="m14">
<mml:mrow>
<mml:munder>
<mml:mi>max</mml:mi>
<mml:mi>D</mml:mi>
</mml:munder>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>It can be seen from <xref ref-type="disp-formula" rid="e2">Formula 2</xref> that the goal of the discriminator is to maximize the discrimination accuracy for the data. In other words, we aim to maximize the discriminant result <inline-formula id="inf13">
<mml:math id="m15">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for the real data <inline-formula id="inf14">
<mml:math id="m16">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and minimize the result <inline-formula id="inf15">
<mml:math id="m17">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of the generated sample <inline-formula id="inf16">
<mml:math id="m18">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> (maximize <inline-formula id="inf17">
<mml:math id="m19">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>).</p>
<p>The purpose of the generator is to generate samples that the discriminator cannot distinguish as false, and its objective function is:<disp-formula id="e3">
<mml:math id="m20">
<mml:mrow>
<mml:munder>
<mml:mi>min</mml:mi>
<mml:mi>G</mml:mi>
</mml:munder>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>The generator is optimized by Eq. <xref ref-type="disp-formula" rid="e3">3</xref>. Specifically, the probability score <inline-formula id="inf18">
<mml:math id="m21">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of the discriminator for the generated sample <inline-formula id="inf19">
<mml:math id="m22">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is maximized (<inline-formula id="inf20">
<mml:math id="m23">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is minimized). During training, the alternate optimization methods are used: fix one side and update the parameters of the other network. In other words, the model updates the discriminator&#x2019;s parameters firstly through the fixed generator so that the discriminator maximizes the discriminant result. Then we fix discriminator&#x2019;s parameters for updating the generator, which minimize the result that discriminator works. Finally, when the probability distribution <inline-formula id="inf21">
<mml:math id="m24">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of the samples generated by the generator <inline-formula id="inf22">
<mml:math id="m25">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is infinitely close to the probability distribution <inline-formula id="inf23">
<mml:math id="m26">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of the real samples (that is, <inline-formula id="inf24">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>), the global optimal solution can be reached.</p>
<p>ACGAN is a variant of GAN [<xref ref-type="bibr" rid="B13">13</xref>], and its structure is illustrated in <xref ref-type="fig" rid="F1">Figure 1B</xref>. By incorporating auxiliary label information c into the generator, the generated samples can be constrained to possess certain characteristics, thus allowing for more precise expression of the samples and the generation of specific samples according to it. Moreover, in order to ensure accurate classification, ACGAN adds a softmax layer to the discriminator network, thus enabling the improved model to not only judge the authenticity of the data, but also classify the input samples.</p>
<p>The loss function of ACGAN consists of two parts: the discriminative loss <inline-formula id="inf25">
<mml:math id="m28">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the classification loss <inline-formula id="inf26">
<mml:math id="m29">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. The role of discriminative loss is to judge the authenticity of the generated samples, thereby improving the quality of the samples generated by the generator. The role of the classification loss is to measure the accuracy of the classification of the sample category. And, the specific calculation of <inline-formula id="inf27">
<mml:math id="m30">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is:<disp-formula id="e4">
<mml:math id="m31">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext> </mml:mtext>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf28">
<mml:math id="m32">
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the cross-entropy loss function, <inline-formula id="inf29">
<mml:math id="m33">
<mml:mrow>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the category label of the real data <inline-formula id="inf30">
<mml:math id="m34">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf31">
<mml:math id="m35">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the category label of the generated data <inline-formula id="inf32">
<mml:math id="m36">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf33">
<mml:math id="m37">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the category label distribution of the sample.</p>
<p>Since a classifier is added to the discriminator <inline-formula id="inf34">
<mml:math id="m38">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, the network can not only distinguish the authenticity of the data, but also classify the data, so its loss function needs to calculate two parts: discriminant loss <inline-formula id="inf35">
<mml:math id="m39">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and classification loss <inline-formula id="inf36">
<mml:math id="m40">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. The specific calculation is as follows.<disp-formula id="e5">
<mml:math id="m41">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext> </mml:mtext>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m42">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>Similarly, the loss function of the generator <inline-formula id="inf37">
<mml:math id="m43">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> also needs to consider the classification loss:<disp-formula id="e7">
<mml:math id="m44">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext> </mml:mtext>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m45">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>
<xref ref-type="disp-formula" rid="e6">Formulas 6</xref>, <xref ref-type="disp-formula" rid="e8">8</xref> ultimately constitute the entire loss function of the ACGAN model. During the training process, the model is continually optimized to enhance the quality of the samples generated by the model and augment the classification accuracy of the model.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<p>Although Generative Adversarial Networks (GANs) and Auxiliary Classifier GANs (ACGANs) can effectively alleviate the few-shot classification problem by generating samples, they still meet the limitation on inadequate information extraction capabilities, gradient vanishing, and pattern collapse. To address these issues, we propose a novel network structure. Specifically, a residual adversarial learning model with Wasserstein divergence based on ACGAN under multi-level semantic guidance is proposed, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Overview of our framework. Given an image <inline-formula id="inf38">
<mml:math id="m46">
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> as input, our framework first extracts rich semantic information through multi-level semantic feature extractor to guide generator. After that, we deliver the noise <inline-formula id="inf39">
<mml:math id="m47">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and label <inline-formula id="inf40">
<mml:math id="m48">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> to generator for generating sample <inline-formula id="inf41">
<mml:math id="m49">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. Finally, we can obtain the classification result <inline-formula id="inf42">
<mml:math id="m50">
<mml:mrow>
<mml:msup>
<mml:mi>c</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and the discriminate result <inline-formula id="inf43">
<mml:math id="m51">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo>?</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> (True or Fake?) of generated sample <inline-formula id="inf44">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> by discriminator. <inline-formula id="inf45">
<mml:math id="m53">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf46">
<mml:math id="m54">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf47">
<mml:math id="m55">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> indicates respectively the classification loss, the discriminant loss, Wasserstein divergence during training.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g002.tif"/>
</fig>
<p>First, the random noise vector <inline-formula id="inf48">
<mml:math id="m56">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and sample label <inline-formula id="inf49">
<mml:math id="m57">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> are input into the generator. The generator generates synthetic samples <inline-formula id="inf50">
<mml:math id="m58">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, expanding the scale of the training data. By utilizing a multi-level semantic feature extractor to process original samples, semantic and contextual information can effectively be captured and used for guiding sample generation of generator. Then, the discriminator takes the generated sample <inline-formula id="inf51">
<mml:math id="m59">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and real sample <inline-formula id="inf52">
<mml:math id="m60">
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> as the inputs, and outputs the discriminant result <inline-formula id="inf53">
<mml:math id="m61">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo>?</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> (True or Fake) and the classification result <inline-formula id="inf54">
<mml:math id="m62">
<mml:mrow>
<mml:msup>
<mml:mi>c</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> of the generated sample. During the adversarial training of model, the Wasserstein divergence <inline-formula id="inf55">
<mml:math id="m63">
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
</inline-formula> is used as the distance measurement between the distributions of the initial data and the distributions of the generated data.</p>
<sec id="s3-1">
<title>3.1 The modification of network</title>
<p>Despite the fact that ACGAN achieved significantly satisfactory results in image generation [<xref ref-type="bibr" rid="B13">13</xref>], it still faces the problem of insufficient feature extraction ability when it is applied to tasks within the few-shot environment, resulting in inadequate acquisition of image information and a consequent decrease in model performance. To address this issue, the overall network structure of ACGAN is optimized, as illustrated in <xref ref-type="fig" rid="F3">Figure 3</xref>. The specific improvements of the network structure are detailed below.<list list-type="simple">
<list-item>
<p>(1) As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the residual module (Residual) is introduced into the network structure of the generator and the discriminator to optimize the feature learning ability of the model, so that the model can extract more valuable features. Meanwhile, it can ensure the quality of the samples generated by the model while optimizing the model&#x2019;s ability to discriminate and classify images. The specific network structure of the introduced residual module is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
</list-item>
<list-item>
<p>(2) When the kernel size of the deconvolution layer cannot be divisible by stride in the actual calculation, uneven overlapping problems will occur. Also, the generated sample images would have some checkerboard-like artifacts [<xref ref-type="bibr" rid="B30">30</xref>]. Therefore, in order to avoid such problems, as shown in <xref ref-type="fig" rid="F3">Figure 3A</xref>, the up-sampling layer (US) and the convolutional layer (Conv) are used to generate sample images in the generator network structure. As shown in <xref ref-type="fig" rid="F3">Figure 3B</xref>, in the discriminator network structure, two convolutional layers are added before the sigmoid and softmax classification layers, which makes the classifier in the discriminator learn more image information and improve the classification performance.</p>
</list-item>
<list-item>
<p>(3) The generator network mainly consists of several residual modules and convolutional layers as well as operating up-sampling layers. The input of the model is the randomly generated 128-dimensional vector <inline-formula id="inf56">
<mml:math id="m64">
<mml:mrow>
<mml:mi mathvariant="normal">z</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and the sample label <inline-formula id="inf57">
<mml:math id="m65">
<mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, which undergoes a fully connected layer (FC) and the reshape (reshape) operation. Before the convolution calculation, the first two convolution layers have both performed the up-sampling operation using the nearest neighbor interpolation, which increases the feature map by two times. At the same time, Batch Normalization (BN) is used to optimize the network throughput the training. There are three residual modules between each two convolution layers to improve the feature learning ability of the model, and the Leaky-ReLU activation function is used between each layer. The discriminator network also includes 6 convolutional layers and 3 residual modules. And the 3 residual modules follow the first convolutional layer. At the same time, a Dropout layer (Dropout) is further introduced to prevent overfitting problems. Furthermore, the Leaky-ReLU activation function is used between each layer.</p>
</list-item>
</list>
</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Improved generator and discriminator network structure. <bold>(A)</bold> The structure of the improved generator network. <bold>(B)</bold> The structure of the improved discriminator network.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The specific network structure diagram of the residual block.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g004.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>3.2 Multi-level semantic feature extractor</title>
<p>Recently, ACGAN has achieved remarkable progress in the field of image generation. Given a category label, ACGAN can map random noise into high-resolution images with abundant texture features and comprehensive shape details. However, satisfactory results depend on training ACGAN with sufficient quantity of samples. When there is an inadequate number of samples, the effectiveness of ACGAN in generating samples close to reality is compromised due to its inability to obtain enough semantic information, which motivates us to design a multi-level semantic feature extractor to facilitate sample generation tasks, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. As illustrated above, the role of the multi-level semantic feature extractor is to extract the semantic and contextual information of defects at different levels in the image. Therefore, the original samples are input into the multi-level semantic feature extractor to obtain the learned hierarchical features, such as texture and shape, which are then incorporated into the generator to serve as guidance for sample generation. Specifically, the sample image <inline-formula id="inf58">
<mml:math id="m66">
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> corresponding to the true sample label <inline-formula id="inf59">
<mml:math id="m67">
<mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is processed by the multi-level semantic feature extractor to obtain rich semantic information, which is then aligned with the different convolutional layers in the generator, facilitating better integration of multi-level semantic features into the process of sample generation. The core of alignment operation mainly relies on a convolutional layer, which adjusts semantic features extracted by MSFE to the size corresponding to different layers of the generator, and adds the adjusted features to the original ones to obtain new features under semantic guidance. The aligned features are added to the convolutional layers of the generator, leveraging diverse levels of semantic features to facilitate the generation of specific defective samples, as depicted in <xref ref-type="fig" rid="F3">Figure 3A</xref>. We use VGG19 pretrained on the ImageNet dataset as a multi-level semantic feature extractor, and use the features extracted from layers 7 to 23 in it to guide the generator.</p>
</sec>
<sec id="s3-3">
<title>3.3 Objective function</title>
<p>The Kullback-Leibler (KL) divergence [<xref ref-type="bibr" rid="B15">15</xref>] is prone to gradient instability in the Generative Adversarial Networks (GANs) training phase, and can also lead to mode collapse. To address these issues, the Wasserstein GAN (WGAN) uses the Wasserstein distance to ensure that the gradient of the model is continuous during the training process [<xref ref-type="bibr" rid="B11">11</xref>]. However, WGAN utilizes weight clipping to restrict the weights within a fixed range strictly, which greatly limits the expressiveness of the network. Consequently, WGAN-GP [<xref ref-type="bibr" rid="B31">31</xref>] adopts gradient penalty to enhance the stability of the network training. According to the research conducted by [<xref ref-type="bibr" rid="B32">32</xref>], in experiments, WGAN-GP typically employs the technique of interpolating between real and fake data to simulate a uniform distribution across the whole space. This approach is somewhat mechanistic and empirical, which makes it challenging to simulate the full spatial distribution using limited sampling.</p>
<p>In order to solve this problem, Wu et al. [<xref ref-type="bibr" rid="B32">32</xref>] proposed Wasserstein divergence to reduce the distance loss function properly between two distributions, as shown in <xref ref-type="disp-formula" rid="e9">Formula 9</xref>. It removes the K-Lipschitz conditional restriction, and changes the penalty term added to the loss function.<disp-formula id="e9">
<mml:math id="m68">
<mml:mrow>
<mml:mtable columnalign="center">
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mi>max</mml:mi>
<mml:mi>D</mml:mi>
</mml:munder>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>u</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;" separators="|">
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mi>p</mml:mi>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where k and <italic>p</italic> are selected empirically. Generally, k &#x3d; 2, <italic>p</italic> &#x3d; 6. <inline-formula id="inf60">
<mml:math id="m69">
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> represents the gradient. <inline-formula id="inf61">
<mml:math id="m70">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> comes from the distribution <inline-formula id="inf62">
<mml:math id="m71">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of the real data; similarly, <inline-formula id="inf63">
<mml:math id="m72">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> comes from the generated sample distribution <inline-formula id="inf64">
<mml:math id="m73">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf65">
<mml:math id="m74">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>u</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is a distribution derived from the real data distribution <inline-formula id="inf66">
<mml:math id="m75">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the generated data distribution <inline-formula id="inf67">
<mml:math id="m76">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf68">
<mml:math id="m77">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the discriminator, and <inline-formula id="inf69">
<mml:math id="m78">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mo>&#x2219;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the calculated expected value. Experiments in [<xref ref-type="bibr" rid="B32">32</xref>] prove that all different distributions have improved performance.</p>
<p>Based on the loss function of ACGAN [<xref ref-type="bibr" rid="B13">13</xref>], we use Wasserstein divergence to address the potential gradient explosion issue in the training process. Hence, the loss function of our method consists of two parts: the loss function <inline-formula id="inf70">
<mml:math id="m79">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of discriminator and the loss function <inline-formula id="inf71">
<mml:math id="m80">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of generator, with each loss function consisting of two components: the adversarial loss function <inline-formula id="inf72">
<mml:math id="m81">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the conditional loss function <inline-formula id="inf73">
<mml:math id="m82">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The purpose of <inline-formula id="inf74">
<mml:math id="m83">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is to ensure that the discriminator can distinguish between real and generated samples and accurately classify them based on their respective conditions, as shown below:<disp-formula id="e10">
<mml:math id="m84">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>u</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="&#x2016;" close="&#x2016;" separators="|">
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mi>p</mml:mi>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<disp-formula id="e11">
<mml:math id="m85">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext> </mml:mtext>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m86">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>where <inline-formula id="inf75">
<mml:math id="m87">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the adversarial loss function that is modified with Wasserstein divergence; <inline-formula id="inf76">
<mml:math id="m88">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the conditional loss function; <inline-formula id="inf77">
<mml:math id="m89">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mo>&#x2219;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes the cross-entropy loss function; <inline-formula id="inf78">
<mml:math id="m90">
<mml:mrow>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> indicates the category label of real data sample <inline-formula id="inf79">
<mml:math id="m91">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf80">
<mml:math id="m92">
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> denotes the category label of generated data <inline-formula id="inf81">
<mml:math id="m93">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf82">
<mml:math id="m94">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the distribution of sample class labels. During the training process of discriminator, our objective is to maximize its loss function <inline-formula id="inf83">
<mml:math id="m95">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Likewise, the purpose of <inline-formula id="inf84">
<mml:math id="m96">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is to generate high-quality data samples such that the discriminator cannot distinguish whether the sample is real or fake, as illustrated below:<disp-formula id="e13">
<mml:math id="m97">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
<disp-formula id="e14">
<mml:math id="m98">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext> </mml:mtext>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="[" close="]" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
<disp-formula id="e15">
<mml:math id="m99">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>where <inline-formula id="inf85">
<mml:math id="m100">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the adversarial loss function for the generator. Similarly, we aim to maximize its loss function <inline-formula id="inf86">
<mml:math id="m101">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> in the training process.</p>
</sec>
<sec id="s3-4">
<title>3.4 Network training</title>
<p>During the training process of the model, the discriminator continuously enhances its capability to distinguish between real samples and generated samples, while the generator continuously improves its ability to generate realistic samples. The discriminator updates its weights by utilizing both real and generated samples, and the generator updates its weights through the error feedback from the discriminator. The training process of the model is a maximization and minimization process. In the adversarial training of the discriminator and the generator, the discriminator minimizes the probability of misclassification, and the generator maximizes the error probability of the discriminator. The iterative training method of the generator and the discriminator is employed to prevent the over-fitting of the generator network. The specific training steps of the model are illustrated in <xref ref-type="statement" rid="Algorithm_1">Algorithm 1</xref>.</p>
<p>
<statement content-type="algorithm" id="Algorithm_1">
<label>Algorithm 1. Residual Adversarial Learning Model with Wasserstein Divergence</label>
<p>
<list list-type="simple">
<list-item>
<p>
<bold>Data:</bold> image dataset</p>
</list-item>
<list-item>
<p>
<bold>Output:</bold> trained Discriminator and Generator, Training Accuracy</p>
</list-item>
<list-item>
<p>1 &#x2003;<bold>for</bold> <italic>epoch&#x3d;0</italic> <bold>to</bold> <italic>n</italic> <bold>do</bold>
</p>
</list-item>
<list-item>
<p>2 &#x2003;&#x2003;randomly sample from real samples and get (<italic>real_images, labels</italic>), and randomly sample from a uniform distribution to obtain noise <italic>z</italic>
</p>
</list-item>
<list-item>
<p>3 &#x2003;&#x2003;input (<italic>z, labels</italic>) into Generator to generate sample <italic>fake_images</italic>
</p>
</list-item>
<list-item>
<p>4 &#x2003;&#x2003;generated sample <italic>fake_images</italic> and real sample <italic>real_images</italic> are fed into discriminator</p>
</list-item>
<list-item>
<p>5 &#x2003;&#x2003;&#x2003;calculate the gradient of the real sample space, calculate the gradient of the generated sample space, and calculate the Wasserstein divergence according to <xref ref-type="disp-formula" rid="e9">Formula 9</xref>
</p>
</list-item>
<list-item>
<p>6 &#x2003;&#x2003;<bold>for</bold> <italic>D_epoch&#x3d;0</italic> <bold>to</bold> <italic>m</italic> <bold>do</bold>
</p>
</list-item>
<list-item>
<p>7 &#x2003;&#x2003;&#x2003;calculate Discriminator&#x2019;s loss by <xref ref-type="disp-formula" rid="e10">Formulas 10</xref>, <xref ref-type="disp-formula" rid="e11">11</xref>, <xref ref-type="disp-formula" rid="e12">12</xref>
</p>
</list-item>
<list-item>
<p>8 &#x2003;&#x2003;&#x2003;update Discriminator parameters</p>
</list-item>
<list-item>
<p>9 &#x2003;&#x2003;<bold>end for</bold>
</p>
</list-item>
<list-item>
<p>10 &#x2003;&#x2003;&#x2003;calculate Generator&#x2019;s loss according to <xref ref-type="disp-formula" rid="e13">Formulas 13</xref>, <xref ref-type="disp-formula" rid="e14">14</xref>, <xref ref-type="disp-formula" rid="e15">15</xref>
</p>
</list-item>
<list-item>
<p>11 &#x2003;&#x2003;&#x2003;update Generator parameters</p>
</list-item>
<list-item>
<p>12 &#x2003;&#x2003;<bold>end for</bold>
</p>
</list-item>
</list>
</p>
</statement>
</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments</title>
<p>In order to verify the effectiveness of the proposed method, experiments are conducted on the NEU-CLS dataset using a Windows 10 system with 16&#xa0;GB of memory, an AMD Ryzen 7 4800HS processor, and an NVIDIA GTX 1660 Ti graphics card. The model is constructed using the PyTorch platform.</p>
<sec id="s4-1">
<title>4.1 Dataset</title>
<p>This paper performs experiments on the NEU-CLS hot-rolled steel surface defect dataset from Northeastern University [<xref ref-type="bibr" rid="B34">34</xref>]. The dataset consists of 6 types of defects, and each category contains 300 grayscale images (200 &#xd7; 200 pixels). These six types of defects are: crazing (Cr), inclusion (In), patches (Pa), pitted surface (PS), rolled-in scale (RS) and scratches (Sc), as illustrated in <xref ref-type="fig" rid="F5">Figure 5</xref>.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Six types of steel surface defects.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g005.tif"/>
</fig>
<p>In the experiment, the NEU-CLS dataset is divided according to a 2:1 ratio, with 1,200 images used as the training set and 600 images used as the test set. It takes 10,000 epochs to train our network with Adam optimizer and a batch of 64 images. The parameter settings of the model are as follows: learning rate of <italic>&#x3b1;</italic> &#x3d; 0.0002, random noise vector dimension of z &#x3d; 128, and Adam optimization parameters of &#x3b2;1 &#x3d; 0.5 and &#x3b2;2 &#x3d; 0.999. In addition, we use VGG19 pretrained on the ImageNet dataset as a multi-level semantic feature extractor, and use the features extracted from layers 7 to 23 in it to guide the generator.</p>
</sec>
<sec id="s4-2">
<title>4.2 Few-shot classification of steel surface defects</title>
<p>Considering the restricted size of the dataset, we conduct experiments with six different training sample sets (200, 150, 100, 50, 30, 10) to evaluate the few-shot classification performance enhancement of the proposed method after training, and to comparatively analyze the impact of the data size on the model. The numbers of test sets are kept constant. The results of the comparison between ACGAN and the method proposed in this paper under different training sample sizes are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Average classification accuracy of different sample sizes.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">Sample size for each category (total sample size)</th>
<th colspan="2" align="center">Average accuracy (%)</th>
<th rowspan="2" align="center">Increase (%)</th>
</tr>
<tr>
<th align="center">ACGAN</th>
<th align="center">Ours</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">200 (1,200)</td>
<td align="center">94.83</td>
<td align="center">
<bold>98.67</bold>
</td>
<td align="center">4.04</td>
</tr>
<tr>
<td align="center">150 (900)</td>
<td align="center">93.67</td>
<td align="center">
<bold>98.33</bold>
</td>
<td align="center">4.66</td>
</tr>
<tr>
<td align="center">100 (600)</td>
<td align="center">90.00</td>
<td align="center">
<bold>95.50</bold>
</td>
<td align="center">5.5</td>
</tr>
<tr>
<td align="center">50 (300)</td>
<td align="center">83.83</td>
<td align="center">
<bold>94.67</bold>
</td>
<td align="center">10.84</td>
</tr>
<tr>
<td align="center">30 (180)</td>
<td align="center">77.83</td>
<td align="center">
<bold>94.00</bold>
</td>
<td align="center">16.17</td>
</tr>
<tr>
<td align="center">10 (60)</td>
<td align="center">66.50</td>
<td align="center">
<bold>89.67</bold>
</td>
<td align="center">23.17</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values mean the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>According to <xref ref-type="table" rid="T1">Table 1</xref>, it can be observed that the classification performance of ACGAN and the proposed model decreases as the training sample size decreases. It is evident that insufficient samples reduce the generalization capability of the model, resulting in a poorer performance on the test set. Furthermore, the decline of our model is more gradual than that of ACGAN, indicating that the method proposed in this paper is more stable and robust when dealing with few-shot issues. As illustrated in <xref ref-type="fig" rid="F6">Figure 6</xref>, the trend of classification results of ACGAN and our model under different training sample sizes can be observed.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Trend chart of classification results under different sample sizes.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g006.tif"/>
</fig>
<p>Observing <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>, it can be seen that the accuracy of our model has a distinct advantage over ACGAN under different training sample sizes. When the sample size is 200, the average accuracy of our model reaches 98.67%. At the same time, when the training sample size is 10, the average accuracy of the model in this paper is 89.67%, while the accuracy of ACGAN drops to 66.5%. This indicates that ACGAN is more reliant on data. Furthermore, as the training sample size decreases, the classification accuracy gap between ACGAN and the model proposed in this paper increases. When the sample size is 10, the accuracy of ACGAN is 23.17% lower than that of the method proposed in this paper, making it evident that ACGAN is far less effective than the model in this paper when dealing with few-shot problems.</p>
<p>To illustrate the classification ability of the proposed model for each type of defect, <xref ref-type="fig" rid="F7">Figure 7</xref> shows the confusion matrix of our model under different sample sizes, where the numbers 0&#x2013;5 in the abscissa and ordinate represent defect types, respectively: Cr, In, Pa, PS, RS, and Sc. It is evident that our method can train an ideal model under different training sample sizes and can accurately classify most of the defects. Moreover, when the sample size is 200, the model can accurately classify all Pa defects. Under different training sample sizes, the cases of classifying Cr as RS and RS as Cr occupy a large proportion in the wrong classification cases. The high similarity between Cr and RS defects and the lack of distinct inter-class features lead to misjudgment of the model. The overall results demonstrate that the method proposed in this paper only misjudges a few fault types under different sample sizes, and the overall accuracy remains high as the sample size decreases.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Confusion matrix of our model under different training sample sizes. <bold>(A)</bold> 200, <bold>(B)</bold> 150, <bold>(C)</bold> 100, <bold>(D)</bold> 50, <bold>(E)</bold> 30, and <bold>(F)</bold> 10.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g007.tif"/>
</fig>
<p>In order to further validate the classification performance of our model, we compare it with the classic ResNet18 and ResNet50 classification methods. To ensure the efficient classification performance of the classic classification models, the ResNet18 and ResNet50 models are pre-trained using the ImageNet dataset. Additionally, we also compared with the latest few-shot deep learning classification models, including: the model proposed by Lian et al. [<xref ref-type="bibr" rid="B28">28</xref>], which combine generative adversarial networks and convolutional neural networks to generate exaggerated defect image samples to ensure the accuracy of micro-surface defect detection; and the model proposed by Li et al. [<xref ref-type="bibr" rid="B35">35</xref>], which replace the fully connected classification layer with an orthogonal SoftMax layer, significantly reducing the complexity of the model and making it suitable for few-shot classification. Moreover, in order to fully demonstrate the impact of MSFE on the final classification results, MSFE is deliberately excluded in the original framework and a corresponding experiment is conducted. The experimental results are presented in <xref ref-type="table" rid="T2">Table 2</xref>, and it can be seen that, in the case of different sample sizes, the methods proposed in this paper have achieved the best results and achieved the highest classification accuracy.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Results of steel surface defects under different methods and sample sizes.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">Methods</th>
<th colspan="6" align="center">Average accuracy (%)</th>
</tr>
<tr>
<th align="center">200</th>
<th align="center">150</th>
<th align="center">100</th>
<th align="center">50</th>
<th align="center">30</th>
<th align="center">10</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">ResNet18</td>
<td align="center">92.33</td>
<td align="center">90.67</td>
<td align="center">85.00</td>
<td align="center">83.33</td>
<td align="center">76.33</td>
<td align="center">59.5</td>
</tr>
<tr>
<td align="center">ResNet50</td>
<td align="center">93.00</td>
<td align="center">92.33</td>
<td align="center">85.33</td>
<td align="center">83.00</td>
<td align="center">77.00</td>
<td align="center">63.17</td>
</tr>
<tr>
<td align="center">Res-ACGAN</td>
<td align="center">96.17</td>
<td align="center">95.00</td>
<td align="center">91.00</td>
<td align="center">84.67</td>
<td align="center">79.00</td>
<td align="center">70.50</td>
</tr>
<tr>
<td align="center">[<xref ref-type="bibr" rid="B28">28</xref>]</td>
<td align="center">96.50</td>
<td align="center">95.50</td>
<td align="center">91.00</td>
<td align="center">89.50</td>
<td align="center">87.50</td>
<td align="center">76.33</td>
</tr>
<tr>
<td align="center">[<xref ref-type="bibr" rid="B35">35</xref>]</td>
<td align="center">96.67</td>
<td align="center">94.67</td>
<td align="center">90.50</td>
<td align="center">85.33</td>
<td align="center">84.83</td>
<td align="center">71.33</td>
</tr>
<tr>
<td align="center">Ours (lack MSFE)</td>
<td align="center">97.00</td>
<td align="center">96.00</td>
<td align="center">94.33</td>
<td align="center">93.67</td>
<td align="center">91.33</td>
<td align="center">86.00</td>
</tr>
<tr>
<td align="center">Ours</td>
<td align="center">
<bold>98.67</bold>
</td>
<td align="center">
<bold>98.33</bold>
</td>
<td align="center">
<bold>95.50</bold>
</td>
<td align="center">
<bold>94.67</bold>
</td>
<td align="center">
<bold>94.00</bold>
</td>
<td align="center">
<bold>89.67</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values mean the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>In order to further verify the importance of introducing various parts in our model, we compare the classification performance of the model after introducing the residual module, Wasserstein distance and penalty weight GP [<xref ref-type="bibr" rid="B30">30</xref>], Wasserstein divergence, and MSFE into the ACGAN model respectively. At the same time, in order to verify the important role played by the residual module in the discriminator network, we replace the residual module in the discriminator of our model with the CBAM attention mechanism module proposed by [<xref ref-type="bibr" rid="B35">35</xref>], and introduce the SENet module to conduct comparative experiments. The training sample size is 200, and the experimental results are shown in <xref ref-type="table" rid="T3">Table 3</xref>. It can be found that after adding the residual module to the original model, the classification accuracy of ACGAN increases by 1.34%, the classification accuracy of ACGAN &#x2b; Wasserstein &#x2b; GP increases by 1.17%, and the classification accuracy of ACGAN &#x2b; SENet &#x2b; Wasserstein &#x2b; GP increases by 0.84%. After adding Wasserstein divergence to the original model, the classification accuracy of ACGAN increases by 1%, and the classification accuracy of ACGAN &#x2b; Res increases by 0.83%, which is higher than that of using Wasserstein distance and penalty weight, showing that introducing the residual module and Wasserstein divergence into the model can improve the feature extraction ability of the model and further improve the model&#x2019;s ability to discriminate and classify sample images. In addition, the introduction of attention mechanism modules SENet and CBAM in the discriminator network can improve the classification ability of the model, but the discriminator network structure proposed in this paper has achieved the best results in experiments. The incorporation of MSFE in the original framework results in a 1.67% increase in classification accuracy. This implies that employing semantic features at varying levels to guide the generator can enhance its efficiency, thereby advancing the classification abilities of the discriminator.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Classification accuracy of introducing different modules.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Method</th>
<th align="center">Accuracy (%)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">ACGAN</td>
<td align="center">94.83</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res</td>
<td align="center">96.17</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Wasserstein &#x2b; GP</td>
<td align="center">95.33</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Wasserstein-div</td>
<td align="center">95.83</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res &#x2b; Wasserstein &#x2b; GP</td>
<td align="center">96.50</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; SENet &#x2b; Wasserstein &#x2b; GP</td>
<td align="center">95.83</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res &#x2b; SENet &#x2b; Wasserstein &#x2b; GP</td>
<td align="center">96.67</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res &#x2b; CBAM &#x2b; Wasserstein &#x2b; GP</td>
<td align="center">96.83</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res &#x2b; Wasserstein-div</td>
<td align="center">97.00</td>
</tr>
<tr>
<td align="center">ACGAN &#x2b; Res &#x2b; Wasserstein-div &#x2b; MSFM (Ours)</td>
<td align="center">
<bold>98.67</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values mean the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-3">
<title>4.3 Quality assessment of generated samples</title>
<p>
<xref ref-type="fig" rid="F8">Figure 8</xref> presents a comparison of steel surface defect samples generated by different models, including ACGAN, the model augmented with SENet module, the model augmented with CBAM module [<xref ref-type="bibr" rid="B35">35</xref>], the proposed method while lacking MSFE, and our model. The training process utilizes 200 samples of each type of defect, with 10,000 iterations and other parameters hold constant.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Sample images generated by different methods. <bold>(A)</bold> ACGAN. <bold>(B)</bold> introduce SENet module. <bold>(C)</bold> introduce CBAM module [<xref ref-type="bibr" rid="B34">34</xref>]. <bold>(D)</bold> Ours (lack MSFE). <bold>(E)</bold> Ours.</p>
</caption>
<graphic xlink:href="fphy-11-1208781-g008.tif"/>
</fig>
<p>It can be observed that, compared to the original sample in <xref ref-type="fig" rid="F5">Figure 5</xref>, the samples generated by the method proposed in this paper are more distinct and the quality of the samples are also much better. For instance, for the defect of scratch, such as the third one in the fifth row and the second one in the sixth row in <xref ref-type="fig" rid="F8">Figure 8E</xref> generated by the method proposed in this paper, when compared to the last one in the second row in (a), the last one in the first row in (b), the third one in the fourth row in (c), and the last one in the last row in (d), its defect features are more discernible, the defect is sharper, and it is also more similar to the original sample image. Although the version of lacking MSFE can also generate high-quality sample images, it is evident that its feature extraction ability is inadequate, leading to blurred images and unclear semantic information, as demonstrated in <xref ref-type="fig" rid="F8">Figure 8D</xref>, specifically in the fifth one of the second row and the second item of the fourth row.</p>
<p>In order to assess the quality of samples generated by different models, the MSE (Mean Square Error) and SSIM (Structural Similarity) metrics are employed to evaluate the sample quality. MSE is a metric that reflects the degree of discrepancy between the estimator and the estimated quantity; SSIM is used to measure the similarity between two images. The results of different models are presented in <xref ref-type="table" rid="T4">Table 4</xref>. The smaller the value of MSE, or the larger the value of SSIM, the larger the similarity between original image and generated image. It can be seen from <xref ref-type="table" rid="T4">Table 4</xref> that the MSE and SSIM of our model are more proximate to the original images than other methods, which demonstrates that the sample data distribution generated by our model is more similar to the original sample distribution, and also shows that MSFE and Wasserstein divergence can improve the quality of samples generated by the model.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Comparison of MSE and SSIM values of different models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Methods</th>
<th align="center">MSE</th>
<th align="center">SSIM</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">ACGAN</td>
<td align="center">347.8543</td>
<td align="center">0.6935</td>
</tr>
<tr>
<td align="center">SENet-ACGAN</td>
<td align="center">272.5821</td>
<td align="center">0.7074</td>
</tr>
<tr>
<td align="center">CBAM-ACGAN</td>
<td align="center">222.4989</td>
<td align="center">0.7583</td>
</tr>
<tr>
<td align="center">Ours (lack MSFE)</td>
<td align="center">193.8484</td>
<td align="center">0.7828</td>
</tr>
<tr>
<td align="center">Ours</td>
<td align="center">
<bold>184.2617</bold>
</td>
<td align="center">
<bold>0.7912</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values mean the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>5 Conclusion</title>
<p>Aiming at the difficulties of steel surface defect few-shot classification, this paper introduces multi-level semantic feature extractor under the residual adversarial learning network framework to generate high-quality samples and achieves promising steel surface defect classification. First, we modify the network structure of the adversarial learning model by the residual module, so that the model can obtain more information during training and generate synthetic data to the original sample. To overcome the challenge of inadequate feature extraction in generator networks which may lead to suboptimal sample quality in small-sample environments, we design a multi-level semantic feature extractor for obtaining diverse semantic information at various levels. By leveraging this comprehensive semantic information, we directed sample generation. At the same time, the Wasserstein divergence is introduced into the loss function to solve the problem of unstable model training and to improve the generation efficiency and classification performance of the model. Experiments are conducted on the steel surface defect dataset NEU-CLS from Northeastern University. The results demonstrate that, under the condition of the restricted number of training samples, the method proposed in this paper achieves the highest classification accuracy. Moreover, when the number of training data is reduced, our method exhibits better stability and robustness than classical classification models and state-of-the-art of deep learning models. Additionally, in terms of the quality of generated samples, the MSE value and SSIM value of the samples generated by the model proposed in this paper are the closest to the original samples, further showing the effectiveness of our proposed method. With the popularization of sensors and lightweight devices, the demand for model compression and lightweight models is becoming increasingly important. Improving the real-time performance of defect detection systems is the main trend for deploying online detection systems in actual industrial production in the future.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>LH: Ideas, methodology, experimental design, formal analysis. PS: Software, validation, data curation. ZP: Supervision, Writing&#x2014;review and editing. YX: Supervision, writing&#x2014;review and editing. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the Establishment of Key Laboratory of Shenzhen Science and Technology Innovation Committee under Grant ZDSYS20190902093015527, the Shenzhen Science and Technology Innovation Committee under Grant JSGG20220831104402004, and Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies (2022B1212010005).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>Authors PS and ZP were employed by HBIS Digital Technology Co., Ltd.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ayarkwa </surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Influence of wood defects on some mechanical properties of two tropical Ghanaian hardwoods</article-title>. <source>J Ghana Sci Assoc</source> (<year>1999</year>) <volume>1</volume>:<fpage>131</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.4314/jgsa.v1i2.17813</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>X</given-names>
</name>
</person-group>. <source>Fully convolutional networks for surface defect inspection in industrial environment</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2017</year>).</citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>EDRNet: Encoder&#x2013;Decoder residual network for salient object detection of strip steel surface defects</article-title>. <source>IEEE Trans Instrumentation Meas</source>, <year>2020</year>, <volume>69</volume>:<fpage>1</fpage>&#x2013;. <pub-id pub-id-type="doi">10.1109/TIM.2020.3002277</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zaghdoudi</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Seridi</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Ziani</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Binary Gabor pattern (BGP) descriptor and principal component analysis (PCA) for steel surface defects classification[C]</article-title>. In: <conf-name>Proceeding of the 2020 International Conference on Advanced Aspects of Software Engineering (ICAASE)</conf-name>; <conf-date>November 2020</conf-date>. <publisher-name>IEEE</publisher-name> (<year>2020</year>). p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>W</given-names>
</name>
</person-group>. <article-title>Classification of defects in steel strip surface based on multiclass support vector machine</article-title>. <source>Multimedia tools Appl</source> (<year>2014</year>) <volume>69</volume>(<issue>1</issue>):<fpage>199</fpage>&#x2013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-012-1248-0</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>Two-stream convolutional neural network based on gradient image for aluminum profile surface defects classification and recognition</article-title>. <source>IEEE Access</source> (<year>2020</year>) <volume>8</volume>:<fpage>172152</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1109/access.2020.3025165</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>He</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Effect of dual-convolutional neural network model fusion for Aluminum profile surface defects classification and recognition</article-title>. <source>Math Biosciences Eng</source> (<year>2022</year>) <volume>19</volume>(<issue>1</issue>):<fpage>997</fpage>&#x2013;<lpage>1025</lpage>. <pub-id pub-id-type="doi">10.3934/mbe.2022046</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mayr</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Hoffmann</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Maier</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Christlein</surname>
<given-names>V</given-names>
</name>
</person-group>. <article-title>Weakly supervised segmentation of cracks on solar cells using normalized L p norm[C]</article-title>. In: <conf-name>Proceeding of the 2019 IEEE International Conference on Image Processing (ICIP)</conf-name>; <conf-date>September 2019</conf-date>. <publisher-name>IEEE</publisher-name> (<year>2019</year>). p. <fpage>1885</fpage>&#x2013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>K</given-names>
</name>
</person-group>. <article-title>Surface defect saliency of magnetic tile</article-title>. <source>Vis Comp</source> (<year>2020</year>) <volume>36</volume>(<issue>1</issue>):<fpage>85</fpage>&#x2013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1007/s00371-018-1588-5</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>W</given-names>
</name>
</person-group>. <article-title>Visual inspection of steel surface defects based on improved auxiliary classification generation adversarial network[J/OL]</article-title>. <source>Chin J Mech Eng</source> (<year>2023</year>) <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://kns.cnki.net/kcms/detail/11.2187.TH.20220526.1827.106.html">http://kns.cnki.net/kcms/detail/11.2187.TH.20220526.1827.106.html</ext-link>
</comment>.</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Panaretos</surname>
<given-names>VM</given-names>
</name>
<name>
<surname>Zemel</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Statistical aspects of Wasserstein distances</article-title>. <source>Annu Rev Stat its Appl</source> (<year>2019</year>) <volume>6</volume>:<fpage>405</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-statistics-030718-104938</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Radford</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Metz</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Chintala</surname>
<given-names>S</given-names>
</name>
</person-group>. <source>Unsupervised representation learning with deep convolutional generative adversarial networks[J]</source>. <comment>arXiv preprint</comment> (<year>2015</year>).</citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Odena</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Olah</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Shlens</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Conditional image synthesis with auxiliary classifier gans[C]</article-title>. In: <source>International conference on machine learning</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>PMLR</publisher-name> (<year>2017</year>). p. <fpage>2642</fpage>&#x2013;<lpage>51</lpage>.</citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dosovitskiy</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Brox</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>Generating images with perceptual similarity metrics based on deep networks</article-title>. <source>Adv Neural Inf Process Syst</source> (<year>2016</year>) <volume>29</volume>.</citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Pouget-Abadie</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Mirza</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Warde-Farley</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Ozair</surname>
<given-names>S</given-names>
</name>
<etal/>
</person-group> <article-title>Generative adversarial networks</article-title>. <source>Commun ACM</source> (<year>2020</year>) <volume>63</volume>(<issue>11</issue>):<fpage>139</fpage>&#x2013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1145/3422622</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>HP</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>CT</given-names>
</name>
</person-group>. <article-title>CNNs combined with a conditional GAN for mura defect classification in TFT-LCDs</article-title>. <source>IEEE Trans Semiconductor Manufacturing</source> (<year>2021</year>) <volume>34</volume>(<issue>1</issue>):<fpage>25</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1109/tsm.2020.3048631</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Q</given-names>
</name>
</person-group>. <article-title>Cross-domain fault diagnosis of rolling element bearings using deep generative neural networks</article-title>. <source>IEEE Trans Ind Elect</source> (<year>2018</year>) <volume>66</volume>(<issue>7</issue>):<fpage>5525</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1109/TIE.2018.2868023</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>BG</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Defect detection of fabrics with generative adversarial network based flaws modeling[C]</article-title>. In: <conf-name>Proceeding of the 2020 Chinese Automation Congress (CAC)</conf-name>; <conf-date>November 2020</conf-date>; <conf-loc>Shanghai, China</conf-loc>. <publisher-name>IEEE</publisher-name> (<year>2020</year>). p. <fpage>3334</fpage>&#x2013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheon</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>CO</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Convolutional neural network for wafer surface defect classification and the detection of unknown defect class</article-title>. <source>IEEE Trans Semiconductor Manufacturing</source> (<year>2019</year>) <volume>32</volume>(<issue>2</issue>):<fpage>163</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1109/tsm.2019.2902657</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nakazawa</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Kulkarni</surname>
<given-names>DV</given-names>
</name>
</person-group>. <article-title>Wafer map defect pattern classification and image retrieval using convolutional neural network</article-title>. <source>IEEE Trans Semiconductor Manufacturing</source> (<year>2018</year>) <volume>31</volume>(<issue>2</issue>):<fpage>309</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1109/tsm.2018.2795466</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Ge</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Deep learning-based classification of weld surface defects</article-title>. <source>Appl Sci</source> (<year>2019</year>) <volume>9</volume>(<issue>16</issue>):<fpage>3312</fpage>. <pub-id pub-id-type="doi">10.3390/app9163312</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wan</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>An improved VGG19 transfer learning strip steel surface defect recognition deep neural network based on few samples and imbalanced datasets</article-title>. <source>Appl Sci</source> (<year>2021</year>) <volume>11</volume>(<issue>6</issue>):<fpage>2606</fpage>. <pub-id pub-id-type="doi">10.3390/app11062606</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>D</given-names>
</name>
</person-group>. <article-title>Deep transfer network with joint distribution adaptation: A new intelligent fault diagnosis framework for industry application</article-title>. <source>ISA Trans</source> (<year>2020</year>) <volume>97</volume>:<fpage>269</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1016/j.isatra.2019.08.012</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Deep learning model for imbalanced multi-label surface defect classification</article-title>. <source>Meas Sci Tech</source> (<year>2021</year>) <volume>33</volume>(<issue>3</issue>):<fpage>035601</fpage>. <pub-id pub-id-type="doi">10.1088/1361-6501/ac41a6</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jain</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Seth</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Paruthi</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>EWR</given-names>
</name>
<name>
<surname>Lwin</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Yeo</surname>
<given-names>TT</given-names>
</name>
<etal/>
</person-group> <article-title>Pseudoaneurysm resulting in rebleeding after evacuation of spontaneous intracerebral hemorrhage</article-title>. <source>J Intell Manufacturing</source> (<year>2020</year>) <volume>143</volume>:<fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1016/j.wneu.2020.07.088</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Semi-supervised defect classification of steel surface based on multi-training and generative adversarial network</article-title>. <source>Opt Lasers Eng</source> (<year>2019</year>) <volume>122</volume>:<fpage>294</fpage>&#x2013;<lpage>302</lpage>. <pub-id pub-id-type="doi">10.1016/j.optlaseng.2019.06.020</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>A surface defect detection method based on positive samples[C]</article-title>. In: <source>Pacific rim international conference on artificial intelligence</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2018</year>). p. <fpage>473</fpage>&#x2013;<lpage>81</lpage>.</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lian</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Zareapoor</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>D</given-names>
</name>
</person-group>. <article-title>Deep-learning-based small surface defect detection via an exaggerated local variation-based generative adversarial network</article-title>. <source>IEEE Trans Ind Inform</source> (<year>2019</year>) <volume>16</volume>(<issue>2</issue>):<fpage>1343</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1109/TII.2019.2945403</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsai</surname>
<given-names>DM</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>SKS</given-names>
</name>
<name>
<surname>Chou</surname>
<given-names>YH</given-names>
</name>
</person-group>. <article-title>Auto-annotated deep segmentation for surface defect detection</article-title>. <source>IEEE Trans Instrumentation Meas</source> (<year>2021</year>) <volume>70</volume>:<fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1109/tim.2021.3087826</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dumoulin</surname>
<given-names>V</given-names>
</name>
<name>
<surname>Visin</surname>
<given-names>F</given-names>
</name>
</person-group>. <source>A guide to convolution arithmetic for deep learning</source> (<year>2016</year>). <comment>arXiv preprint arXiv:1603.07285</comment>.</citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Arjovsky</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Chintala</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Bottou</surname>
<given-names>L</given-names>
</name>
</person-group>. <source>Wasserstein GAN</source> (<year>2017</year>).</citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Thoma</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Acharya</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Gool</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Wasserstein divergence for gans[C]</article-title>. In: <source>Proceedings of the European conference on computer vision</source>. <publisher-loc>Switzerland</publisher-loc>: <publisher-name>ECCV</publisher-name> (<year>2018</year>). p. <fpage>653</fpage>&#x2013;<lpage>68</lpage>.</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>K</given-names>
</name>
<name>
<surname>He</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>Q</given-names>
</name>
</person-group>. <article-title>PGA-Net: Pyramid feature fusion and global context attention network for automated surface defect detection</article-title>. <source>IEEE Trans Ind Inform</source> (<year>2019</year>) <volume>16</volume>(<issue>12</issue>):<fpage>7448</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1109/TII.2019.2958826</pub-id>
</citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meng</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>F</given-names>
</name>
</person-group>. <article-title>An intelligent fault diagnosis method of small sample bearing based on improved auxiliary classification generative adversarial network</article-title>. <source>IEEE Sensors J</source> (<year>2022</year>) <volume>22</volume>(<issue>20</issue>):<fpage>19543</fpage>&#x2013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1109/jsen.2022.3200691</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>ZH</given-names>
</name>
<name>
<surname>Xue</surname>
<given-names>JH</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>J</given-names>
</name>
<etal/>
</person-group> <article-title>OSLNet: Deep small-sample classification with an orthogonal softmax layer</article-title>. <source>IEEE Trans Image Process</source> (<year>2020</year>) <volume>29</volume>:<fpage>6482</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2020.2990277</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>