<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2021.773142</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Style-Consistent Image Translation: A Novel Data Augmentation Paradigm to Improve Plant Disease Recognition</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Mingle</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1472694/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Yoon</surname> <given-names>Sook</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/595546/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Fuentes</surname> <given-names>Alvaro</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/551374/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yang</surname> <given-names>Jucheng</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1441808/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Park</surname> <given-names>Dong Sun</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/567101/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Electronics Engineering, Jeonbuk National University</institution>, <addr-line>Jeonbuk</addr-line>, <country>South Korea</country></aff>
<aff id="aff2"><sup>2</sup><institution>Core Research Institute of Intelligent Robots, Jeonbuk National University</institution>, <addr-line>Jeonbuk</addr-line>, <country>South Korea</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Computer Engineering, Mokpo National University</institution>, <addr-line>Jeonnam</addr-line>, <country>South Korea</country></aff>
<aff id="aff4"><sup>4</sup><institution>College of Artificial Intelligence, Tianjin University of Science and Technology</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yuzhen Lu, Mississippi State University, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Borja Espejo-Garca, Agricultural University of Athens, Greece; Dong Chen, Michigan State University, United States; Junfeng Gao, University of Lincoln, United Kingdom; Ebenezer Olaniyi, Mississippi State University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Sook Yoon <email>syoon&#x00040;mokpo.ac.kr</email></corresp>
<corresp id="c002">Dong Sun Park <email>dspark&#x00040;jbnu.ac.kr</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Sustainable and Intelligent Phytoprotection, a section of the journal Frontiers in Plant Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>02</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>773142</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>12</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Xu, Yoon, Fuentes, Yang and Park.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Xu, Yoon, Fuentes, Yang and Park</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Deep learning shows its advantages and potentials in plant disease recognition and has witnessed a profound development in recent years. To obtain a competing performance with a deep learning algorithm, enough amount of annotated data is requested but in the natural world, scarce or imbalanced data are common, and annotated data is expensive or hard to collect. Data augmentation, aiming to create variations for training data, has shown its power for this issue. But there are still two challenges: creating more desirable variations for scarce and imbalanced data, and designing a data augmentation to ease object detection and instance segmentation. First, current algorithms made variations only inside one specific class, but more desirable variations can further promote performance. To address this issue, we propose a novel data augmentation paradigm that can adapt variations from one class to another. In the novel paradigm, an image in the source domain is translated into the target domain, while the variations unrelated to the domain are maintained. For example, an image with a healthy tomato leaf is translated into a powdery mildew image but the variations of the healthy leaf are maintained and transferred into the powdery mildew class, such as types of tomato leaf, sizes, and viewpoints. Second, current data augmentation is suitable to promote the image classification model but may not be appropriate to alleviate object detection and instance segmentation model, mainly because the necessary annotations can not be obtained. In this study, we leverage a prior mask as input to tell the area we are interested in and reuse the original annotations. In this way, our proposed algorithm can be utilized to do the three tasks simultaneously. Further, We collect 1,258 images of tomato leaves with 1,429 instance segmentation annotations as there is more than one instance in one single image, including five diseases and healthy leaves. Extensive experimental results on the collected images validate that our new data augmentation algorithm makes useful variations and contributes to improving performance for diverse deep learning-based methods.</p></abstract>
<kwd-group>
<kwd>tomato disease recognition</kwd>
<kwd>data augmentation</kwd>
<kwd>image translation</kwd>
<kwd>image classification</kwd>
<kwd>instance segmentation</kwd>
<kwd>image style</kwd>
</kwd-group>
<contract-num rid="cn001">2019R1A6A1A09031717</contract-num>
<contract-num rid="cn002">421027-04</contract-num>
<contract-sponsor id="cn001">National Research Foundation of Korea<named-content content-type="fundref-id">10.13039/501100003725</named-content></contract-sponsor>
<contract-sponsor id="cn002">Ministry of Agriculture, Food and Rural Affairs<named-content content-type="fundref-id">10.13039/501100003624</named-content></contract-sponsor>
<counts>
<fig-count count="9"/>
<table-count count="7"/>
<equation-count count="11"/>
<ref-count count="39"/>
<page-count count="16"/>
<word-count count="9700"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Food security has been raised to a high level in many countries partly because food distribution is not compatible with the distribution of the population in the world, and the number of laborers related is becoming less. While enough amount of food is required to feed our humans, many factors may harm plant growth and hence threaten the food supply. Controlling disease is one of the key challenges to keep plants healthy toward obtaining the expected yield. Artificial intelligence has recently witnessed a booming development with a decent disease recognition performance as intelligent machines deployed in farms can reduce workload. Deep learning, a core technique of artificial intelligence, has been successfully adopted to recognize diseases or abnormalities, such as tomato (Fuentes et al., <xref ref-type="bibr" rid="B7">2017</xref>; Liu and Wang, <xref ref-type="bibr" rid="B21">2020</xref>; Wang et al., <xref ref-type="bibr" rid="B34">2021</xref>), banana (Lin et al., <xref ref-type="bibr" rid="B20">2021</xref>), potato (Gao et al., <xref ref-type="bibr" rid="B9">2021</xref>), corn and apple (Zhong and Zhao, <xref ref-type="bibr" rid="B37">2020</xref>), and many other plants (Gao et al., <xref ref-type="bibr" rid="B8">2020</xref>; Liu and Wang, <xref ref-type="bibr" rid="B22">2021</xref>). Recent studies (Martineau et al., <xref ref-type="bibr" rid="B23">2017</xref>; Liu and Wang, <xref ref-type="bibr" rid="B22">2021</xref>; Saranya et al., <xref ref-type="bibr" rid="B29">2021</xref>) show the advantages and potentialities of deep learning methods compared to other methods, such as handcraft feature, in recognizing plant diseases and related tasks.</p>
<p>To obtain a competing performance with a deep learning algorithm, enough amount of annotated data is requested but in the natural world, scarce or imbalanced data are common, and annotated data is expensive or hard to collect. For example, some diseases rarely appear or even never appear on one farm but the healthy plant is more common, which can not lead to a convincing disease recognition performance. To address this challenge, data augmentation is one of the most potential solutions and has been utilized in the agricultural field (Zhu et al., <xref ref-type="bibr" rid="B39">2018</xref>; Nazki et al., <xref ref-type="bibr" rid="B24">2020</xref>; Abbas et al., <xref ref-type="bibr" rid="B1">2021</xref>). Data augmentation aims to generate more data with collected <italic>training</italic> data to improve the deep learning model&#x00027;s performance in <italic>testing</italic> data. Previous studies (Pawara et al., <xref ref-type="bibr" rid="B26">2017</xref>; Douarre et al., <xref ref-type="bibr" rid="B5">2019</xref>; Pinto Sampaio Gomes and Zheng, <xref ref-type="bibr" rid="B27">2020</xref>; Liu and Wang, <xref ref-type="bibr" rid="B22">2021</xref>; Saranya et al., <xref ref-type="bibr" rid="B29">2021</xref>) have validated that data augmentation plays a significant role to improve the performance of deep learning in the agricultural area. In this study, we are interested in image-based recognition for tomato diseases, and hence data augmentation defaults with image-based.</p>
<p>Traditional data augmentation methods generate new augmented data within one specific class, <italic>within-class data augmentation</italic>, where the appearance of the image can be changed but the corresponding class remains, such as rotating or translating an image (Hu et al., <xref ref-type="bibr" rid="B14">2020</xref>; Gorad and Kotrappa, <xref ref-type="bibr" rid="B11">2021</xref>). In contrast, <italic>cross-class data augmentation</italic> methods can translate one image from the class to another class <italic>via image translation</italic> that aims to translate images from a source domain to a target domain, by which the variations are desired to be borrowed from one class to another class. For example, a healthy tomato image, source domain, can be translated into a powdery mildew image, target domain. In general, healthy tomato leaf images are easy to collect with large variations, such as background, viewpoint, size of the leaf, and type of tomato, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. As there is a high variation for the healthy tomato images, we refer the healthy to a <italic>variation-majority class</italic>. On the other hand, the class, hard to obtain images or enough variation, is referred to as a <italic>variation-minority class</italic>, such as some disease tomato leaves. To achieve the cross-class data augmentation, CycleGAN (Zhu et al., <xref ref-type="bibr" rid="B38">2017</xref>), one of the state-of-the-art methods to do image translation, is plausible to be utilized. Based on CycleGAN, Nazki et al. (<xref ref-type="bibr" rid="B24">2020</xref>) proposed an activation reconstruction loss to improve the quality of the generated image. Except for low image quality, CycleGAN tends to change the undesired content such as background, and an attention module was proposed in LeafGAN (Cap et al., <xref ref-type="bibr" rid="B4">2020</xref>) to detect the area that we are interested in. Although they achieved better results than the original CycleGAN, the following two new challenges are addressed in this study.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Data augmentation from healthy tomato leaves to powdery mildew leaves using the proposed style-consistent image translation (SCIT) model. Healthy leaves are in the source domain, belonging to the variation-majority class, while powdery mildew leaves are in the target domain, belonging to the variation-minority class. Ours SCIT model is leveraged to translate the images in the source domain into the target domain, which can take the variations from the variation-majority class to the variation-minority class. Healthy leaves include several variations, background, type of tomato, viewpoint, shape, illumination, size. SCIT model can <italic>only</italic> translate one given leaf in one image as shown right three images with red boundary, compared to other generative adversarial networks (GAN)-based data augmentation methods which translate all leaves in one image (Cap et al., <xref ref-type="bibr" rid="B4">2020</xref>; Nazki et al., <xref ref-type="bibr" rid="B24">2020</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0001.tif"/>
</fig>
<p><italic>First, how can we adapt the majority of variations from the source domain to the target domain?</italic> Image translation is expected to keep the variations but there are no guarantees to keep the variations in current algorithms. To ease this issue, we propose a new paradigm to combine style consistency and image translation to maintain the variations during the image translation process, and hence, we call our algorithm <underline>s</underline>tyle-<underline>c</underline>onsistent <underline>i</underline>mage <underline>t</underline>ranslation (SCIT). Specifically, we follow the hypothesis that images can be factorized into two parts, <italic>label-related</italic> and <italic>style-related</italic> (Gonzalez-Garcia et al., <xref ref-type="bibr" rid="B10">2018</xref>; Lee et al., <xref ref-type="bibr" rid="B18">2018</xref>). While the label-related are the characteristics or patterns of specific classes such as healthy and powdery mildew, the style-related is independent of the labels, such as illumination, viewpoints, and background. In the translation process, we aim to translate the label-related but keep the style-related, which contributes to adapting the styles from the source domain to the target domain. In <xref ref-type="fig" rid="F1">Figure 1</xref>, we can see that the selected healthy tomato leaves are translated into powdery mildew leaves by our method, but the styles in healthy tomato leaves are maintained, and the variations of powdery mildew domain are augmented from the healthy domain.</p>
<p><italic>Second, how can we use image translation as data augmentation to ease object detection and instance segmentation?</italic> Traditional data augmentation and current image translation-based algorithms can be leveraged to improve the image classification but may not be appropriate to alleviate object detection and instance segmentation which are closer to our practical applications. The first reason is that object detection and instance segmentation require more annotations than image classification but current algorithms can not make those necessary annotations. Generally, we only need class information of images to train the image classification model, but the exact locations of each class are necessary to train the object detection algorithm, and both location and instance identity are required to train the instance segmentation model, <bold>Figure 9</bold> giving examples about the two tasks. Furthermore, classification is an image-level task but object detection and instance segmentation are at the instance level. Elusively, the input image undergoes a preprocessing to be one leaf or even part of one leaf to do image classification by which classification model is easier to be trained (Nazki et al., <xref ref-type="bibr" rid="B24">2020</xref>). In contrast, preprocessing is not necessary for object detection and instance segmentation. In fact, the leaves are normally in different scales, as shown in the red boundary images in <xref ref-type="fig" rid="F1">Figure 1</xref>. Therefore, we are requested to translate leaf instances separately in diverse scales instead of translating all leaves as other algorithms have been doing (Cap et al., <xref ref-type="bibr" rid="B4">2020</xref>; Nazki et al., <xref ref-type="bibr" rid="B24">2020</xref>). To address the two issues, we employ a mask as prior knowledge in order to split an image into the <underline>r</underline>egion <underline>o</underline>f <underline>i</underline>nterest (ROI) and <underline>b</underline>ack<underline>g</underline>round (BG). We aim to translate its ROI part but reuse its BG part, by which the original annotations can be reused for the produced images. Further, we design a new framework based on CycleGAN (Zhu et al., <xref ref-type="bibr" rid="B38">2017</xref>). First of all, a mask encoder is designed to be incorporated with the image encoder, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, by which our generator knows where is interested. Although a similar idea appears in RBGAN (Xu et al., <xref ref-type="bibr" rid="B35">2021</xref>), we aim to translate part of the image yet keep the other part and maintain the style during the image translation but RBGAN aims to perform instance-level image translation with a decent translated instance boundary. Besides, our discriminator absorbs both real or fake images and corresponding masks and is pushed to know where is the translated area and whether the area is real or fake. Therefore, with the new generator and discriminator along with the input mask, our algorithm can translate given leaves in diverse scales with the reused annotations, which contributes to ease object detection and instance segmentation as a data augmentation method.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>(A)</bold> Proposed SCIT model, the general structure of the proposed data augmentation model for image classification, object detection, and instance segmentation. <italic>x</italic><sup><italic>s</italic></sup> and <italic>x</italic>&#x02032; denote the <underline>s</underline>ource image and the generated image by our model. While <italic>m</italic><sup><italic>s</italic></sup> denotes instance segmentation <underline>m</underline>ask aligning <underline>s</underline>ource image <italic>x</italic><sup><italic>s</italic></sup> and expecting to align <italic>x</italic>&#x02032;, <italic>m</italic><sup><italic>t</italic></sup> is the instance segmentation <underline>m</underline>ask of real images in the <underline>t</underline>arget domain. The generator consists of an image encoder <italic>E</italic><sub><italic>I</italic></sub>, a mask encoder <italic>E</italic><sub><italic>M</italic></sub>, and a decoder <italic>Dec</italic>. Discriminator <italic>Dis</italic> pushes the region of interest (ROI) of the generated image to have the same label with the ROI of the real image in the target domain, while the pre-trained VGG pushes ROIs to share the same style. <bold>(B)</bold> Flow chart to compute identity loss and cycle-consistent loss. <italic>G</italic><sub><italic>T</italic></sub> and <italic>G</italic><sub><italic>S</italic></sub> is the generator to produce images in the domain <italic>T</italic> and <italic>S</italic>, respectively. The masks to generate the image are omitted.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0002.tif"/>
</fig>
<p>To summarize, our contributions are as follows:</p>
<list list-type="order">
<list-item><p>We propose a data augmentation paradigm, SCIT, which can increase the data variations for the variation-minority class by leveraging the images from the variation-majority class.</p></list-item>
<list-item><p>We propose a framework to perform image translation in diverse scales with necessary annotations as output to ease object detection and instance segmentation, which is out of current data augmentation methods.</p></list-item>
<list-item><p>Taking tomato as an example, we perform extensive experiments on three tasks, image classification, object detection, and instance segmentation. The experimental results suggest that our proposed algorithm improves the performances for diverse deep learning-based methods and outperforms the state-of-the-art data augmentation methods.</p></list-item>
</list>
<p>The remainder of this study is organized as follows. Related studies and our basic idea are introduced in the preliminary section. The proposed method to do data augmentation is instantiated in Section 3, including the framework and loss function. In the experiments section, we show the details about our dataset, implementation to train and test our model, ablation study to understand our algorithm, comparison to other methods in three tasks. Finally, we conclude our studies and future study in the last section.</p></sec>
<sec id="s2">
<title>2. Preliminary</title>
<p>Data augmentation based on the image can be categorized into two main parts, basic image manipulations and deep learning-based algorithms. In this section, we try to highlight the difference between other methods and our method to achieve data augmentation.</p>
<p><bold>Image manipulations</bold>. Image manipulations make use of image processing methods, such as pixel-wise conversion and geometrical transformations. Formally, let <italic>x</italic><sup><italic>s</italic></sup> and <italic>x</italic><sup><italic>a</italic></sup> denote the <underline>s</underline>ource image and the <underline>a</underline>ugmented image. Similarly, <italic>y</italic><sup><italic>s</italic></sup> and <italic>y</italic><sup><italic>a</italic></sup> are corresponding labels. The formulation of basic image manipulation-based data augmentation refers to Equation 1.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M55"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>y</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>f</italic><sub><italic>m</italic></sub> is one of the basic image <underline>m</underline>anipulation functions. The equation suggests that the label of the input image is not changed when the input image undergoes manipulations. The retainment degree of the label is called safety of data augmentation because manipulations can not always keep the label (Shorten and Khoshgoftaar, <xref ref-type="bibr" rid="B31">2019</xref>). For example, after adding too much random noise, some images could not be recognized as before. With these image manipulations, prior works achieved better performance in the missions with small datasets (Hu et al., <xref ref-type="bibr" rid="B14">2020</xref>; Gorad and Kotrappa, <xref ref-type="bibr" rid="B11">2021</xref>). Besides, these basic image manipulation can be adapted from one image to more than one image (Dwibedi et al., <xref ref-type="bibr" rid="B6">2017</xref>). Kuznichov et al. (<xref ref-type="bibr" rid="B17">2019</xref>) copied the leaves from different images to form a new augmented image to promote leaf segmentation and counting. Similarly, Gao et al. (<xref ref-type="bibr" rid="B8">2020</xref>) produced a synthetic image by combining specific objects from different images, in which two different classes can appear in a single image. Their experiments validated that the performance can be also improved by their methods. In this study, we aim to do data augmentation from another viewpoint by using one class to augment another class, hoping that more variations can be produced.</p>
<p><bold>Deep learning-based algorithms</bold>. Different from image manipulations, deep learning-based algorithms employ deep neural networks to generate new images. According to the condition to generate new images, deep learning-based algorithms can be split into label-condition and image-condition. Label-condition algorithms generate images from given labels by using generative adversarial networks (GANs) (Valerio Giuffrida et al., <xref ref-type="bibr" rid="B33">2017</xref>; Pandian et al., <xref ref-type="bibr" rid="B25">2019</xref>; Bi and Hu, <xref ref-type="bibr" rid="B2">2020</xref>; Abbas et al., <xref ref-type="bibr" rid="B1">2021</xref>). In contrast, image-condition algorithms produce images from given images. Style transfer is one of the possible methods (Li et al., <xref ref-type="bibr" rid="B19">2017</xref>; Huang et al., <xref ref-type="bibr" rid="B15">2021</xref>; Shen et al., <xref ref-type="bibr" rid="B30">2021</xref>). Mathematically, it can be formalized as Equation 2.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M56"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>y</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>f</italic><sub><italic>st</italic></sub> is the <underline>s</underline>tyle <underline>t</underline>ransfer function. Because of the function, style transfer algorithms can introduce more variations into the dataset. For instance, images captured in sunny style can be transferred into night style (Shen et al., <xref ref-type="bibr" rid="B30">2021</xref>), which is beyond the basic image manipulations. However, the labels of the transferred images are considered as same as the source, which is the same as the basic image manipulations. Besides, the style transferring-based algorithms try to maintain the content of the image. Take the tomato leaves as an example, it aims to maintain the size, shape of leaves, viewpoints and, hence, making more variations about them is still challenging.</p>
<p>To deal with the challenge mentioned above, we propose a novel paradigm to achieve data augmentation, following the disentangled idea that an image can be factorized into two factors: style-related and label-related (Gonzalez-Garcia et al., <xref ref-type="bibr" rid="B10">2018</xref>; Lee et al., <xref ref-type="bibr" rid="B18">2018</xref>). By using image translation, label-related factors of an image in the source domain can be translated into images in the target domain. Simultaneously, the style-related factors of the source image are desired to be kept in the translated image. In the style-consistent image translation, the variations of the style in the source domain are borrowed into the target domain. Mathematically, this kind of data augmentation algorithm can be formulated as Equation 3.</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M58"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>y</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>&#x02260;</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>f</italic><sub><italic>it</italic></sub> symbolizes <underline>i</underline>mage <underline>t</underline>ranslation function. <italic>y</italic><sup><italic>a</italic></sup> and <italic>y</italic><sup><italic>s</italic></sup> denote the label corresponding to image <italic>x</italic><sup><italic>s</italic></sup> and <italic>x</italic><sup><italic>a</italic></sup>, respectively. The function <italic>g</italic> symbolizes the style extracting function.</p>
<p>We argue that this kind of data augmentation eases practical applications. When it is hard to collect the data or the collected data suffer from a lack of variations in one domain but easier to collect data in another domain, we can use this data augmentation to leverage the variations in the easier domain to promote the variations in the harder domain. For example, images of healthy tomato leaves can be easily collected from farms but images with specific diseases of abnormalities like powdery mildew could not be collected easily, mainly because farmers must do necessary measures to prevent before their appearance or make a fast remedy after their appearance to reduce financial loss. We can augment data for powdery mildew effectively by using image translation from healthy domain to powdery mildew domain. Besides, we emphasize that the proposed data augmentation method can be employed with any kind of image translation model <italic>f</italic><sub><italic>it</italic></sub> and domain-invariant function <italic>g</italic>.</p>
<p>Moreover, we also notice that data augmentation for image classification attracts much more attention than for object detection or instance segmentation. A label is globally assigned to a whole image for the image classification task but a label is locally assigned to a region of an image for object detection or instance segmentation task. Therefore, one of the challenges for object detection and instance segmentation is mainly that dealing with a bounding box or instance mask for a local region is required. To address this issue, we further spatially split one image into two parts, <underline>r</underline>egion <underline>o</underline>f <underline>i</underline>nterest (ROI) <italic>x</italic><sub><italic>roi</italic></sub> and <underline>b</underline>ack<underline>g</underline>round <italic>x</italic><sub><italic>bg</italic></sub>. Formally, input image <italic>x</italic> is split into <italic>x</italic><sub><italic>roi</italic></sub> and <italic>x</italic><sub><italic>bg</italic></sub> according to the binary instance segmentation <underline>m</underline>ask from <underline>s</underline>ource domain <italic>m</italic><sup><italic>s</italic></sup>. Then our proposed generative adversarial network takes the <italic>x</italic><sup><italic>s</italic></sup> and <italic>m</italic><sup><italic>s</italic></sup> as input and aims to translate the <italic>x</italic><sub><italic>roi</italic></sub> into the target domain. Simultaneously, we reuse the background of the source image since the translation model is not interested in the background. In this way, the augmented image <italic>x</italic><sup><italic>a</italic></sup> shares the same bounding box or instance segmentation with the source image <italic>x</italic><sup><italic>s</italic></sup>, but the label is changed. Formally, the SCIT for object detection and instance segmentation can be formalized as Equation 4.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M66"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>m</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>*</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>m</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>*</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mi>a</mml:mi></mml:msubsup><mml:mo>&#x02260;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>a</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M5"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the label corresponding to the ROI of <italic>x</italic><sup><italic>a</italic></sup> and <italic>x</italic><sup><italic>s</italic></sup>. We found this kind of method suitable for image classification.</p></sec>
<sec id="s3">
<title>3. Style-Consistent Image Translation</title>
<p>In this section, we instantiate <italic>f</italic><sub><italic>it</italic></sub> and <italic>g</italic> as a GAN model and VGG network and deploy them to do image classification (Simonyan and Zisserman, <xref ref-type="bibr" rid="B32">2014</xref>), object detection, and instance segmentation for tomato leaves. Specifically, an updated CycleGAN (Zhu et al., <xref ref-type="bibr" rid="B38">2017</xref>) is leveraged in our experiments to translate images from the source domain to the target domain. To keep the style consistent, VGG loss is employed (Huang and Belongie, <xref ref-type="bibr" rid="B16">2017</xref>; Li et al., <xref ref-type="bibr" rid="B19">2017</xref>). But we emphasize that other kinds of image translation methods and style losses are possible and encouraged. Since our method aims to keep the style consistent when doing image translation, it is called SCIT.</p>
<sec>
<title>3.1. Framework</title>
<p>Style-consistent image translation consists of three parts functionally, as shown in <xref ref-type="fig" rid="F2">Figure 2A</xref>. The Generator, <italic>G</italic> for short, is expected to translate the image, while the discriminator <italic>Dis</italic> is assumed to push the translated image similar to the real image in the target domain, and a pre-trained VGG19 is utilized to extract the style, class-unrelated characters.</p>
<p>The generator <italic>G</italic> absorbs source image <italic>x</italic><sup><italic>s</italic></sup> and instance segmentation <underline>m</underline>ask from <underline>s</underline>ource image <italic>m</italic><sup><italic>s</italic></sup> as input, in which two specific encoders, <italic>E</italic><sub><italic>I</italic></sub> and <italic>E</italic><sub><italic>M</italic></sub>, are leveraged to extract features from image and mask, respectively. The outputs of the two encoders are concatenated, followed by a decoder to produce an output image <italic>x</italic>&#x02032;. Formally, the generating process can be formalized as Equation 5.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02295;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x02295; denotes concatenation channel-wise. The instance segmentation mask <italic>m</italic><sup><italic>s</italic></sup> is used to let the generator know where to put attention. Similarly, our discriminator takes in real image <italic>x</italic><sup><italic>t</italic></sup> in the target domain or generated image <italic>x</italic>&#x02032;, along with its instance mask <italic>m</italic><sup><italic>s</italic></sup>, to recognize whether it is fake or real.</p>
<p>In the inference time, the image <italic>x</italic><sup><italic>a</italic></sup> to be used as augmented data is a fusion of the source image <italic>x</italic><sup><italic>s</italic></sup> and the generated image <italic>x</italic>&#x02032;. Equation 6 shows the inference process to get augmented data. Intuitively, the augmented image has the same background as the source image but has the same foreground as the translated image.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>*</mml:mo><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02295;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>*</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></sec>
<sec>
<title>3.2. Loss Functions</title>
<p>The loss functions employed to train our SCIT model are explained in this subsection. To push the generated image toward the real image in the target domain, GAN loss is used as shown in Equation 7.</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>D</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02295;</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>D</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02295;</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Except for being like a real image in the target domain, the generated image <italic>x</italic>&#x02032; is hypothesized to have the same style as the source image <italic>x</italic><sup><italic>s</italic></sup>, which is realized by a pre-trained VGG network, as Equation 8.</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>*</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>*</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>k</italic> denotes a kernel function, and &#x003D5; is an image feature extractor. A linear kernel function is employed in this study, as it gives a decent performance and requires less computation (Li et al., <xref ref-type="bibr" rid="B19">2017</xref>). We apply a pretrained VGG-19 network as the feature extractor without finetuning. Thus, &#x003D5;<sub><italic>i</italic></sub> means a layer in VGG-19, which is similar to the literature. In our experiments, we use <italic>relu</italic>1_1, <italic>relu</italic>2_1, <italic>relu</italic>3_1, <italic>relu</italic>4_1, and <italic>relu</italic>5_1 layers with equal weights and, thus, <italic>n</italic> &#x0003D; 5. Intuitively, the style loss pushes the two images to share the same feature distribution. Specifically, the linear kernel function-based loss encourages them to have the same sample mean in feature distribution space. We refer to Li et al. (<xref ref-type="bibr" rid="B19">2017</xref>) to check the detail about the style loss and its related kernel functions. The deep learning-based style loss with a linear kernel function is adopted in our experiment, but other methods, such as different kernel functions and new style loss are theoretically possible.</p>
<p>To ease the training of generator model <italic>G</italic>, identity loss and cycle-consistency loss are leveraged in our experiments, as shown in <xref ref-type="fig" rid="F2">Figure 2B</xref> (Zhu et al., <xref ref-type="bibr" rid="B38">2017</xref>). To describe clearly the cycle-consistency loss, we use subscript with <italic>S</italic> and <italic>T</italic> to denote the source domain and target domain. For instance, <italic>G</italic><sub><italic>S</italic></sub> means the generator which aims to translate an image into the <underline>s</underline>ource domain while <italic>G</italic><sub><italic>T</italic></sub> denotes the generator which aims to translate an image into the <underline>t</underline>arget domain. Identity loss, defined in Equation 9, comes from that when one instance in the target domain is given to generator <italic>G</italic><sub><italic>T</italic></sub> as input, the generator needs to output the same as its input without any change. Equation 10 shows the way to calculate the cycle-consistent loss. When we translate an image in the source domain into the target domain and translate the result back into the source domain, we want to get the same result as the original input.</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the end, our model is trained with four loss functions described before. Mathematically, Equation 11 shows the sum of the four losses, where &#x003BB;<sub>&#x0002A;</sub> balances each loss.</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></sec></sec>
<sec id="s4">
<title>4. Experiments</title>
<sec>
<title>4.1. Dataset and Implementation Details</title>
<p><bold>Dataset</bold>. We aim to recognize tomato diseases among different farms and collect data in real farms with many variations, such as the type of tomato, the distance between the camera and tomato leaf, weather, and illumination. <xref ref-type="fig" rid="F3">Figure 3</xref> gives examples of the collected images. A total of 1,258 images of tomato leaves are collected, called original data, which covers five types of disease and healthy leaves. <xref ref-type="table" rid="T1">Table 1</xref> displays the number of images for each class. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the original data are first split into 40% testing and 60% training data, respectively. We adopt the training data to train data augmentation models and utilize all healthy leaves images as testing to get the augmented data for the other five diseases. <xref ref-type="table" rid="T1">Table 1</xref> shows the number of augmented data for each class. Since we are not interested in healthy leaves, we do not do data augmentation for the healthy class. A total of 314 generated images and 358 instances are generated for each disease class. Finally, the augmented data and the training data are leveraged to train task models (classification, object detection, and instance segmentation). While object detection and instance segmentation can give more than one label or one instance to one image, image classification requires one holistic label for one image. To ease image classification, we crop the original image to get a single leaf in one image.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Original tomato leaf dataset. From the first to the last row are Healthy, Powdery (powdery mildew), Canker, LMold (leaf mold), ToCV, MagDef (magnesium deficiency). We collect the dataset from different farms at different times. The variations among the dataset include background, type of tomato, the severity of disease, illumination condition, the distance between the camera and interesting leaf, viewpoint to take the picture.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0003.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Dataset used in the experiment.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="6" style="border-bottom: thin solid #000000;"><bold>Original data</bold></th>
<th valign="top" align="center" colspan="2"><bold>Augmented data</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>All</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Testing data</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Training data</bold></th>
<th/>
</tr>
<tr>
<th valign="top" align="left"><bold>Type</bold></th>
<th valign="top" align="center"><bold>Images</bold></th>
<th valign="top" align="center"><bold>Masks</bold></th>
<th valign="top" align="center"><bold>Images</bold></th>
<th valign="top" align="center"><bold>Masks</bold></th>
<th valign="top" align="center"><bold>Images</bold></th>
<th valign="top" align="center"><bold>Masks</bold></th>
<th valign="top" align="center" style="border-top: thin solid #000000;"><bold>Images</bold></th>
<th valign="top" align="center" style="border-top: thin solid #000000;"><bold>Masks</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Healthy</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">Powdery</td>
<td valign="top" align="center">170</td>
<td valign="top" align="center">185</td>
<td valign="top" align="center">71</td>
<td valign="top" align="center">73</td>
<td valign="top" align="center">106</td>
<td valign="top" align="center">112</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
</tr>
<tr>
<td valign="top" align="left">Canker</td>
<td valign="top" align="center">298</td>
<td valign="top" align="center">309</td>
<td valign="top" align="center">119</td>
<td valign="top" align="center">122</td>
<td valign="top" align="center">185</td>
<td valign="top" align="center">187</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
</tr>
<tr>
<td valign="top" align="left">Leaf mold</td>
<td valign="top" align="center">166</td>
<td valign="top" align="center">198</td>
<td valign="top" align="center">74</td>
<td valign="top" align="center">78</td>
<td valign="top" align="center">105</td>
<td valign="top" align="center">120</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
</tr>
<tr>
<td valign="top" align="left">ToCV</td>
<td valign="top" align="center">172</td>
<td valign="top" align="center">227</td>
<td valign="top" align="center">81</td>
<td valign="top" align="center">89</td>
<td valign="top" align="center">115</td>
<td valign="top" align="center">138</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
</tr>
<tr>
<td valign="top" align="left">Mag Def</td>
<td valign="top" align="center">138</td>
<td valign="top" align="center">152</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">59</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">93</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">358</td>
</tr>
<tr>
<td valign="top" align="left">All</td>
<td valign="top" align="center">999</td>
<td valign="top" align="center">1,071</td>
<td valign="top" align="center">402</td>
<td valign="top" align="center">421</td>
<td valign="top" align="center">597</td>
<td valign="top" align="center">650</td>
<td valign="top" align="center">1,570</td>
<td valign="top" align="center">1,790</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>It includes healthy tomato leaves and five kinds of diseases of abnormalities. To do the tasks (image classification, object detection, and instance segmentation), the original data are split into two parts according to the number of masks, 60% training and 40% as testing. The healthy images are only used to do image translation</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Data utilization process. The originally collected data are split into testing and training data, 0.4 and 0.6, respectively. The training data is firstly adopted to train our SCIT model that then generates the augmented data. Next, the training data along with the augmented data are leveraged to train the task model (image classification, object detection, or instance segmentation). The dotted lines suggest the data augmentation process. The SCIT model is one of our main contributions, which introduces new variations for the original data and, thus, encourages the task model to have better performance on the testing data.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0004.tif"/>
</fig>
<p><bold>Implementation details</bold>. The original training dataset is employed to train our data augmentation model. Meanwhile, basic image manipulation is adopted to enlarge the original data since the collected data are not enough to train the SCIT model. Specifically, we use three times random brightening or darkening, and three times random cropping. Hence, the training data are enlarged to six times in total. During the training process of our data augmentation model, random flip in horizontal and vertical are employed and every image is resized to 256 in both height and width. Each type of disease employs one specific SCIT model.</p>
<p>We use Adam optimizer to train our model for 100 epochs with a learning rate of 0.0002. The batch size is set as 6 with three TITAN V GPUs (12 GB memory). After training, the trained models are adopted to generate disease images, which are later taken as augmented data to train task models. Every training process for each class and &#x003BB;<sub><italic>sty</italic></sub> roughly spends 5 h. Therefore, all translation models require about 26 days (5 translation models for each class and 5 &#x003BB;<sub><italic>sty</italic></sub> settings, each setting is executed 5 times.)</p>
<p><bold>Architectures</bold>. The proposed SCIT model consists of three sub-models. First, the generator consists of an image encoder, mask encoder, and decoder. The image encoder, aiming to extract necessary information from the input images, leverages several stacks of convolution-ReLU-BatchNorm layers and nine residual blocks, while the mask encoder only adopts the same number of stacks of convolution-ReLU-BatchNorm layers without residual blocks since the mask is much simpler than images. In contrast, the decoder, aiming to produce bigger size images from the smaller size of the feature map, employs several stacks of deconvolution-ReLU-BatchNorm. Second, the discriminator also applies several stacks of convolution-LeakyReLU-BatchNorm. The details of our generator and discriminator are referred to in <xref ref-type="table" rid="T2">Table 2</xref> and our codes will be public soon<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. In terms of the style computation model, a pretrained VGG19 model<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> is leveraged.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>The architecture details adopted in our algorithm.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Network</bold></th>
<th valign="top" align="center"><bold>Input size</bold></th>
<th valign="top" align="left"><bold>Operation</bold></th>
<th valign="top" align="left"><bold>Normalization</bold></th>
<th valign="top" align="left"><bold>Active function</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>E</italic><sub><italic>I</italic></sub></td>
<td valign="top" align="center">(256, 256, 3)</td>
<td valign="top" align="left">Conv7-C64-S1-P3</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(256, 256, 64)</td>
<td valign="top" align="left">Conv3-C128-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(128, 128, 128)</td>
<td valign="top" align="left">Conv3-C256-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(64, 64, 256)</td>
<td valign="top" align="left" colspan="3">Residual block * 9</td>
</tr>
<tr>
<td valign="top" align="left"><italic>E</italic><sub><italic>M</italic></sub></td>
<td valign="top" align="center">(256, 256, 3)</td>
<td valign="top" align="left">Conv7-C64-S1-P3</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(256, 256, 64)</td>
<td valign="top" align="left">Conv3-C128-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(128, 128, 128)</td>
<td valign="top" align="left">Conv3-C256-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Dec</italic></td>
<td valign="top" align="center">(64, 64, 512)</td>
<td valign="top" align="left">Conv1-C256-S1-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(64, 64, 256)</td>
<td valign="top" align="left">DeConv3-C128-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(128, 128, 128)</td>
<td valign="top" align="left">DeConv3-C64-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">ReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(256, 256, 64)</td>
<td valign="top" align="left">Conv7-C3-S1-P0</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">Tanh</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Dis</italic></td>
<td valign="top" align="center">(256, 256, 3)</td>
<td valign="top" align="left">Conv4-C64-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">LeakyReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(128, 128, 64)</td>
<td valign="top" align="left">Conv4-C128-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">LeakyReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(64, 64, 128)</td>
<td valign="top" align="left">Conv4-C256-S2-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">LeakyReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(32, 32, 256)</td>
<td valign="top" align="left">Conv4-C512-S1-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">LeakyReLU</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">(31, 31, 256)</td>
<td valign="top" align="left">Conv4-C1-S1-P1</td>
<td valign="top" align="left">InstNorm</td>
<td valign="top" align="left">Sigmoid</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The input size is height, width, and channel. In the operation, Convk is a convolution layer with kernel size as k, DeConv is the deconvolution layer. Ck, Sk, and Pk denote the number of channels, stride, and padding, respectively. Instance normalization (InstNorm) is used. We utilize nine residual blocks (He et al., <xref ref-type="bibr" rid="B12">2016</xref>) to extract necessary information from the image while no residual block is used in the mask encoder as the mask is much simpler than the image</italic>.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.2. Ablation Study</title>
<p><bold>FIDs and Visualization</bold>. In this subsection, we analyze the impact of the style-consistent loss by changing the value of &#x003BB;<sub><italic>sty</italic></sub> in Equation 11. We use Fr&#x000E9;chet inception distance (FID) (Heusel et al., <xref ref-type="bibr" rid="B13">2017</xref>) to show its impact, one of the popular methods to access the quality of the generated images by computing the distance between two images distributions, real images, and the translated images. In general, the lower the FID value, the closer distance between the distributions. We point out that FID is not suitable to access the generated images by our SCIT model since we assume that the real images in the target domain are not available. But the FID can be used to show the tendency between our generated images and all available data that we have when &#x003BB;<sub><italic>sty</italic></sub> changes.</p>
<p>To compute the FID, all original images including training and testing images are leveraged as real images while the generated images are taken as fake images. We borrow code<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> to compute FID. <xref ref-type="table" rid="T3">Table 3</xref> shows the FID values for each class, &#x003BB;<sub><italic>sty</italic></sub> ranging from 0 to 4. From the table, we observe that FID tends to be larger as &#x003BB;<sub><italic>sty</italic></sub> ranges from 0.5 to 4, which proves that the generated images are farther from the real images when we try to keep the style in the image translation process. Moreover, the performance of &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0 differs in different classes. It shows lower FIDs in LMold and MagDef but higher FIDs in translated powdery mildew.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>The impact of &#x003BB;<sub><italic>sty</italic></sub> on FIDs. For each &#x003BB;<sub><italic>sty</italic></sub>, we execute five times and report the mean and SD.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0</bold></th>
<th valign="top" align="center"><bold>&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0.5</bold></th>
<th valign="top" align="center"><bold>&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 1</bold></th>
<th valign="top" align="center"><bold>&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 2</bold></th>
<th valign="top" align="center"><bold>&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 4</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Powdery</td>
<td valign="top" align="center">70.3 &#x000B1; 1.35</td>
<td valign="top" align="center">60.2 &#x000B1; 1.41</td>
<td valign="top" align="center">63.0 &#x000B1; 1.40</td>
<td valign="top" align="center">65.7 &#x000B1; 0.60</td>
<td valign="top" align="center">67.3 &#x000B1; 0.84</td>
</tr>
<tr>
<td valign="top" align="left">Canker</td>
<td valign="top" align="center">97.5 &#x000B1; 2.36</td>
<td valign="top" align="center">93.4 &#x000B1; 2.90</td>
<td valign="top" align="center">93.5 &#x000B1; 1.73</td>
<td valign="top" align="center">98.0 &#x000B1; 2.81</td>
<td valign="top" align="center">99.3 &#x000B1; 1.93</td>
</tr>
<tr>
<td valign="top" align="left">LMold</td>
<td valign="top" align="center">97.9 &#x000B1; 3.09</td>
<td valign="top" align="center">98.2 &#x000B1; 2.97</td>
<td valign="top" align="center">100.8 &#x000B1; 2.34</td>
<td valign="top" align="center">103.1 &#x000B1; 3.32</td>
<td valign="top" align="center">112.6 &#x000B1; 5.14</td>
</tr>
<tr>
<td valign="top" align="left">ToCV</td>
<td valign="top" align="center">79.6 &#x000B1; 5.95</td>
<td valign="top" align="center">72.7 &#x000B1; 6.68</td>
<td valign="top" align="center">75.7 &#x000B1; 5.35</td>
<td valign="top" align="center">86.6 &#x000B1; 3.52</td>
<td valign="top" align="center">93.0 &#x000B1; 2.14</td>
</tr>
<tr>
<td valign="top" align="left">MagDef</td>
<td valign="top" align="center">83.2 &#x000B1; 4.10</td>
<td valign="top" align="center">83.1 &#x000B1; 1.71</td>
<td valign="top" align="center">85.1 &#x000B1; 1.77</td>
<td valign="top" align="center">90.2 &#x000B1; 2.83</td>
<td valign="top" align="center">89.5 &#x000B1; 3.12</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="F5">Figure 5</xref> shows the generated image with different &#x003BB;<sub><italic>sty</italic></sub>. The visual comparisons in the figure comply with the FID values in <xref ref-type="table" rid="T3">Table 3</xref>. First, the generated images with &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0 are clear to show the corresponding class, but the style is far from the input images. Further, the style is better to be maintained when &#x003BB;<sub><italic>sty</italic></sub> becomes larger while the abnormal severity tends to be less. We argue that the variation of the severity contributes to improving task performance. As the collected data are limited to variations, the FID in big &#x003BB;<sub><italic>sty</italic></sub> tends to be worst for some classes, such as in Lmold and MagDef. But when the variations in the collected data are bigger, such as powdery mildew, our SCIT model tends to be better.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Some generated images for image translation with different &#x003BB;<sub><italic>sty</italic></sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0005.tif"/>
</fig>
<p><bold>Image classification</bold>. Except for FID and visualization, we conduct the ablation study on the image classification task. To compare the accuracy in image classification, three categories of classification models are utilized for our tomato leaves. VGG, ResNet, and DenseNet are often deployed in applications with big scale datasets, while MobileNet and ShuffleNet aim to save computations for mobile devices. MNASNet is designed to find the optimal model setting. As our dataset is not big, smaller architecture is the default for all models. On the other hand, since our main objective is data augmentation, we use an open code<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> to produce it. For each model, we execute five times independently for each augmented data. All models are trained for 400 epochs and the best performance is recorded. The initial learning rate is set as 0.02 and decreased to 0.01 and 0.005 at epoch 50 and 200 respectively. SGD is the optimizer with 0.9 as the momentum and batch size is 64 using one GPU. <xref ref-type="table" rid="T4">Table 4</xref> displays the comparison results. The table shows that the performance varies with different &#x003BB;<sub><italic>sty</italic></sub>, &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 2 showing its superiority to other values except in MNASNet, and controlling the style tends to be better than without controlling the style, such as &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0 resulting in less accuracy than other values of &#x003BB;<sub><italic>sty</italic></sub> with model ResNet, DenseNet, MobileNet, and ShuffleNet. It validates that our model, style controlling, is reliable to dedicate the classification in the applications with a small dataset. Moreover, the classification accuracy changes with different models. We argue that it is related to the model itself. As ResNet has fewer parameters and more powerful architecture than VGG, the performances in ResNet are better than VGG. DenseNet, an advanced version of ResNet, also obtains decent results. While MobileNet receives competing results, MNASNet behaves much lower than other methods. We guess that the optimal model setting of MNASNet learned from other datasets is not suitable for our tomato dataset.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Ablation study of &#x003BB;<sub><italic>sty</italic></sub> in image classification.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>VGG11</bold></th>
<th valign="top" align="center"><bold>ResNet18</bold></th>
<th valign="top" align="center"><bold>DenseNet121</bold></th>
<th valign="top" align="center"><bold>MobileNet v2</bold></th>
<th valign="top" align="center"><bold>ShuffleNet v2</bold></th>
<th valign="top" align="center"><bold>MNASNet</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0</td>
<td valign="top" align="center">89.04 &#x000B1; 2.12</td>
<td valign="top" align="center">92.68 &#x000B1; 0.63</td>
<td valign="top" align="center">96.01 &#x000B1; 0.32</td>
<td valign="top" align="center">94.14 &#x000B1; 0.23</td>
<td valign="top" align="center">79.56 &#x000B1; 0.42</td>
<td valign="top" align="center">67.65 &#x000B1; 7.01</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 0.5</td>
<td valign="top" align="center">88.03 &#x000B1; 3.75</td>
<td valign="top" align="center">94.24 &#x000B1; 0.55</td>
<td valign="top" align="center">96.25 &#x000B1; 0.46</td>
<td valign="top" align="center">95.56 &#x000B1; 0.32</td>
<td valign="top" align="center">80.37 &#x000B1; 0.19</td>
<td valign="top" align="center">65.08 &#x000B1; 6.65</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 1</td>
<td valign="top" align="center">93.82 &#x000B1; 2.34</td>
<td valign="top" align="center">94.10 &#x000B1; 0.64</td>
<td valign="top" align="center">96.24 &#x000B1; 0.48</td>
<td valign="top" align="center">96.20 &#x000B1; 0.28</td>
<td valign="top" align="center">86.46 &#x000B1; 0.69</td>
<td valign="top" align="center" style="color:#ee1c23">73.11 &#x000B1; 4.52</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 2</td>
<td valign="top" align="center" style="color:#ee1c23">94.36 &#x000B1; 2.63</td>
<td valign="top" align="center" style="color:#ee1c23">94.61 &#x000B1; 0.38</td>
<td valign="top" align="center" style="color:#ee1c23">96.48 &#x000B1; 0.44</td>
<td valign="top" align="center" style="color:#ee1c23">96.28 &#x000B1; 0.32</td>
<td valign="top" align="center" style="color:#ee1c23">87.44 &#x000B1; 0.69</td>
<td valign="top" align="center">69.45 &#x000B1; 7.48</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 4</td>
<td valign="top" align="center">86.17 &#x000B1; 3.12</td>
<td valign="top" align="center">93.96 &#x000B1; 0.76</td>
<td valign="top" align="center">95.32 &#x000B1; 0.31</td>
<td valign="top" align="center">95.57 &#x000B1; 0.49</td>
<td valign="top" align="center">82.35 &#x000B1; 0.79</td>
<td valign="top" align="center">64.52 &#x000B1; 6.79</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The red font shows the best accuracy for each classification model</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p><bold>Verification of SCIT</bold>. As shown in <xref ref-type="table" rid="T5">Table 5</xref>, performance is compared using different training datasets in three popular networks for image classification to check how the augmented data by our SCIT model work as training data. For this experiment, we use the combinations of three datasets as training datasets. Let <inline-formula><mml:math id="M21"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the <underline>o</underline>riginal training data, <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the training data by our image <underline>t</underline>ranslation SCIT method, and <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the augmented training data by basic image <underline>m</underline>anipulation. <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> individually shows potential with an average accuracy of 79.84 on the testing data, even though its performance is worse than <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. One of the reasons is that there are some translation failures in our SCIT model, as shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, which is common with deep learning-based image generation, such as BigGAN (Brock et al., <xref ref-type="bibr" rid="B3">2018</xref>). However, <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> coupled with <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> improves performance averagely by 3.58% over three classification models. Hence, we conclude that our SCIT model can ease downstream applications to be a data augmentation method.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Verification of SCIT.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Training data</bold></th>
<th valign="top" align="center"><bold>ResNet18</bold></th>
<th valign="top" align="center"><bold>DenseNet121</bold></th>
<th valign="top" align="center"><bold>MobileNet v2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M14"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="center">86.94 &#x000B1; 0.69</td>
<td valign="top" align="center">91.21 &#x000B1; 0.72</td>
<td valign="top" align="center">85.27 &#x000B1; 0.47</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="center">78.24 &#x000B1; 1.05</td>
<td valign="top" align="center">81.38 &#x000B1; 1.06</td>
<td valign="top" align="center">79.90 &#x000B1; 1.51</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="center">90.01 &#x000B1; 0.42</td>
<td valign="top" align="center">94.06 &#x000B1; 0.41</td>
<td valign="top" align="center">92.37 &#x000B1; 0.54</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="center">94.61 &#x000B1; 0.38</td>
<td valign="top" align="center">96.48 &#x000B1; 0.44</td>
<td valign="top" align="center">96.28 &#x000B1; 0.32</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>We show the performance of three popular classifier models according to differently combined training datasets with averages and standard deviations. <inline-formula><mml:math id="M18"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the <underline>o</underline>riginal training dataset, <inline-formula><mml:math id="M19"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the augmented training data by basic image <underline>m</underline>anipulation, and <inline-formula><mml:math id="M20"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the training data by our image <underline>t</underline>ranslation method</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Some translation failure examples from our algorithm.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0006.tif"/>
</fig></sec>
<sec>
<title>4.3. Image Classification</title>
<p><bold>Compared algorithms</bold>. To validate our algorithm, we compare it to other image generation methods using GANs, as GANs have good reputations to produce clear images. The following methods are compared:</p>
<list list-type="bullet">
<list-item><p><italic>Baseline</italic>. We adopt the image manipulation-based data augmentation with the original training dataset. We argue that our data augmentation is complementary to other data augmentation methods. In other words, taking more advanced data augmentation methods such as mixup (Zhang et al., <xref ref-type="bibr" rid="B36">2018</xref>) and cut and paste (Dwibedi et al., <xref ref-type="bibr" rid="B6">2017</xref>) could be a stronger baseline for our future study.</p></list-item>
<list-item><p><italic>DCGAN</italic> (Radford et al., <xref ref-type="bibr" rid="B28">2015</xref>). DCGAN-based (Pandian et al., <xref ref-type="bibr" rid="B25">2019</xref>) or label-condition GANs (Valerio Giuffrida et al., <xref ref-type="bibr" rid="B33">2017</xref>; Pandian et al., <xref ref-type="bibr" rid="B25">2019</xref>; Bi and Hu, <xref ref-type="bibr" rid="B2">2020</xref>; Abbas et al., <xref ref-type="bibr" rid="B1">2021</xref>) are GAN-based algorithms to do data augmentation in which the generator produces images from random noises or given labels. We adopt the original DCGAN to do data augmentation and to produce a higher resolution, two more upsampling layers and convolution layers are added to the original DCGAN.</p></list-item>
<list-item><p><italic>LeafGAN</italic> (Cap et al., <xref ref-type="bibr" rid="B4">2020</xref>). LeafGAN aims to keep the background, one of the challenges of CycleGAN, and introduces an attention module to distinguish the foreground and background. Performance in the cucumber dataset shows its superiority over than original CycleGAN.</p></list-item>
<list-item><p><italic>CycleGAN*</italic>. To get a stronger comparison, we updated CycleGAN. CycleGAN* directly reuses the background from the input image with a mask as input. We emphasize that CycleGAN* is the same as the LeafGAN when the mask is given. Simultaneously, our algorithm degrades to CycleGAN* without keeping the style of the input image during the image translation process.</p></list-item>
</list>
<p><bold>Quantitive comparisons</bold>. For each method, five independent training processes are performed. Except for the baseline, all methods are utilized to generate images with resolution 256 in width and height. The number of generated images is also the same except for the baseline. <xref ref-type="table" rid="T6">Table 6</xref> shows the performances of image classification for tomato leaves with different deep-learning models. From the table, CycleGAN* and our model can significantly improve the classification accuracy. Simultaneously, our data augmentation method achieves the best accuracy and F1 score overall models, which suggests that choosing a good data augmentation method is the way to obtain a better result. In contrast, DCGAN and LeafGAN can not always boost performance. Moreover, each data augmentation method shows the best performance with DenseNet and the worst with ShuffleNet, which suggests that the classification model is also essential and we should choose a better model in our own applications.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Comparison results to other methods to perform image classification for tomato leaves with multiple deep learning-based models.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="left"><bold>Training data</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>F1 Score</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">VGG11</td>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="center">78.72 &#x000B1; 3.78</td>
<td valign="top" align="center">98.95 &#x000B1; 0.12</td>
<td valign="top" align="center">86.31 &#x000B1; 5.21</td>
<td valign="top" align="center">87.42 &#x000B1; 1.51</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DCGAN</td>
<td valign="top" align="center">78.34 &#x000B1; 6.36</td>
<td valign="top" align="center">98.30 &#x000B1; 0.78</td>
<td valign="top" align="center">86.30 &#x000B1; 4.67</td>
<td valign="top" align="center">87.46 &#x000B1; 1.31</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">LeafGAN</td>
<td valign="top" align="center">87.73 &#x000B1; 2.47</td>
<td valign="top" align="center">99.06 &#x000B1; 0.05</td>
<td valign="top" align="center">94.03 &#x000B1; 1.37</td>
<td valign="top" align="center">89.28 &#x000B1; 0.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CycleGAN*</td>
<td valign="top" align="center">89.04 &#x000B1; 2.12</td>
<td valign="top" align="center">99.21 &#x000B1; 0.33</td>
<td valign="top" align="center">95.37 &#x000B1; 1.55</td>
<td valign="top" align="center">89.81 &#x000B1; 0.55</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>94.36</bold> <bold>&#x000B1;2.63</bold></td>
<td valign="top" align="center"><bold>99.46</bold> <bold>&#x000B1;0.13</bold></td>
<td valign="top" align="center"><bold>96.50</bold> <bold>&#x000B1;2.03</bold></td>
<td valign="top" align="center"><bold>90.14</bold> <bold>&#x000B1;0.51</bold></td>
</tr>
<tr>
<td valign="top" align="left">ResNet18</td>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="center">86.75 &#x000B1; 0.81</td>
<td valign="top" align="center">99.07 &#x000B1; 0.43</td>
<td valign="top" align="center">92.06 &#x000B1; 0.61</td>
<td valign="top" align="center">88.80 &#x000B1; 0.83</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DCGAN</td>
<td valign="top" align="center">82.52 &#x000B1; 0.99</td>
<td valign="top" align="center">98.77 &#x000B1; 0.29</td>
<td valign="top" align="center">89.44 &#x000B1; 0.74</td>
<td valign="top" align="center">88.17 &#x000B1; 0.59</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">LeafGAN</td>
<td valign="top" align="center">90.36 &#x000B1; 0.82</td>
<td valign="top" align="center">99.27 &#x000B1; 0.31</td>
<td valign="top" align="center">94.49 &#x000B1; 0.56</td>
<td valign="top" align="center">89.93 &#x000B1; 0.39</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CycleGAN*</td>
<td valign="top" align="center">92.68 &#x000B1; 0.63</td>
<td valign="top" align="center">99.35 &#x000B1; 0.12</td>
<td valign="top" align="center">95.50 &#x000B1; 0.33</td>
<td valign="top" align="center">90.33 &#x000B1; 0.22</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>94.61</bold> <bold>&#x000B1;0.38</bold></td>
<td valign="top" align="center"><bold>99.68</bold> <bold>&#x000B1;0.07</bold></td>
<td valign="top" align="center"><bold>96.59</bold> <bold>&#x000B1;0.12</bold></td>
<td valign="top" align="center"><bold>90.65</bold> <bold>&#x000B1;0.19</bold></td>
</tr>
<tr>
<td valign="top" align="left">DenseNet121</td>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="center">90.97 &#x000B1; 0.40</td>
<td valign="top" align="center">99.59 &#x000B1; 0.27</td>
<td valign="top" align="center">94.94 &#x000B1; 0.27</td>
<td valign="top" align="center">90.16 &#x000B1; 0.40</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DCGAN</td>
<td valign="top" align="center">90.45 &#x000B1; 1.28</td>
<td valign="top" align="center">99.63 &#x000B1; 0.22</td>
<td valign="top" align="center">94.52 &#x000B1; 0.74</td>
<td valign="top" align="center">89.97 &#x000B1; 0.39</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">LeafGAN</td>
<td valign="top" align="center">92.02 &#x000B1; 0.75</td>
<td valign="top" align="center"><bold>99.77</bold> <bold>&#x000B1;0.15</bold></td>
<td valign="top" align="center">95.57 &#x000B1; 0.59</td>
<td valign="top" align="center">90.37 &#x000B1; 0.50</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CycleGAN*</td>
<td valign="top" align="center">96.01 &#x000B1; 0.32</td>
<td valign="top" align="center">99.44 &#x000B1; 0.22</td>
<td valign="top" align="center">97.14 &#x000B1; 0.22</td>
<td valign="top" align="center">90.16 &#x000B1; 0.62</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>96.48</bold> <bold>&#x000B1;0.44</bold></td>
<td valign="top" align="center">99.61 &#x000B1; 0.14</td>
<td valign="top" align="center"><bold>97.87</bold> <bold>&#x000B1;0.40</bold></td>
<td valign="top" align="center"><bold>90.76</bold> <bold>&#x000B1;0.31</bold></td>
</tr>
<tr>
<td valign="top" align="left">MobileNet v2</td>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="center">84.37 &#x000B1; 2.10</td>
<td valign="top" align="center">99.05 &#x000B1; 0.35</td>
<td valign="top" align="center">90.79 &#x000B1; 1.34</td>
<td valign="top" align="center">88.35 &#x000B1; 0.90</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DCGAN</td>
<td valign="top" align="center">82.80 &#x000B1; 1.69</td>
<td valign="top" align="center">98.94 &#x000B1; 0.37</td>
<td valign="top" align="center">89.68 &#x000B1; 1.12</td>
<td valign="top" align="center">88.33 &#x000B1; 0.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">LeafGAN</td>
<td valign="top" align="center">91.31 &#x000B1; 0.53</td>
<td valign="top" align="center">99.31 &#x000B1; 0.19</td>
<td valign="top" align="center">95.17 &#x000B1; 0.30</td>
<td valign="top" align="center">89.94 &#x000B1; 0.39</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CycleGAN*</td>
<td valign="top" align="center">94.14 &#x000B1; 0.23</td>
<td valign="top" align="center"><bold>99.44</bold> <bold>&#x000B1;0.39</bold></td>
<td valign="top" align="center">96.59 &#x000B1; 0.77</td>
<td valign="top" align="center"><bold>90.22</bold> <bold>&#x000B1;0.42</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>96.28</bold> <bold>&#x000B1;0.32</bold></td>
<td valign="top" align="center">99.36 &#x000B1; 0.22</td>
<td valign="top" align="center"><bold>97.53</bold> <bold>&#x000B1;0.30</bold></td>
<td valign="top" align="center">90.10 &#x000B1; 0.52</td>
</tr>
<tr>
<td valign="top" align="left">ShuffleNet v2</td>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="center">71.82 &#x000B1; 0.66</td>
<td valign="top" align="center">98.98 &#x000B1; 0.35</td>
<td valign="top" align="center">84.76 &#x000B1; 0.31</td>
<td valign="top" align="center">88.12 &#x000B1; 0.38</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DCGAN</td>
<td valign="top" align="center">78.05 &#x000B1; 1.82</td>
<td valign="top" align="center">98.80 &#x000B1; 0.18</td>
<td valign="top" align="center">86.43 &#x000B1; 1.10</td>
<td valign="top" align="center">87.47 &#x000B1; 0.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">LeafGAN</td>
<td valign="top" align="center">77.84 &#x000B1; 2.26</td>
<td valign="top" align="center"><bold>99.15</bold> <bold>&#x000B1;0.31</bold></td>
<td valign="top" align="center">88.89 &#x000B1; 1.45</td>
<td valign="top" align="center">88.97 &#x000B1; 0.70</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CycleGAN*</td>
<td valign="top" align="center">79.56 &#x000B1; 0.42</td>
<td valign="top" align="center">99.12 &#x000B1; 0.11</td>
<td valign="top" align="center">89.54 &#x000B1; 0.48</td>
<td valign="top" align="center"><bold>89.92</bold> <bold>&#x000B1;0.17</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center"><bold>87.44</bold> <bold>&#x000B1;0.69</bold></td>
<td valign="top" align="center"><bold>99.15</bold> <bold>&#x000B1;0.46</bold></td>
<td valign="top" align="center"><bold>92.43</bold> <bold>&#x000B1;0.87</bold></td>
<td valign="top" align="center">89.76 &#x000B1; 0.69</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The boldfaces show the best average evaluation metric for each classification model</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p><bold>Qualitative results</bold>. <xref ref-type="fig" rid="F7">Figure 7</xref> shows several generated samples from each algorithm. First, DCGAN can learn similar patterns such as the white part for the powdery and yellow part for LMold, but produces poor images and, in ToCV, fails. The visual results verify its impact in <xref ref-type="table" rid="T6">Table 6</xref>. Moreover, LeafGAN gives plausible images but its attention module is hard to find decent objects to be translated. Hence, it tends to change the background and fails to do image translation such as the canker image. In contrast, CycleGAN*, an advanced LeafGAN with a perfect attention module, achieves much better results and hence boosts the classification performances. Furthermore, our method, adopting a style loss to maintain the style during the image translation and hence taking the variations from the source domain to the target domain, obtains decent visual images.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Qualitative results of different algorithms to do image translation for image classification.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0007.tif"/>
</fig></sec>
<sec>
<title>4.4. Object Detection and Instance Segmentation</title>
<p>As discussed before, our algorithm reuses the annotation for object detection and instance segmentation. As DCGAN is not image translation-based and LeafGAN can not maintain the annotations, we compare our algorithm to CycleGAN* in this subsection. To do object detection and instance segmentation, mmDetection<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> is borrowed as it supports many models. For object detection, we leveraged FasterRCNN, MaskRCNN, and PointRend. In a different paradigm, YOLO aiming to achieve high-speed object detection is also used. Except for object detection, MaskRCNN and PointRend are deployed to do instance segmentation. As &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 2 shows its superiority in classification, we use it as the default value and compare it to baseline and CycleGAN*. For each model, we also execute it five times separately and compute the average performance.</p>
<p><xref ref-type="table" rid="T7">Table 7</xref> displays the comparison. We observe that while the CycleGAN* improves the performance, our method gives more improvements. Except for AP50 in YOLO-v3, our model achieves the best mAP and AP50. Besides, we find that FasterRCNN, MaskRCNN, and PointRend obtain similar results and better results than YOLO-v3 to do object detection. <xref ref-type="fig" rid="F8">Figure 8</xref> displays some generated samples as data augmentation for object detection and instance segmentation, in which one image includes more than one healthy leaf but we choose one of them to do data augmentation with necessary annotations as output. <xref ref-type="fig" rid="F9">Figure 9</xref> illustrates several samples of instance segmentation results using our SCIT model as a data augmentation method. As shown in the figure, the predicted results are highly competent to the ground truth.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Performance of object detection and instance segmentation for tomato leaves in multiple deep learning-based models.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Baseline</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>CycleGAN*</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Ours</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>mAP</bold></th>
<th valign="top" align="center"><bold>AP50</bold></th>
<th valign="top" align="center"><bold>mAP</bold></th>
<th valign="top" align="center"><bold>AP50</bold></th>
<th valign="top" align="center"><bold>mAP</bold></th>
<th valign="top" align="center"><bold>AP50</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">FasterRCNN</td>
<td valign="top" align="center">49.5 &#x000B1; 0.007</td>
<td valign="top" align="center">76.8 &#x000B1; 0.010</td>
<td valign="top" align="center">50.7 &#x000B1; 0.015</td>
<td valign="top" align="center">77.2 &#x000B1; 0.012</td>
<td valign="top" align="center" style="color:#ee1c23">51.5 &#x000B1; 0.019</td>
<td valign="top" align="center" style="color:#2e3092">77.4 &#x000B1; 0.015</td>
</tr>
<tr>
<td valign="top" align="left">MaskRCNN</td>
<td valign="top" align="center">52.4 &#x000B1; 0.014</td>
<td valign="top" align="center">79.6 &#x000B1; 0.015</td>
<td valign="top" align="center">55.6 &#x000B1; 0.010</td>
<td valign="top" align="center">80.4 &#x000B1; 0.011</td>
<td valign="top" align="center" style="color:#ee1c23">56.6 &#x000B1; 0.005</td>
<td valign="top" align="center" style="color:#2e3092">80.5 &#x000B1; 0.007</td>
</tr>
<tr>
<td valign="top" align="left">PointRend</td>
<td valign="top" align="center">51.7 &#x000B1; 0.009</td>
<td valign="top" align="center">79.4 &#x000B1; 0.006</td>
<td valign="top" align="center">52.8 &#x000B1; 0.007</td>
<td valign="top" align="center">80.9 &#x000B1; 0.011</td>
<td valign="top" align="center" style="color:#ee1c23">53.4 &#x000B1; 0.007</td>
<td valign="top" align="center" style="color:#2e3092">81.1 &#x000B1; 0.011</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-v3</td>
<td valign="top" align="center">29.5 &#x000B1; 0.007</td>
<td valign="top" align="center">58.7 &#x000B1; 0.013</td>
<td valign="top" align="center">31.2 &#x000B1; 0.015</td>
<td valign="top" align="center">63.2 &#x000B1; 0.025</td>
<td valign="top" align="center" style="color:#ee1c23">32.6 &#x000B1; 0.013</td>
<td valign="top" align="center" style="color:#2e3092">65.6 &#x000B1; 0.024</td>
</tr>
<tr>
<td valign="top" align="left">MaskRCNN</td>
<td valign="top" align="center">62.6 &#x000B1; 0.015</td>
<td valign="top" align="center" style="color:#2e3092">80.1 &#x000B1; 0.009</td>
<td valign="top" align="center">66.6 &#x000B1; 0.006</td>
<td valign="top" align="center">80.0 &#x000B1; 0.012</td>
<td valign="top" align="center" style="color:#ee1c23">67.1 &#x000B1; 0.010</td>
<td valign="top" align="center">79.9 &#x000B1; 0.007</td>
</tr>
<tr>
<td valign="top" align="left">PointRend</td>
<td valign="top" align="center">56.1 &#x000B1; 0.023</td>
<td valign="top" align="center">80.6 &#x000B1; 0.007</td>
<td valign="top" align="center">67.6 &#x000B1; 0.007</td>
<td valign="top" align="center">81.0 &#x000B1; 0.010</td>
<td valign="top" align="center" style="color:#ee1c23">68.3 &#x000B1; 0.006</td>
<td valign="top" align="center" style="color:#2e3092">81.3 &#x000B1; 0.008</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Four models have been used to do object detection, while two to do instance segmentation. Red font is utilized to show the best average mAP for each model while <inline-formula><mml:math id="M71"><mml:mrow><mml:mstyle mathcolor="#2e3092"><mml:mi>blue</mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> is adopted to show the best average AP50. In general, higher mAP and higher AP50 are better</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Some generated samples to do data augmentation for object detection and instance segmentation in which one image could include more than one leaves but we can just translate the desired one and maintain the others.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0008.tif"/>
</fig>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Result samples of instance segmentation in different diseases with our SCIT model as data augmentation model and &#x003BB;<sub><italic>sty</italic></sub> &#x0003D; 2. The predicted results are decent compared to the ground truth. Zoom in to see the bounding boxes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-773142-g0009.tif"/>
</fig></sec></sec>
<sec id="s5">
<title>5. Conclusion and Future Work</title>
<p>In this study, we introduced a new data augmentation method to improve the abnormality recognition for tomato leaves, termed SCIT which aims to keep the image style from the source image when doing image translation. Armed with this data augmentation paradigm, the data variation in the variation-minority classes is enlarged by the variation-majority class. Simultaneously, we extended the data augmentation from image classification to object detection and instance segmentation which is more competing to do downstream applications. Experimental results validated that the proposed data augmentation method outperforms the baseline and popular methods, in image classification, object detection, and instance segmentation. Although our algorithm was verified to be useful, our future study includes how to integrate different types of data augmentation methods to facilitate the data-hungry deep learning methods, such as advanced image manipulations (mixup Zhang et al., <xref ref-type="bibr" rid="B36">2018</xref> and cut and paste Dwibedi et al., <xref ref-type="bibr" rid="B6">2017</xref>). We hope that our study can stimulate the community to use a more powerful data augmentation method to improve the recognition performance for diseases or other abnormalities in the agricultural field where data are hard or expensive to collect.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding authors.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>MX conceived the idea and designed the algorithm, conducted the experiments, and wrote the manuscript. SY supervised the project, analyzed the algorithm, and contributed to part of the writing and overall improvement of the manuscript. AF collected the original images and performed the preliminary experiment on object detection and revised the manuscript. JY enriched the idea, contributed to data annotation, and improved the manuscript. DP conceptualized the paper, supervised the project, and got funding. All authors read and approved the manuscript.</p></sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This research was supported by the Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (No. 2019R1A6A1A09031717); by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (NRF-2021R1A2C1012174); and by Korea Institute of Planning and Evaluation for Technology in Food, Agriculture and Forestry (IPET) and Korea Smart Farm R&#x00026;D Foundation (KosFarm) through Smart Farm Innovation Technology Development Program, funded by Ministry of Agriculture, Food and Rural Affairs (MAFRA) and Ministry of Science and ICT (MSIT), Rural Development Administration (RDA) (421027-04).</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec> </body>
<back>
<ack><p>We thank Yao Meng for the discussion on instance segmentation and preliminary results about instance segmentation. We also appreciate Ruihan Ma for her suggestions and image annotations in healthy tomato leaves.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abbas</surname> <given-names>A.</given-names></name> <name><surname>Jain</surname> <given-names>S.</given-names></name> <name><surname>Gour</surname> <given-names>M.</given-names></name> <name><surname>Vankudothu</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Tomato plant disease detection using transfer learning with c-gan synthetic images</article-title>. <source>Comput. Electron. Agric</source>. <volume>187</volume>:<fpage>106279</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.106279</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bi</surname> <given-names>L.</given-names></name> <name><surname>Hu</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Improving image-based plant disease classification with generative adversarial network under limited training set</article-title>. <source>Front. Plant Sci</source>. <volume>11</volume>:<fpage>583438</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.583438</pub-id><pub-id pub-id-type="pmid">33343595</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brock</surname> <given-names>A.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Simonyan</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Large scale gan training for high fidelity natural image synthesis,&#x0201D;</article-title> in <source>International Conference on Learning Representations</source> (<publisher-loc>Stockholm</publisher-loc>).</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cap</surname> <given-names>Q. H.</given-names></name> <name><surname>Uga</surname> <given-names>H.</given-names></name> <name><surname>Kagiwada</surname> <given-names>S.</given-names></name> <name><surname>Iyatomi</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Leafgan: an effective data augmentation method for practical plant disease diagnosis</article-title>. <source>IEEE Trans. Autom. Sci. Eng</source>. <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1109/TASE.2020.3041499</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Douarre</surname> <given-names>C.</given-names></name> <name><surname>Crispim-Junior</surname> <given-names>C. F.</given-names></name> <name><surname>Gelibert</surname> <given-names>A.</given-names></name> <name><surname>Tougne</surname> <given-names>L.</given-names></name> <name><surname>Rousseau</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Novel data augmentation strategies to boost supervised segmentation of plant disease</article-title>. <source>Comput. Electron. Agric</source>. <volume>165</volume>:<fpage>104967</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.104967</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dwibedi</surname> <given-names>D.</given-names></name> <name><surname>Misra</surname> <given-names>I.</given-names></name> <name><surname>Hebert</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Cut, paste and learn: Surprisingly easy synthesis for instance detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1301</fpage>&#x02013;<lpage>1310</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuentes</surname> <given-names>A.</given-names></name> <name><surname>Yoon</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>S. C.</given-names></name> <name><surname>Park</surname> <given-names>D. S.</given-names></name></person-group> (<year>2017</year>). <article-title>A robust deep-learning-based detector for real-time tomato plant diseases and pests recognition</article-title>. <source>Sensors</source> <volume>17</volume>, <fpage>2022</fpage>. <pub-id pub-id-type="doi">10.3390/s17092022</pub-id><pub-id pub-id-type="pmid">28869539</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>French</surname> <given-names>A. P.</given-names></name> <name><surname>Pound</surname> <given-names>M. P.</given-names></name> <name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Pridmore</surname> <given-names>T. P.</given-names></name> <name><surname>Pieters</surname> <given-names>J. G.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep convolutional neural networks for image-based convolvulus sepium detection in sugar beet fields</article-title>. <source>Plant Methods</source> <volume>16</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1186/s13007-020-00570-z</pub-id><pub-id pub-id-type="pmid">32165909</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Westergaard</surname> <given-names>J. C.</given-names></name> <name><surname>Sundmark</surname> <given-names>E. H. R.</given-names></name> <name><surname>Bagge</surname> <given-names>M.</given-names></name> <name><surname>Liljeroth</surname> <given-names>E.</given-names></name> <name><surname>Alexandersson</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>Automatic late blight lesion recognition and severity quantification based on field imagery of diverse potato genotypes by deep learning</article-title>. <source>Knowl Based Syst</source>. <volume>214</volume>:<fpage>106723</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106723</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gonzalez-Garcia</surname> <given-names>A.</given-names></name> <name><surname>Weijer</surname> <given-names>J. v. D.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Image-to-image translation for cross-domain disentanglement,&#x0201D;</article-title> in <source>Proceedings of the 32nd International Conference on Neural Information Processing Systems</source> (<publisher-loc>Montr&#x000E9;al Canada</publisher-loc>), <fpage>1294</fpage>&#x02013;<lpage>1305</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gorad</surname> <given-names>B.</given-names></name> <name><surname>Kotrappa</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Novel dataset generation for indian brinjal plant using image data augmentation,&#x0201D;</article-title> in <source>IOP Conference Series: Materials Science and Engineering, Vol. 1065</source> (<publisher-loc>Bristol</publisher-loc>: <publisher-name>IOP Publishing</publisher-name>), <fpage>012041</fpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep residual learning for image recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>.<pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heusel</surname> <given-names>M.</given-names></name> <name><surname>Ramsauer</surname> <given-names>H.</given-names></name> <name><surname>Unterthiner</surname> <given-names>T.</given-names></name> <name><surname>Nessler</surname> <given-names>B.</given-names></name> <name><surname>Hochreiter</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;GANs trained by a two time-scale update rule converge to a local nash equilibrium,&#x0201D;</article-title> in <source>Proceedings of the 31st International Conference on Neural Information Processing Systems</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>NIPS</publisher-name>), <fpage>6629</fpage>&#x02013;<lpage>6640</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Xu</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Qian</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;The identification of corn leaf diseases based on transfer learning and data augmentation,&#x0201D;</article-title> in <source>Proceedings of the 2020 3rd International Conference on Computer Science and Software Engineering</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>58</fpage>&#x02013;<lpage>65</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>A.</given-names></name> <name><surname>Tang</surname> <given-names>Y.</given-names></name> <name><surname>Zhuang</surname> <given-names>J.</given-names></name> <name><surname>Hou</surname> <given-names>C.</given-names></name> <name><surname>Tan</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Deep color calibration for uav imagery in crop monitoring using semantic style transfer with local to global attention</article-title>. <source>Int. J. Appl. Earth Observat. Geoinform</source>. <volume>104</volume>:<fpage>102590</fpage>. <pub-id pub-id-type="doi">10.1016/j.jag.2021.102590</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Arbitrary style transfer in real-time with adaptive instance normalization,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1501</fpage>&#x02013;<lpage>1510</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuznichov</surname> <given-names>D.</given-names></name> <name><surname>Zvirin</surname> <given-names>A.</given-names></name> <name><surname>Honen</surname> <given-names>Y.</given-names></name> <name><surname>Kimmel</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Data augmentation for leaf segmentation and counting tasks in rosette plants,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>Long Beach, CA</publisher-loc>).</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>H.-Y.</given-names></name> <name><surname>Tseng</surname> <given-names>H.-Y.</given-names></name> <name><surname>Huang</surname> <given-names>J.-B.</given-names></name> <name><surname>Singh</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>M.-H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Diverse image-to-image translation via disentangled representations,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>35</fpage>&#x02013;<lpage>51</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>N.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Hou</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Demystifying neural style transfer,&#x0201D;</article-title> in <source>Proceedings of the 26th International Joint Conference on Artificial Intelligence</source> (<publisher-loc>Melbourne</publisher-loc>), <fpage>2230</fpage>&#x02013;<lpage>2236</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>H.</given-names></name> <name><surname>Zhou</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Em-ernet for image-based banana disease recognition</article-title>. <source>J. Food Measur. Characterizat</source>. <volume>15</volume>, <fpage>4696</fpage>&#x02013;<lpage>4710</lpage>. <pub-id pub-id-type="doi">10.1007/s11694-021-01043-0</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Tomato diseases and pests detection based on improved yolo v3 convolutional neural network</article-title>. <source>Front. Plant Sci</source>. <volume>11</volume>:<fpage>898</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.00898</pub-id><pub-id pub-id-type="pmid">32612632</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Plant diseases and pests detection based on deep learning: a review</article-title>. <source>Plant Methods</source> <volume>17</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1186/s13007-021-00722-9</pub-id><pub-id pub-id-type="pmid">33627131</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martineau</surname> <given-names>M.</given-names></name> <name><surname>Conte</surname> <given-names>D.</given-names></name> <name><surname>Raveaux</surname> <given-names>R.</given-names></name> <name><surname>Arnault</surname> <given-names>I.</given-names></name> <name><surname>Munier</surname> <given-names>D.</given-names></name> <name><surname>Venturini</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>A survey on image-based insect classification</article-title>. <source>Pattern. Recognit</source>. <volume>65</volume>, <fpage>273</fpage>&#x02013;<lpage>284</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2016.12.020</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nazki</surname> <given-names>H.</given-names></name> <name><surname>Yoon</surname> <given-names>S.</given-names></name> <name><surname>Fuentes</surname> <given-names>A.</given-names></name> <name><surname>Park</surname> <given-names>D. S.</given-names></name></person-group> (<year>2020</year>). <article-title>Unsupervised image translation using adversarial networks for improved plant disease recognition</article-title>. <source>Comput. Electron. Agric</source>. <volume>168</volume>:<fpage>105</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.105117</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pandian</surname> <given-names>J. A.</given-names></name> <name><surname>Geetharamani</surname> <given-names>G.</given-names></name> <name><surname>Annette</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Data augmentation on plant leaf disease image dataset using image manipulation and deep learning techniques,&#x0201D;</article-title> in <source>2019 IEEE 9th International Conference on Advanced Computing (IACC)</source> (<publisher-loc>Tiruchirappalli</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>199</fpage>&#x02013;<lpage>204</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pawara</surname> <given-names>P.</given-names></name> <name><surname>Okafor</surname> <given-names>E.</given-names></name> <name><surname>Schomaker</surname> <given-names>L.</given-names></name> <name><surname>Wiering</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Data augmentation for plant classification,&#x0201D;</article-title> in <source>International Conference on Advanced Concepts for Intelligent Vision Systems</source> (<publisher-loc>Antwerp</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>615</fpage>&#x02013;<lpage>626</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pinto Sampaio Gomes</surname> <given-names>D.</given-names></name> <name><surname>Zheng</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Recent data augmentation strategies for deep learning in plant phenotyping and their significance,&#x0201D;</article-title> in <source>2020 Digital Image Computing: Techniques and Applications (DICTA)</source> (<publisher-loc>Melbourne</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Metz</surname> <given-names>L.</given-names></name> <name><surname>Chintala</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Unsupervised representation learning with deep convolutional generative adversarial networks</article-title>. <source>arXiv preprint</source> arXiv:1511.06434.<pub-id pub-id-type="pmid">33873122</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saranya</surname> <given-names>S. M.</given-names></name> <name><surname>Rajalaxmi</surname> <given-names>R.</given-names></name> <name><surname>Prabavathi</surname> <given-names>R.</given-names></name> <name><surname>Suganya</surname> <given-names>T.</given-names></name> <name><surname>Mohanapriya</surname> <given-names>S.</given-names></name> <name><surname>Tamilselvi</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep learning techniques in tomato plant-a review</article-title>. <source>J. Phys</source>. <volume>1767</volume>, <fpage>012010</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/1767/1/012010</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>M.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Maheshwari</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Cdtd: A large-scale cross-domain benchmark for instance-level image-to-image translation and domain adaptive object detection</article-title>. <source>Int. J. Comput. Vis</source>. <volume>129</volume>, <fpage>761</fpage>&#x02013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-020-01394-z</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shorten</surname> <given-names>C.</given-names></name> <name><surname>Khoshgoftaar</surname> <given-names>T. M.</given-names></name></person-group> (<year>2019</year>). <article-title>A survey on image data augmentation for deep learning</article-title>. <source>J. Big Data</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-019-0197-0</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv[Preprint]</source>. arXiv:1409.1556.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Valerio Giuffrida</surname> <given-names>M.</given-names></name> <name><surname>Scharr</surname> <given-names>H.</given-names></name> <name><surname>Tsaftaris</surname> <given-names>S. A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Arigan: Synthetic arabidopsis plants using generative adversarial network,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision Workshops</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2064</fpage>&#x02013;<lpage>2071</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Early real-time detection algorithm of tomato diseases and pests in the natural environment</article-title>. <source>Plant Methods</source> <volume>17</volume>, <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1186/s13007-021-00745-2</pub-id><pub-id pub-id-type="pmid">33892765</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Fuentes</surname> <given-names>A.</given-names></name> <name><surname>Park</surname> <given-names>D. S.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yoon</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Instance-level image translation with a local discriminator</article-title>. <source>IEEE Access</source>. <volume>9</volume>, <fpage>111802</fpage>&#x02013;<lpage>111813</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3102263</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Cisse</surname> <given-names>M.</given-names></name> <name><surname>Dauphin</surname> <given-names>Y. N.</given-names></name> <name><surname>Lopez-Paz</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;mixup: Beyond empirical risk minimization,&#x0201D;</article-title> in <source>International Conference on Learning Representations</source> (<publisher-loc>Stockholm</publisher-loc>).</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Research on deep learning in apple leaf disease recognition</article-title>. <source>Comput. Electron. Agric</source>. <volume>168</volume>:<fpage>105146</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.105146</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>J.-Y.</given-names></name> <name><surname>Park</surname> <given-names>T.</given-names></name> <name><surname>Isola</surname> <given-names>P.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Unpaired image-to-image translation using cycle-consistent adversarial networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2223</fpage>&#x02013;<lpage>2232</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Aoun</surname> <given-names>M.</given-names></name> <name><surname>Krijn</surname> <given-names>M.</given-names></name> <name><surname>Vanschoren</surname> <given-names>J.</given-names></name> <name><surname>Campus</surname> <given-names>H. T.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Data augmentation using conditional generative adversarial networks for leaf counting in arabidopsis plants,&#x0201D;</article-title> in <source>British Machine Vision Conference</source>, <volume>324</volume>.</citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/xml94/SCIT">https://github.com/xml94/SCIT</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://pytorch.org/vision/stable/models.html&#x00023;torchvision.models.vgg19">https://pytorch.org/vision/stable/models.html&#x00023;torchvision.models.vgg19</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/mseitzer/pytorch-fid">https://github.com/mseitzer/pytorch-fid</ext-link></p></fn>
<fn id="fn0004"><p><sup>4</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/bearpaw/pytorch-classification">https://github.com/bearpaw/pytorch-classification</ext-link></p></fn>
<fn id="fn0005"><p><sup>5</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/open-mmlab/mmdetection">https://github.com/open-mmlab/mmdetection</ext-link></p></fn>
</fn-group>
</back>
</article> 